Software Engineer - Data & Cloud
The Opportunity
Burbio is seeking a talented, growth-minded Software Engineer to help drive the data-collection infrastructure at the center of our business. You will take meaningful ownership of the systems that gather and process millions of pages of public documents, including complex web crawlers, document-processing pipelines, PostgreSQL/Pinecone databases, cloud infrastructure, and APIs.
This is a broad, hands-on role for someone who enjoys solving difficult data-collection problems and wants responsibility beyond a narrowly defined development lane. You will work closely with our development, data, and product teams, contribute to technical decisions, and grow toward a Technical Lead role.
Why Burbio
Burbio is a fast-moving, AI-forward company with an established and respected position in the K-12 market, working some of the largest suppliers of curriculum, equipment, and services in the country. Our technical team continually creates new data-collection methods, technical processes, and client-facing products - from the Construction Tracker and District Contacts to APIs and our MCP integration. We can move from an initial concept to a product in the market in as little as 45 days.
Engineers at Burbio do more than execute specifications; they are expected to bring ideas and technical judgment to product discussions, work closely with company leadership, data teams, and customers, and help determine what we build and how it should work. You will see your ideas move quickly into products that clients purchase while gaining broad experience across data collection, cloud systems, databases, APIs, and emerging AI delivery methods.
Our Technology Ecosystem
- Languages: Python and SQL
- Cloud: Google Cloud Platform, including Cloud Run, virtual machines, cloud jobs, and IAM
- Database: PostgreSQL and Pinecone(Vector Database)
- Data delivery and AI: RESTful APIs, LLM APIs, RAG workflows, and MCP integrations
Key Responsibilities
Web Scraping, Data Collection & Processing
- Own and improve web data collection: Maintain and extend large-scale crawlers and document-processing pipelines that ingest millions of pages from varied public websites.
- Solve collection challenges: Adapt to dynamic sites, rate limits, anti-bot measures, and inconsistent document structures balancing maximum speed and coverage with respect for site terms of service and crawling ethics.
- Strengthen data quality: Build monitoring, validation, deduplication, and recovery processes that keep collection pipelines accurate and dependable.
- Transform unstructured information: Develop pipelines that extract, normalize, and structure web documents for downstream products, databases, and client workflows.
Databases, APIs & Cloud Infrastructure
- Design and optimize Postgres systems: Organize relational schemas, protect data integrity, minimize duplication, and tune complex queries and joins for performance at scale.
- Build APIs and integrations: Develop clean, scalable RESTful endpoints and backend pipelines that deliver structured Burbio data to products, clients, and external systems.
- Operate cloud infrastructure: Manage workloads across GCP virtual machines and Cloud Run, structure serverless jobs, monitor compute performance, and maintain secure IAM permissions.
- Create internal tools: Build lightweight utilities, webhooks, and automation that improve human-in-the-loop quality assurance and team productivity.
Technical Ownership & AI Engineering
- Help shape and launch new products: Contribute technical and product ideas from the earliest stages, evaluate what is possible, and help turn new data capabilities into client-ready products.
- Build with AI: Use LLM APIs, retrieval workflows, agentic workflows and AI-assisted development tools to improve document processing, product capabilities, and engineering velocity.
- Advance engineering practices: Help shape architecture, break down specifications, conduct structured GitHub code reviews, and balance rapid delivery with sustainable code quality.
- Document and share knowledge: Maintain useful technical documentation and architecture diagrams while collaborating with other developers to strengthen the team collectively.
Qualifications
Required
- 2-3 years of professional software engineering experience, preferably in data-heavy or cloud-based applications.
- Strong Python and SQL skills.
- Hands-on experience building or maintaining web scrapers, crawlers, document-processing systems, or comparable data-ingestion pipelines.
- Strong relational database fundamentals, including schema design, data mapping, complex joins, and query optimization.
- Experience with cloud infrastructure on GCP or AWS, serverless deployments, and RESTful API development.
- Working knowledge of Git, GitHub pull-request workflows, and automated testing.
- A systems mindset: the ability to solve immediate problems while improving how code and infrastructure are organized over time.
- Indiana or Kentucky residency.
Helpful, but Not Required
- Experience with PostgreSQL and Google Cloud Platform.
- Experience integrating LLM APIs or building RAG-based document and data workflows.
- Familiarity with Model Context Protocol, function calling, or tool definitions, which will be helpful when collaborating with Burbio team members responsible for MCP integrations.
- Experience using AI-assisted development tools such as GitHub Copilot, Cursor, or agentic coding workflows.
- Familiarity with the structure of the U.S. K-12 education system or the ability to learn a new data domain quickly.
Application Process
Apply through Handshake or email dennis@burbio.com directly.