Staff Data Engineer - AI Platform in Abbeyville, Colorado at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Data Engineer - AI Platform based in the United States.
This is a staff-level engineering opportunity focused on the architecture, reliability, and performance of critical data infrastructure supporting a high-availability government cloud environment. You’ll work on a distributed serving layer that transforms complex datasets into fast, dependable answers for mission-critical investigations. The role combines deep data engineering, distributed systems, database optimization, production operations, and AI-assisted development. You’ll have broad technical ownership and play a key role in strengthening infrastructure resilience, scalability, and compliance. The environment is fast-moving and highly collaborative, with engineers empowered to make technical decisions close to the systems they operate. You’ll partner across data platform, product, and forward-deployed engineering teams while tackling technically challenging problems with real-world impact. This role is ideal for an experienced engineer who thrives on autonomy, ambiguity, high standards, and meaningful technical challenges.
- Lead data infrastructure strategy: Provide technical leadership for the distributed serving layer and help shape its architecture, scalability, performance, reliability, and long-term evolution.
- Own database performance: Drive performance optimization for the serving layer, using advanced query profiling and AI-assisted tooling to identify bottlenecks and resolve inefficient query patterns before they affect customers.
- Build and harden data pipelines: Design, develop, and improve reliable pipelines supporting government cloud investigations, with strong attention to scalability, correctness, maintainability, and compliance.
- Strengthen infrastructure resilience: Reduce single points of failure by becoming a senior independent owner of critical infrastructure and improving system redundancy, operational readiness, and incident response.
- Lead production troubleshooting: Apply sophisticated debugging, log analysis, AI-assisted research, and code exploration to diagnose complex production problems and deliver rapid, durable fixes.
- Drive operational excellence: Participate in on-call responsibilities, improve runbooks and monitoring, lead incident retrospectives, and translate operational learnings into lasting infrastructure improvements.
- Deliver critical infrastructure initiatives: Take complex projects from technical discovery and architectural design through implementation, deployment, validation, and ongoing ownership.
- Support regulatory and compliance needs: Build infrastructure capabilities that satisfy evolving government cloud requirements around availability, auditing, data retention, backups, security, and operational controls.
- Champion AI-enabled engineering: Use AI tools to accelerate development, debugging, code reviews, documentation, configuration, research, and other workflows while maintaining rigorous technical standards.
- Influence technical direction: Make evidence-based architectural and engineering decisions, communicate trade-offs clearly, and contribute technical perspective to broader Data Platform initiatives.
- Mentor and raise engineering standards: Share expertise, improve engineering practices, and help create a culture of strong ownership, craftsmanship, collaboration, and continuous improvement.
- Collaborate across functions: Work closely with Data Platform, Product, Forward Deployed Engineering, and other teams to ensure reliable capabilities and alignment across critical environments.
- U.S. citizenship is required due to government cloud data access requirements.
- Extensive hands-on experience designing, operating, and scaling distributed OLAP, analytical database, or serving-layer systems, including technologies such as StarRocks, Trino, ClickHouse, or comparable platforms.
- Deep expertise in query optimization, database performance, distributed systems, and large-scale data infrastructure.
- Strong track record owning data pipeline reliability, production infrastructure, incident response, and operational excellence.
- Demonstrated ability to take end-to-end ownership of complex infrastructure, from architecture and implementation through production operations.
- Experience independently troubleshooting unfamiliar systems and making sound technical decisions in high-pressure production environments.
- Strong experience with on-call operations, incident management, observability, and reliability engineering practices.
- Advanced practical fluency with AI engineering tools such as Claude, Cursor, or similar platforms, using them to accelerate research, debugging, code review, development, documentation, and problem-solving.
- Ability to use AI strategically to improve engineering quality, speed, leverage, and decision-making rather than simply automate repetitive tasks.
- Strong architectural judgment and the ability to evaluate technical trade-offs across performance, reliability, security, compliance, and maintainability.
- Excellent communication and collaboration skills, with the ability to influence technical direction across teams and explain complex concepts clearly.
- Demonstrated leadership through technical influence, mentorship, and the ability to raise engineering standards without relying solely on formal authority.
- Comfortable operating in a high-velocity, high-ownership environment where priorities evolve and ambiguity is part of the work.
- Strong bias toward action, experimentation, continuous learning, and measurable outcomes.
- Fully remote work for eligible US-based employees.
- Opportunity to provide technical leadership over sophisticated data infrastructure supporting mission-critical government cloud applications.
- Significant autonomy and influence over architecture, infrastructure strategy, technical standards, and operational practices.
- Hands-on exposure to distributed data systems, AI-assisted engineering, production infrastructure, and highly regulated cloud environments.
- Opportunity to solve complex technical problems at the intersection of AI, security, public safety, and mission-critical technology.
- Collaborative distributed-first culture with strong asynchronous communication and close cross-functional partnership.
- Environment that values technical craftsmanship, ownership, experimentation, speed, and continuous improvement.
- Opportunities to mentor other engineers and shape engineering practices across the broader organization.
- Meaningful work with tangible real-world impact and opportunities for continued technical growth.