Senior Software Engineer, Data Infrastructure - AI Platform in Canada Creek, Nova Scotia at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineer, Data Infrastructure - AI Platform based in the Canada.
This is a senior-level engineering role focused on building, optimizing, and operating critical data infrastructure in a highly regulated government cloud environment. You’ll work on a distributed serving layer that delivers fast, reliable data access for mission-critical investigations and applications. The role combines database performance engineering, data pipeline reliability, production operations, and modern AI-assisted development practices. You’ll take meaningful ownership of systems where availability, compliance, and operational resilience are essential. As part of a distributed and highly collaborative engineering team, you’ll work closely with data platform, product, and forward-deployed engineering partners. The environment values technical judgment, speed, experimentation, and end-to-end accountability. This is an opportunity to deepen your expertise in distributed systems while contributing to technology with significant real-world impact.
- Own serving-layer performance: Lead performance tuning for a distributed OLAP serving layer, using query profiling and AI-assisted tools to identify and resolve inefficient query patterns before they become customer-facing issues.
- Build and strengthen data pipelines: Develop and harden reliable pipelines supporting government cloud investigations, emphasizing correctness, scalability, maintainability, and compliance.
- Improve infrastructure resilience: Become an independent operator of critical serving infrastructure, reducing single points of failure and strengthening the team’s ability to respond quickly to production incidents.
- Lead production troubleshooting: Apply AI-assisted debugging, log analysis, code exploration, and systematic investigation techniques to identify root causes and resolve complex production issues efficiently.
- Own operational reliability: Participate in on-call rotations, respond to incidents, improve operational procedures, and ensure lessons from production events are incorporated into runbooks and infrastructure improvements.
- Deliver high-impact infrastructure projects: Take initiatives from technical investigation and design through implementation, testing, deployment, and ongoing operational ownership.
- Support compliance and security requirements: Build and maintain infrastructure capabilities that meet evolving government cloud requirements, including audit, data retention, backup, and reliability standards.
- Collaborate across teams: Work closely with Data Platform, Forward Deployed Engineering, Product, and other engineering teams to maintain reliable capabilities and alignment across government and commercial environments.
- Apply AI to engineering workflows: Use AI tools to accelerate debugging, code review, research, documentation, configuration, and other development workflows while maintaining strong engineering standards.
- Contribute to architecture and technical strategy: Make evidence-based decisions around system architecture, performance, reliability, and trade-offs, while sharing expertise with the broader engineering organization.
- Drive continuous improvement: Participate in sprint planning, team discussions, production retrospectives, and asynchronous collaboration to improve systems, processes, and operational practices.
- U.S. citizenship is required due to government cloud data access requirements.
- Strong hands-on experience operating distributed OLAP, analytical database, or serving-layer systems, such as StarRocks, Trino, ClickHouse, or similar technologies.
- Demonstrated expertise in query tuning, database performance optimization, and distributed systems at scale.
- Experience owning data pipeline reliability, production infrastructure, incident response, and operational troubleshooting.
- Strong understanding of production systems and the ability to independently take ownership of unfamiliar infrastructure with limited oversight.
- Experience with on-call responsibilities and a demonstrated commitment to operational excellence, reliability, and effective incident management.
- Proven ability to use AI-powered engineering tools such as Claude, Cursor, or comparable platforms to accelerate debugging, code review, research, documentation, and development.
- Strong applied AI fluency, with the ability to use AI not only for automation but also to structure problems, improve output quality, and increase engineering leverage.
- Excellent technical judgment and a strong ownership mindset, with the ability to take problems from discovery through resolution and production deployment.
- Strong communication and collaboration skills, particularly in distributed, cross-functional engineering environments.
- Ability to operate effectively in a fast-moving, high-ownership environment where priorities may shift and ambiguity is common.
- Ability to balance rapid execution with high standards for reliability, security, compliance, maintainability, and technical quality.
- Fully remote work for eligible Canada-based employees.
- Opportunity to work on sophisticated data infrastructure supporting mission-critical government cloud applications.
- Hands-on exposure to distributed database technologies, AI-assisted engineering, production systems, and highly regulated cloud environments.
- High-autonomy engineering culture where system operators have significant influence over architecture and technical trade-offs.
- Close collaboration with experienced engineers and cross-functional teams across data platform, product, and forward-deployed engineering.
- Opportunity to work on technically challenging problems with meaningful real-world impact.
- Fast-paced environment emphasizing ownership, experimentation, technical craftsmanship, and continuous learning.
- Exposure to evolving AI engineering practices and an organizational culture where applied AI fluency is considered an important part of technical excellence.
- Distributed-first working environment with strong asynchronous communication and clear ownership across projects.