Senior Site Reliability Engineer in Canada Creek, Nova Scotia at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in Canada.
Join a growing AI/ML organization operating at the intersection of geospatial intelligence and environmental technology. As a Senior Site Reliability Engineer, you will take ownership of cloud infrastructure and help strengthen reliability, observability, and operational excellence across engineering and product teams. You will design and evolve scalable infrastructure on Google Cloud Platform while enabling developers through automation and self-service tooling. Your work will directly influence system availability, incident response, deployment performance, and cloud efficiency. You will champion modern reliability practices, from SLOs and error budgets to DORA metrics and observability. This is a fully remote opportunity offering significant technical ownership in a collaborative, high-impact engineering environment.
- Design, build, and continuously evolve scalable and reliable cloud infrastructure on Google Cloud Platform.
- Develop internal tooling, automation, and self-service capabilities that increase engineering efficiency and reduce operational dependencies.
- Strengthen the observability platform across metrics, logging, distributed tracing, and monitoring to improve system visibility and reduce mean time to recovery.
- Establish infrastructure cost visibility, governance, and optimization initiatives to improve cloud efficiency and manage spending responsibly.
- Define and champion reliability practices including Service Level Objectives (SLOs), Service Level Indicators (SLIs), error budgets, and DORA metrics.
- Lead incident management activities, coordinate response efforts, facilitate post-incident reviews, and drive meaningful improvements based on incident learnings.
- Participate in the on-call rotation and help ensure production systems remain stable, available, and resilient.
- Partner with Product and Engineering teams to improve operational practices, deployment reliability, and overall software delivery performance.
- 3+ years of professional experience in Site Reliability Engineering, production engineering, or a closely related infrastructure role.
- Strong hands-on experience with Google Cloud Platform, including cloud cost optimization, governance, and infrastructure management.
- Proven experience managing Kubernetes clusters and workloads in production environments.
- Experience with Infrastructure as Code tools such as Terraform or Google Cloud Deployment Manager.
- Strong scripting and automation capabilities using Python, Bash, Go, or comparable languages.
- Extensive experience with observability technologies such as Prometheus, Grafana, OpenTelemetry, centralized logging, and distributed tracing.
- Practical experience with incident management, post-incident analysis, production troubleshooting, and on-call operations.
- Ability to define, implement, and monitor SLOs, SLIs, and error budgets.
- Strong understanding of reliability engineering principles and a proactive approach to identifying and resolving operational risks.
- Familiarity with DORA metrics and their application to engineering workflows is an asset.
- Experience working in AI/ML, geospatial technology, or other data-intensive technical environments is considered a plus.
- Strong communication and collaboration skills, with the ability to work effectively across distributed Product and Engineering teams.
- Candidates must be authorized to work in their country of residence; visa sponsorship is not available.
- Fully remote work environment.
- Opportunity to work with modern cloud infrastructure, Kubernetes, Infrastructure as Code, and advanced observability technologies.
- Significant ownership over infrastructure reliability, operational excellence, and cloud optimization initiatives.
- Opportunity to influence engineering practices through SLOs, error budgets, DORA metrics, and incident management.
- Collaboration with cross-functional and distributed engineering teams across North America, Europe, and the UK.
- Exposure to innovative AI/ML and geospatial technology applications.
- Compensation will be discussed during the interview process and will be aligned with experience and geographic location.