Staff Site Reliability Engineer in Canada Creek, Nova Scotia at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer based in Canada.
This role offers the opportunity to shape the reliability and scalability of a modern cybersecurity SaaS platform.
You will act as a technical leader, designing resilient systems across cloud and on-premises environments while driving platform engineering excellence.
The position sits at the intersection of software engineering, infrastructure, automation, and operational reliability.
You will help build highly available systems, improve developer velocity, and establish best practices around automation, observability, and security.
As a Staff-level engineer, you will influence engineering strategy, mentor teammates, and solve complex technical challenges at scale.
This is an opportunity to make a meaningful impact in a collaborative environment focused on innovation, continuous improvement, and AI-driven engineering practices.
As a Staff Site Reliability Engineer, you will own the evolution of infrastructure, deployment systems, and reliability practices. You will lead platform initiatives, improve operational excellence, and partner with engineering teams to create scalable, secure, and efficient solutions.
- Design, build, and maintain highly available, secure, and resilient systems across cloud and on-premises environments.
- Lead platform engineering initiatives that simplify development workflows and improve engineering productivity.
- Own and optimize foundational services, including API gateways, service meshes, caching systems, configuration management, and secrets management.
- Establish and promote an “Everything as Code” approach, ensuring infrastructure, configurations, and deployment processes are automated, version-controlled, and managed through GitOps workflows.
- Develop, standardize, and improve CI/CD pipelines to enable secure, reliable, and efficient software delivery.
- Implement infrastructure-as-code practices using tools such as Terraform, OpenTofu, or Ansible to prevent configuration drift.
- Design reliability strategies, including automated testing, chaos engineering practices, and disaster recovery simulations.
- Build and enhance observability capabilities through metrics, logs, traces, and monitoring solutions.
- Define and maintain Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for critical applications.
- Partner with engineering leadership to shape long-term reliability and platform strategies.
- Mentor engineers, promote reliability best practices, and contribute to technical documentation and knowledge sharing.
- Lead incident reviews and help teams continuously improve system resilience.
The ideal candidate is an experienced infrastructure and reliability engineer with strong software engineering knowledge and a passion for building scalable, automated systems. You should have experience leading technical initiatives, improving engineering practices, and operating complex distributed environments.
- 7+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or related disciplines.
- At least 2 years of experience operating at a Senior or Staff engineering level.
- Proven experience designing and managing infrastructure across cloud and on-premises environments.
- Strong experience with cloud platforms such as AWS and/or Azure.
- Hands-on expertise with Docker and Kubernetes, including cluster administration, networking, and security concepts.
- Experience designing release orchestration strategies such as Canary or Blue-Green deployments.
- Proficiency in at least one systems programming language such as Go, Java, or C# for automation, tooling, and internal services.
- Strong understanding of observability practices, including telemetry, monitoring, logging, tracing, and application performance management.
- Experience with GitOps workflows and Infrastructure as Code / Configuration as Code principles.
- Strong knowledge of Linux and Windows systems internals.
- Experience working with continuous integration environments and modern DevSecOps practices.
- Ability to mentor engineers, influence technical direction, and collaborate effectively across teams.
Nice to have:
- Experience with UI automation testing.
- Background designing and supporting microservice-based architectures.
- Experience managing virtual machines and test environments.
- Experience migrating workloads between on-premises and cloud environments.
- Familiarity with advanced security and reliability practices.
- Flexible work environment designed around trust, autonomy, and collaboration.
- Opportunity to work on impactful cybersecurity solutions used by organizations worldwide.
- Competitive compensation package aligned with experience and expertise.
- Professional growth opportunities through continuous learning and technical development.
- Ability to influence engineering strategy and build solutions from the ground up.
- Supportive culture focused on diversity, inclusion, and employee growth.
- Opportunity to collaborate with talented engineers across global teams.
- Environment that encourages innovation, knowledge sharing, and technical leadership.