Manager, Cloud Services and Site Reliability in Canada Creek, Nova Scotia at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Manager, Cloud Services and Site Reliability based in Canada.
This leadership role is responsible for the reliability, availability, scalability, and operational excellence of high-volume, business-critical SaaS applications.
You will lead and develop an SRE team while providing technical direction across cloud infrastructure and production operations.
The role partners closely with engineering, product, platform, and security teams to strengthen resilient and scalable services.
You’ll drive improvements in incident response, automation, monitoring, service health, and operational readiness.
Using data and operational insights, you’ll identify reliability risks and turn them into measurable improvements.
This is an opportunity to shape engineering practices, reduce operational toil, and build a culture of ownership and continuous improvement.
You’ll work in a collaborative, remote-friendly environment where technical expertise and people leadership are equally valued.
- Lead, coach, mentor, and develop a high-performing SRE team, establishing clear expectations and fostering ownership, collaboration, accountability, and continuous improvement.
- Drive reliability engineering practices across critical cloud services, including SLOs, SLIs, monitoring, alerting, capacity planning, and service health reporting.
- Partner with engineering and platform teams to improve the architecture, scalability, resilience, maintainability, and operational performance of cloud-based systems.
- Own and continuously improve incident management practices, including major incident coordination, post-incident reviews, root-cause analysis, and follow-up actions.
- Champion automation, tooling, and engineering practices that reduce manual operational work and improve consistency and scalability.
- Analyze operational data, service metrics, and risk indicators to identify reliability gaps, prioritize improvements, and communicate progress to technical and business stakeholders.
- Partner with security and engineering teams to support secure, compliant, and operationally mature production environments.
- Contribute to disaster recovery, infrastructure automation, CI/CD, cost optimization, and multi-cloud initiatives where relevant.
- Improve documentation, operational readiness, and service management practices across teams.
- Evaluate tools, technologies, and vendors that can strengthen service reliability and operational effectiveness.
- Influence cross-functional stakeholders and help establish a culture focused on resilient systems, measurable service health, and continuous improvement.
- 5+ years of experience in SRE, DevOps, infrastructure, cloud operations, or a related technical operations discipline, including experience leading or managing technical teams.
- Strong understanding of cloud platforms, distributed systems, production operations, and modern site reliability engineering principles.
- Hands-on experience implementing or improving SLOs, SLIs, monitoring, alerting, incident response, and post-incident review processes.
- Demonstrated ability to hire, mentor, coach, and develop engineers while creating a healthy, inclusive, and accountable team culture.
- Strong communication and stakeholder-management skills, with the ability to explain complex technical concepts to engineers, product teams, and business leaders.
- Proven track record of using operational data, structured problem-solving, and technical insight to improve reliability and team effectiveness.
- Experience with infrastructure automation, CI/CD, disaster recovery, cost optimization, or multi-cloud environments.
- Ability to influence operational change across teams and drive adoption of improved processes, documentation, and reliability practices.
- Strong judgment and prioritization skills, with the ability to balance immediate operational needs against longer-term reliability and engineering improvements.
- Comfortable working in a fast-paced environment where collaboration, ownership, and continuous learning are essential.
- Anticipated salary range of $151,000–$200,000, with actual compensation based on skills, experience, qualifications, location, budget, and applicable employment requirements.
- Equity in the form of non-qualifying stock options.
- High-quality health benefits.
- Retirement plan with employer matching.
- Flexible Time Off and Paid Time Off benefits.
- Career development and growth opportunities.
- Internal mobility and cross-training opportunities.
- Volunteer opportunities.
- Remote-friendly work environment.
- Opportunity to make a meaningful impact on the reliability and resilience of business-critical SaaS services.