Senior Software Engineering Manager, Managed Gateways SREs in Canada Creek, Nova Scotia at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineering Manager, Managed Gateways SREs based in Canada.
This is an exciting opportunity to lead the growth of a high-impact Site Reliability Engineering team responsible for ensuring the reliability, scalability, and performance of mission-critical managed gateway services. You will play a key role in shaping technical strategy, building a strong engineering culture, and driving operational excellence across globally distributed cloud infrastructure. In the early stages, you will combine technical leadership with hands-on engineering, collaborating closely with cross-functional teams to deliver resilient, cloud-native solutions. This role is ideal for an experienced engineering leader who thrives in fast-paced environments and is passionate about automation, reliability, and empowering teams to build highly available platforms at scale.
- Build and lead a high-performing Site Reliability Engineering team, establishing engineering standards, best practices, and a culture of operational excellence.
- Contribute directly to critical infrastructure implementations during the team's growth phase, combining leadership with hands-on technical execution.
- Design, deploy, and optimize highly available, scalable cloud-native infrastructure using Kubernetes and modern cloud technologies.
- Drive reliability initiatives by managing monitoring, alerting, incident response, root cause analysis, and continuous service improvements.
- Develop automation, self-service tooling, and operational processes that improve deployment efficiency and reduce manual effort.
- Define, monitor, and improve service level objectives (SLOs), service level indicators (SLIs), and overall platform reliability.
- Collaborate with engineering, product, and support teams to ensure new features are operationally ready and aligned with long-term platform goals.
- Mentor engineers, support career development, and foster a collaborative, high-performance team environment.
- Proven experience leading Site Reliability Engineering, DevOps, or Infrastructure Engineering teams in fast-paced, high-growth organizations.
- Strong expertise designing, deploying, and operating highly available distributed systems on AWS, Azure, GCP, or similar cloud platforms.
- Extensive hands-on experience with Kubernetes and container orchestration in production environments.
- Proficiency in Go (Golang) or another modern programming language used for infrastructure automation and platform development.
- Solid understanding of observability practices, including experience with tools such as Prometheus, Grafana, OpenTelemetry, or comparable platforms.
- Demonstrated experience managing production incidents, performing root cause analysis, and implementing long-term reliability improvements.
- Familiarity with API gateways, service mesh technologies, or network proxy solutions is considered an asset.
- Strong leadership, communication, collaboration, and stakeholder management skills, with the ability to balance strategic vision and technical execution.
- Cloud or Kubernetes certifications (such as AWS Certified DevOps Engineer, CKA, or CKAD), open-source contributions, or experience in developer infrastructure organizations are advantageous.
- Competitive compensation package.
- Remote-friendly work environment with flexibility.
- Opportunity to build and lead a new engineering team with significant ownership and influence.
- Work on large-scale, cloud-native infrastructure supporting mission-critical services.
- Career growth opportunities within a rapidly scaling global technology organization.
- Collaborative engineering culture focused on innovation, reliability, and continuous improvement.
- Exposure to cutting-edge technologies across Kubernetes, cloud platforms, automation, and distributed systems.