Staff Software Engineer, Infrastructure in Haciendas del Canada, Nuevo León at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Software Engineer, Infrastructure based in Canada.
This is a Staff-level infrastructure engineering role focused on building the platform foundations that enable engineering teams to move faster and operate more reliably.
You’ll combine hands-on software engineering with technical leadership, shaping self-service infrastructure, cloud platforms, deployment systems, and operational tooling.
A major focus will be replacing expert-driven provisioning and operational workflows with paved roads that provide clear ownership, safe defaults, strong guardrails, and measurable adoption.
You’ll help evolve multi-region, cross-account infrastructure and Kubernetes foundations while improving reliability, security, scalability, and cost efficiency.
The role also includes developing AI-assisted operational workflows that reduce toil while keeping production changes safe, auditable, and human-reviewed.
You’ll influence architecture across teams through RFCs, design reviews, technical standards, and pragmatic engineering decisions.
Working in a small, growing team, you’ll have substantial ownership and the opportunity to drive infrastructure investments through real production adoption.
- Turn ambiguous infrastructure challenges into clear technical proposals and drive them through RFCs, architecture reviews, and cross-team alignment.
- Design and build self-service platform capabilities and APIs, primarily in Go, covering onboarding, provisioning, deployment, observability defaults, and day-2 operations.
- Establish reliable delivery standards using Terraform, GitOps with Argo CD, progressive delivery, automated testing, and continuous deployment practices.
- Evolve multi-tenant EKS infrastructure to improve reliability, security, scalability, and cost efficiency.
- Develop and improve ingress and traffic-routing capabilities, including Envoy Gateway and multi-region, cross-account connectivity.
- Strengthen SLOs, alerting, incident response, and operational follow-up through Grafana Cloud and improved observability practices.
- Measure success through outcomes for consuming engineering teams, including faster provisioning and deployment, greater self-service, and improved operational reliability.
- Develop AI-assisted operational workflows such as alert enrichment, incident context gathering, runbook-assisted diagnosis, remediation recommendations, and onboarding assistants.
- Maintain appropriate human oversight for AI-assisted operational actions, with an emphasis on safety, auditability, and responsible automation.
- Participate in the on-call rotation after onboarding and shadowing, while helping improve the overall health of on-call through better alerts, runbooks, automation, and blameless postmortems.
- Lead strategic platform initiatives from initial design through production adoption and establish durable technical patterns across engineering teams.
- 8+ years of professional, hands-on, full-time software engineering experience in backend, infrastructure, or platform engineering.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Strong software engineering expertise in Go or a comparable programming language, including system design, testing, debugging, code review, and long-term maintainability.
- Proven experience designing, delivering, and operating cloud services or infrastructure platforms in production.
- Deep expertise in at least one area such as Kubernetes, networking, cloud platforms, reliability engineering, or developer platforms.
- Strong Linux, networking, and production operations fundamentals.
- Experience setting technical direction and leading initiatives that require alignment across multiple engineering teams.
- Strong written and verbal communication skills, particularly in remote environments and through RFCs, design documents, and incident writeups.
- Experience with EKS, ingress, CNI, or service-mesh technologies is valuable.
- Familiarity with OpenTelemetry, Prometheus, Grafana, CI/CD, progressive delivery, GitHub Actions, Argo CD, and canary deployments is a plus.
- Experience leading large-scale migrations, platform adoption programs, or cross-team infrastructure initiatives is beneficial.
- Strong systems judgment, curiosity, pragmatic decision-making, and the ability to develop deep expertise while navigating adjacent technical domains.
- Willingness to participate in an operational on-call rotation and contribute to improving its effectiveness.
- Visa sponsorship may be considered on a case-by-case basis depending on business needs.
- CA$238,250–CA$382,250 + equity for Canada-based candidates.
- Remote-first work arrangement.
- Flexible scheduling and autonomy in managing your working hours.
- Generous paid time off, quarterly wellness days, and an end-of-year wellness break.
- Home-office support to help create an effective remote workspace.
- Technology stipend equivalent to US$100 net per month.
- Annual learning and development stipend covering conferences, courses, certifications, and continued professional learning.
- 16 weeks of paid parental leave after six months of employment.
- Equity participation for full-time employees.
- Medical, retirement, and paid-holiday benefits, with details varying by country.
- Opportunities to lead major infrastructure initiatives involving self-service provisioning, multi-region networking, continuous deployment, Kubernetes, and platform engineering.
- Significant technical ownership within a small, growing infrastructure team.
- Exposure to AI-assisted and agentic operational workflows, with a focus on safe and auditable automation.
- Fully remote collaboration with distributed engineering teams.
- Offices available in Seattle and Paris for connection and collaboration.