Senior Staff DevOps Engineer in India at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Staff DevOps Engineer based in India.
This is a high-impact technical leadership role within a global Infrastructure Platform team responsible for operating large-scale, cloud-native SaaS infrastructure.
You will shape the architecture, reliability, security, and scalability of platforms running across AWS, Azure, and GCP.
A major focus will be Kubernetes leadership, including EKS, AKS, service mesh, GitOps, and cloud-native deployment patterns.
You will also help advance AI-assisted and agentic automation across DevOps and SRE workflows.
The role offers significant organizational influence through architecture decisions, technical standards, mentoring, and cross-functional leadership.
You’ll work with globally distributed engineering teams across India, the US, EMEA, and APAC in a fast-paced 24×7 environment.
This is an individual-contributor leadership position with substantial scope to improve reliability, automation, security, and engineering productivity.
- Lead cloud infrastructure strategy: Design, build, and evolve resilient, secure, scalable, and cost-efficient multi-account cloud infrastructure, primarily across AWS while supporting Azure and GCP environments.
- Drive Kubernetes excellence: Serve as a technical authority for production Kubernetes, including EKS and AKS cluster architecture, upgrades, networking, storage, capacity planning, security, multi-tenancy, and troubleshooting.
- Own service mesh capabilities: Design and operate production service mesh solutions, implementing secure service-to-service communication, mTLS, traffic management, observability, resilience, and progressive delivery.
- Establish platform standards: Define and promote best practices for Kubernetes, cloud networking, IAM, secrets management, infrastructure governance, workload design, Helm, deployment patterns, and security controls.
- Advance Infrastructure as Code and GitOps: Automate infrastructure provisioning, deployment, monitoring, incident response, and capacity management using Terraform, CI/CD, GitOps, and related platform engineering practices.
- Strengthen reliability and operations: Improve SLIs, SLOs, error budgets, observability, runbooks, incident response, on-call practices, post-incident reviews, and systemic remediation across a 24×7 SaaS environment.
- Support security and compliance: Help implement secure platform controls and maintain infrastructure aligned with requirements such as PCI DSS, including audit readiness, governance, and secure-by-design practices.
- Champion AI-enabled operations: Introduce LLM-based tooling, AI coding assistants, and agentic workflows to improve infrastructure development, incident triage, root-cause analysis, deployment validation, compliance checks, and operational efficiency.
- Build safe AI automation: Establish appropriate guardrails, observability, cost controls, and human-in-the-loop practices for production AI and agentic workflows, including solutions built with services such as Amazon Bedrock.
- Provide technical leadership: Influence engineers across teams and geographies without direct management responsibility, drive consensus on architecture, contribute to critical escalations, and raise engineering standards.
- Mentor engineering talent: Coach Staff, Senior, and mid-level engineers while documenting reusable patterns, sharing technical knowledge, and encouraging stronger engineering practices.
- Lead strategic initiatives: Drive cross-functional projects that improve uptime, deployment velocity, cloud consistency, cost efficiency, operational toil, and engineering productivity.
- Experience: 12+ years working in 24×7 production operations and highly available SaaS or cloud environments, with prior experience as a technical lead or Staff+ individual contributor in a global engineering organization.
- Cloud expertise: 5+ years of hands-on experience with multi-account AWS infrastructure, including AWS Organizations, Account Factory, guardrails, SCPs, landing zones, networking, IAM, and cross-account connectivity.
- Kubernetes: 5+ years of production Kubernetes experience at scale, with deep expertise in EKS and/or AKS, cluster operations, networking, storage, security, workload management, and performance optimization.
- Infrastructure as Code: 5+ years of Terraform experience managing infrastructure across multiple AWS accounts and regions.
- CI/CD & GitOps: 5+ years designing and implementing CI/CD pipelines for Terraform, Kubernetes, and microservices, plus practical experience with GitOps platforms such as ArgoCD, Kargo, or Flux.
- Programming & systems: Strong Python, Go, or similar programming skills combined with advanced shell scripting and solid knowledge of Linux, networking, distributed systems, and production troubleshooting.
- Service mesh: Hands-on experience implementing and operating a production service mesh such as Istio, Linkerd, or AWS App Mesh.
- Observability: Experience with monitoring and logging technologies such as Prometheus, Grafana, OpenSearch, or equivalent platforms.
- SRE practices: Strong understanding of SLIs, SLOs, error budgets, incident management, observability, reliability engineering, and operational excellence.
- AI & automation: Experience applying AI tools on AWS or equivalent platforms to improve engineering productivity, automation, or operational efficiency; experience with Amazon Bedrock, LLM agents, or agentic workflows is highly valuable.
- Security & compliance: Experience designing secure cloud platforms and familiarity with regulated enterprise environments; knowledge of PCI DSS and related compliance practices is preferred.
- Regional infrastructure: Experience supporting data sovereignty or regional cloud deployments is a plus, particularly across markets with specific residency requirements.
- Leadership: Strong interpersonal and communication skills, with the ability to influence engineers and stakeholders across teams, time zones, and organizational boundaries.
- Education: Bachelor’s or Master’s degree in Computer Science or a related technical discipline, or equivalent practical experience.
- Working style: Self-directed, collaborative, adaptable, quality-focused, and comfortable operating in a complex, fast-moving environment where technical decisions have broad organizational impact.
- Fully remote position for candidates based in India.
- Opportunity to work on large-scale, cloud-native infrastructure supporting a global enterprise SaaS platform.
- Significant technical influence and ownership without requiring people management.
- Exposure to AWS, Azure, GCP, Kubernetes, service mesh, GitOps, and Infrastructure as Code at scale.
- Opportunity to lead the adoption of AI-assisted and agentic DevOps workflows, including LLM-powered automation.
- Collaboration with globally distributed engineering teams across India, US, EMEA, and APAC.
- Scope to shape platform architecture, engineering standards, reliability practices, and long-term infrastructure strategy.
- Opportunities to mentor experienced engineers and establish yourself as a subject-matter expert across the broader engineering organization.
- Participation in strategic initiatives focused on automation, reliability, security, cost efficiency, and engineering productivity.
- Inclusive workplace committed to equal opportunity and reasonable accommodations throughout the hiring process.
- Opportunity to contribute to a modern engineering environment focused on innovation, operational excellence, and continuous improvement.