Staff Site Reliability Engineer (AWS, Terraform, Distributed Systems) in India at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer (AWS, Terraform, Distributed Systems) based in India.
This is a high-impact engineering role focused on building and operating highly available, resilient, and scalable distributed systems.
You will architect cloud-native solutions that support mission-critical services operating at significant global scale.
The role combines site reliability engineering, cloud infrastructure, automation, observability, and complex incident management.
You will work across engineering teams to modernize applications, improve operational efficiency, and influence technical architecture.
The position offers opportunities to work with AWS, Terraform, containerized microservices, monitoring platforms, and modern automation technologies.
You will also contribute to ML deployment infrastructure and help establish reliable, repeatable engineering practices.
As a senior technical contributor, you will drive continuous improvement, knowledge sharing, and reliability across a fast-paced global environment.
- Architect, implement, and maintain solutions designed to deliver highly available and reliable application services, with a strong focus on resilient distributed systems.
- Automate monitoring, deployment, operational processes, and incident-response activities to improve reliability, efficiency, and consistency while reducing manual errors.
- Lead complex troubleshooting and technical investigations, identify root causes, implement corrective actions, and help prevent recurring incidents.
- Collaborate with global engineering and cross-functional teams to build, deploy, maintain, and continuously improve application services.
- Drive the modernization of traditional applications toward cloud-native architectures and contribute to key technical and architectural decisions.
- Design and manage scalable infrastructure using cloud platforms and infrastructure-as-code practices, particularly AWS, Terraform, and Ansible.
- Implement and optimize observability solutions covering monitoring, logging, alerting, system performance, and service health.
- Support and optimize ML model deployment pipelines and their associated monitoring systems.
- Participate in an on-call rotation, including occasional off-hours support, to maintain operational continuity in a 24x7 environment.
- Develop and maintain documentation for infrastructure, operational procedures, troubleshooting practices, and knowledge transfer across engineering teams.
- Promote continuous improvement, automation, reliability engineering practices, and effective collaboration within Agile teams.
- Bachelor’s degree with at least 5 years of relevant experience, or an advanced degree with the corresponding professional experience; candidates with extensive equivalent experience may also be considered.
- Strong professional experience in Site Reliability Engineering, DevOps, cloud infrastructure, distributed systems, or a closely related engineering discipline.
- Hands-on experience with cloud platforms such as AWS, with additional exposure to Azure or GCP and hybrid cloud architectures.
- Proven expertise with infrastructure-as-code technologies, particularly Terraform and/or Ansible.
- Experience designing, deploying, and supporting containerized microservices and high-performance distributed applications.
- Strong knowledge of observability and monitoring technologies such as Prometheus, Grafana, Datadog, and ELK.
- Programming or automation experience with languages such as Python, Go, and Java.
- Familiarity with modern application and messaging technologies such as Spring Boot, SQS, IBM MQ, Kafka-like streaming systems, Flink, Hazelcast, or comparable platforms.
- Experience with distributed-system design, system architecture, performance optimization, troubleshooting, and reliability engineering.
- Experience supporting ML model deployment pipelines and associated monitoring is highly valuable.
- Strong understanding of automation, incident management, cloud-native architecture, and operational best practices.
- Experience working in Agile teams and fast-paced 24x7 production environments.
- Strong communication, collaboration, documentation, and knowledge-sharing skills, with the ability to influence technical decisions across teams.
- Comfortable adopting emerging technologies, including generative AI tools, to improve productivity and everyday engineering workflows.
- Fully remote working arrangement within India, with occasional office presence potentially required with advance notice.
- Opportunity to work on large-scale, mission-critical distributed systems and globally significant technology platforms.
- Exposure to modern cloud, infrastructure-as-code, observability, automation, containerization, and ML technologies.
- Strong opportunities for technical leadership, architectural influence, continuous learning, and professional development.
- Collaborative environment with global engineering teams and cross-functional stakeholders.
- Experience working with modern SRE and cloud-native engineering practices in a fast-paced, 24x7 environment.
- Inclusive workplace focused on equal opportunity, professional growth, and meaningful technical impact.