Lead Software Engineer, Cloud Site Reliability (SRE) in India at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead Software Engineer, Cloud Site Reliability (SRE) based in India.
This role leads cloud reliability and 24x7 site reliability operations for critical technology environments, with a strong focus on Azure infrastructure and cloud-native platforms. You will take ownership of major incidents, drive operational excellence, and ensure high availability and SLA adherence across production systems. The position combines hands-on cloud engineering with observability, automation, incident management, and reliability improvements. You will work extensively with Azure, AKS, Kubernetes, Docker, Datadog, and infrastructure-as-code technologies to build resilient and scalable environments. The role also provides an opportunity to advance proactive monitoring, anomaly detection, AIOps, and self-healing capabilities. As a technical leader, you will mentor engineers, collaborate with development teams, and communicate operational performance to stakeholders and leadership.
- Lead 24x7 NOC and site reliability operations through rotational shifts, ensuring system availability, operational stability, and adherence to service-level agreements.
- Act as Major Incident Manager for P1 and P2 incidents, coordinating triage activities, war rooms, technical teams, and stakeholder communications through resolution.
- Manage and troubleshoot Azure infrastructure, including virtual machines, networking, storage, and related cloud services.
- Administer Azure Kubernetes Service (AKS), Kubernetes, and Docker environments, including scaling, troubleshooting, performance optimization, and reliability improvements.
- Establish and enhance observability practices across logs, metrics, and traces using platforms such as Datadog and Azure Monitor.
- Drive proactive monitoring, alert optimization, anomaly detection, and AIOps initiatives to identify and address potential reliability issues before they impact users.
- Build automation and self-healing workflows using Terraform, ARM templates, Helm, Power Automate, PowerShell, Python, Bash, and other appropriate technologies.
- Collaborate with engineering teams to strengthen deployment pipelines, improve system reliability, and advance cloud-native architecture and operational practices.
- Develop operational dashboards and reports using Power BI and ServiceNow to provide visibility into incidents, reliability, and service performance.
- Lead monthly business reviews and provide clear operational reporting and insights to leadership and stakeholders.
- Mentor team members, promote knowledge sharing, standardize operational processes, and drive continuous improvements across the reliability function.
- Support broader initiatives involving multi-cloud environments, predictive monitoring, and self-healing systems where appropriate.
- 7–12 years of professional experience in CloudOps, Site Reliability Engineering, NOC, or comparable 24x7 operations environments.
- Strong hands-on expertise with Azure infrastructure, particularly virtual machines, networking, storage, and related IaaS services.
- Extensive experience with Azure Kubernetes Service (AKS), Kubernetes, and Docker, including troubleshooting, scaling, and performance tuning.
- Strong experience with monitoring and observability platforms such as Datadog and Azure Monitor, including integrations, alerting, dashboards, and operational analysis.
- Proven experience in incident management and major incident handling, including P1/P2 coordination, stakeholder communication, root-cause analysis, and operational reporting.
- Experience with Infrastructure as Code technologies such as Terraform, ARM templates, and Helm.
- Strong scripting capabilities using PowerShell, Python, Bash, or similar automation technologies.
- Experience working with ServiceNow, particularly Incident, Problem, and Change Management modules and associated dashboards.
- Good understanding of distributed systems, cloud-native architectures, and reliability engineering principles.
- Excellent communication, leadership, coordination, and problem-solving skills, with the ability to work effectively across technical and business teams.
- Experience working in multi-cloud environments, particularly Azure and AWS, is a plus.
- Exposure to AIOps, predictive monitoring, anomaly detection, or self-healing systems is desirable.
- Relevant certifications in Azure, Datadog, Kubernetes, or related cloud and reliability technologies are advantageous.
- Bachelor’s degree or equivalent technical education and professional experience.
- Fully remote opportunity based in India, with a rotational shift structure supporting 24x7 operations.
- Leadership responsibility across cloud reliability, infrastructure operations, incident management, and operational excellence.
- Hands-on exposure to Azure IaaS, AKS, Kubernetes, Docker, Datadog, Azure Monitor, and Infrastructure as Code.
- Opportunity to develop advanced observability, AIOps, predictive monitoring, automation, and self-healing capabilities.
- Collaboration with engineering teams on cloud-native architecture, deployment pipelines, scalability, and reliability initiatives.
- Opportunities to mentor team members and influence operational standards and engineering practices.
- Exposure to multi-cloud technologies and large-scale distributed systems.
- An inclusive work environment focused on teamwork, openness, respect, fairness, and continuous improvement.
- Support for professional development through exposure to modern cloud and reliability technologies and relevant certification paths.
- Opportunities to participate in leadership reporting, business reviews, and cross-functional initiatives with broad organizational visibility.