JobTarget Logo

Staff Site Reliability Engineer-Observability in India at Jobgether

NewJob Function: Engineering
Jobgether
India, India
Posted on
New job! Apply early to increase your chances of getting hired.

Explore Related Opportunities

Job Description

Staff Site Reliability Engineer-Observability

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer - Observability based in India.

We are seeking an experienced Staff Site Reliability Engineer specializing in Observability to build and enhance the reliability, performance, and scalability of large-scale cloud infrastructure. In this senior technical role, you will lead initiatives across observability, automation, cloud operations, and platform engineering while partnering with development teams to improve system resilience. You will play a key role in designing modern monitoring solutions, driving cloud migration efforts, strengthening security and compliance, and reducing operational toil through AI-powered automation. This is an exciting opportunity to influence engineering best practices, mentor fellow engineers, and shape the future of highly available distributed systems in a fully remote environment.

Accountabilities:
  • Design, implement, and continuously improve enterprise observability solutions using modern monitoring, logging, and tracing technologies, including Prometheus, Grafana, OpenTelemetry, Elastic, and related platforms.
  • Lead initiatives that improve the reliability, scalability, and performance of cloud-native and hybrid infrastructure across AWS and Kubernetes environments.
  • Manage vulnerability remediation, patch compliance, and infrastructure health across large-scale on-premises and cloud deployments while ensuring service-level objectives are achieved.
  • Operate, enhance, and optimize Kubernetes platforms and GitOps workflows, including ArgoCD and multi-tenant containerized environments.
  • Develop automation solutions and AI-assisted operational tooling that streamline incident response, reduce manual effort, and improve operational efficiency.
  • Partner with software engineering teams to establish monitoring standards, define meaningful service metrics, optimize alerting strategies, and improve production readiness.
  • Participate in incident management, root cause analysis, and continuous improvement initiatives while mentoring engineers and promoting Site Reliability Engineering best practices.
Requirements
  • 8+ years of experience designing, operating, and supporting large-scale cloud infrastructure, with significant expertise in AWS environments.
  • 5+ years of hands-on experience with Kubernetes platforms such as EKS, AKS, GKE, Fargate, or similar container orchestration technologies.
  • 4+ years of Linux systems administration experience in production environments.
  • Strong programming skills in Go, Python, Ruby, or comparable languages used for automation and platform engineering.
  • Proven expertise with Infrastructure as Code technologies such as Terraform, Ansible, AWS CDK, or equivalent tools.
  • Strong understanding of observability platforms, distributed systems monitoring, logging, tracing, and performance optimization.
  • Experience with CI/CD pipelines, Git-based version control, automated testing, and modern software delivery practices.
  • Knowledge of SQL databases such as MySQL or PostgreSQL, networking fundamentals, distributed infrastructure, and cybersecurity best practices.
  • Excellent problem-solving, communication, collaboration, and mentoring skills with the ability to thrive in high-pressure production environments.
Benefits
  • Fully remote position based in India with flexibility to work from home.
  • Competitive compensation package with performance-based bonus and incentive opportunities.
  • Employee equity grants and participation in an employee stock purchase plan, where applicable.
  • Comprehensive health and wellness benefits for employees and eligible dependents.
  • Retirement savings and statutory benefits in accordance with applicable local policies.
  • Generous paid time off, company holidays, parental leave, and family-friendly benefits.
  • Opportunity to work on cutting-edge cloud infrastructure, AI-powered automation, and large-scale distributed systems.
  • Collaborative engineering culture focused on continuous learning, innovation, mentorship, and professional growth.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Job Location

India, India

Frequently asked questions about this position

Continue to apply
Enter your email to continue. You’ll be redirected to the employer’s application.
By clicking Continue, you understand and agree to JobTarget's Terms of Use and Privacy Policy.