Senior Observability Engineer / Platform Engineer in India at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Observability Engineer / Platform Engineer based in India.
This is a remote opportunity for an experienced engineer to shape and operate enterprise-grade observability platforms across cloud-native environments. You will work across metrics, logs, traces, events, alerting, and service health to strengthen the reliability of critical production systems. The role combines hands-on engineering with platform enablement, automation, and operational excellence. You will support Kubernetes and public cloud environments while helping engineering and SRE teams gain better visibility into system performance and reliability. Your work will directly contribute to proactive incident detection, faster troubleshooting, and improved availability. You will also have the opportunity to contribute to evolving practices around distributed tracing, SLOs, intelligent alerting, and AIOps. This is a permanent remote role designed for professionals who enjoy solving complex platform and reliability challenges.
- Design, implement, and maintain enterprise observability platforms covering metrics, logs, traces, and events.
- Build and manage observability solutions using platforms such as Prometheus, Grafana, OpenSearch, Splunk, Elastic, Datadog, Dynatrace, New Relic, or equivalent technologies.
- Develop dashboards, SLOs, SLIs, alerting rules, and service health monitoring frameworks to provide actionable visibility into production environments.
- Integrate monitoring and observability capabilities across Kubernetes, containerized workloads, and public cloud platforms.
- Enable effective incident management, root cause analysis, and production troubleshooting through robust observability practices.
- Automate monitoring configuration, platform onboarding, and operational processes using Infrastructure-as-Code and CI/CD pipelines.
- Partner with engineering, platform, SRE, DevOps, and operations teams to improve system reliability, performance, scalability, and availability.
- Contribute to observability maturity initiatives, including distributed tracing, OpenTelemetry, AIOps, intelligent alerting, and automated remediation.
- 6+ years of experience in Platform Engineering, SRE, DevOps, Cloud Operations, Observability Engineering, or a closely related discipline.
- Strong hands-on expertise with observability technologies such as Prometheus, Grafana, Splunk, OpenSearch, Elastic, or equivalent platforms.
- Experience with distributed tracing solutions such as Jaeger, Tempo, and OpenTelemetry, with a solid understanding of modern observability practices.
- Strong knowledge of Kubernetes, Docker, and containerized workloads, ideally gained through enterprise-scale production environments.
- Practical experience working with AWS, Azure, or GCP cloud environments.
- Demonstrated ability to manage incidents, tune alerts, troubleshoot complex production issues, and improve operational reliability.
- Strong scripting and automation capabilities using Python, Shell, or similar languages.
- Familiarity with CI/CD and Infrastructure-as-Code tools such as Terraform, Jenkins, GitHub Actions, or ArgoCD.
- Understanding of SRE practices, including SLO/SLI frameworks and error budgets, is preferred.
- Exposure to AIOps, intelligent incident response, automated remediation, and OpenTelemetry implementations is an advantage.
- Experience working with enterprise-scale production platforms and an awareness of security, compliance, and governance considerations in cloud-native environments.
- Strong collaboration, communication, analytical, and problem-solving skills, with the ability to work effectively across engineering and operations teams.
- Permanent remote working model, with the role based in India.
- Flexible working arrangements designed to accommodate employee, customer, and business needs.
- Opportunities to work on enterprise-scale cloud, platform engineering, and observability initiatives.
- Exposure to modern technologies and practices across observability, Kubernetes, cloud platforms, SRE, automation, and AIOps.
- An inclusive and diverse working environment that values different perspectives, skills, and experiences.
- Flexibility around working hours and arrangements where business and customer requirements allow.
- Supportive initiatives for professionals returning to work after an extended career break due to health or family circumstances.
- Opportunities for professional development, meaningful ownership, and long-term career growth.
- Salary: Competitive compensation aligned with experience and market standards.
- Healthcare and additional perks: Details to be confirmed during the hiring process.