JobTarget Logo

Technical Lead, Platform Engineering (Observability) in India at Jobgether

NewJob Function: Admin/Clerical/Secretarial
Jobgether
India, India
Posted on
New job! Apply early to increase your chances of getting hired.

Explore Related Opportunities

Job Description

Technical Lead, Platform Engineering (Observability)

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Technical Lead, Platform Engineering (Observability) based in India.

This is a senior technical leadership opportunity focused on building observability capabilities for complex, large-scale distributed systems.
You will shape the technical direction of a platform that provides engineers with the visibility needed to build reliable, high-performing services.
The role spans metrics, logs, traces, alerting, automation, developer tooling, and AI-assisted approaches to incident management.
You will tackle challenges across hundreds of microservices, improving signal quality while reducing incident detection and mitigation times.
The position combines hands-on engineering with technical leadership, mentoring, architecture, and cross-functional collaboration.
You will work closely with platform engineering, SRE, security, and product engineering teams to establish scalable reliability practices.
This role is ideal for an experienced platform engineer who enjoys solving complex infrastructure problems and making large-scale systems easier to understand and operate.

Accountabilities:
  • Observability Platform Engineering: Design, build, operate, and continuously improve an observability platform covering metrics, logs, distributed traces, dashboards, and alerting at scale.
  • Technical Leadership: Drive the technical direction of observability capabilities, make architectural decisions, establish engineering standards, and lead initiatives from design through implementation and operation.
  • Reliability Improvement: Deliver measurable improvements in Mean Time to Detect (MTTD) and Mean Time to Mitigate (MTTM), while improving service reliability and operational resilience.
  • AI-Powered Observability: Design and implement intelligent solutions for automated anomaly detection, alert correlation, root cause analysis, and AI-assisted incident response.
  • Self-Service Tooling: Build developer-focused tooling that enables product engineering teams to independently instrument services, create dashboards, configure alerts, and adopt observability best practices.
  • Standards & SLOs: Define and champion observability standards, instrumentation practices, Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error-budget frameworks across the engineering organization.
  • Platform Optimization: Improve data collection, processing pipelines, alert quality, and operational workflows while reducing noise, unnecessary toil, and observability-related costs.
  • Cross-Functional Collaboration: Partner with SRE, Security, platform teams, and product engineering teams to ensure comprehensive system visibility and consistent reliability practices.
  • Automation: Automate recurring operational processes and workflows to increase engineering efficiency and improve the team's ability to respond to incidents.
  • Mentorship & Culture: Lead by example, mentor engineers, contribute to technical discussions, and help foster a strong engineering culture centered on collaboration, ownership, and continuous improvement.
  • Documentation: Produce high-quality technical and architectural documentation and facilitate design discussions with relevant engineering stakeholders.
Requirements
  • Professional Experience: 9+ years of experience building, operating, and maintaining scalable production systems, with significant experience in platform engineering, infrastructure, SRE, or related disciplines.
  • Observability Expertise: Strong hands-on experience with production observability and monitoring platforms such as Datadog, Prometheus, Grafana, or equivalent technologies.
  • Distributed Systems: Deep understanding of metrics, logging, distributed tracing, instrumentation patterns, and observability data pipeline architecture.
  • Kubernetes: Proven experience deploying, operating, and troubleshooting Kubernetes-based production environments and containerized workloads.
  • Programming: Strong proficiency in Go or Python for developing infrastructure tooling, platform services, and automation.
  • Cloud & Infrastructure as Code: Experience with GCP and/or AWS and Infrastructure as Code technologies such as Terraform.
  • Alerting & Incident Detection: Demonstrated ability to design and tune alerting systems to improve signal quality, reduce alert fatigue, and accelerate incident detection.
  • Reliability Engineering: Strong understanding of SLIs, SLOs, error budgets, and modern reliability engineering practices.
  • Developer Platforms: Experience building internal platforms, tools, or self-service capabilities that improve developer productivity and engineering workflows.
  • Communication: Strong written and verbal communication skills, with the ability to produce design documentation, explain complex technical concepts, and lead technical discussions.
  • Technical Leadership: Proven ability to influence technical direction, mentor engineers, make sound architectural decisions, and drive initiatives across teams.
  • AI & Advanced Observability: Experience applying AI or machine learning to observability use cases such as anomaly detection, alert correlation, or root cause analysis is preferred.
  • Large-Scale Systems: Experience supporting observability across large distributed environments, ideally involving hundreds of microservices.
  • OpenTelemetry: Hands-on experience with OpenTelemetry for instrumentation and telemetry collection is an advantage.
  • Observability Cost Optimization: Experience with sampling strategies, data tiering, pipeline optimization, or other approaches to controlling observability costs at scale is preferred.
  • Developer Experience: Passion for improving engineering productivity through better platform tooling, automation, and self-service capabilities.
  • Open Source: Contributions to or active participation in open-source observability communities are a plus.
  • Incident Management: Experience with incident management processes, tooling, and operational response practices is desirable.
Benefits
  • Full-time employment with the opportunity to work on large-scale platform and observability challenges.
  • Hybrid work model with 2 days per week working from the office and 3 days remotely.
  • Bengaluru-based office environment with opportunities for collaboration and team engagement.
  • Opportunity to shape observability practices across a complex, distributed engineering environment.
  • Exposure to advanced technologies spanning cloud infrastructure, Kubernetes, AI-powered operations, distributed systems, and developer platforms.
  • Strong opportunities for technical leadership, mentoring, and career development.
  • Collaborative environment involving platform engineering, SRE, security, and product engineering teams.
  • Opportunity to directly improve engineering productivity, system reliability, and incident response at scale.
  • Professional environment focused on innovation, continuous improvement, collaboration, and high engineering standards.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Job Location

India, India

Frequently asked questions about this position

Continue to apply
Enter your email to continue. You’ll be redirected to the employer’s application.
By clicking Continue, you understand and agree to JobTarget's Terms of Use and Privacy Policy.