Sr. DevOps Engineer in India at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Sr. DevOps Engineer based in India.
We are seeking an experienced DevOps professional to drive the reliability, scalability, and automation of large-scale production infrastructure. This role focuses on building and maintaining resilient Linux-based environments that support critical backup and disaster recovery platforms. You will collaborate closely with engineering, SRE, and platform teams to improve operational excellence and deliver highly available systems. The position requires strong expertise in cloud infrastructure, Kubernetes, automation, and infrastructure-as-code practices. You will have the opportunity to solve complex technical challenges, optimize performance, and shape reliable engineering processes. This is a high-impact role for someone passionate about platform engineering, automation, and modern infrastructure.
The Senior DevOps Engineer will own the design, operation, and continuous improvement of production infrastructure, ensuring reliability, security, and scalability across complex environments. You will contribute to automation initiatives, incident management, and platform modernization while partnering with multiple technical teams.
- Design, build, and maintain highly available Linux-based production infrastructure supporting critical services.
- Manage and optimize large-scale Linux environments, including performance tuning, kernel configuration, storage, networking, and troubleshooting.
- Administer Kubernetes clusters with a focus on reliability, scalability, security, and efficient resource utilization.
- Develop and maintain Infrastructure as Code solutions using Terraform or Pulumi with reusable and modular approaches.
- Build and improve CI/CD pipelines using tools such as GitHub Actions, Jenkins, or ArgoCD.
- Create automation solutions using Shell scripting, Python, or Go to reduce manual operational efforts.
- Implement monitoring, logging, and observability solutions using platforms such as Prometheus, Grafana, Datadog, or ELK.
- Lead incident response activities, root cause analysis, postmortems, and continuous improvement initiatives.
- Collaborate with software engineering teams to improve deployment processes, reliability, and production readiness.
- Implement infrastructure security practices including access controls, secrets management, vulnerability scanning, and audit logging.
- Troubleshoot complex production issues involving infrastructure, networking, storage, and operating systems.
- Support platform evolution by improving resilience, automation, and operational efficiency.
The ideal candidate is a senior DevOps or Site Reliability professional with extensive experience managing enterprise-scale infrastructure and building reliable automation solutions.
- 8–12 years of experience in DevOps, Platform Engineering, Linux Administration, or Site Reliability Engineering.
- Strong expertise in Linux system administration, including:
- Performance tuning and optimization.
- Kernel configuration and troubleshooting.
- Storage and filesystem management.
- Process management and system diagnostics.
- Networking fundamentals.
- Hands-on experience managing Kubernetes clusters in production environments.
- Strong knowledge of Infrastructure as Code tools such as Terraform or Pulumi.
- Experience designing and maintaining CI/CD pipelines using GitHub Actions, Jenkins, ArgoCD, or similar technologies.
- Proficiency with monitoring and observability tools including Prometheus, Grafana, Datadog, or ELK.
- Strong scripting skills with Shell, Python, or Go for infrastructure automation.
- Experience supporting production incidents, on-call operations, alert management, and reliability improvements.
- Strong understanding of DNS, load balancers, firewalls, VPCs, and hybrid networking concepts.
- Experience implementing security best practices such as RBAC, secrets management, Vault, and audit logging.
- Ability to analyze complex technical problems and work effectively in collaborative environments.
Preferred qualifications:
- Experience with OpenStack technologies such as Nova, Swift, Neutron, or Cinder.
- Experience managing workloads across AWS, Azure, GCP, and private cloud environments.
- Experience with large-scale distributed infrastructure environments.
- Knowledge of storage platforms, backup solutions, and disaster recovery concepts, including RPO and RTO.
- Experience with FinOps practices, infrastructure capacity planning, and cost optimization.
- Exposure to Chaos Engineering and resilience testing.
- Familiarity with bare-metal provisioning technologies such as PXE, MaaS, or Ironic.
- Ability to troubleshoot Go-based services and contribute to infrastructure tooling.
- Hybrid/onsite work environment in Pune.
- Opportunity to work on large-scale infrastructure and mission-critical platforms.
- Exposure to modern DevOps practices, cloud technologies, and automation frameworks.
- Collaboration with experienced engineering, SRE, and platform teams.
- Opportunity to contribute to high-impact reliability and scalability initiatives.
- Career growth opportunities within a global technology organization.
- Inclusive workplace environment focused on innovation, accountability, and professional development.