Platform Engineer in India at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Platform Engineer based in India.
This role focuses on maintaining reliable, scalable, and highly available technology platforms in a fast-paced operational environment.
You will work across Kubernetes, containerization, Linux, DevOps, CI/CD, Python, and cloud technologies to support system stability.
The position combines proactive monitoring, incident response, troubleshooting, automation, and continuous improvement.
You will help reduce operational dependencies by strengthening runbooks, automation, monitoring, and self-service capabilities.
The role requires close collaboration with technical partners while maintaining service levels and responding quickly to production incidents.
You will work with distributed systems and modern deployment strategies, with opportunities to improve platform resilience and delivery processes.
The position is suited to engineers who enjoy hands-on technical problem solving and are comfortable working rotational shifts.
- Work with a proactive platform operations team to maintain system stability, minimize incidents, consistently achieve service-level targets, and improve capacity and operational performance through metrics and best practices.
- Resolve Tier-1 incidents using established runbooks, including system restarts, network connectivity checks, model resets, data-flow checks, alert reprocessing, environment verification, and other standard operational fixes.
- Monitor systems and alerting channels continuously, review operational metrics, provide immediate incident response, and prepare daily and weekly reports covering system performance, incidents, and key indicators.
- Review root-cause analysis and resolution requests, execute documented runbook procedures, and escalate issues to L2/L3 support when they involve model, hardware, security, or other issues beyond the defined operational scope.
- Maintain strong adherence to SLAs by prioritizing incidents according to severity and ensuring timely response, resolution, escalation, and communication throughout the incident lifecycle.
- Perform regular operational checks, including server environments, real-time overlays, security desk systems, PagerDuty alerts, Redis database cron jobs, and other platform health indicators.
- Schedule and coordinate system downtime for updates and maintenance activities while minimizing disruption to ongoing operations.
- Improve platform operations by reducing manual and people-dependent processes, strengthening automation, and moving recurring workflows toward more reliable and self-sufficient operation.
- Maintain consistent communication and collaboration with internal and external technical partners, ensuring operational status, incidents, escalations, and required actions are clearly communicated.
- Have 1–6 years of relevant professional experience in platform engineering, DevOps, backend development, infrastructure operations, or a related technical discipline.
- Demonstrate strong Kubernetes expertise, ideally supported by the Certified Kubernetes Application Developer (CKAD) certification or comparable practical knowledge.
- Have hands-on experience creating and maintaining container infrastructure and scalable Kubernetes environments, including Deployments, Pods, Jobs, StatefulSets, ConfigMaps, Services, NodePort, Ingress, Volumes, and Custom Resource Definitions.
- Be highly proficient with Linux and Docker, including container maintenance, troubleshooting, container logs, health checks, probes, and Linux shell operations.
- Have experience with Helm Charts, multi-container pod designs using sidecar and init containers, distributed-system design patterns, and Kubernetes deployment strategies such as blue/green, canary, and rolling updates.
- Demonstrate strong DevOps and CI/CD knowledge, including continuous integration, deployment automation, monitoring, alerting, troubleshooting, and production reliability practices.
- Possess strong Python development skills, including iterators, exception handling, file handling, data types and structures, object-oriented programming, and production-grade software development rather than scripting alone.
- Have practical experience with web frameworks such as Flask and Gunicorn, along with strong understanding of software design patterns, database integration, streaming pipelines, and multiprocessing architectures.
- Demonstrate strong Git and GitHub knowledge and experience working with technologies such as Redis and Kafka for data integration or streaming workflows.
- Experience with GStreamer, FFmpeg, and OpenCV for video streaming or image processing is highly valued, along with knowledge of cloud services and hybrid cloud environments.
- Be comfortable troubleshooting complex technical issues, working with operational metrics, following structured incident procedures, and communicating clearly with technical stakeholders.
- Be available to work rotational shifts: Shift A from 6 AM–2 PM IST, Shift B from 2 PM–10 PM IST, or Shift C from 10 PM–6 AM IST, with two consecutive days off each week.
- Be comfortable joining immediately and working remotely from a base location associated with Mumbai, Bangalore, or Trivandrum.
- Full-time remote work opportunity based in India, with Mumbai, Bangalore, and Trivandrum identified as base locations.
- Exposure to modern platform engineering technologies including Kubernetes, Docker, Python, CI/CD, cloud services, distributed systems, and production monitoring.
- Opportunity to work on highly scalable infrastructure and contribute to platform stability, automation, reliability, and operational efficiency.
- Hands-on experience with incident management, deployment strategies, observability, and production-grade engineering practices.
- A collaborative environment focused on transparency, diversity, integrity, continuous learning, and professional growth.
- Rotational shift schedule with two consecutive days off each week.
- Opportunities to work alongside proactive technical teams and collaborate with partners across a dynamic technology environment.
- The source description does not specify a salary range, healthcare coverage, or additional financial benefits.