Senior HPC Cluster Engineer - AI, ML in India at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior HPC Cluster Engineer - AI, ML based in India.
This role offers the opportunity to engineer and operate large-scale infrastructure powering advanced AI and high-performance computing workloads. You will help build and manage heterogeneous GPU-accelerated clusters across on-premises and cloud environments. The position combines production operations, automation, performance engineering, reliability, and direct collaboration with AI/ML researchers. You will play a key role in improving resource utilization, system stability, and the overall experience of users running demanding workloads. Working with globally distributed engineering teams, you will influence infrastructure strategy and continuously improve operational practices. The environment is highly technical, fast-moving, and focused on enabling the next generation of AI and scientific computing.
- Provide technical leadership for systems administration and service delivery across large-scale AI/HPC infrastructure, coordinating upgrades, incident response, and reliability improvements.
- Own the day-to-day operation of production AI/HPC clusters, monitoring system health, user experience, resource utilization, and adherence to internal service-level targets.
- Build, maintain, and scale heterogeneous AI/ML clusters across on-premises and cloud environments, covering compute, networking, storage, and GPU-accelerated infrastructure.
- Develop scalable automation and tooling to improve the deployment, configuration, management, and operational efficiency of AI/HPC environments.
- Collaborate with global engineering teams and internal users to understand evolving research and computing requirements and deliver reliable infrastructure solutions.
- Support researchers in running complex AI/HPC workloads, including performance analysis, troubleshooting, optimization, and workload tuning.
- Analyze cluster efficiency, job fragmentation, and GPU utilization to identify opportunities to reduce resource waste and improve overall capacity.
- Conduct root cause analysis for infrastructure and service issues, proactively identifying risks and implementing corrective actions before they affect users.
- Lead SEV triage, incident response, and postmortems for reliability events affecting production infrastructure or users.
- Build strong relationships with customers, researchers, and cross-functional engineering teams to improve service delivery and anticipate infrastructure needs.
- Participate in an on-call rotation and provide timely support for critical production GPU clusters.
- Bachelor's degree in Computer Science, Electrical Engineering, or a related discipline, or equivalent practical experience.
- 5+ years of experience designing, deploying, and operating large-scale compute or infrastructure environments.
- Strong experience with AI/HPC job schedulers such as Slurm, Kubernetes, PBS, RTDA, BCM, or LSF.
- Proficiency administering Linux distributions such as CentOS/RHEL and/or Ubuntu.
- Hands-on experience with cluster configuration and infrastructure management tools including BCM, Terraform, Ansible, Puppet, Salt, or similar technologies.
- Strong knowledge of container technologies such as Docker, Singularity, Podman, Shifter, or Charliecloud.
- Proficiency with Python programming and Bash scripting for infrastructure automation and operational tooling.
- Practical experience supporting AI/HPC workflows that use MPI.
- Experience analyzing and tuning the performance of diverse AI and HPC workloads.
- Strong troubleshooting, analytical, and root-cause analysis skills, with the ability to resolve complex distributed infrastructure problems.
- Excellent collaboration and communication skills, particularly when working with researchers, infrastructure engineers, and globally distributed teams.
- Passion for continuous learning and for staying current with emerging technologies and best practices in HPC and AI/ML infrastructure.
Nice-to-have experience:
- Experience working with NVIDIA GPUs, CUDA programming, NCCL, or MLPerf benchmarking.
- Understanding of AI/ML concepts, algorithms, models, and frameworks such as PyTorch or TensorFlow.
- Experience with InfiniBand, IPoIB, and RDMA technologies.
- Knowledge of high-performance distributed storage platforms such as Lustre or GPFS.
- Experience with advanced GPU cluster optimization and large-scale AI/ML infrastructure.
- Opportunity to work on large-scale AI and high-performance computing infrastructure supporting advanced research and engineering workloads.
- Exposure to GPU-accelerated computing, distributed systems, cloud infrastructure, HPC, and emerging AI/ML technologies.
- Collaboration with highly skilled engineering and research teams across global locations.
- Opportunities to influence infrastructure architecture, automation, reliability, and operational best practices.
- Continuous learning and exposure to rapidly evolving technologies in AI, ML, and high-performance computing.
- Flexible work arrangements may be available, with opportunities based in Bengaluru, Pune, or remote within India.
- Full-time role within a technically advanced and innovation-focused environment.
- Opportunity to contribute to infrastructure that enables next-generation AI and scientific computing.