Intelligent Infrastructure Engineer in United States Embassy at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an Intelligent Infrastructure Engineer based in the United States.
This role offers the opportunity to build and operate the foundational technology powering advanced AI training and inference workloads at scale.
The Intelligent Infrastructure Engineer will design, optimize, and maintain high-performance platforms that enable machine learning teams to develop and deploy next-generation solutions.
Working at the intersection of infrastructure engineering, distributed systems, and artificial intelligence, this position focuses on reliability, scalability, and efficiency.
The ideal candidate will bring deep technical expertise in GPU infrastructure, cloud platforms, and large-scale computing environments.
This role provides the chance to solve complex engineering challenges while improving developer experience and accelerating AI innovation.
You will collaborate with technical teams to create robust infrastructure capable of supporting demanding AI workloads in a fast-evolving environment.
The Intelligent Infrastructure Engineer will be responsible for designing, developing, and operating scalable infrastructure platforms that support AI and machine learning workloads while ensuring reliability, performance, and cost efficiency.
- Design, build, and operate infrastructure platforms supporting large-scale AI model training and inference workloads.
- Manage and optimize GPU clusters, distributed training environments, and scheduling systems for machine learning applications.
- Improve platform reliability, performance, scalability, and operational efficiency across AI infrastructure.
- Develop software solutions and automation tools using Python and systems programming languages such as Go or C++.
- Work with distributed training frameworks, accelerator architectures, and high-performance computing environments.
- Configure and optimize Kubernetes, Slurm, Ray, or similar orchestration platforms for ML workloads.
- Troubleshoot and improve Linux-based infrastructure, networking systems, and high-performance storage solutions.
- Partner with machine learning engineers, researchers, and infrastructure teams to improve developer workflows and platform capabilities.
- Implement strong software engineering practices, including testing, CI/CD processes, and code reviews.
- Contribute to cost optimization initiatives and operational improvements for AI workloads.
The ideal candidate will have extensive experience in infrastructure engineering, distributed systems, and AI platform operations, with strong software development capabilities.
- Bachelor’s or Master’s degree in Computer Science, Engineering, or a related technical discipline.
- 6+ years of experience in infrastructure, platform engineering, HPC engineering, or related fields.
- Hands-on experience operating GPU clusters or large-scale machine learning infrastructure.
- Strong proficiency in Python and experience with a systems programming language such as Go or C++.
- Deep understanding of distributed training, accelerator architectures, and communication frameworks.
- Experience with Kubernetes, Slurm, Ray, or comparable workload scheduling systems.
- Strong knowledge of Linux internals, networking concepts, and high-performance storage technologies.
- Experience working with major cloud providers and their machine learning infrastructure offerings.
- Strong software engineering skills, including testing, automation, CI/CD, and collaborative development practices.
- Excellent communication skills and ability to collaborate effectively across technical teams.
- Preferred experience with InfiniBand or RDMA networking, open-source ML infrastructure contributions, custom orchestration systems, frontier model training, or AI workload FinOps.
- Competitive annual salary range of $100,000–$150,000, based on experience and qualifications.
- Fully remote work opportunity within the United States.
- Full-time direct employment opportunity.
- Opportunity to work on advanced AI infrastructure and large-scale machine learning systems.
- Career growth opportunities within a technology-focused environment.
- Exposure to cutting-edge cloud, AI, and distributed computing technologies.
- Collaborative culture focused on innovation, engineering excellence, and continuous learning.