Senior Software Engineer — Infra Agent Systems in India at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineer — Infra Agent Systems based in India.
This is a high-impact engineering role focused on building production AI agents that operate and automate large-scale GPU infrastructure. You’ll design systems that diagnose hardware failures, investigate incidents, gather evidence, and support remediation across complex infrastructure environments. The role spans AI agent systems, distributed services, knowledge graphs, retrieval, orchestration, and developer tooling. You’ll own systems end to end, from architecture and implementation through deployment, observability, and production operations. You’ll collaborate across infrastructure, datacenter, and engineering teams to turn operational knowledge into reliable automation. This is an opportunity to help shape how autonomous AI systems can safely and intelligently operate real-world infrastructure at massive scale.
- Design and build production AI agent systems capable of diagnosing, investigating, and supporting remediation of infrastructure issues across large-scale GPU environments.
- Develop the distributed services, orchestration frameworks, knowledge graphs, retrieval systems, and supporting infrastructure that power AI agents.
- Build fleet intelligence capabilities that combine telemetry, infrastructure state, operational knowledge, and historical incidents to improve agent decision-making.
- Integrate agent systems with observability, incident management, ticketing, fleet inventory, source control, communication platforms, and internal infrastructure through reliable APIs.
- Own services throughout their lifecycle, including architecture, implementation, testing, deployment, monitoring, reliability, and production support.
- Improve agent quality and reliability through evaluations, retrieval optimization, better tools, and continuous feedback from production environments.
- Convert insights and knowledge generated through production use into reliable, reviewed software, workflows, and automation.
- Contribute to the evolution of platform architecture and engineering practices as autonomous infrastructure capabilities scale.
- Bachelor’s degree or equivalent professional experience in Computer Science, Engineering, or a related technical field.
- 5+ years of professional experience building production backend systems, distributed systems, infrastructure platforms, or similarly complex software.
- Strong systems design capabilities and demonstrated experience taking significant systems from initial architecture through production.
- Deep expertise in at least one relevant area, such as AI agent systems, orchestration, tool use, evaluation, grounding, knowledge graphs, graph data modeling, search, retrieval, ranking, RAG, or semantic search.
- Strong backend engineering skills, including API design, service boundaries, data modeling, and integrations across complex technical environments.
- Experience with Kubernetes, GitOps practices such as ArgoCD, infrastructure-as-code, and cloud platforms.
- Proficiency in one or more relevant programming languages, such as Go, TypeScript, Python, or Rust, with the ability to work across multiple languages when required.
- Strong analytical and problem-solving abilities, with an interest in solving ambiguous and technically challenging infrastructure problems.
- Ability to own systems in production, balancing engineering quality, reliability, operational requirements, and delivery speed.
- Experience with GPU infrastructure, datacenters, bare-metal environments, hardware failure modes, BMC/IPMI, or cluster schedulers is a plus.
- Experience with graph databases, event-driven systems, messaging platforms such as NATS or Kafka, or observability tools such as Prometheus and Grafana is advantageous.
- Experience building evaluation frameworks or improving the reliability and quality of LLM-powered systems is also a plus.
- Comfortable working in a remote, highly collaborative engineering environment with strong ownership and autonomy.
- Fully remote position based in India.
- Opportunity to work on AI agents operating at significant infrastructure scale.
- Exposure to the intersection of artificial intelligence, distributed systems, GPU infrastructure, knowledge systems, and automation.
- Opportunity to build foundational systems and influence the architecture of emerging autonomous infrastructure technologies.
- End-to-end ownership of production software, from design and development through deployment and operations.
- Collaborative environment with technically ambitious engineers and researchers working on challenging AI infrastructure problems.
- Significant opportunities for learning, experimentation, and professional growth in a rapidly evolving technology space.
- Opportunity to contribute to systems that have direct, measurable production impact.