Senior Site Reliability Engineer — Token Factory (Inference Platform) in Châtenay-en-France, Île-de-France at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer — Token Factory (Inference Platform) based in France.
This is a senior engineering role focused on the reliability, performance, and observability of a large-scale AI inference platform.
You will help operate infrastructure serving foundation models across text, vision, audio, and emerging multimodal workloads.
The role combines Kubernetes, infrastructure-as-code, observability, automation, and production incident management at significant scale.
You will optimize GPU-heavy workloads, strengthen resilience, and ensure high-throughput APIs meet demanding reliability and cost targets.
You will work closely with software engineers and infrastructure teams to build self-healing systems and robust operational processes.
The environment is fast-moving, highly technical, international, and focused on solving complex infrastructure challenges for the AI ecosystem.
This is an opportunity to have a direct impact on the infrastructure powering next-generation AI applications.
- Own the reliability, performance, and observability of the inference platform and its supporting infrastructure.
- Design, implement, and continuously improve telemetry pipelines covering metrics, logs, and traces.
- Build monitoring and observability solutions capable of processing large volumes of production signals and converting them into actionable insights.
- Configure and optimize Kubernetes infrastructure for high availability, scalability, and efficient resource utilization.
- Tune Kubernetes autoscaling mechanisms to improve the efficiency and utilization of GPU resources.
- Develop and maintain Terraform modules and infrastructure-as-code patterns that embed resilience and reliability into new clusters and services.
- Design and improve request-routing, retry, and failure-handling mechanisms to minimize the impact of transient infrastructure or service failures.
- Develop automation and operational tooling to detect, isolate, and remediate incidents quickly.
- Create, maintain, and improve runbooks for incident response and operational procedures.
- Participate in production incident management, troubleshooting issues and restoring services within demanding reliability objectives.
- Lead or contribute to post-mortem processes and implement corrective actions to prevent recurring incidents.
- Define and improve reliability practices for high-throughput APIs, including alerting strategies and Service Level Objectives (SLOs).
- Investigate distributed-system failures and performance issues across infrastructure and application layers.
- Optimize systems from the kernel and infrastructure layer through to the application layer.
- Support and improve the operation of GPU-intensive inference workloads and accelerator-based infrastructure.
- Contribute to scaling the inference platform while balancing performance, reliability, and infrastructure costs.
- Collaborate closely with software engineers to incorporate reliability and operational excellence into product and platform development.
- Promote automation, self-healing capabilities, and engineering practices that reduce operational overhead and improve system resilience.
- Significant experience in Site Reliability Engineering, Production Engineering, DevOps, or a closely related infrastructure discipline.
- Deep practical knowledge of Kubernetes in production environments.
- Strong experience with Prometheus and Grafana for monitoring, metrics, dashboards, and observability.
- Advanced experience with Terraform and infrastructure-as-code practices.
- Strong scripting and automation skills using Python and/or Bash.
- Solid understanding of distributed systems and the ways production backends can fail under real-world conditions.
- Experience designing effective alerts, monitoring strategies, and SLOs for high-throughput services or APIs.
- Strong troubleshooting and debugging skills across infrastructure, networking, operating systems, and application layers.
- Experience designing systems for high availability, resilience, scalability, and graceful failure recovery.
- Hands-on experience with GPU-heavy workloads or accelerator-based infrastructure is highly valuable.
- Familiarity with GPU inference technologies such as vLLM, Triton, Ray, or comparable accelerator and model-serving stacks.
- Experience with MLOps, model hosting, AI infrastructure, or machine-learning platforms is advantageous.
- Strong understanding of infrastructure automation, deployment, configuration management, and operational tooling.
- Ability to analyze complex performance and reliability problems and translate findings into practical engineering improvements.
- Strong incident-management and root-cause-analysis capabilities.
- Ability to collaborate effectively with software engineers and other technical teams to integrate reliability into platform development.
- Proactive mindset with a strong focus on automation, self-healing systems, and continuous improvement.
- Comfortable working independently, taking ownership of critical infrastructure, and operating effectively in a fast-paced technical environment.
- Competitive compensation.
- Career growth and continuous learning opportunities.
- Flexibility and significant ownership in your work.
- Collaborative and innovative international working environment.
- Opportunity to work on high-impact AI infrastructure and inference technologies.
- Exposure to large-scale GPU infrastructure and complex distributed systems.
- Opportunity to contribute to infrastructure supporting next-generation multimodal AI applications.
- Diverse and highly technical international teams.
- Inclusive workplace committed to equal employment opportunities.
- Workplace accommodations available throughout the application process where required.
- Employment is subject to authorization to work in the country where the position is based.