Research Engineer (Reinforcement Learning) in Roma, Lazio at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Research Engineer (Reinforcement Learning) based in Italy.
Join a small, senior engineering team building the next generation of voice- and text-driven AI agents.
You’ll focus on post-training models to make agents more capable, reliable, and effective over long-running interactions.
Your work will span environments, verifiers, synthetic data, training experiments, evaluations, and production deployment.
You’ll tackle challenging problems such as persistent context, reliable tool use, and multi-turn agent behavior.
The role combines hands-on research and engineering, with a strong emphasis on measurable improvements in model performance.
You’ll work closely with experienced engineers in a remote, collaborative environment where technical craft and creativity are highly valued.
Your contributions will directly shape AI systems operating at significant production scale.
- Build training environments, verifiers, and supporting infrastructure for post-training models.
- Own the synthetic data pipeline from data generation through quality assurance and validation.
- Run end-to-end training experiments, analyze results, and clearly identify the factors driving model improvements.
- Design and maintain evaluations that models must pass before production releases.
- Select and adapt suitable open-weight foundation models for specific agent and product requirements.
- Develop trained behaviors that perform consistently across both voice and text-based agents.
- Deploy trained models to production and continuously improve them based on real-world usage and feedback.
- Develop robust approaches to long-horizon interactions, accumulated context, and reliable tool use during live conversations.
- Strong Python engineering skills and the ability to build reliable, production-quality systems.
- Demonstrated experience taking a machine learning model from raw data through experimentation and into production.
- A strong data-centric mindset, with attention to coverage, diversity, quality, and data leakage.
- The ability to anticipate reward exploitation and design robust rewards, verifiers, and evaluation mechanisms.
- Practical experience working with GPUs and a realistic understanding of their capabilities and limitations.
- Strong judgment around when model training is the right solution—and when a simpler approach is preferable.
- Ability to collaborate effectively within a remote, distributed, and highly autonomous team.
- Experience with post-training techniques such as fine-tuning, reward design, or reinforcement learning, including approaches such as GRPO, is highly desirable.
- Familiarity with RL and fine-tuning frameworks such as TRL, verl, OpenRLHF, or custom training loops is a plus.
- Experience with technologies such as vLLM or SGLang for fast rollouts and FSDP for multi-GPU training is advantageous.
- Experience training tool-using or multi-turn agents, as well as building execution sandboxes, verifiers, evaluation harnesses, or developer tooling, is valuable.
- Familiarity with open-weight model families such as Qwen or Llama and techniques such as LoRA is a plus.
- Opportunity to make a significant impact on a fast-growing developer platform and help shape its future.
- Collaboration with a small, highly experienced team that values technical excellence, creativity, and ownership.
- Competitive salary and equity package.
- Health, dental, and vision benefits.
- Flexible vacation policy.
- Remote-friendly working environment with flexibility and autonomy.