JobTarget Logo

Senior Software Engineer, AI Benchmarking in United States Embassy at Jobgether

NewJob Function: Information Technology
Jobgether
United States Embassy, 0930, Philippines
Posted on
New job! Apply early to increase your chances of getting hired.

Explore Related Opportunities

Job Description

Senior Software Engineer, AI Benchmarking

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Software Engineer, AI Benchmarking based in United States.

This role offers the opportunity to build advanced software systems that help evaluate and improve the safety of frontier AI models.
You will design, scale, and maintain AI benchmarking tools focused on understanding biological capabilities and dual-use risks.
Your work will directly support critical decisions made by researchers, policymakers, and technology organizations working on AI safety.
The position combines software engineering excellence with applied AI research in a mission-driven environment.
You will collaborate with a small, highly skilled team building impactful tools at the intersection of artificial intelligence and biosecurity.
This is an ideal opportunity for an engineer who enjoys solving complex technical challenges while contributing to globally important safety initiatives

Accountabilities:

The Senior Software Engineer, AI Benchmarking will be responsible for developing, operating, and continuously improving evaluation infrastructure used to assess advanced AI systems. This role requires strong engineering skills, curiosity about emerging AI technologies, and the ability to independently drive projects from concept to implementation.

  • Conduct pre-release and post-release assessments of advanced and open-source AI models using specialized biosecurity evaluation frameworks.
  • Build and maintain software tools that enable scalable AI capability assessments and research workflows.
  • Develop model adapters, troubleshoot unexpected AI agent behaviors, and improve evaluation reliability and methodology.
  • Expand evaluation systems using modern agent frameworks, advanced prompting strategies, jailbreaking techniques, and elicitation approaches.
  • Design, implement, and maintain infrastructure required to run large-scale AI evaluations and analyses.
  • Contribute to cloud infrastructure development, deployment workflows, and operational improvements.
  • Analyze evaluation results and support the development of public-facing insights and reporting on AI capability trends.
  • Improve engineering practices, code quality, documentation, and maintainability across internal tools and systems.
  • Collaborate with researchers and technical stakeholders to translate evaluation needs into robust software solutions.
Requirements:

The ideal candidate is an experienced software engineer with a strong interest in AI systems, evaluation methodologies, and technology safety. They should be comfortable working independently in a fast-paced environment and have experience building reliable software solutions involving modern AI technologies.

  • 5+ years of professional software engineering experience or significant open-source contribution experience.
  • Strong proficiency in Python and experience developing production-quality software.
  • Experience building applications, tools, or infrastructure involving large language models or AI agents.
  • Experience with Docker and cloud platforms, particularly AWS services such as ECS.
  • Strong understanding of software engineering best practices, including code quality, readability, testing, and maintainability.
  • Ability to work autonomously, prioritize effectively, and contribute within a small, rapidly evolving team.
  • Interest in AI safety, responsible technology development, and reducing risks associated with advanced AI systems.
  • Commitment to maintaining strong human oversight when using AI-assisted coding tools.

Preferred qualifications include:

  • Experience designing or running AI evaluations, agent systems, scoring frameworks, model assessments, red-teaming, or safety research.
  • Experience with infrastructure-as-code tools such as Terraform, OpenTofu, or AWS CDK.
  • Experience developing modern web applications using technologies such as React and TypeScript.
  • Familiarity with AI evaluation frameworks, LLM APIs, or agent development tools.
  • Experience working with technologies such as Inspect AI, OpenAI-compatible APIs, Anthropic APIs, Together APIs, Codex, or Claude Code.
Benefits:
  • Fully remote work opportunity with access to office locations in Cambridge, MA and Berkeley, CA.
  • Opportunity to work on high-impact projects focused on improving the safety and reliability of advanced AI systems.
  • Collaborative environment with a mission-driven, international team of engineers, researchers, and experts.
  • Ability to influence the development of tools used to evaluate frontier AI capabilities.
  • Competitive compensation package based on experience and qualifications.
  • Opportunity to contribute to research and initiatives addressing major global technology and biosecurity challenges.
  • Flexible and inclusive workplace culture that values diverse perspectives and backgrounds.
  • Supportive hiring process including opportunities for technical discussions, coding assessments, and paid trial work where appropriate.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Job Location

United States Embassy, 0930, Philippines

Frequently asked questions about this position

Similar Jobs In United States Embassy, Other

New

Full Stack Engineer (Go/Rust)

Jobgether
United States Embassy, Other
New
New

Storage Systems Engineer

Jobgether
United States Embassy, Other
New

Dealer Services Document Review Specialist

Jobgether
United States Embassy, Other
New

Principal Software Solutions Engineer

Jobgether
United States Embassy, Other
Continue to apply
Enter your email to continue. You’ll be redirected to the employer’s application.
By clicking Continue, you understand and agree to JobTarget's Terms of Use and Privacy Policy.