JobTarget Logo

AI Systems Performance Specialist in United States Embassy at Jobgether

NewJob Function: Human Resources
Jobgether
United States Embassy, 0930, Philippines
Posted on
New job! Apply early to increase your chances of getting hired.

Explore Related Opportunities

Job Description

AI Systems Performance Specialist

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for an AI Systems Performance Specialist based in United States.

The AI Systems Performance Specialist will optimize large-scale artificial intelligence systems by improving performance, efficiency, and scalability across training and inference workloads.
This role focuses on maximizing throughput, reducing latency, and lowering infrastructure costs through advanced optimization techniques.
You will work across the AI technology stack, from GPU-level optimization and distributed computing to model efficiency and production deployment.
The ideal candidate combines deep machine learning systems expertise with strong engineering discipline and a passion for measurable performance improvements.
This position offers the opportunity to solve complex challenges in AI infrastructure while collaborating with engineering teams building next-generation intelligent systems.
You will contribute to performance standards, optimization strategies, and technical innovations that directly impact production AI capabilities.

Accountabilities:

The AI Systems Performance Specialist will lead efforts to improve the efficiency and reliability of advanced AI workloads through profiling, optimization, and engineering best practices. This role requires strong technical ownership, analytical thinking, and the ability to collaborate across machine learning and infrastructure teams.

  • Profile and optimize end-to-end AI training and inference pipelines to improve throughput, latency, and cost efficiency.
  • Identify performance bottlenecks across data pipelines, model execution, memory usage, communication layers, and infrastructure components.
  • Implement optimization strategies including quantization, sparsity, pruning, and other model efficiency techniques.
  • Optimize distributed training systems using approaches such as tensor parallelism, pipeline parallelism, FSDP, and ZeRO-style sharding.
  • Improve large language model serving performance through techniques such as KV cache optimization, continuous batching, and speculative decoding.
  • Develop and apply compiler-level optimizations using technologies such as Triton, XLA, TorchInductor, or TVM.
  • Optimize data loading, storage access patterns, and dataset sharding strategies for high-performance AI workloads.
  • Build and maintain benchmarking frameworks, regression testing systems, and performance measurement tools.
  • Collaborate with machine learning and platform engineering teams to integrate optimization best practices into production workflows.
  • Drive cost optimization initiatives through improvements in model architecture, hardware utilization, and workload scheduling.
  • Evaluate emerging AI hardware and software technologies and recommend adoption strategies.
  • Create technical documentation, optimization playbooks, and knowledge-sharing materials for engineering teams.
  • Stay current with AI systems research and translate new developments into practical production improvements.
Requirements:

The successful candidate will bring extensive experience in AI systems, performance engineering, or high-performance computing, with a strong ability to analyze and optimize complex machine learning workloads. The ideal profile combines software engineering expertise, deep understanding of modern AI infrastructure, and strong problem-solving skills.

  • Bachelor’s or Master’s degree in Computer Science, Computer Engineering, or a related technical field.
  • 6+ years of experience in performance engineering, machine learning systems, distributed computing, or high-performance computing environments.
  • Strong programming skills in Python and C++.
  • Hands-on experience optimizing deep learning workloads on modern GPU architectures.
  • Deep understanding of distributed training and inference architectures.
  • Experience using profiling and performance analysis tools across CPU, GPU, and distributed systems.
  • Strong knowledge of memory hierarchies, communication primitives, and parallel computing strategies.
  • Familiarity with model compression techniques and understanding their impact on accuracy and performance.
  • Excellent measurement, debugging, and analytical reasoning abilities.
  • Strong communication and collaboration skills with the ability to work effectively across engineering teams.

Preferred qualifications include:

  • Experience optimizing large language model inference systems at production scale.
  • Contributions to AI infrastructure projects such as vLLM, TensorRT-LLM, DeepSpeed, or similar technologies.
  • Experience developing custom GPU kernels using Triton, CUTLASS, or related frameworks.
  • Familiarity with FinOps practices for managing AI infrastructure costs.
  • Technical publications, conference presentations, or community contributions related to AI systems performance.
Benefits:
  • Fully remote work opportunity within the Continental United States.
  • Competitive annual salary range of approximately $100,000 - $150,000, depending on experience and qualifications.
  • Full-time direct employment opportunity.
  • Opportunity to work on advanced AI systems and large-scale machine learning infrastructure.
  • Career growth opportunities within an innovative technology environment.
  • Exposure to cutting-edge AI optimization techniques and emerging technologies.
  • Collaborative culture focused on engineering excellence, learning, and continuous improvement.
  • Opportunity to contribute to impactful cloud, AI, and enterprise technology solutions.
  • Inclusive workplace committed to equal opportunity and professional development.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Job Location

United States Embassy, 0930, Philippines

Frequently asked questions about this position

Similar Jobs In United States Embassy, Other

New

Data Engineering Specialist – AI

Jobgether
United States Embassy, Other

AI Trainer - Mandarin/Cantonese

Wing Assistant
Singapore, Other

AI Trainer - Swahili

Wing Assistant
Uganda, Other
Continue to apply
Enter your email to continue. You’ll be redirected to the employer’s application.
By clicking Continue, you understand and agree to JobTarget's Terms of Use and Privacy Policy.