JobTarget Logo

Staff ML Engineer – AWS Trainium & SageMaker in New York at Jobgether

NewJob Function: Human Resources
Jobgether
New York, 10455, United States
Posted on
New job! Apply early to increase your chances of getting hired.

Explore Related Opportunities

Job Description

Staff ML Engineer AWS Trainium & SageMaker

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff ML Engineer – AWS Trainium & SageMaker based in United States.

As a Staff ML Engineer, you will design, train, optimize, and operate production machine learning workloads on AWS Trainium and Amazon SageMaker.
You will work deeply across the ML stack, from PyTorch training code and distributed execution to accelerator hardware and compiler behavior.
The role requires understanding how training workloads behave on custom silicon rather than treating infrastructure as a black box.
You will diagnose complex training issues, optimize throughput and cost, and build reliable pipelines for real production workloads.
You will work directly with engineering teams to translate business and technical requirements into scalable training solutions.
The environment is highly hands-on, technical, and client-facing, with a focus on solving specialized problems that require deep engineering expertise.
This is an opportunity to work at the intersection of machine learning, cloud infrastructure, distributed systems, and custom AI acceleration.

Accountabilities:
  • Train and operate machine learning models using Amazon SageMaker with AWS Trainium as the underlying compute infrastructure.
  • Develop and optimize PyTorch training workloads for execution on AWS Trainium.
  • Analyze how PyTorch code compiles and executes across Trainium and NeuronCore architecture.
  • Optimize memory utilization, throughput, and other performance characteristics of training workloads.
  • Diagnose training failures and performance issues caused by hardware, compiler behavior, device configuration, data, or model code.
  • Distinguish model- and data-level problems from accelerator- and compiler-level issues during troubleshooting.
  • Translate high-level requirements for Trainium workloads into complete, production-ready training pipelines.
  • Design training workflows that balance performance, scalability, reliability, and cloud infrastructure costs.
  • Tune distributed and multi-device training workloads for throughput and cost efficiency.
  • Operate production training workloads and help ensure their reliability throughout the ML lifecycle.
  • Work directly with client engineering teams to scope, design, and deliver specialized machine learning workloads.
  • Collaborate with internal engineering teams to develop production-grade solutions rather than isolated prototypes or notebook experiments.
  • Investigate complex technical issues across the ML software and hardware stack.
  • Apply a deep understanding of accelerator behavior to improve training architecture and implementation decisions.
  • Contribute to production engineering practices around deployment, monitoring, troubleshooting, and operational reliability.
  • Help translate emerging AI infrastructure capabilities into practical production solutions.
  • Communicate technical findings and tradeoffs clearly with both technical stakeholders and client teams.
Requirements:
  • Strong hands-on experience with PyTorch, including a deep understanding of training workflows.
  • Experience with distributed or multi-device model training is highly desirable.
  • Production experience using Amazon SageMaker for model training and/or inference.
  • Strong understanding of how machine learning workloads interact with accelerator hardware and device-specific compilation.
  • Ability to debug issues that originate at the hardware, accelerator, compiler, or runtime layer rather than solely within model or data code.
  • Experience optimizing training workloads for performance, throughput, memory utilization, or cost.
  • Strong Python programming fundamentals.
  • Understanding of distributed training concepts and production ML infrastructure.
  • Ability to design and operate end-to-end managed training pipelines.
  • AWS experience and familiarity with cloud-based machine learning infrastructure.
  • AWS Trainium or AWS Inferentia experience and familiarity with the AWS Neuron SDK is strongly preferred.
  • Candidates without direct Trainium or Inferentia experience may also be considered if they have deep PyTorch expertise and demonstrated ability to quickly learn new hardware targets.
  • Strong analytical and problem-solving skills, particularly when diagnosing complex system-level issues.
  • Ability to reason across multiple layers of the technology stack, from model code through frameworks, compilers, accelerators, and cloud infrastructure.
  • Experience working in a production engineering environment with high standards for reliability and delivery.
  • Strong communication and collaboration skills for working directly with client and internal engineering teams.
  • Comfortable operating in ambiguous environments where requirements and technical challenges may evolve.
  • Willingness and ability to travel approximately 20% within the United States.
  • Must be legally authorized to work in the United States.
Benefits:
  • Opportunity to work on production AI systems using AWS Trainium and Amazon SageMaker.
  • Exposure to custom AI accelerator hardware, compiler behavior, distributed training, and advanced ML infrastructure.
  • Hands-on work with complex machine learning workloads for enterprise clients.
  • Direct collaboration with client and internal engineering teams.
  • Opportunity to solve specialized technical problems that extend beyond conventional ML application development.
  • Production-focused engineering environment emphasizing systems that ship and operate at scale.
  • Approximately 20% U.S.-based travel associated with the role.
  • Final interview and onboarding may require onsite participation.
  • Professional growth through work across machine learning, cloud infrastructure, hardware acceleration, and production engineering.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Job Location

New York, 10455, United States

Frequently asked questions about this position

Similar Jobs In Other / Non-US, New York

Hot Job

Production Supervisor

Cafe Spice
Beacon, New York

Slime Production Associate

Sloomoo Institute LLC
New York, New York

Production Professional [Part Time]

Bread Alone Bakery
Lake Katrine, New York

1st Shift Supervisor

Dimar Manufacturing
Clarence, New York

GENERAL WAREHOUSE ASSOCIATE

Dival Safety Equipment
victor, New York
Continue to apply
Enter your email to continue. You’ll be redirected to the employer’s application.
By clicking Continue, you understand and agree to JobTarget's Terms of Use and Privacy Policy.