JobTarget Logo

Senior MLOps Platform Engineer {S} in New York at Jobgether

NewJob Function: Engineering
Jobgether
New York, 10455, United States
Posted on
New job! Apply early to increase your chances of getting hired.

Explore Related Opportunities

Job Description

Senior MLOps Platform Engineer {S}

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior MLOps Platform Engineer based in United States.

This role is an opportunity to build and operate the infrastructure powering next-generation Agentic AI services at scale.
You will own the end-to-end MLOps foundation across on-premise Kubernetes environments and AWS.
Working alongside data scientists, product engineers, and SREs, you will turn experimental AI models into reliable production services.
Your work will shape architecture, automation, observability, security, and performance for mission-critical applications.
You will help enable thousands of concurrent AI agents through highly available and governed platform capabilities.
The environment emphasizes strong engineering practices, reusable tooling, and continuous improvement.
A flexible schedule is also available within the pay period to support work-life balance.

Accountabilities:
  • Design, implement, and operate a unified MLOps platform spanning on-premise Kubernetes clusters and AWS, enabling rapid onboarding and consistent governance for Agentic AI services.
  • Develop reusable GitLab CI and CI/CD pipelines covering model packaging, containerization, automated testing, canary deployments, and rollbacks.
  • Build and maintain observability, monitoring, and alerting capabilities using tools such as Prometheus, Grafana, OpenTelemetry, and CloudWatch to track latency, throughput, resource utilization, and data drift.
  • Create self-service CLIs, SDKs, and dashboards that allow data science and product teams to register models, configure inference endpoints, and manage versions with minimal DevOps support.
  • Architect and maintain robust data pipelines for training data, model artifacts, and inference logs across S3 and on-premise object storage.
  • Partner with research and product engineering teams to transform AI prototypes into production-grade services with strong reproducibility, security, compliance, and maintainability.
  • Optimize inference performance through GPU/CPU scaling, model quantization, batching strategies, and other techniques for high-throughput workloads.
  • Champion security, cost efficiency, disaster recovery, and operational best practices across hybrid infrastructure, including IAM, network policies, and secret management.
  • Mentor junior engineers and contribute to technical documentation, knowledge sharing, skills development, and engineering reviews.

Requirements:

  • Bachelor’s degree in Computer Science, Engineering, or a related technical field.
  • 5+ years of experience building and operating production-grade software infrastructure, preferably across hybrid on-premise and cloud environments.
  • Deep expertise with Kubernetes, including cluster provisioning, Helm, operators, custom resources, and container runtimes such as Docker and OCI.
  • Hands-on experience with AWS services including EKS, SageMaker, S3, IAM, CloudWatch, and Step Functions, with the ability to connect on-premise resources to AWS through VPN or Direct Connect.
  • Strong Python software engineering skills plus proficiency in at least one compiled language such as Go, Rust, or Java for developing platform components and SDKs.
  • Experience with CI/CD and GitOps tooling such as Argo CD, Flux, GitLab, or GitHub Actions.
  • Strong understanding of distributed systems, including consensus, fault tolerance, load balancing, and performance tuning for high-throughput, low-latency inference pipelines.
  • Experience with data engineering technologies such as Airflow, Prefect, Kafka, Spark, or Flink and the ability to build robust, versioned data pipelines.
  • Familiarity with observability platforms including Prometheus, Grafana, OpenTelemetry, or ELK, along with experience defining meaningful SLIs and SLOs for AI services.
  • Demonstrated ability to collaborate with research and product teams to transition experimental code into scalable, maintainable production services.
  • Strong problem-solving skills, excellent written and verbal communication, and enthusiasm for building scalable AI infrastructure.
  • Working knowledge of Scrum and Agile software development methodologies is preferred.
  • Ability to work remotely from within the United States. No visa sponsorship is available, and access to export-controlled information may require U.S. Person status or an applicable special license.
  • Successful completion of required pre-employment screening, including credit, background, and drug screening.

Benefits:

  • Remote position within the United States, supporting teams in Aurora, Colorado and King of Prussia, Pennsylvania.
  • Flexible scheduling options within the pay period to support work-life balance.
  • Comprehensive medical, dental, and vision insurance.
  • Employer contributions to eligible HSA accounts.
  • 401(k) retirement plan with strong employer contributions.
  • Three weeks of vacation accrual per year, plus sick leave and time off for unscheduled life events.
  • 13 paid holidays.
  • Upfront tuition assistance for approved degree programs.
  • Annual bonus program based on company and employee performance.
  • Company-paid life insurance, AD&D, short-term disability, and long-term disability coverage.
  • Four weeks of paid parental leave.
  • Employee Assistance Program (EAP).
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Job Location

New York, 10455, United States

Frequently asked questions about this position

Similar Jobs In Other / Non-US, New York

Hot Job

Document Control Specialist / Project Officer Associate

The LiRo Group
Long Island City, New York
Hot Job

IT Solutions Engineer

LMT Technology Solutions
Rochester, New York
New

SAP Security & GRC Engineer

Jobgether
Other / Non-US, New York
New

Application Integration Engineer

Jobgether
Other / Non-US, New York
New

Staff Machine Learning Systems Engineer

Jobgether
Other / Non-US, New York
Continue to apply
Enter your email to continue. You’ll be redirected to the employer’s application.
By clicking Continue, you understand and agree to JobTarget's Terms of Use and Privacy Policy.