Senior AI/ML Operations Engineer in New York at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior AI/ML Operations Engineer based in United States.
This senior individual contributor role is responsible for the infrastructure, pipelines, and operational reliability behind production AI and machine learning systems.
You will work across classical ML, generative AI, and agentic applications, helping move solutions from prototypes into dependable production environments.
The role combines platform engineering, MLOps, model lifecycle management, data pipelines, and AI infrastructure.
You will build scalable deployment patterns, operate RAG and agentic systems, and strengthen observability, evaluation, and governance practices.
Working closely with engineering, data science, security, DevOps, and business teams, you will solve complex operational challenges and enable new AI/ML use cases.
The position offers significant hands-on ownership in a technology-driven environment where emerging AI capabilities are applied to real-world healthcare data challenges.
You will also mentor junior engineers and contribute to customer-specific implementations when needed.
- Deploy, promote, and maintain ML and GenAI models, pipelines, and code across environments using established CI/CD infrastructure.
- Develop reusable deployment patterns, tooling, and operational practices that reduce the effort required to launch new AI/ML use cases and client-specific solutions.
- Build and maintain data pipelines supporting classical ML and GenAI workloads, covering ingestion, feature engineering, processing, and serving.
- Operate within Databricks and Snowflake governance frameworks, including access controls and environment boundaries, to support secure and compliant promotion of data, code, and models.
- Independently diagnose and resolve production issues across data pipelines, infrastructure, model-serving systems, and related AI/ML components.
- Automate and monitor production ML inference and feature-engineering workflows, including alerting, incident response, and operational reliability.
- Manage model lifecycles through platforms such as MLflow and Unity Catalog, including experiment tracking, registration, versioning, and controlled promotion across environments.
- Build and maintain infrastructure for retrieval-augmented generation (RAG), including vector search indexing and retrieval pipelines.
- Deploy, host, and maintain MCP servers and tool integrations supporting agentic applications.
- Build evaluation infrastructure for AI systems and collaborate with AI engineering teams on evaluation methodologies and quality standards.
- Implement observability for agents and models through logging, tracing, monitoring, and analysis of production behavior.
- Partner with business stakeholders to understand data and feature requirements and coordinate infrastructure needs with Software Engineering, Data Engineering, Data Science, Security, and DevOps teams.
- Mentor junior engineers on platform practices, operational standards, and reliable AI/ML engineering.
- Contribute to customer-specific implementations when required, including semantic layer configuration and domain-specific analytics initiatives.
Requirements:
- 5+ years of professional experience in AI/ML engineering, MLOps, or a closely related discipline.
- Extensive hands-on AI/ML experience with meaningful depth in either generative AI/agentic systems or classical machine learning, alongside working exposure to the other area.
- Experience with AI/GenAI technologies such as agentic frameworks, RAG systems, vector search, MCP or comparable tool-integration protocols, model-serving or gateway layers, and AI evaluation design.
- Alternatively, strong classical ML experience covering common algorithm families such as XGBoost, gradient boosting, random forests, and neural networks, along with feature engineering, training pipelines, and production deployment.
- Deep hands-on experience with Databricks and/or Snowflake, including AI/ML pipeline development, governance and access-control frameworks such as Unity Catalog, and integration with CI/CD infrastructure.
- Strong experience with model lifecycle and registry platforms such as MLflow, including experiment tracking, model registration, versioning, and promotion across environments.
- Strong proficiency in Python and SQL.
- Demonstrated ability to independently investigate, troubleshoot, and resolve production infrastructure and operational issues.
- Excellent communication skills and the ability to collaborate directly with both technical and non-technical stakeholders.
- A proactive approach to learning new technologies, tools, and frameworks and applying them effectively to real-world projects.
- Experience in healthcare, health insurance, or regulated data environments is a plus.
- Experience building or operating multi-agent systems is a plus.
- Experience with Mosaic AI Gateway or comparable model-serving and gateway platforms is a plus.
Benefits:
- Compensation based on experience, skills, and location, including base salary plus eligibility for performance bonuses and equity grants.
- Unlimited paid time off.
- Work-from-anywhere flexibility.
- Comprehensive health coverage with multiple plan options.
- Equity grants for all employees.
- Growth-focused environment with opportunities for professional development.
- One-time home office setup allowance.
- Monthly cell phone allowance.