MLOps Engineer in Boston, Massachusetts at Applisights.com
Explore Related Opportunities
Job Description
We’re hiring an MLOps Engineer to own the infrastructure our AI runs on.
Our AI systems are already in production and already touching real loans. The next chapter is making them scale, deploy safely, and change quickly — infrastructure that absorbs real production load, a promotion path that makes shipping boring, and internal tooling that lets the team improve AI behaviour without an infrastructure engineer in the loop for every change.
This is not a role where you inherit a mature platform and keep the lights on. You’ll take AI services that were built to prove the idea and turn them into systems that hold up under load, deploy predictably across environments, and give the rest of engineering a safe, fast path to production. You own the arc from “this works on one machine” to “this is a platform the whole company builds on.”
This role is right for you if:
You’ve taken AI or ML systems from a single deployment to infrastructure that scales — and you’ve been on call for the result
You treat infrastructure as a product with users, not a ticket queue
You think deploys should be boring, reversible, and frequent — and you’ve built the systems that make that true
You’d rather remove yourself from the critical path than be the person everyone has to ask
What You’ll Build
Scalable AI Serving Infrastructure
Our AI workloads are moving from early-stage deployment to production scale. You’ll design what they run on:
Re-architect how AI services are deployed and run — from single-host setups to horizontally scalable, orchestrated infrastructure that absorbs traffic spikes without degrading
Design for the specific realities of LLM-backed workloads: long-running requests, streaming responses, bursty concurrency, expensive downstream calls, and upstream rate limits
Own capacity, autoscaling, and unit economics — you should be able to say what we spend per unit of work, and why
Deployment & Environment Promotion
Shipping AI changes should be a routine, reviewable event — not a coordinated risk:
Build the promotion path from development through pre-production to production, with environments that are consistent and reproducible rather than each one a special case
Make everything version-controlled and reviewable — infrastructure, configuration, and application logic on the same rails
Ship CI/CD that gives engineers fast deploys with real rollback, staged rollout, and change history you can audit
Internal Platform & Self-Serve Tooling
The highest-leverage thing you’ll build is the thing that lets other people ship without you:
Build internal tooling that lets engineers — and technically-minded teammates outside engineering — define, modify, and test AI workflow logic without touching deployment plumbing
Design the guardrails that make that safe: validation, versioning, review, staged promotion, and a clear path back when something is wrong
Treat internal users as customers. Success is measured in how many changes ship correctly without you being involved
Reliability, Observability & Operations
You’ll care about whether the system actually holds — not just whether it deployed:
Instrument the AI stack end to end: latency, throughput, failure modes, cost, and output-quality signals
Build alerting that catches real degradation rather than noise, and incident practice that changes the system instead of assigning blame
Own secrets, access control, and environment isolation in a regulated industry where those things carry real consequences