JobTarget Logo

Staff Software Engineer, IDSM in United States at Jobgether

NewJob Function: Information Technology
Jobgether
United States, United States
Posted on
New job! Apply early to increase your chances of getting hired.

Explore Related Opportunities

Job Description

Staff Software Engineer, IDSM

This position is listed on behalf of a partner company, which manages all applications and next steps. Our partner is looking for a Staff Software Engineer, IDSM based in United States .

We are looking for an experienced Staff Software Engineer to help advance an enterprise platform for developing, deploying, and managing machine learning and AI solutions at scale. In this senior individual contributor role, you will design scalable backend systems and build capabilities that simplify the machine learning lifecycle, from model development and registration to deployment and monitoring. You will work closely with customers and cross-functional engineering teams to deliver reliable integrations with cloud-based machine learning platforms. Your work will also support improved resource discovery, model observability, and large language model (LLM) hosting. This opportunity is ideal for a technically strong engineer who enjoys solving complex distributed systems challenges and shaping architecture in a collaborative, innovation-driven environment. You will have meaningful ownership over platform capabilities that help organizations operationalize AI securely, efficiently, and reliably.

Accountabilities
  • Scalable Platform Engineering: Design, develop, and maintain high-performance backend services that support enterprise-scale data science and machine learning workflows.

  • Cloud Model Deployment: Collaborate with customers and internal teams to design and implement model deployment solutions integrating with platforms such as AWS SageMaker and Azure Machine Learning.

  • Model Lifecycle Management: Enhance capabilities across model development, training, registration, deployment, and ongoing management to improve reliability and efficiency.

  • Data Science Resource Discovery: Help design and launch a centralized data science catalog that enables users to discover, explore, and summarize resources across the platform.

  • Model Monitoring and Observability: Integrate monitoring capabilities that provide a comprehensive view of deployed model health, performance, and operational status.

  • Metadata and Discoverability: Expand tagging and metadata functionality across platform entities to improve organization, tracking, reuse, and resource discovery.

  • LLM Hosting and Infrastructure: Extend large language model hosting capabilities to meet customer requirements for scalability, performance, and operational logging.

  • API Design and Integration: Build secure, reliable, and scalable APIs, including RESTful APIs and gRPC services, and integrate backend systems with frontend interfaces and third-party services.

  • Performance Optimization: Profile, troubleshoot, and optimize backend applications and distributed workloads across cloud environments and containerized infrastructure.

  • Testing and Delivery Automation: Implement comprehensive unit, integration, and end-to-end testing practices, and contribute to robust continuous integration and continuous delivery (CI/CD) pipelines.

  • Cross-Functional Collaboration: Partner with engineering, product, and customer-facing teams to translate technical and business requirements into practical, maintainable solutions.

  • Technical Leadership: Contribute to architectural decisions, share engineering best practices, support team learning, and continuously improve platform quality and performance.

Requirements:

  • At least 8 years of experience in software engineering, primarily in an individual contributor capacity.

  • Strong knowledge of artificial intelligence and machine learning systems, including model development, training, optimization, deployment, and lifecycle management.

  • Demonstrated experience designing, building, and maintaining scalable backend systems in distributed computing environments.

  • Proficiency in designing and implementing secure, high-performance APIs, including RESTful APIs and/or gRPC.

  • Experience with cloud platforms such as AWS, Microsoft Azure, or Google Cloud Platform.

  • Strong understanding of backend performance profiling, troubleshooting, and optimization in cloud-based or distributed environments.

  • Familiarity with containerization and orchestration technologies such as Docker and Kubernetes.

  • Experience implementing automated testing strategies, including unit, integration, and end-to-end tests, as well as maintaining CI/CD pipelines.

  • Ability to collaborate effectively across engineering teams and integrate backend services with frontend applications and external systems.

  • Strong analytical and problem-solving skills, with the ability to address complex technical challenges and deliver maintainable, reliable solutions.

  • Excellent communication skills and the ability to work collaboratively with technical stakeholders and customers.

  • A growth mindset, curiosity, and a willingness to learn, share knowledge, and continuously improve engineering practices.

  • Experience with distributed computing frameworks such as Apache Spark, or machine learning platforms such as AWS SageMaker and Azure Machine Learning, is an advantage.

  • Familiarity with model registries, model monitoring, enterprise MLOps, and LLM hosting infrastructure is beneficial.

  • Ability to work remotely from Canada.

Benefits:

  • Competitive compensation: Annual US base salary range of USD $200,000–$235,000. The final range may vary based on experience, qualifications, and location.

  • Equity opportunities: Potential eligibility for equity compensation.

  • Performance incentives: Additional company bonus opportunities may be available.

  • Retirement benefits: Access to a 401(k) plan, subject to applicable eligibility and location requirements.

  • Health coverage: Medical, dental, and vision benefits, subject to the applicable benefits package.

  • Wellness support: Potential wellness stipends to support employee well-being.

  • Remote work flexibility: Opportunity to work remotely from Canada.

  • Technical ownership: Meaningful responsibility for designing and improving enterprise AI and machine learning platform capabilities.

  • Complex engineering challenges: Work on distributed systems, cloud infrastructure, model monitoring, and LLM hosting at enterprise scale.

  • Collaborative culture: Join a team that values continuous learning, open communication, diverse perspectives, and ongoing improvement.

  • Professional development: Opportunities to deepen expertise in AI/ML infrastructure, cloud engineering, and modern software architecture.

How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Job Location

United States, United States

Frequently asked questions about this position

Apply For This Position

Apply Now