Principal Staff Engineer, AI Platform Research, Data Science in Abbeyville, Colorado at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Principal Staff Engineer, AI Platform Research, Data Science based in United States.
This is a highly technical leadership opportunity focused on building the data and AI infrastructure behind next-generation cybersecurity products. You will architect and develop large-scale platforms supporting LLMs, neural networks, RAG, and sophisticated agentic systems at Exabyte scale. The role combines deep hands-on engineering with technical leadership, research enablement, and mentorship. You will help turn advanced AI research into reliable, production-grade capabilities while establishing modern MLOps and DataOps practices. Working across data science, product, and engineering teams, you will influence platform strategy while remaining deeply involved in implementation and delivery. This is an environment where rapid innovation, engineering rigor, responsible AI adoption, and large-scale distributed systems come together.
- Architect, build, and optimize highly scalable data platforms and pipelines supporting LLMs, neural networks, Retrieval-Augmented Generation (RAG), and AI agentic systems at Exabyte scale.
- Design and deploy agentic workflows and agent-harnessing capabilities that enable autonomous, data-driven security solutions.
- Remain hands-on in software development, writing elegant, production-ready code with a strong focus on performance, maintainability, testing, reliability, and rapid delivery.
- Design fault-tolerant, cost-effective distributed systems using advanced approaches to sharding, partitioning, concurrency, and large-scale data processing.
- Provide technical leadership across data modeling, normalization, semantic cataloging, and platform architecture for AI/ML workloads.
- Establish and advance MLOps and DataOps standards for LLM-powered systems, including monitoring, observability, automation, and zero-touch recovery.
- Own the complete lifecycle of critical AI and data services, from architecture and development through testing, deployment, production monitoring, and continuous improvement.
- Partner with Data Scientists, Product Managers, and engineering teams to transform research prototypes into scalable, secure, production-grade services.
- Drive adoption of modern AI platform technologies and continuously evaluate emerging tools, frameworks, and development practices.
- Champion secure, AI-assisted software development practices while maintaining strong engineering and cybersecurity standards.
- Mentor engineers through technical workshops, architecture discussions, design reviews, and hands-on guidance, strengthening the team's expertise in AI platform engineering.
- Lead by example in engineering excellence, fostering a culture of high-quality execution, technical curiosity, collaboration, and continuous learning.
- Bachelor's, Master's, or PhD in Computer Science, Data Engineering, or a related STEM discipline, or equivalent practical experience.
- 5+ years of progressive experience in Data Engineering or Platform Engineering, including at least 3 years architecting and building AI/ML or Data Science platforms at significant scale.
- Previous experience operating at a Staff-level engineering capacity, with demonstrated technical leadership, architecture ownership, and mentorship responsibilities.
- Strong hands-on expertise in LLM engineering, including fine-tuning, prompt engineering, deployment, Retrieval-Augmented Generation, and agentic workflow development.
- Proven experience designing and delivering large-scale distributed systems, including expertise in sharding, partitioning, concurrency, fault tolerance, and performance optimization.
- Expert-level proficiency in at least one programming language such as Python, Go, Rust, or JVM-based technologies.
- Strong experience with AI/ML and MLOps technologies such as MLflow, SageMaker, Vertex AI, LangChain, or LlamaIndex.
- Experience with distributed processing frameworks such as Spark, Dask, or Flink.
- Strong knowledge of cloud environments including AWS, GCP, or OCI and associated data services.
- Experience with Docker and Kubernetes in production environments.
- Familiarity with streaming technologies such as Kafka or Pulsar.
- Experience with data warehousing and orchestration platforms such as Snowflake, BigQuery, Airflow, or Kubeflow.
- Strong understanding of software engineering best practices, including peer code reviews, resilient architecture, automated testing, observability, and secure development.
- Demonstrated ability to use AI technologies to improve decision-making, streamline workflows, increase efficiency, and drive measurable business outcomes.
- Ability to work effectively across technical and non-technical teams, communicate complex concepts clearly, and operate successfully in a rapidly evolving environment.
- Prior experience in cybersecurity, intelligence, or highly regulated/compliance-focused industries is a plus.
- Contributions to open-source data or AI/ML projects are a plus.
- Base salary: $195,000–$290,000 per year for U.S. candidates.
- Eligibility for bonuses and equity grants.
- Comprehensive health insurance and wellness benefits.
- 401(k) benefits.
- Paid time off, vacation, and holidays to support time away and recharge.
- Paid parental and adoption leave.
- Comprehensive physical and mental wellness programs.
- Professional development opportunities available across career levels and roles.
- Employee networks, geographic communities, and volunteer opportunities that support connection and engagement.
- Flexible, remote-first work environment for U.S.-based employees.
- Opportunity to work on large-scale AI, data, and distributed-systems challenges with significant real-world impact.
- Inclusive workplace culture committed to belonging, equal opportunity, accessibility, and supporting veterans and individuals with disabilities.