Data Engineer/Senior Data Engineer in Santa Clara, California at PlusAI Inc
Explore Related Opportunities
Job Description
Knowing how well our virtual driver drives — and why it fell short — is what lets us ship with confidence. In this role, you will own that loop end to end: the metrics that quantify driving performance, the pipelines that compute them at fleet scale, the analysis that turns them into judgments about autonomy behavior — including where on the map that behavior changes — the data and tooling that gate software releases, and the agentic workflows that take a detected issue from triage to a proposed fix. You will work primarily in Python across large-scale data processing, geospatial analytics, evaluation frameworks, and LLM-powered automation. We welcome engineers from data engineering, analytics, geospatial, evaluation, or robotics backgrounds; prior autonomous-vehicle experience is helpful but not required.
We are open to candidates at either the Engineer or Software Engineer level. Level will be determined by experience, technical depth, scope of ownership, and demonstrated impact. You do not need experience with every technology in our stack; we value strong fundamentals, ownership, and the ability to learn.
Responsibilities:Define and compute driving performance metrics — covering safety, comfort, progress, interventions, and compliance — and build the scalable pipelines that evaluate them consistently across fleet and simulation data
Analyze road-test and simulation data in depth to identify trends, regressions, and anomalies in autonomy behavior, and turn them into clear findings that engineering teams act on
Build geospatial analytics over fleet driving map-matched metrics, route and corridor performance, location-based clustering of events and issues, and geographic coverage analysis that shows where the virtual driver performs well and where it struggles
Build the data foundations and tooling for release management, including release-over-release comparisons, readiness and gating criteria, and traceable evidence supporting release decisions
Build AI agentic workflows that triage detected issues at scale — clustering and deduplicating failures, attributing root cause, routing to the right owners, and proposing fixes with supporting evidence for engineering review
Ensure that your work is performed in accordance with the company's Quality Management System (QMS) requirements and contribute to continuous improvement efforts
BS, MS, or PhD in Computer Science, engineering, or a related technical field, or equivalent practical experience
Proficiency in Python, with experience building scalable data processing systems or evaluation frameworks
Experience developing metrics and analyzing large-scale time-series, event, or geospatial data, including principled metric definitions, validation, and error analysis
Experience building LLM-powered or agentic workflows for data analysis, evaluation, or automation
Ability to solve open-ended technical challenges and communicate findings clearly to engineering and program stakeholders
Self-driven with a strong sense of ownership: a quick learner who is eager to take responsibility and drive projects forward end to end
Experience with distributed data processing such as Apache Spark, and workflow orchestration such as Airflow or Argo Workflows
Familiarity with LLM agent frameworks (e.g., LangChain, LangGraph, or similar) and prompt/tool-orchestration patterns
Experience with geospatial data and tooling, such as GIS formats, map matching, spatial indexing and joins, PostGIS, GeoPandas, or map-based visualization libraries
Experience with release engineering, quality gating, or automated regression detection
Experience with autonomous vehicles, robotics, or other safety-critical systems
Experience building dashboards or analytics interfaces that expose metrics to users
$130,000 - $200,000 a year