Software Engineer, Early Career (Copy) in Palo Alto, California at Deep Infra Inc.
NewEmployment Type: Full-Time
Deep Infra Inc.
Palo Alto, California, United States
Posted on
New job! Apply early to increase your chances of getting hired.
Explore Related Opportunities
Miscellaneous Computer Occupations jobs near me in CaliforniaJobs near me in CaliforniaMiscellaneous Computer Occupations jobs
Job Description
About DeepInfra
Why this role matters
What You'll Do
What You Bring
Bonus
Why DeepInfra
How we work
DeepInfra is building the foundation for companies to use modern AI in production — simply, reliably, and at scale. Our team has deep experience building large systems that serve hundreds of millions of users, and we're bringing that same level of rigor to a rapidly evolving AI inference space. Our mission is to make advanced AI available to people and teams everywhere.
We're an early, tight-knit team where you can influence product direction, try bold ideas, and drive meaningful work forward quickly. If you want to join a fast-growing company at a defining moment, we'd love to talk.
DeepInfra is backed by leading investors including A.Capital, Felicis, 500 Global, Georges Harik, Samsung Next, Supermicro, Upper90, Peak6, SVAngel and Nvidia.
As DeepInfra's enterprise pipeline grows, our customers need a technical partner who can run rigorous evals, defend benchmarks, and speak fluently to both engineering and procurement — someone who can own the technical win from first call through production.
This is a pioneering role. You'll work closely with Sales, our co-founders, and the engineering team on the deals that matter most. You'll own the technical win end to end: running head-to-head bake-offs against leading AI providers, tuning deployments on the latest hardware, and turning what you learn into reusable assets that make every future deal faster to close. As our first FDE, you'll also define what the function looks like as GTM scales.
- Own the technical win and the POC timeline, working closely with Sales and Engineering, from call one.
- Design and run reproducible benchmark harnesses (TTFT, ITL, throughput/GPU, p95/p99) and quality-parity evals.
- Run head-to-head bake-offs against leading AI providers — and win them.
- Tune model-to-hardware deployments on B200/B300/GB300 NVL72.
- Build cost-per-token models and write migration plans.
- Handle enterprise security and compliance review, and get deployments to launch readiness.
- Own account health post-signature, driving usage reviews and expansion.
- Turn what you learn into reusable benchmark reports, reference architectures, and AE enablement material.
- Customer-facing engineering with an owned technical outcome at an infrastructure or ML platform company.
- Strong Python skills.
- Dual-audience presence with commercial instinct — credible with a skeptical staff engineer, clear with a CFO, and able to tell a technical objection from a procurement one.
- Hands-on experience with inference internals: vLLM, SGLang, or TRT-LLM, batching, KV cache math, quantization.
- Experience with agentic or coding-assistant workloads at scale.
- Prefix-cache-heavy long context workloads.
- Diffusion image/video, ASR/TTS, or multi-LoRA serving.
- Open-source contributions to vLLM or SGLang.
- Deep NVLink/InfiniBand topology knowledge.
- Define DeepInfra's Forward Deployed Engineering function from day one and have a direct impact on its direction.
- Work directly with co-founders and the inference team on the deals that matter most.
- Join a small, high-performing team where your work ships quickly and reaches customers around the world.
- Help shape how enterprises adopt some of the world's leading open-source AI models.
Three traits define the people who thrive here, and this role leans on all three.
Initiative. We take ownership and step in where we can add value. Whether it’s starting something new, improving what exists, or helping move ideas forward, we aim to be proactive and thoughtful in how we contribute.
Drive. We’re energized by hard problems. Building AI infrastructure is complex, and we lean into that. We care about doing things well, moving fast, and continuously improving — because solving meaningful challenges is what motivates us.
Grit. Things don’t always work on the first try — and that’s expected. We stay persistent, adapt quickly, and learn as we go. We take setbacks seriously, but not personally, and use them to get better.
Scan to Apply
Just scan this QR code to apply from your phone.
Job Location
Palo Alto, California, United States
Frequently asked questions about this position
Similar Jobs In Palo Alto, California
Hot Job
Print Production Supervisor - Riot Color
ARC Document Solutions
San Jose, California
New
Distributed Systems Architect
Bright Vision Technologies
Santa Clara, California
New
Multi-Cloud Solutions Architect
Bright Vision Technologies
Milpitas, California
New
Storage Solutions Engineer
Bright Vision Technologies
Milpitas, California
New
Lead Backend Engineer
Bright Vision Technologies
Milpitas, California
Continue to apply
Enter your email to continue. You’ll be redirected to the employer’s application.By clicking Continue, you understand and agree to JobTarget's Terms of Use and Privacy Policy.