JobTarget Logo

HPC Storage Engineer - West Coast in New York at Jobgether

NewJob Function: Information Technology
Jobgether
New York, 10455, United States
Posted on
New job! Apply early to increase your chances of getting hired.

Explore Related Opportunities

Job Description

HPC Storage Engineer - West Coast

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a HPC Storage Engineer - West Coast based in the United States.

This is a senior, hands-on infrastructure engineering role focused on building and scaling a multi-region storage platform for demanding AI workloads.
You’ll own critical storage systems spanning network volumes, local NVMe, and S3-compatible object storage at petabyte scale.
Your work will directly influence training, fine-tuning, and inference performance, including cold-start speed, data streaming, and workload reliability.
You’ll operate at the intersection of storage, networking, hardware, and software, with substantial ownership from architecture through production operations.
The role offers significant latitude to automate manual processes, establish SLOs, optimize performance, and shape long-term storage strategy.
You’ll collaborate closely with SRE, network engineering, supply chain teams, and infrastructure partners in a fast-moving remote environment.
This opportunity is ideal for an engineer who enjoys solving complex production problems and building infrastructure that serves millions of developers.

Accountabilities
  • Own the capacity, durability, availability, and performance of network volumes, local NVMe, and S3-compatible object storage.
  • Tune the complete I/O path, including device and filesystem configuration, caching, read-ahead strategies, replication, erasure coding, and client-side mount behavior.
  • Diagnose complex storage and performance issues end to end, identifying root causes and implementing durable solutions.
  • Lead capacity expansions, hardware refreshes, migrations, and data rebalancing while minimizing or eliminating customer-visible disruption.
  • Design and optimize the networking infrastructure supporting storage workloads, including high-throughput east-west fabrics, MTU and jumbo-frame configuration, congestion and flow control, multipath, and NIC/offload settings.
  • Optimize storage traffic across RDMA/RoCE and high-speed InfiniBand or Ethernet environments, collaborating with network engineering on topology, oversubscription, and cross-region data movement.
  • Develop and ship production software in Go, Python, Rust, or similar languages for storage control-plane services, provisioning, data movement, and monitoring.
  • Build and extend integrations with internal control-plane services, S3-compatible interfaces, CSI drivers, Kubernetes APIs, vendor platforms, and cloud-provider APIs.
  • Replace manual operational procedures with reliable automation and infrastructure-as-code, while participating fully in code reviews, testing, and CI.
  • Instrument storage infrastructure with meaningful metrics covering IOPS, throughput, latency, errors, retries, capacity utilization, and tenant consumption.
  • Build dashboards, SLOs, and alerts that identify degradation proactively and support reliable production operations.
  • Participate in an on-call rotation and lead blameless post-incident follow-through, ensuring lessons learned translate into measurable system improvements.
Requirements
  • 8+ years of experience in infrastructure, storage, or systems engineering, including substantial ownership of production storage environments at scale.
  • Deep practical experience with at least one distributed storage platform such as Ceph, MinIO, Lustre, GPFS/Spectrum Scale, MooseFS, WekaFS, VAST, ZFS-based systems, or a comparable technology.
  • Strong knowledge of Linux internals and the storage stack, including block devices, filesystems, NVMe, page cache, I/O schedulers, NFS/SMB, iSCSI, and NVMe-oF.
  • Hands-on experience building or operating S3-compatible object storage services.
  • Strong networking fundamentals and demonstrated experience tuning networks specifically for storage workloads.
  • Proven ability to write and ship production-quality software using Go, Python, Rust, or a similar programming language, beyond scripting alone.
  • Experience with observability platforms such as Prometheus, Grafana, Datadog, or equivalent, including designing meaningful metrics and monitoring strategies.
  • Demonstrated ability to analyze and resolve performance problems under real production pressure.
  • Self-directed approach, with the ability to take broad infrastructure goals, develop an options analysis, diagnose problems, and execute solutions with minimal supervision.
  • Strong continuous-improvement mindset, with a track record of eliminating operational toil and replacing recurring manual work with automation.
  • High ownership and accountability, including the willingness to follow problems across team boundaries through to resolution.
  • Collaborative, low-ego communication style combined with confidence in technical decision-making.
  • Experience with AI/ML storage workloads, including checkpointing, dataset streaming, model-weight distribution, GPU-adjacent data locality, or GPUDirect Storage, is preferred.
  • Familiarity with Kubernetes storage internals, including CSI drivers, PV/PVC lifecycles, StatefulSets, and local persistent volumes, is a plus.
  • Bare-metal or colocation experience, including hardware selection, vendor management, firmware, and physical failure domains, is beneficial.
  • Experience operating multi-tenant environments where isolation, fairness, and quality of service are critical is preferred.
  • Background in a rapidly scaling cloud or infrastructure provider is advantageous.
Benefits
  • Base salary: $180,000–$260,000, with the final range determined based on career level, experience, qualifications, and location.
  • Meaningful equity through stock options, giving employees an opportunity to share in the company’s growth.
  • Generous medical, dental, and vision coverage.
  • Flexible paid time off.
  • Remote-first work environment with collaborative teams and Slack as a primary internal communication channel.
  • $1,200 home office and equipment stipend to help create an effective remote workspace.
  • Opportunity to work on cutting-edge AI infrastructure with a strong emphasis on ownership, learning, and technical impact.
  • Inclusive workplace committed to equal opportunity and respect for people from diverse backgrounds.
  • Candidates must be legally authorized to work in the United States; employment visa sponsorship is not currently available.
How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Job Location

New York, 10455, United States

Frequently asked questions about this position

Similar Jobs In Other / Non-US, New York

NewUrgently Hiring

3rd Shift Material Handler

Alro Steel Corporation
Buffalo, New York
Hot Job

Laser Operator

Dimar Manufacturing
Clarence, New York
New

Port Chester, NY XPS Cargo Van Contractors Needed

R.A.S. Logistics
Port Chester, New York
New

CNC Machine Operator

Graham Manufacturing
Batavia, New York

Lead Data Center NOC Engineer

Jobgether
Other / Non-US, New York
Continue to apply
Enter your email to continue. You’ll be redirected to the employer’s application.
By clicking Continue, you understand and agree to JobTarget's Terms of Use and Privacy Policy.