Staff Site Reliability Engineer, Ads in New York at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer, Ads based in United States.
This is a senior technical leadership role focused on strengthening the reliability of a large-scale advertising ecosystem.
You’ll shape reliability strategy across critical systems spanning ad serving, auctions, targeting, reporting, measurement, and billing.
The role combines hands-on engineering with architecture, automation, incident leadership, and long-term platform resilience.
You’ll partner with engineering leaders and multiple technical teams to influence roadmaps and improve operational excellence.
Your work will directly support highly available, low-latency systems where reliability and performance have meaningful business impact.
You’ll also mentor engineers and help establish measurable reliability practices across a broad, distributed technology environment.
The position offers significant ownership and the opportunity to define how reliability engineering evolves at organizational scale.
- Lead reliability initiatives across critical advertising domains, including ad serving, auctions, targeting, reporting, measurement, attribution, and billing.
- Partner with engineering leadership to establish and execute roadmaps focused on reliability, scalability, operational excellence, and developer productivity.
- Design and build scalable platforms, tooling, automation, and infrastructure capabilities that improve system resilience and engineering efficiency.
- Lead architecture reviews and influence technical decisions for high-traffic, revenue-critical distributed systems.
- Establish and monitor reliability metrics and SLOs around critical advertiser and platform journeys, using data to identify risks and prioritize improvements.
- Participate in on-call rotations, lead complex incident investigations, and coordinate cross-functional responses to major production events.
- Identify systemic reliability risks and implement durable solutions that improve availability, performance, scalability, and operational maturity.
- Drive automation, observability, incident management, performance optimization, and other practices that strengthen production reliability.
- Mentor engineers and provide technical leadership across multiple teams, helping raise engineering standards and reliability expertise.
- Collaborate with Product, Data Science, Infrastructure, and Engineering stakeholders to ensure reliability considerations are embedded into product and infrastructure investments.
- 8+ years of experience in Site Reliability Engineering, Infrastructure Engineering, or a related discipline, with experience operating large-scale distributed systems.
- Proven experience evolving and supporting high-traffic, user-facing production environments with demanding availability and performance requirements.
- Deep expertise in distributed systems, scalability engineering, cloud-native architectures, and highly available system design.
- Strong software engineering capabilities, ideally with experience in backend programming languages such as Go.
- Extensive knowledge of observability practices and technologies, including metrics, logging, tracing, alerting, and performance monitoring.
- Demonstrated experience improving reliability through SLOs, automation, incident management, performance optimization, and systematic operational practices.
- Strong troubleshooting and problem-solving abilities across complex, modern distributed technology stacks.
- Excellent cross-functional communication and collaboration skills, with the ability to influence technical direction and align teams around shared reliability goals.
- Experience supporting advertising technology or other large-scale, revenue-critical platforms is highly desirable.
- Familiarity with reliability challenges involving ad serving, real-time auctions, budget pacing, campaign delivery, measurement, attribution, or billing is a strong advantage.
- Experience operating high-QPS, low-latency services where system performance directly affects business outcomes is preferred.
- Experience establishing reliability programs with measurable business and operational results is a plus.
- Hands-on experience with Kubernetes, cloud infrastructure, and large-scale distributed systems is beneficial.
- Familiarity with technologies such as Kafka, ClickHouse, Spark, Flink, or BigQuery is advantageous.
- Experience partnering with Product, Data Science, and advertising engineering teams, or supporting machine learning inference and recommendation systems at scale, is a plus.
- Base salary range of $217,000–$303,900 USD, with final compensation determined by factors such as skills, experience, credentials, and role level.
- Eligibility for equity in the form of restricted stock units.
- Comprehensive health benefits, including medical, dental, and vision coverage.
- 401(k) program with employer matching.
- Workspace benefits and support for a home office.
- Personal and professional development funds.
- Family planning support.
- Flexible vacation and global days off.
- 4+ months of paid parental leave.
- Paid volunteer time off.
- Flexible-first remote work environment within the United States.