Staff Site Reliability Engineer in United States Embassy at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer based in the United States.
This is a high-impact leadership opportunity for an experienced reliability engineering professional who thrives at the intersection of technology strategy, operational excellence, and platform innovation. In this role, you will help shape the future of large-scale cloud infrastructure, observability practices, and production reliability across a rapidly evolving technology environment. Working closely with engineering leaders and technical teams, you will drive long-term reliability initiatives, improve operational maturity, and champion automation at scale. The position offers the chance to influence engineering culture, mentor talented professionals, and implement forward-thinking approaches leveraging AI and machine learning. If you enjoy solving complex distributed systems challenges and building resilient platforms that directly support business success, this role provides a meaningful opportunity to make a lasting impact.
- Define and execute the technical strategy for observability, alerting, platform infrastructure, and overall operational excellence.
- Lead the design and evolution of scalable, secure, reliable, and cost-efficient cloud-native platforms and distributed systems.
- Establish and champion reliability best practices, including SLIs, SLOs, error budgets, capacity planning, operational readiness reviews, and automation initiatives.
- Guide teams through complex production incidents, driving effective response processes and ensuring lessons learned translate into lasting engineering improvements.
- Build and promote self-service platform capabilities that reduce operational toil, improve developer productivity, and strengthen system reliability.
- Drive adoption of AI and machine learning technologies for observability, anomaly detection, incident management, automated remediation, and resource optimization.
- Partner with engineering leadership on critical architectural decisions and long-term platform strategy.
- Mentor engineers across multiple levels, providing technical guidance and fostering a culture of reliability ownership and continuous improvement.
- Serve as a trusted advisor on production risk, operational health, and infrastructure scalability across the organization.
- 12+ years of experience in software engineering, infrastructure engineering, platform engineering, or Site Reliability Engineering.
- At least 6 years of dedicated SRE experience and 3+ years leading complex, cross-functional technical initiatives involving large-scale distributed systems.
- Deep expertise in observability, infrastructure reliability, incident response, capacity management, automation, and operational excellence.
- Advanced experience with container orchestration technologies, particularly Kubernetes, and modern observability platforms such as Datadog, New Relic, or similar solutions.
- Strong software engineering skills with proficiency in languages such as Python, Go, Bash, or other general-purpose programming languages.
- Proven experience building automation frameworks, internal tooling, platform services, or Infrastructure-as-Code solutions.
- Solid understanding of cloud-native architectures, distributed systems design, and production-scale operational challenges.
- Demonstrated ability to mentor engineers, influence technical direction, and communicate effectively with technical and non-technical stakeholders, including executive leadership.
- Experience applying AI/ML concepts to operational workflows, observability, or platform management is highly desirable.
- Familiarity with compliance-focused or regulated environments such as FedRAMP, CJIS, HIPAA, SOC 2, or PCI is preferred.
- Competitive and equitable compensation package.
- Comprehensive medical, dental, and vision insurance coverage for eligible employees.
- Maternity and paternity leave programs.
- Short-term and long-term disability coverage.
- Opportunity to work within a fast-growing and innovative technology environment.
- Access to experienced leadership and mentorship from highly skilled professionals.
- Strong focus on professional development and continuous learning.
- High-quality company merchandise and employee perks.
- Collaborative culture centered on innovation, ownership, and impact.