Senior Site Reliability Engineer in United States Embassy at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer based in United States.
This role offers the opportunity to build and maintain reliable infrastructure that supports innovative healthcare technology solutions.
You will help design scalable systems, improve operational efficiency, and ensure high availability across complex technology environments.
Working closely with engineering teams, data professionals, and technical leaders, you will drive automation and modern infrastructure practices.
The position focuses on creating resilient platforms through cloud technologies, container orchestration, and advanced reliability engineering principles.
You will play a key role in reducing operational complexity, improving developer productivity, and shaping infrastructure strategy.
This is an ideal opportunity for an experienced SRE professional who enjoys solving complex technical challenges and building systems that make a meaningful impact.
As a Senior Site Reliability Engineer, you will be responsible for designing, automating, and maintaining scalable infrastructure while improving reliability, performance, and developer experience. You will collaborate across engineering disciplines to build robust platforms, streamline operations, and support mission-critical workloads.
- Design and implement systems for application and infrastructure lifecycle management, including CI/CD pipelines, continuous deployment, and Kubernetes cluster operations.
- Develop automation solutions to reduce manual processes, eliminate operational bottlenecks, and improve engineering efficiency.
- Monitor, troubleshoot, and resolve infrastructure issues while minimizing downtime and maintaining service reliability.
- Support and improve containerized workloads using technologies such as Kubernetes and modern cloud-native tooling.
- Contribute to the evolution of site reliability engineering practices, standards, and team objectives.
- Improve infrastructure processes, including deployment workflows, database changes, monitoring, and operational tooling.
- Collaborate with engineering, data, and technical teams to optimize system performance, scalability, and reliability.
- Participate in incident response and alert management activities to maintain production availability.
- Promote a collaborative engineering culture focused on innovation, ownership, and continuous improvement.
- Evaluate and implement solutions that enhance infrastructure security, automation, and operational maturity.
The ideal candidate brings strong experience in site reliability engineering, cloud infrastructure, and software development practices. You should be comfortable working independently, solving complex technical problems, and collaborating with cross-functional teams in a fast-changing environment.
- 5+ years of programming experience with proficiency in languages such as Python, Go, or Shell scripting.
- Strong experience with containerization and orchestration technologies, including Docker, Containerd, and Kubernetes.
- Hands-on experience with cloud platforms such as AWS, GCP, or Azure.
- Knowledge of cloud-native technologies and tools such as Helm, gRPC, Prometheus, and related CNCF solutions.
- Solid understanding of networking concepts, including TCP/IP, UDP, DNS, firewalls, routing, and load balancing.
- Experience with Linux system administration and Linux architecture principles.
- Strong understanding of SRE concepts, including monitoring, automation, performance optimization, and reliability engineering.
- Experience building and maintaining CI/CD pipelines and modern deployment workflows.
- Ability to proactively identify problems, work with limited guidance, and drive solutions independently.
- Strong communication and collaboration skills with the ability to work effectively across technical teams.
- Adaptability and willingness to learn new technologies in a rapidly evolving environment.
- Competitive base salary range of $160,000–$208,000 USD.
- Equity opportunities and performance-based compensation programs.
- Comprehensive medical, dental, and vision coverage.
- 401(k) matching and retirement support programs.
- Flexible paid time off policy and initiatives supporting work-life balance.
- Remote-first work environment with flexibility to collaborate from anywhere.
- Mental health resources and employee wellbeing programs.
- Professional development opportunities, including learning programs, mentorship, and career growth support.
- Employee stock purchase opportunities with discounted equity options.
- Home office setup reimbursement and monthly internet/cell phone support.
- Paid parental leave and additional family-focused benefits.