Site Reliability Engineer in United States Embassy at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer based in United States.
This role offers the opportunity to build and operate the critical infrastructure behind a rapidly scaling, AI-powered software platform.
You will own systems that support thousands of customers and ensure reliability, scalability, and security across complex environments.
The position focuses on automation, cloud infrastructure, observability, disaster recovery, and continuous improvement.
You will work as a senior technical contributor, solving challenging infrastructure problems and shaping engineering best practices.
The role combines hands-on technical execution with architectural decision-making and cross-functional collaboration.
You will join a fully remote engineering environment that values ownership, transparency, experimentation, and long-term thinking.
This is an opportunity to make a direct impact on the systems powering a modern technology platform.
As a Site Reliability Engineer, you will be responsible for designing, maintaining, and improving the infrastructure, automation, and operational systems that support mission-critical applications. You will help ensure reliability, security, and performance while collaborating with engineering teams to continuously improve how systems are built and operated.
- Automate database lifecycle management, including provisioning, scaling, failover processes, and retirement strategies across large-scale database environments.
- Improve infrastructure security by reducing reliance on static credentials and implementing identity-based authentication approaches.
- Enhance system reliability by minimizing downtime, improving deployment processes, and strengthening disaster recovery capabilities.
- Design and maintain multi-region recovery strategies to ensure fast and reliable failover during service disruptions.
- Own and improve observability systems, including telemetry pipelines, monitoring platforms, and operational visibility tools.
- Maintain and optimize CI/CD workflows to enable efficient, reliable, and secure software delivery.
- Manage cloud infrastructure and automation using modern tools such as Kubernetes, Terraform, Ansible, and AWS services.
- Act as a senior escalation point for production incidents, troubleshooting complex issues and driving long-term solutions.
- Collaborate with engineering teams to improve architecture, operational practices, and platform scalability.
- Leverage AI-powered tools and automation techniques to increase productivity, improve problem-solving, and accelerate engineering workflows.
The ideal candidate is an experienced infrastructure professional who has successfully operated large-scale production systems and thrives in a remote, high-autonomy environment. You bring strong technical expertise, a reliability-focused mindset, and the ability to communicate effectively across distributed teams.
- 5+ years of experience building and operating modern infrastructure systems for senior-level candidates, or 8+ years for staff-level candidates.
- Strong experience with cloud platforms, particularly AWS, and containerized environments such as Kubernetes.
- Hands-on expertise with infrastructure automation tools including Terraform, Ansible, and CI/CD technologies such as GitHub Actions or ArgoCD.
- Experience managing production databases and data platforms, including MongoDB, PostgreSQL, Elasticsearch, or similar technologies.
- Strong understanding of observability and monitoring tools such as Grafana, Prometheus/Mimir, Loki, Tempo, OpenTelemetry, or equivalent platforms.
- Experience supporting mission-critical production environments, including incident response, troubleshooting, and disaster recovery.
- Solid understanding of networking concepts and protocols, including DNS, HTTP, and TCP.
- Ability to design simple, maintainable, and resilient systems using scalable engineering practices.
- Experience working effectively in fully distributed teams with strong written and verbal communication skills in English.
- Based in the United States and able to work effectively across relevant time zones.
- Competitive compensation package with an organization-wide goal-based bonus program.
- Fully remote work environment with flexibility and autonomy.
- Approximately five weeks of paid time off at the start, plus company holidays and a winter holiday break.
- Additional paid time off earned through tenure milestones.
- Flexible work options, including the ability to choose a standard schedule or an 80% work arrangement with adjusted compensation.
- Paid parental leave for primary and secondary caregivers.
- One-month paid sabbatical after five years with the team.
- Comprehensive healthcare benefits for US employees, including medical coverage with most premiums covered, dental, vision, HSA, FSA, and long-term disability insurance.
- 401(k) plan with employer matching contributions and immediate vesting.
- Support for professional growth, experimentation, and the use of modern engineering tools.
- Collaborative remote culture focused on transparency, ownership, continuous improvement, and innovation.