Systems Reliability Engineer in United States Embassy at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Systems Reliability Engineer based in United States.
This role offers the opportunity to design, automate, and operate highly reliable distributed systems in a modern engineering environment.
You will work at the intersection of software development and infrastructure operations to improve platform stability, scalability, and performance.
The position focuses on reducing operational complexity through automation, observability, and engineering best practices.
You will help ensure critical systems remain available and resilient while proactively addressing reliability challenges.
The ideal candidate will bring strong systems expertise, programming skills, and a passion for building dependable technology platforms.
You will collaborate with engineering teams to solve complex production challenges and continuously improve operational excellence.
This is a remote opportunity for a technically driven professional who values ownership, innovation, and measurable impact.
The Systems Reliability Engineer will be responsible for maintaining and improving the reliability, performance, and scalability of large-scale production systems. The role requires a balance of software engineering, infrastructure expertise, automation, and operational leadership.
- Design, build, and maintain reliable distributed systems while improving availability and performance.
- Apply software engineering principles to infrastructure and operations challenges to reduce manual effort and operational complexity.
- Develop automation tools and solutions using programming languages such as Python, Go, or Java.
- Operate and troubleshoot Linux-based systems at scale, including networking, performance optimization, and system-level issues.
- Manage Kubernetes and containerized workloads in production environments.
- Build, maintain, and improve CI/CD pipelines supporting both applications and infrastructure.
- Implement and enhance observability solutions using tools such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or similar platforms.
- Monitor system health, define reliability improvements, and contribute to proactive performance management.
- Lead incident response activities, conduct post-incident reviews, and implement preventative improvements.
- Support the development and adoption of reliability practices including automation, service ownership, and operational excellence.
- Collaborate with engineering and cross-functional teams to improve system design and resilience.
The ideal candidate combines strong software engineering capabilities with deep infrastructure knowledge and experience operating complex production environments.
- Bachelor’s degree in Computer Science, Engineering, or a related technical discipline.
- 5+ years of experience in Site Reliability Engineering, DevOps, or production engineering roles supporting large-scale distributed systems.
- Strong programming experience in at least one of the following: Python, Go, or Java.
- Deep hands-on experience managing Linux environments, including networking, troubleshooting, and performance tuning.
- Production experience with Kubernetes and container-based architectures.
- Strong knowledge of observability and monitoring platforms such as Prometheus, Grafana, OpenTelemetry, ELK/EFK, or equivalent solutions.
- Experience designing and managing CI/CD pipelines for applications and infrastructure.
- Understanding of distributed system concepts including consistency models, partitioning, and failure handling.
- Proven experience leading incident response and conducting effective post-mortem reviews.
- Excellent communication, collaboration, and technical documentation skills.
- Preferred experience with SLOs, error budgets, chaos engineering practices, and reliability frameworks.
- Preferred familiarity with cloud platforms such as AWS, Azure, or GCP.
- Preferred background in capacity planning, performance engineering, load testing, or service mesh technologies.
- Competitive annual salary range of $100,000–$150,000.
- Fully remote work environment within the United States.
- Full-time employment opportunity with career growth potential.
- Opportunity to work on large-scale distributed systems and advanced technology projects.
- Collaborative environment focused on engineering excellence and continuous improvement.
- Exposure to modern cloud, automation, observability, and reliability practices.
- Inclusive workplace committed to equal employment opportunities.