Reliability Monitoring Engineer in United States Embassy at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Reliability Monitoring Engineer based in the United States.
The Reliability Monitoring Engineer will design, build, and maintain observability solutions that help engineering teams understand, operate, and improve complex systems.
This role focuses on creating reliable monitoring strategies across metrics, logs, traces, dashboards, and alerting workflows.
The ideal candidate will transform large volumes of operational data into actionable insights that improve system reliability and performance.
You will work across modern cloud-native environments, supporting scalable platforms while optimizing visibility and operational efficiency.
This position requires strong technical expertise, a proactive approach to problem-solving, and the ability to collaborate with engineering teams and stakeholders.
It is an opportunity to shape observability practices, improve incident response capabilities, and strengthen the reliability of critical technology platforms.
The Reliability Monitoring Engineer will own the development and operation of observability platforms, ensuring systems provide accurate, actionable, and efficient operational insights. This role will focus on improving monitoring maturity, reducing operational noise, and enabling teams to make data-driven reliability decisions.
- Design, implement, and maintain observability solutions covering metrics, logging, tracing, dashboards, and alerting workflows.
- Build and operate scalable telemetry collection pipelines and monitoring infrastructure.
- Configure and optimize platforms such as Prometheus, Grafana, and commercial observability solutions.
- Develop monitoring strategies that improve system visibility, reliability, and operational efficiency.
- Implement distributed tracing, structured logging, and OpenTelemetry-based solutions.
- Manage high-volume, high-cardinality metrics and log environments while ensuring performance and scalability.
- Create and maintain dashboards, alerts, and reporting solutions that provide actionable insights for engineering teams.
- Support Site Reliability Engineering (SRE) practices through service level objectives (SLOs), error budgets, and reliability improvements.
- Integrate observability tools with CI/CD pipelines, incident management systems, and operational workflows.
- Analyze monitoring data to identify trends, improve performance, and optimize infrastructure costs.
- Collaborate with engineering and business teams to translate technical signals into meaningful operational outcomes.
The ideal candidate will have strong experience in reliability engineering, observability platforms, and cloud-native infrastructure. They should be comfortable designing scalable monitoring solutions, troubleshooting complex environments, and communicating technical insights across teams.
- Bachelor’s degree in Computer Science, Engineering, or a related technical field.
- 5+ years of experience in SRE, platform engineering, reliability engineering, or observability-focused roles.
- Hands-on experience with Prometheus, Grafana, and at least one major observability platform such as Datadog, New Relic, or Splunk.
- Strong understanding of OpenTelemetry, distributed tracing, telemetry pipelines, and structured logging practices.
- Proficiency in at least one programming language such as Go, Python, or Java.
- Experience managing high-throughput metrics and log processing systems.
- Solid understanding of SRE principles, including SLOs, SLIs, and error budgets.
- Experience integrating observability solutions with CI/CD platforms and incident response processes.
- Strong knowledge of Linux systems, networking concepts, and container technologies.
- Excellent troubleshooting, analytical, communication, and collaboration skills.
- Experience with observability technologies such as Thanos, Mimir, Cortex, Loki, or Tempo is a plus.
- Familiarity with eBPF-based monitoring tools, open-source observability projects, and cost optimization strategies is preferred.
- Exposure to regulated environments requiring audit-ready logging and monitoring is beneficial.
- Competitive annual salary range of $100,000–$150,000, depending on experience and qualifications.
- Fully remote work opportunity within the United States.
- Full-time employment with opportunities for professional growth and career advancement.
- Opportunity to work on modern reliability, monitoring, and cloud technology initiatives.
- Collaborative environment focused on innovation and technical excellence.
- Exposure to large-scale systems, distributed platforms, and advanced observability practices.
- Benefits package and employee support programs.
- Opportunity to contribute to impactful technology solutions across enterprise environments.