Site Reliability Specialist (Observability & Kubernetes) in United States at Everbridge
Explore Related Opportunities
Job Description
At Everbridge, we’re building a resilient, scalable, and secure cloud platform that powers critical services used around the world. We’re looking for a Senior Platform Site Reliability Specialist to own, operate, and evolve our enterprise observability platform.
In this role, you will be responsible for the up-keep, reliability, scalability, and strategic growth of Everbridge’s observability stack, EKS, and supporting services, ensuring our engineering teams have deep visibility into system health, performance, and reliability across a large-scale, cloud-native environment. You will also be working with other cloud technologies within the AWS and GCP areas.
We’re looking for someone who shows up for the team, not just themselves. This role works best for a person who communicates clearly, collaborates easily, and treats interactions with other teams with respect and professionalism. You should be comfortable being involved, offering support, and helping move work forward without ego. We value people who build trust, keep things running smoothly, and make the teams around them better.
- Head the design, operation, and evolution of Everbridge’s observability stack
- Build and maintain a highly available, scalable observability platform
- Standardize instrumentation, dashboards, alerts, and SLOs
- Support incident response, root cause analysis, and capacity planning
- Operate and scale Grafana and technology
- Grafana Loki (logs)
- Grafana Mimir (metrics)
- Grafana Tempo (tracing)
- Grafana Alerting
- Maintain reliability and security of EKS clusters running observability
- Manage cluster lifecycle and upgrades
- Terraform for infrastructure provisioning
- HashiCorp Packer
- Gitlab CI/CD at Scale
- 6+ years in SRE / Platform Engineering
- Strong Grafana ecosystem experience
- Kubernetes and Amazon EKS expertise
- Terraform proficiency
- OpenTelemetry experience
- Large-scale observability systems
- Cost optimization experience
$118,700 - $145,000 a year