Senior Site Reliability Engineer – Telephony & Communications Platform (AWS) in United States Embassy at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Senior Site Reliability Engineer – Telephony & Communications Platform (AWS) based in United States.
This role offers the opportunity to engineer highly reliable, scalable, and secure cloud-based systems that power critical business operations.
You will join a reliability-focused engineering team responsible for building autonomous systems that improve platform performance, resilience, and operational excellence.
The position combines software engineering, infrastructure automation, and cloud architecture to solve complex reliability challenges at scale.
You will design and maintain robust AWS environments, improve observability, and help teams deliver dependable customer experiences.
Working in a collaborative and innovative environment, you will contribute to continuous improvement initiatives and modern engineering practices.
This is an impactful opportunity for an experienced engineer passionate about automation, distributed systems, and building platforms designed for growth.
As a Senior Site Reliability Engineer, you will be responsible for ensuring platform reliability, availability, and performance while driving improvements across infrastructure, automation, and operational processes. You will collaborate with cross-functional teams to build resilient systems, improve engineering standards, and support critical production environments.
- Own the reliability, scalability, and performance of key platform services, ensuring highly available production systems.
- Design, build, and maintain AWS infrastructure using Terraform and Infrastructure as Code principles.
- Develop and improve CI/CD pipelines, automation frameworks, monitoring solutions, and operational tooling.
- Define and maintain reliability practices including Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets.
- Lead incident response activities, perform root cause analysis, and implement long-term reliability improvements.
- Improve observability through logging, metrics, alerting, and proactive system monitoring.
- Collaborate with engineering teams to identify operational risks and implement scalable solutions.
- Participate in a 24/7 on-call rotation and contribute to production support excellence.
- Mentor engineers and promote best practices in reliability engineering, automation, and infrastructure management.
- Support communications and telephony platform reliability, including improvements related to voice infrastructure when applicable.
The ideal candidate is an experienced reliability-focused engineer with deep expertise in cloud infrastructure, automation, and highly available systems. You bring strong technical skills, a problem-solving mindset, and the ability to collaborate across engineering teams to deliver reliable solutions.
- 8+ years of experience in software engineering, infrastructure engineering, or operations roles.
- 4+ years of hands-on Site Reliability Engineering experience supporting production environments.
- Strong experience with AWS services, including EC2, ECS/EKS, VPC, Route 53, IAM, Lambda, S3, and CloudWatch.
- Advanced knowledge of Terraform, Infrastructure as Code, CI/CD pipelines, observability practices, and automation.
- Strong scripting skills with languages such as Python, Bash, or PowerShell.
- Proven experience supporting highly available, low-latency, and scalable production systems.
- Strong understanding of distributed systems, cloud architecture, and reliability best practices.
- Ability to troubleshoot complex technical issues and drive solutions independently.
- Strong communication skills with the ability to collaborate effectively with engineering teams.
- Experience with telephony or communication platforms such as Amazon Connect, Twilio, SIP/RTP, or other VoIP technologies is a plus.
- Familiarity with voice quality metrics such as MOS, jitter, and packet loss is preferred.
- Experience working in regulated environments such as SOC 2, HIPAA, or FedRAMP is a plus.
- Exposure to Azure or Google Cloud Platform environments is beneficial.
- Competitive compensation package with a salary range of $175,000 – $195,000.
- Comprehensive benefits package including medical, dental, and vision insurance for eligible employees.
- Paid maternity and paternity leave.
- Short-term and long-term disability coverage.
- Flexible paid time off policy.
- Fully remote work environment for engineering teams.
- Opportunity to learn from an experienced leadership team and collaborate with talented engineers.
- Access to innovation initiatives such as hackathons and technology-focused projects.
- Company-provided equipment and employee perks.
- Opportunity to contribute to a rapidly growing technology environment focused on improving professional workflows.