Site Reliability Engineer - FedRAMP in United States Embassy at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer - FedRAMP based in United States.
This role offers the opportunity to strengthen the reliability and operational excellence of a large-scale SaaS platform supporting government and sovereign cloud environments.
You will work at the intersection of cloud infrastructure, security, automation, and software engineering to improve system resilience.
The position focuses on building reliable services, enhancing observability, and supporting incident response in compliance-driven environments.
You will collaborate with experienced engineers across engineering, security, and operations teams to solve complex reliability challenges.
The role provides hands-on ownership of infrastructure improvements while contributing to long-term reliability strategies.
This is an ideal opportunity for an engineer who enjoys solving ambiguous problems, improving systems, and working in a high-impact cloud environment.
The Site Reliability Engineer will help maintain and improve the operational foundation of a secure, cloud-based platform. This role requires strong technical execution, proactive problem-solving, and close collaboration across engineering and operational teams.
- Gain deep understanding of platform workloads, dependencies, and operational workflows through documentation, code analysis, and collaboration with subject matter experts.
- Create and maintain operational documentation, including runbooks, incident guides, onboarding resources, and knowledge-sharing materials.
- Participate in incident response activities, including investigation, mitigation, root cause analysis, and post-incident improvements.
- Support the implementation and maintenance of reliability practices, including SLIs, SLOs, error budgets, and availability improvements.
- Improve system observability by developing monitoring, alerting, dashboards, and instrumentation strategies.
- Reduce operational complexity through automation, tooling improvements, and elimination of repetitive tasks.
- Support infrastructure delivery through infrastructure-as-code, CI/CD pipelines, deployment workflows, and configuration management.
- Contribute to secure and compliant infrastructure changes within regulated environments.
- Collaborate with engineering, security, compliance, and operations teams to improve reliability, communicate risks, and resolve technical challenges.
- Participate in on-call rotations and help ensure platform stability and resilience.
The ideal candidate brings strong software engineering fundamentals, cloud infrastructure experience, and the ability to operate effectively in regulated environments. They are comfortable investigating complex systems, improving reliability practices, and collaborating across technical teams.
- 3+ years of experience in software engineering, including at least 1 year working in Site Reliability Engineering, Platform Engineering, or DevOps roles supporting cloud-hosted services.
- Experience with cloud infrastructure platforms such as Azure or similar cloud providers.
- Familiarity with compliance-focused environments such as government, FedRAMP, CMMC, financial services, or healthcare industries.
- Ability to understand and troubleshoot application code to investigate system behavior independently.
- Experience with observability and monitoring tools such as Prometheus, Grafana, OpenTelemetry, or ELK stack.
- Hands-on experience with infrastructure-as-code tools such as Terraform, Terragrunt, or Pulumi.
- Experience with container orchestration platforms, particularly Kubernetes.
- Experience managing CI/CD workflows using tools such as GitHub Actions, Azure DevOps, GitLab CI, or ArgoCD.
- Strong programming skills in languages such as TypeScript, JavaScript, Go, Java, C#, or similar.
- Understanding of distributed systems concepts, networking fundamentals, and cloud reliability principles.
- Strong written and verbal communication skills with the ability to explain technical concepts clearly.
- Experience with government or sovereign cloud environments, SaaS platforms, multi-tenant systems, resilience testing, or chaos engineering is a plus.
- Familiarity with AI-assisted development workflows and LLM-powered tools for automation, documentation, or engineering productivity is beneficial.
- Competitive compensation package with geographic-based salary ranges:
- Zone 1: $151,500 - $252,500 USD
- Zone 2: $138,900 - $231,400 USD
- Zone 3: $126,300 - $210,400 USD
- Zone 4: $109,800 - $183,000 USD
- Unlimited paid time off and 12 paid holidays, including dedicated company wellness days.
- Paid volunteer time and community engagement opportunities.
- Paid parental leave programs.
- Medical, dental, and vision coverage starting from the first day.
- Mental health support, therapy resources, and digital wellness programs.
- 401(k) retirement plan with company matching contributions.
- Fertility, adoption, and surrogacy support.
- Virtual veterinary care benefits.
- Legal services, identity protection, and supplemental insurance options.
- Healthcare, dependent care, and commuting spending accounts.
- Access to professional development resources, learning platforms, mentoring, workshops, and training opportunities.
- Opportunity to work remotely while collaborating with global engineering teams.
- Exposure to impactful cloud reliability projects within secure and regulated environments.