Senior Site Reliability Engineer - Dublin in Dublin 2, Dublin at Rapid Ratings International, Inc.
Explore Related Opportunities
Job Description
About RapidRatings
RapidRatings® sets the standard for financial health transparency between business partners, transforming the way the world’s leading companies manage enterprise and financial risk. RapidRatings provides the most sophisticated analysis of the financial health of public and private companies in over 140 countries worldwide. The company’s predictive analytics provide insights into how suppliers, vendors, and other third parties are likely to perform. For more information, visit rapidratings.com.
We are hiring a hands-on Senior Site Reliability Engineer to keep the RapidRatings platform reliable, available, and performant as it scales. Reporting to the Head of Platform & Cloud Engineering and working within the Platform & Cloud Engineering team, you implement and run the reliability practice for our services: SLOs and error budgets, telemetry, incident response, and self-healing automation. You keep the paved-road AWS and Kubernetes infrastructure that our AI systems, agents, and product squads depend on healthy, observable, and cost-aware.
What You'll OwnReliability & SLOs. Implement and maintain SLO and SLI frameworks and error budgets across critical services, and use error-budget signals to guide reliability work and release risk.
Observability & telemetry. Build and run logging, metrics, tracing, and alerting across Datadog, Prometheus, Grafana, and OpenTelemetry, keeping platform health visible and alert noise low.
Incident response. Take a lead role in on-call and incident response, drive down mean-time-to-recovery, and run blameless post-incident reviews with root-cause analysis and follow-through on actions.
Automation & toil reduction. Replace manual, repetitive operations with self-healing automation and tooling, so reliability scales without adding operational overhead.
AWS, Kubernetes & IaC. Operate high-availability, multi-region AWS infrastructure on Kubernetes (EKS) with Terraform or OpenTofu, keeping it resilient, performant, and cost-aware.
Paved-road platform. Maintain and improve the self-service infrastructure patterns and guardrails that let product squads ship securely on top of the platform.
AI infrastructure support. Help operate and keep reliable the model-serving, routing, and governance layer that our AI agents and copilots run on.
What This Role Is Not
This is not application development. Product squads own their own build, test, and deployment on the paved road, using AI tooling on top of the platform you keep reliable. You strengthen the foundational platform and its reliability, not application-level features or scripts for product teams.
Qualifications & Technical Bar
Kubernetes & AWS. Hands-on experience running production Kubernetes (EKS) and the native AWS stack at scale, with infrastructure as code (Terraform or OpenTofu).
Reliability practice. Practical experience with SLOs, SLIs, error budgets, capacity planning, and incident management.
Observability. Expert with logging and monitoring tooling such as Datadog, Prometheus, Grafana, and OpenTelemetry, with a clear grasp of why monitoring and alerting matter.
Systems engineering. Strong Linux background and solid networking (TCP/IP, DNS, routing), running entirely in the cloud, with experience maintaining high-availability systems.
Automation & software. Proficiency in Python, Go, or Bash for building tooling, custom telemetry, and automated remediation, with configuration kept in source control under a GitOps approach.
Security-aware. Comfortable interpreting and remediating dependency, container, and DAST findings (for example ZAP), and working within SOC 2 and ISO 27001 controls.
Ways of working. Experience in an agile or scrum process, with clear written and verbal communication across technical and non-technical colleagues.
AI-assisted operations. Exposure to AIOps or AI-assisted detection and remediation, or readiness to adopt AI tooling as part of day-to-day reliability work.
Who You Are
- A natural problem solver who stays curious, works logically, and digs past symptoms to root cause.
- You treat AI as a strong collaborator, using the platform and guardrails the function provides to lift reliability across the group, not only your own output.
- You take ownership and carry work through to completion, working to the standards, designs, and reviews set across the function.
- You communicate clearly across every tier and keep a calm, positive attitude under pressure, including during production incidents and against tight deadlines.
Role Shape
Balance: a hands-on individual-contributor role focused on reliability engineering and operations, including a share of the on-call rotation, rather than people management.
Scope: part of the Platform & Cloud Engineering team, reporting to the Head of Platform & Cloud Engineering, with no direct reports.
Our Values
Integrity, Innovation, Accountability, Resilience, Community.
Some of our Benefits:
• Health insurance via RapidRatings with a generous allowance.
• Workplace pension with matched contributions up to 5%.
• 25 days of paid time off (PTO)
Pay Transparency Statement (Ireland)
In line with the EU Pay Transparency Directive, we are committed to openness about compensation.
Salary Range: €80,000 to €100,000 base salary (Dublin, Ireland)
The salary offered to the successful candidate is determined by objective criteria including relevant experience and skills, level of responsibility, demands of the role, and business need. As part of the total compensation package, this role may also be eligible for performance bonus and other benefits.