Engineering Manager, SRE in Soest, Utrecht at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Engineering Manager, SRE based in Netherlands.
This is a hands-on engineering leadership role responsible for building a highly reliable foundation for a globally distributed technology platform.
You will lead a Site Reliability Engineering team while remaining deeply involved in technical direction and complex infrastructure challenges.
The role combines people leadership with expertise across Kubernetes, AWS, PostgreSQL, CI/CD, observability, infrastructure as code, and reliability engineering.
You will shape how the team balances operational excellence, incident response, reliability improvements, and longer-term engineering initiatives.
A key focus will be maturing SLOs, error budgets, observability, and reliability practices across the wider engineering organization.
You will also act as a trusted technical and organizational partner to engineering, security, and senior leadership.
The environment is fully remote and asynchronous, offering significant autonomy in a fast-growing, globally distributed organization.
- Lead and develop a Site Reliability Engineering team, owning the full career lifecycle of direct reports including onboarding, feedback, performance management, progression, coaching, and hiring.
- Establish a clear team direction and priorities aligned with broader company goals, balancing operational commitments with project delivery and protecting the team’s focus.
- Serve as the team's spokesperson across engineering and with senior leadership, communicating priorities, progress, risks, and technical challenges clearly.
- Own SRE delivery goals, deciding what the team commits to, how work is prioritized, and how operational responsibilities are managed.
- Design and maintain effective support rotations and on-call processes while strengthening incident response practices.
- Provide technical leadership across Kubernetes, AWS, PostgreSQL, DNS and TLS, CI infrastructure, and the broader infrastructure platform.
- Guide the development of reliability practices including SLOs, error budgets, observability, incident response, and post-incident improvements.
- Partner closely with Security on infrastructure threats, patching, controls, audits, and compliance obligations.
- Manage relationships with infrastructure and platform vendors, including renewals and commercial discussions with support from senior leadership.
- Remain hands-on enough to review technical work, challenge architectural decisions, participate credibly in incidents, and identify emerging reliability issues before they escalate.
- Build strong relationships across engineering and encourage teams to bring operational and reliability challenges forward early.
- Continuously improve team health, collaboration, conflict resolution, and retrospective practices.
- Proven experience leading an SRE, infrastructure, platform engineering, DevOps, or similarly focused technical team, with direct responsibility for performance and career development.
- Strong hands-on background in site reliability, DevOps, or cloud infrastructure engineering, with sufficient technical depth to review designs, challenge implementation decisions, and contribute during production incidents.
- Production experience with Kubernetes and AWS at meaningful scale, including the operational realities of running cloud infrastructure.
- Hands-on experience building, enabling, or scaling AI infrastructure and working with AI-related engineering workloads.
- Strong understanding of observability principles and practices, infrastructure as code with Terraform, and CI/CD platforms such as GitLab CI, GitHub Actions, or Jenkins.
- Experience with Docker, shell scripting, and production infrastructure operations.
- Proven ownership of reliability practices including incident response, on-call operations, SLOs, error budgets, and turning incidents into lasting engineering improvements.
- Experience working in regulated environments, with an understanding of infrastructure controls, compliance, and security requirements.
- Exceptional prioritization skills, particularly when operational workloads compete with project commitments.
- Excellent written communication and documentation skills, with the ability to lead effectively in a highly distributed and asynchronous environment.
- Strong relationship-building, collaboration, conflict-resolution, and stakeholder-management capabilities.
- A coaching-oriented leadership style, with evidence of developing engineers both technically and professionally.
- Strong judgment, accountability, adaptability, curiosity, and commitment to high-quality execution.
- Nice-to-have experience with Elixir, Java, Clojure, Node.js, Python, or another backend programming language.
- Additional desirable experience includes OpenTelemetry, distributed tracing, Honeycomb, PostgreSQL or Aurora performance optimization, connection pool management, query tuning, Linux systems administration, security, FinOps, and cloud cost management.
- Experience growing an engineering team from a small base and establishing a strong hiring bar is advantageous.
- Ability to work effectively across global teams and time zones.
- Annual salary range of USD $75,450–$169,700, with actual compensation determined by location, experience, skills, training, business needs, and market conditions.
- Fully remote, work-from-anywhere environment.
- Flexible working hours within an asynchronous work culture.
- Flexible paid time off.
- 16 weeks of paid parental leave.
- Budget for coworking spaces, learning, and wellness activities, including gym memberships.
- Mental health support services.
- Stock options.
- Home office budget and IT equipment.
- Global exposure through collaboration with colleagues across multiple continents.
- Opportunities to travel internationally and meet colleagues at company events.
- A high-autonomy environment where employees are encouraged to organize their schedules around their lives and personal commitments.
- Opportunity to influence the maturity of reliability engineering practices while working on complex infrastructure and platform challenges.