Staff Site Reliability Engineer in Winit Germany GmbH, Bremen at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer based in Germany.
This is a fully remote opportunity for a highly experienced SRE leader to shape reliability across complex, AI-driven production environments.
You will act as a senior technical authority, designing resilient infrastructure and establishing SRE practices that scale with rapid business growth.
The role spans cloud infrastructure, Kubernetes, observability, CI/CD, data platforms, and machine learning systems.
You’ll work across platform, product, data, and ML engineering teams to improve availability, performance, security, and operational efficiency.
A key focus will be productionizing AI workloads, standardizing customer environments, and strengthening infrastructure for enterprise-scale deployments.
You’ll tackle complex reliability challenges hands-on while influencing architecture and technology direction across engineering.
The environment values ownership, technical excellence, automation, continuous improvement, and the ability to influence without relying on formal authority.
- Architect, deploy, operate, and continuously improve scalable, secure production environments, with a strong preference for AWS-based infrastructure.
- Lead reliability initiatives across multiple engineering streams and establish consistent SRE practices throughout the organization.
- Design, evolve, migrate, and optimize Kubernetes-based infrastructure, including production hardening and scaling.
- Establish and enforce robust Infrastructure-as-Code standards using Terraform or equivalent technologies.
- Define, implement, and operationalize SLIs, SLOs, error budgets, and other reliability practices.
- Strengthen observability across applications, infrastructure, data pipelines, and ML systems to improve visibility into system health and performance.
- Collaborate with product and data teams to incorporate product telemetry, model analytics, and operational data into reliability insights.
- Design and optimize CI/CD pipelines across the complete software lifecycle, from build and testing through deployment and rollback.
- Improve release safety, deployment frequency, operational predictability, and adherence to service-level objectives.
- Lead incident response for complex, cross-system failures and drive thorough post-incident reviews and corrective actions.
- Reduce operational toil through automation, platform engineering, and improved tooling and processes.
- Design scalable processes and infrastructure for absorbing, standardizing, monitoring, and troubleshooting customer environments.
- Support and productionize ML workloads by implementing MLOps practices for model deployment, monitoring, and retraining workflows.
- Ensure infrastructure and operational practices meet enterprise-grade security, compliance, and regulatory requirements.
- Mentor engineers, share best practices, and raise the overall reliability and engineering standards across teams.
- Collaborate with Staff Engineers and Architects to influence global product architecture and long-term technology strategy.
- Extensive hands-on experience in Site Reliability Engineering, Production Engineering, or a closely related infrastructure role.
- Proven experience establishing or scaling SRE practices within high-growth, complex, or highly distributed technical environments.
- Deep expertise with AWS or Azure cloud infrastructure and modern cloud-native architectures.
- Strong production experience with Kubernetes, including migration, scaling, optimization, and security hardening.
- Advanced Infrastructure-as-Code expertise using Terraform or an equivalent technology.
- Demonstrated experience designing, implementing, and optimizing end-to-end CI/CD pipelines.
- Strong knowledge of observability practices and tooling across distributed applications and infrastructure.
- Experience troubleshooting complex multi-tenant, customer-hosted, or enterprise environments.
- Experience supporting production data platforms and machine learning systems.
- Practical MLOps experience, including model deployment, monitoring, and operational lifecycle management.
- Strong understanding of distributed systems, scalability, resilience, fault tolerance, and failure modes.
- Ability to think across systems and understand the interactions between infrastructure, applications, data, and ML workloads.
- Strong communication and collaboration skills, with the ability to work effectively across engineering and business functions.
- Experience with large-scale global B2B or B2C products is desirable.
- Experience working with AI/ML, NLP, or LLM-based products is a strong advantage.
- Familiarity with integrating product analytics and model performance metrics into operational monitoring is beneficial.
- Experience operating in enterprise environments with stringent security, compliance, and regulatory requirements is preferred.
- Experience implementing regulatory controls within cloud infrastructure is a plus.
- Experience scaling infrastructure during periods of rapid growth is advantageous.
- Experience evaluating infrastructure tools, platforms, and vendors is desirable.
- Experience deploying and operating solutions within large enterprise customer accounts or VPCs is a strong plus.
- Strong problem-solving skills, high ownership, and accountability.
- Ability to anticipate failure modes, operate across multiple engineering streams, and influence technical decisions without formal authority.
- Continuous-learning mindset with a strong commitment to improving systems, processes, and engineering practices.
- Full-time, permanent employment.
- Fully remote position within European time zones.
- Opportunity to work on infrastructure supporting advanced AI and agent-based workloads.
- Significant technical ownership and influence over reliability practices and architecture.
- Opportunity to work across cloud infrastructure, Kubernetes, distributed systems, data platforms, and ML operations.
- Exposure to complex enterprise environments and large-scale production deployments.
- Collaboration with highly experienced engineers, architects, product teams, and AI/ML specialists.
- Opportunity to shape reliability standards and technology strategy as the organization scales.
- Strong culture of ownership, continuous improvement, and technical excellence.
- International and distributed working environment.
- Career growth and opportunities to expand technical leadership and organizational impact.