Devops/SRE Tech Lead I Observabilidade in Brazil at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a DevOps/SRE Tech Lead I Observabilidade based in Brazil .
This is a technical leadership role focused on building reliable, scalable, and highly observable cloud infrastructure.
You will own the architecture and evolution of the observability platform, enabling engineering teams to detect, understand, and resolve issues faster.
The role combines hands-on infrastructure engineering with team mentorship, delivery management, and cross-functional technical leadership.
You will work extensively with AWS, Kubernetes, infrastructure as code, CI/CD, telemetry pipelines, and modern monitoring technologies.
A major focus will be creating an Observability-as-a-Service model that delivers actionable insights while controlling telemetry noise and costs.
You will also help shape developer experience, incident response, security, and platform standards across the engineering organization.
This is an opportunity to make a broad technical impact while guiding a small infrastructure team in a fast-moving and collaborative environment.
- Design, implement, and maintain highly available observability infrastructure, including OpenTelemetry Collectors, metrics/logs/traces agents, telemetry aggregators, and visualization platforms.
- Build and govern the observability ecosystem for applications and infrastructure, defining actionable alerting standards with tools such as Alertmanager or Incident.io while reducing alert fatigue and operational noise.
- Develop and evolve AWS cloud infrastructure using Infrastructure as Code, ensuring resilient, scalable, secure, and cost-efficient architectures.
- Administer and continuously improve production Kubernetes clusters, maintaining high standards for scalability, security, reliability, and cloud cost management.
- Build, optimize, and maintain fast and reliable CI/CD pipelines while automating repetitive operational work and reducing engineering toil.
- Own telemetry architecture end to end, ensuring logs, metrics, and traces are efficiently generated, collected, processed, retained, and stored.
- Establish and improve Observability-as-a-Service capabilities that enable development teams to adopt effective monitoring and incident detection practices.
- Lead and mentor the infrastructure/platform engineering team by organizing the backlog, distributing work, supporting technical decisions, and helping engineers develop their capabilities.
- Facilitate collaboration with development and product teams, translating infrastructure risks, technical debt, and platform priorities into business-oriented discussions and negotiating appropriate backlog capacity.
- Support incident response processes and provide calm, decisive technical leadership during production outages and other high-pressure situations.
- Improve the overall developer experience by identifying opportunities across the software development lifecycle and introducing automation, standards, and tooling that help engineering teams deliver more effectively.
- Contribute to security and platform maturity through practices such as secrets management, policy as code, and secure software supply-chain controls.
- Deep professional experience with public cloud platforms, preferably AWS, and the ability to design resilient and scalable cloud architectures.
- Advanced Kubernetes expertise and strong knowledge of container ecosystems such as Docker and containerd.
- Mature experience operating and supporting distributed, highly available production systems.
- Proven hands-on experience deploying and managing observability infrastructure through code and IaC/Helm.
- Strong practical knowledge of observability technologies such as Prometheus Operator, OpenTelemetry Collector, Grafana Stack including Loki, Tempo, and Mimir, or comparable commercial platforms such as Datadog, New Relic, or Dynatrace.
- Deep understanding of telemetry engineering, including how logs, metrics, and traces are generated, collected, processed, routed, stored, and retained.
- Strong understanding of telemetry retention, ingestion, sampling, aggregation, and cost optimization.
- Proficiency with Infrastructure as Code tools such as Terraform, Pulumi, or equivalent technologies.
- Strong programming and scripting capabilities for automation using Python, Go, or advanced Bash.
- Proven experience designing and maintaining CI/CD pipelines with technologies such as GitLab CI, GitHub Actions, Jenkins, or similar.
- Experience with incident response processes and the ability to make effective technical decisions under production pressure.
- Strong understanding of the software development lifecycle and a clear focus on improving developer experience and engineering productivity.
- Experience facilitating agile ceremonies and coordinating delivery for a small technical team, or a strong aptitude for doing so.
- Strong communication and negotiation skills, including the ability to translate technical debt, infrastructure risks, and reliability concerns into clear business language.
- Servant-leadership mindset with a genuine interest in mentoring engineers, removing blockers, and developing technical talent.
- Pragmatic approach to engineering, balancing ideal architecture with business priorities, delivery timelines, and real-world constraints.
- Experience with GitOps and service mesh technologies such as Argo CD, Flux, Istio, or Linkerd is a plus.
- Practical FinOps experience and a track record of optimizing cloud infrastructure costs is desirable.
- Experience with advanced observability and eBPF technologies such as Cilium or Pixie is advantageous.
- Experience optimizing observability costs through trace sampling, log aggregation, filtering, and collector-level data management is a strong plus.
- Knowledge of DevSecOps practices, including secrets management, policy as code such as OPA/Gatekeeper, and software supply-chain security is desirable.
- Relevant certifications such as CKA, AWS Solutions Architect, or equivalent credentials are beneficial.
- Advanced technical English, including the ability to read, communicate, and participate in discussions with global teams or external providers.
- 100% remote work with the flexibility to work from wherever you are.
- Access to offices in Rio de Janeiro and São Paulo for employees who prefer an in-person workspace.
- Opportunities for professional growth, learning, and career development.
- Flexible working hours, including while working remotely.
- Flexible vacation arrangements in addition to the applicable employment framework.
- Bradesco health and dental insurance.
- Group life insurance.
- Meal allowance through Caju.
- Monthly home-office support for furniture and internet expenses.
- Dedicated home-office allowance.
- OrienteMe psychological support.
- Conexa Saúde psychological and nutritional support.
- Wellhub membership.
- Birthday day off.
- Six months of maternity leave.
- One month of paternity leave.
- Childcare assistance.
- SESC partnership and benefits.
- Team game nights, both in-person and online.
- Learning and development programs supporting training such as courses, talks, books, and other professional development resources.
- Collaborative and supportive team environment.