JobTarget Logo

Director, Site Reliability Engineering in Randburg, Northern Cape at Jobgether

NewJob Function: Information Technology
Jobgether
Randburg, Northern Cape, 2161, South Africa
Posted on
New job! Apply early to increase your chances of getting hired.

Explore Related Opportunities

Job Description

Director, Site Reliability Engineering

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Director, Site Reliability Engineering based in Africa.

As Director, Site Reliability Engineering, you will lead the teams and technical strategy responsible for keeping critical infrastructure reliable, scalable, and secure for millions of users. You will work across software, systems, automation, cloud infrastructure, and operational processes to solve complex reliability challenges at scale. The role combines strategic leadership with hands-on technical depth, including troubleshooting production systems and partnering closely with software engineering teams. You will help shape the future architecture and deployment practices of a large-scale, privacy-focused technology environment. You will also drive improvements in automation, observability, incident response, and engineering efficiency. This is a remote-first leadership opportunity with significant ownership, autonomy, and impact.

Accountabilities:
  • Lead and develop Site Reliability Engineering teams responsible for the reliability, scalability, performance, and operational health of large-scale systems.

  • Define and execute the technical direction for infrastructure, deployment, reliability engineering, automation, and operational practices.

  • Lead high-impact and complex initiatives from initial proposal and planning through implementation, measurement, and postmortem.

  • Investigate and resolve sources of instability across high-traffic, distributed systems, identifying root causes and implementing sustainable remediation.

  • Establish and improve tools, services, monitoring, alerts, incident-response processes, and operational practices that identify and mitigate reliability risks.

  • Partner closely with software engineers to troubleshoot production issues, evaluate performance considerations, and implement appropriate code-level or infrastructure-level solutions.

  • Drive automation for infrastructure provisioning and configuration management to improve efficiency, scalability, consistency, and reliability.

  • Leverage cloud-native architectures and services to strengthen system resilience and support continued growth.

  • Help ensure products and infrastructure meet established reliability standards while minimizing user impact during failures and incidents.

  • Identify emerging technical needs and opportunities to guide the long-term evolution of deployment and infrastructure architecture.

  • Support a culture of ownership, continuous improvement, measurable outcomes, and effective post-incident learning.

Requirements:
  • 10+ years of relevant professional experience in Site Reliability Engineering, platform engineering, infrastructure engineering, software engineering, or related fields.

  • 4+ years of experience leading SRE or comparable engineering teams.

  • Experience participating in or managing 24/7 on-call operations for large-scale production environments.

  • Advanced programming experience and the ability to read, write, troubleshoot, and deploy software across production systems.

  • Strong experience with Linux administration and troubleshooting, web technologies, distributed systems, and high-traffic production environments.

  • Demonstrated ability to lead complex technical projects from ambiguous initial requirements through execution and postmortem.

  • Experience developing effective reliability tooling, services, monitoring, alerting, and incident-response capabilities.

  • Strong investigative and root-cause analysis skills, particularly within distributed and high-scale systems.

  • Experience designing and implementing infrastructure automation, provisioning, and configuration-management solutions.

  • Hands-on experience with cloud-native services and architectures, including application packaging and deployment using Docker and Docker Compose.

  • Experience with high-level programming languages such as Go, Perl, TypeScript, Python, or comparable technologies.

  • Experience with AI-driven software development, including the design and implementation of agentic workflows.

  • Strong ability to turn ambiguous or complex problems into practical, innovative solutions with measurable outcomes.

  • Strategic thinking and technical foresight, with the ability to anticipate future infrastructure and reliability requirements.

  • Excellent communication and collaboration skills, with the ability to work effectively across engineering teams and technical disciplines.

  • Strong sense of ownership, autonomy, and accountability in a remote-first working environment.

Benefits:
  • Annual compensation of $243,800 USD, plus stock options.

  • Transparent compensation structure, with team members at the same professional level and within the same global region receiving the same compensation regardless of functional team, location, gender, educational background, or years of experience.

  • Fully remote, flexible working arrangement with no core working hours.

  • Average full-time commitment of approximately 40 hours per week.

  • Company-sponsored health benefits for eligible team members based in the United States; these benefits do not extend to team members based in Canada or other countries.

  • Paid parental leave.

  • Support for home-office setup.

  • Co-working allowances.

  • Opportunities to participate in company-wide and team gatherings, with travel expected at least twice per year for an all-hands meeting and a team retreat.

  • Remote-first environment centered on trust, inclusivity, ownership, and empowered project management.

  • Equal employment opportunities and a commitment to an accessible, inclusive hiring process.

  • Reasonable accommodations are available for candidates who require support during the application process.

  • Successful candidates must complete a background check as a condition of employment.

  • The role requires participation in video meetings with cameras enabled.

How Jobgether works:
We use an AI-powered matching process to ensure your application is reviewed quickly, objectively, and fairly against the role's core requirements. Our system identifies the top-fitting candidates, and this shortlist is then shared directly with the hiring company. The final decision and next steps (interviews, assessments) are managed by their internal team.
We appreciate your interest and wish you the best!
Data Privacy Notice: By submitting your application, you acknowledge that Jobgether will process your personal data to evaluate your candidacy and share relevant information with the hiring employer. This processing is based on legitimate interest and pre-contractual measures under applicable data protection laws (including GDPR). You may exercise your rights (access, rectification, erasure, objection) at any time.
#LI-CL1

Job Location

Randburg, Northern Cape, 2161, South Africa

Frequently asked questions about this position

Similar Jobs In Randburg, Northern Cape

New

Head of UX/UI

Jobgether
Randburg, Northern Cape
New

Senior Middleware Administrator

Jobgether
Randburg, Northern Cape
New

Technical Support Team Lead (SaaS | SQL)

Jobgether
Randburg, Northern Cape
New

Project Manager (Art Team)

Jobgether
Randburg, Northern Cape
Continue to apply
Enter your email to continue. You’ll be redirected to the employer’s application.
By clicking Continue, you understand and agree to JobTarget's Terms of Use and Privacy Policy.