Site Reliability Engineer in Randburg, Northern Cape at Jobgether
Explore Related Opportunities
Job Description
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Site Reliability Engineer based in South Africa.
Join a technology-driven team responsible for maintaining the reliability, scalability, and security of mission-critical infrastructure supporting high-availability platforms. In this role, you will combine infrastructure engineering, automation, and operational excellence to ensure seamless system performance across cloud and on-premises environments. Working alongside cross-functional engineering teams, you will proactively optimize production systems, strengthen monitoring capabilities, and respond to complex operational challenges. This is an excellent opportunity for an experienced infrastructure professional who enjoys solving technical problems, improving reliability through automation, and contributing to resilient, enterprise-grade services in a fast-paced environment.
- Design, implement, maintain, and optimize highly available infrastructure supporting mission-critical applications and services.
- Monitor production environments, analyze system performance, and proactively identify opportunities to improve stability, scalability, and operational efficiency.
- Respond to technical escalations, troubleshoot infrastructure, networking, hardware, and software issues, and lead resolution of critical incidents.
- Develop and maintain monitoring, alerting, backup, recovery, and disaster recovery procedures to maximize uptime and minimize business impact.
- Manage cloud platforms, virtualization technologies, network infrastructure, and remote monitoring systems to ensure secure and reliable operations.
- Build and maintain infrastructure automation using configuration management, scripting, and Infrastructure-as-Code tools.
- Participate in post-incident reviews, document operational improvements, and contribute to continuous reliability and security enhancements.
- Collaborate with engineering and operations teams to strengthen CI/CD pipelines, system resilience, and infrastructure best practices.
- Bachelor's degree in Computer Science, Information Technology, or a related field; relevant professional certifications are advantageous.
- At least 3 years of experience as a Site Reliability Engineer, DevOps Engineer, Systems Administrator, or in a similar infrastructure role.
- Strong experience with cloud platforms such as AWS or Oracle Cloud.
- Proficiency with Linux and Windows server administration, virtualization technologies, and enterprise infrastructure management.
- Experience with Docker, Kubernetes, Terraform, Git, GitLab CI/CD, ELK Stack, Prometheus, and Grafana.
- Knowledge of MySQL, PostgreSQL, networking concepts (LAN/WAN, HTTP, TCP/IP), system security, and backup/recovery strategies.
- Experience with automation and scripting tools such as Ansible, Bash, Rundeck, or Puppet.
- Familiarity with Nginx, PHP-FPM, SSL, DNS, and Cloudflare is considered an advantage.
- Strong analytical thinking, troubleshooting skills, attention to detail, and the ability to work independently and collaboratively.
- Availability to respond to critical production incidents outside standard business hours when required.
- Competitive salary package.
- Private health insurance.
- Annual wellness allowance.
- Birthday leave.
- Company-sponsored team-building events and social activities.
- Relocation support, where applicable.
- Opportunity to work with modern cloud technologies, automation tools, and large-scale infrastructure.
- Professional development opportunities within a collaborative engineering environment.