Site Reliability Engineer

Posted 4 Days Ago
Be an Early Applicant
Cape Town, Western Cape, ZAF
In-Office
Mid level
Cloud • Information Technology • Internet of Things • Security
The Role
The Site Reliability Engineer will ensure the reliability, scalability, performance, and availability of platforms, network services, infrastructure, and customer-facing systems. Responsibilities include infrastructure automation, CI/CD improvement, cloud and on-premises environment management, container orchestration, observability, incident response, security, compliance, and on-call support. The role collaborates with network, software development, DevOps, and NOC teams to reduce operational toil and improve system resilience.
Summary Generated by Built In
Established in 2001, RSAWEB is South Africa’s fastest growing internet service provider (ISP) with a focus on providing connectivity to home customers, and a wide array of technology solutions to businesses. We are obsessed about ensuring all our customers receive the best possible digital experience and exceptional customer service. Thousands of customers have given RSAWEB a 5-star rating, with an average rating of 4.7 out of 5 on Google – the best-rated ISP in South Africa. We are extremely proud of winning KFM’s Best of the Cape Awards: Best ISP in 2021 and 2022 being one of the fastest streaming ISPs on Netflix and a consistently top-rated ISP on MyBroadband. These accolades are not for nothing, as we constantly strive to improve our products, services, and solutions to enhance each customer’s experience. Having invested heavily in infrastructure, RSAWEB has built a strong presence in South Africa with Data Centres in Johannesburg and Cape Town.

Our Products and Services:
•Fibre-to-the-Home (FTTH)
•Fibre-to-the-Business (FTTB)
•Enterprise connectivity
•Mobile connectivity and data management
•Cloud infrastructure and more!

At RSAWEB, we are passionate about using our creativity, to provide innovative solutions and services, that allow our customers to succeed in all areas of life. We believe that we are in the business of connecting customers and businesses with each other and a world of infinite possibility and opportunity, through technology. Our mission transcends our values through every customer, every interaction, every connection, every day.

Our values:
•We Build Trust and Ownership
•We Honour & Respect People
•We Cultivate Passion & Creativity
•We Innovate Feverishly
•We Go the Extra Mile
•We Believe in Humility
•We Communicate Openly & Honestly
•We Make it Fun
•We Teach, Grow & Learn
•We Do More, With Less

Role Purpose:

The Site Reliability Engineer (SRE) is responsible for ensuring the reliability, performance, scalability, and availability of RSAWEB's platforms, network services, and customer-facing systems. This role blends software engineering, infrastructure automation, and operations to deliver highly reliable services and improve the efficiency of technical teams.

Key Responsibilities
1. Reliability & System Performance
  • Maintain high availability and performance across platforms, services, and infrastructure.

  • Define, measure, and improve SLIs/SLOs/SLAs for critical systems.

  • Troubleshoot system and network reliability issues proactively.

2. Automation & DevOps Enablement
  • Build automation for deployments, monitoring, configuration, and operational tasks.

  • Improve CI/CD pipelines and assist engineers with release engineering.

  • Reduce manual work (toil) by implementing self-service tools and automation workflows.

3. Infrastructure Engineering
  • Deploy, manage, and optimise cloud and on-prem infrastructure (Linux servers, virtualisation, containers).

  • Work with network teams to ensure resilient integration between systems and ISP network elements.

  • Manage and scale containerised platforms (Docker, Kubernetes).

4. Observability & Monitoring
  • Implement and maintain monitoring, alerting, and logging solutions (e.g., Prometheus, Grafana, ELK, Datadog).

  • Ensure actionable, low-noise alerting and system dashboards.

  • Use metrics to identify performance bottlenecks and reliability risks.

5. Incident Management
  • Participate in incident response, including root cause analysis and corrective actions.

  • Improve monitoring and automation to prevent repeated issues.

  • Assist with on-call rotations to support critical services.

6. Security & Compliance
  • Implement security best practices across systems and deployments.

  • Support vulnerability scanning, patching, and secure configurations.

  • Ensure compliance with internal and industry standards (ISO, POPIA, etc).

7. Collaboration & Support
  • Work closely with Network Engineering, DevOps, Software Development, and NOC teams.

  • Provide technical guidance in system design, scalability, and reliability improvements.

  • Improve operational processes through documentation and automation.



Requirements
Minimum Qualifications
  • Diploma or degree in Computer Science, Engineering, Information Technology, or related field.

  • Relevant certifications (AWS/Azure/GCP, Linux, Kubernetes, Terraform) are beneficial.

Experience Requirements
  • 3–5+ years in SRE, DevOps, Systems Engineering, or Infrastructure roles.

  • Experience supporting large-scale, mission-critical environments (preferably ISP or telecom).

  • Strong background in Linux (CentOS, Ubuntu, Debian) administration.

  • Experience with container orchestration and Infrastructure as Code.

Technical Skills

  • Strong scripting skills (Python, Bash, Go preferred).

  • CI/CD tools: GitHub Actions, GitLab CI, Jenkins, ArgoCD, etc.

  • IaC: Terraform, Ansible, Pulumi, CloudFormation.

  • Cloud platforms: AWS / Azure / GCP (or private cloud / OpenStack).

  • Monitoring: Prometheus, Grafana, Zabbix, ELK, Datadog.

  • Networking fundamentals: DNS, DHCP, firewalls, load balancing, routing.

  • Databases: SQL and NoSQL basics.

  • Knowledge of ISP infrastructure such as BNGs, RADIUS, DNS clusters (advantage).


Benefits
•Medical Aid (Discovery)
•Reduced Gap Cover Rates (Turnberry Premier)
•Retirement Annuity Contribution (Allan Gray)
•Medical Insurance (Momentum - Health4Me)
•Discounted Internet Connectivity
•Free Employee Wellness Programme (Lyra Wellbeing, formerly ICAS)
•Exposure to latest industry technologies and standards
•Lastly, a work environment that rivals the very best!

If you have not heard from us within 2 weeks of submitting your application, please consider your application as unsuccessful.

Skills Required

  • Diploma or degree in Computer Science, Engineering, Information Technology, or a related field
  • 3-5+ years of experience in SRE, DevOps, Systems Engineering, or Infrastructure roles
  • Experience supporting large-scale, mission-critical environments
  • Strong Linux administration experience, including CentOS, Ubuntu, or Debian
  • Experience with container orchestration and Infrastructure as Code
  • Strong scripting skills in Python, Bash, or Go
  • Experience with CI/CD tools such as GitHub Actions, GitLab CI, Jenkins, or ArgoCD
  • Experience with Infrastructure as Code tools such as Terraform, Ansible, Pulumi, or CloudFormation
  • Experience with AWS, Azure, GCP, private cloud, or OpenStack
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Zabbix, ELK, or Datadog
  • Knowledge of networking fundamentals, including DNS, DHCP, firewalls, load balancing, and routing
  • Basic SQL and NoSQL database knowledge
  • Relevant AWS, Azure, GCP, Linux, Kubernetes, or Terraform certifications
  • Experience with ISP infrastructure such as BNGs, RADIUS, and DNS clusters
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
21 Employees
Year Founded: 2001

What We Do

RSAWEB is a South African internet service provider and technology company serving residential and business customers. It delivers fibre and other connectivity services alongside cloud infrastructure, mobile data management, security, IoT, and related digital solutions. The company positions itself as a technology partner, using its network and data-centre infrastructure to help customers increase revenue, manage risk, and control costs.

Similar Jobs

RAMP Group Pty Ltd Logo RAMP Group Pty Ltd

Site Reliability Engineer

Cloud • Information Technology • Internet of Things • Security
In-Office
Cape Town, Western Cape, ZAF
21 Employees

Invisible Technologies Logo Invisible Technologies

Site Reliability Engineer

Artificial Intelligence • Information Technology • Machine Learning • Professional Services • Software • Analytics • Consulting
Remote or Hybrid
3 Locations
370 Employees

RELX Logo RELX

Site Reliability Engineer

Information Technology • Legal Tech • Analytics
In-Office or Remote
7 Locations
10001 Employees

impact.com Logo impact.com

Site Reliability Engineer

Marketing Tech • Software
In-Office
Cape Town, Western Cape, ZAF
1247 Employees

Similar Companies Hiring

Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account