Site Reliability Engineer

Posted 8 Days Ago
Be an Early Applicant
Sydney, New South Wales, AUS
In-Office
Mid level
Fintech • Software • Financial Services
The Role
Support reliable, scalable production environments through incident management, observability, cloud operations, system administration, container support, vulnerability remediation, and automation. The role uses SRE practices such as SLIs, SLOs, error budgets, incident reviews, and on-call support across AWS, GCP, Linux, Windows, Docker, and Kubernetes environments.
Summary Generated by Built In

About HUB24


At HUB24, we’re rethinking the way wealth management works, combining platform, technology and data to create better outcomes for financial professionals and their clients.


Our purpose is simple:  Empower better financial futures, together.


What sets us apart is how we work. We back bold thinking, move with pace, and turn ideas into action. You’ll have the opportunity to make a real impact across your team, the business, and for the clients we support every day.


HUB24 Limited is an ASX-listed company (ASX: HUB) and part of the ASX100. We have over 1,100 employees across Australia, with offices in Sydney, Melbourne, Brisbane, Perth and the Gold Coast. 


Why you’ll enjoy working here


We create an environment where you can do your best work and see the impact of it.

  • Work with smart, collaborative people who get things done.
  • Your ideas won’t sit in a backlog, they’ll be heard, tested and actioned.
  • Grow your career your way, with support to learn, stretch and explore new opportunities.

We also offer benefits to support you inside and outside of work:

  • Genuinely flexible and hybrid ways of working.
  • Employee Share Scheme.
  • Additional leave and wellbeing support.
  • Enhanced parental leave and support through different life stages.
  • Everyday benefits, including discounts and financial wellbeing support.

Why this is an exciting opportunity


HUB24 is expanding its Site Reliability Engineering function and investing in Dynatrace as our core monitoring and observability platform. This is an opportunity to be part of that team which will play a critical role in ensuring the reliability, performance and scalability across our technology platforms.
Working closely with engineering, infrastructure and operations teams, you will embed SRE best practices, proactively manage system health and help design resilient, high-availability services that support our customers and business growth.

What you’ll be doing


  • Demonstrated hands-on experience in Site Reliability Engineering (SRE), DevOps, or Infrastructure Engineering, supporting reliable, high-performing, and scalable production environments.
  • Provide technical support for complex production incidents, ensuring timely resolution and minimising customer impact.
  • Lead incident triage, troubleshooting, root cause analysis, and Post Incident Reviews (PIRs), driving permanent fixes and continuous service improvements.
  • Monitor system performance, availability, and security across HUB platforms, proactively identifying and escalating risks.
  • Apply SRE best practices, including SLIs, SLOs, error budgets, and reliability engineering principles, to improve service resilience and operational excellence.
  • Support and troubleshoot AWS and GCP cloud environments, ensuring stability, performance, and operational efficiency.
  • Administer and maintain Linux and Windows servers, including performance tuning, configuration management, and operational support.
  • Support containerised environments using Docker and Kubernetes, helping maintain scalable and resilient workloads.
  • Manage server patching and vulnerability remediation activities in line with compliance, security, and risk management requirements.
  • Contribute to automation initiatives, developing operational scripts and runbooks to reduce manual effort and improve consistency.
  • Participate in an on-call roster, providing support for critical systems, platforms, and services as required.

What you bring


You don’t need to tick every box, but experience in the below will set you up for success.


Experience & Skills


  • 3–5 years of hands-on experience in Site Reliability Engineering, Observability Engineering, DevOps, or Infrastructure Engineering, including support for production systems and incident response.
  • Practical experience in incident management, including triage, troubleshooting, root cause analysis, and participation in on-call support for high-availability environments.
  • Working knowledge of cloud platforms such as AWS and/or GCP, including infrastructure troubleshooting, operational support, and optimisation activities.
  • Experience using observability tools, preferably Dynatrace, for monitoring, alerting, dashboarding, and performance analysis.
  • Solid system administration skills across Linux and Windows environments, including troubleshooting and basic performance tuning.
  • Exposure to containerisation and orchestration technologies such as Docker and Kubernetes is preferred.
  • Experience supporting server patching and vulnerability remediation, including prioritisation based on risk and compliance requirements.
  • Experience with automation and Infrastructure as Code tools such as Terraform, Ansible, or similar is desirable.
  • Strong communication skills, with the ability to document findings, explain technical issues clearly, and collaborate effectively across teams.

Our process


We aim to keep the process simple and respectful of your time:

  • You’ll receive an acknowledgement after applying.
  • Our Talent team will review your application and keep you updated.
  • If shortlisted, we’ll connect to learn more about you.
  • Interviews may be virtual or in person.
  • You’ll receive an outcome and feedback.

If you need any adjustments, please let us know - we’re here to support you.



Our commitment

We’re committed to building an inclusive environment where everyone feels valued and supported to do their best work. We welcome applications from people of all backgrounds, identities and experiences.


Agencies, we work with a panel of preferred suppliers and are not accepting any unsolicited CVs.

Skills Required

  • 3-5 years of hands-on experience in Site Reliability Engineering, Observability Engineering, DevOps, or Infrastructure Engineering
  • Experience supporting production systems and incident response
  • Practical experience in incident management, including triage, troubleshooting, root cause analysis, and on-call support
  • Working knowledge of AWS and/or GCP, including infrastructure troubleshooting, operational support, and optimization
  • Experience using observability tools, preferably Dynatrace, for monitoring, alerting, dashboarding, and performance analysis
  • System administration skills across Linux and Windows environments
  • Experience supporting server patching and vulnerability remediation
  • Strong communication, documentation, and cross-team collaboration skills
  • Exposure to Docker and Kubernetes
  • Experience with automation and Infrastructure as Code tools such as Terraform or Ansible
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Sydney, NSW
859 Employees

What We Do

HUB24 Group (ASX:HUB) leads the wealth industry as the best provider of integrated platform, technology and data solutions, and we’re not done yet. At HUB24, we believe in the value of advice and by collaborating with the industry and leveraging our technology and data expertise, we’re helping to solve key challenges to enable the delivery of accessible financial advice, and empower better financial futures for more Australians. Our solutions include Australia’s best platform HUB24, leading SMSF software Class, and myprosperity’s innovative client portal technology. Recognised by advisers and the industry as Australia’s best platform, HUB24 leverages data and technology to deliver choice, flexibility, efficiencies and value for advisers and their clients. We’re committed to continued investment in innovation that delivers, making the complex simple and enabling advisers to empower better financial futures for more Australians. As a leading provider of SMSF software, Class, together with NowInfinity, delivers trust accounting, portfolio management, legal documentation and corporate compliance solutions that transform the way finance professionals do business. We champion automation, simplicity and connectively to drive business profitability and enhanced client experience. myprosperity is a leading provider of client portals for accountants and financial advisers, enabling streamlined service delivery and increased productivity for their business and enhanced customer experience for their clients. We make it easier for Australians to share, communicate and collaborate with their financial professional across all aspects of their financial lives.

Similar Jobs

Axon Logo Axon

Site Reliability Engineer

Artificial Intelligence • Cloud • Social Impact • Software • Wearables
In-Office or Remote
3 Locations
2700 Employees

Leidos Logo Leidos

Site Reliability Engineer

Information Technology • Software
In-Office or Remote
4 Locations
27104 Employees

Social Discovery Group Logo Social Discovery Group

Site Reliability Engineer

Artificial Intelligence • Fintech • Machine Learning • Software • App development • Conversational AI • Generative AI
In-Office or Remote
7 Locations
In-Office
3 Locations
3366 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account