Senior Site Reliability Engineer (Platform Engineering/DevOps)

Posted 6 Days Ago
Be an Early Applicant
Sydney, New South Wales, AUS
Hybrid
Senior level
eCommerce • Fintech • Information Technology • Insurance • Software
Cover Genius protects millions of customers of the world’s largest online companies. Our goal is to protect all of them.
The Role
Lead infrastructure and reliability strategy across the organization. Architect multi-region cloud platforms, define observability and SLOs, lead incident response, build automation and self-service platforms, mentor engineers, apply AI-assisted tools, and drive cloud cost optimisation and disaster recovery practices.
Summary Generated by Built In

About the Company

Cover Genius is a Series E Insurtech that protects the global customers of the world’s largest digital companies including Booking Holdings, owner of Priceline, Kayak and Booking.com, Intuit, Hopper, Skyscanner, Ryanair, Turkish Airlines, Descartes ShipRush, Zip and SeatGeek. We’re also available at Amazon, Flipkart, eBay, Wayfair and SE Asia’s largest company, Shopee.

Our partners integrate with XCover, our award-winning insurance distribution platform, to embed protection for millions of customers worldwide each year. Our team and products have been recognized with dozens of awards including by the Financial Times who ranked Cover Genius as the #1 fastest growing company in APAC in 2020. Our diverse team across 20+ countries and many language groups commits itself to diverse cultural programs, in particular “CG Gives” which makes social entrepreneurs out of us all and funds development initiatives in global communities.

Our People are Bold, Authentic, Purposeful and Inspired  

Our People are not Perfect, Traditional, Complacent or Cautious 

About the Role

As a Senior Site Reliability Engineer, you'll own reliability and infrastructure strategy across the organisation — not just for a single team. Decisions you make on system design, tooling, and process will directly affect every engineering team's ability to operate and ship at scale.

To drive success in this role, you will have a strong background in cloud infrastructure and platform engineering, with experience across infrastructure-as-code, CI/CD and release automation, observability, security, and disaster recovery. You should possess strong technical judgement, the ability to set standards other engineers adopt, and a proactive approach to eliminating operational risk before it becomes a problem.

Regular collaboration with software engineering teams, security teams, and other relevant stakeholders will be key in ensuring the reliability and efficiency of our production systems are achieved.

Key Responsibilities

  • Drive infrastructure and reliability strategy that connects directly to business outcomes - revenue protection, customer trust, and developer velocity

  • Analyze, test, and evolve systems to improve reliability and performance at an architectural/infrastructure level, setting technical direction other engineers build against

  • Architect multi-region, multi-AZ infrastructure with clear failover and disaster recovery strategies, applying deep AWS and GCP expertise to govern cloud infrastructure across multiple teams and projects

  • Define observability strategy and standards, and develop the tooling and dashboards other teams build on

  • Define and own SLOs and error budgets for services in your area, and use them to prioritise reliability work against feature velocity

  • Take a leading role in major incidents and lead troubleshooting on the most complex production issues, driving deployment safety and process improvements while building a strong postmortem and continuous-improvement culture across the organisation

  • Reduce operational toil by building automation and self-service platforms, rather than absorbing repetitive work yourself

  • Develop and maintain design, troubleshooting, and runbook standards that other engineers can follow without tribal knowledge

  • Mentor other engineers and raise the bar on production ownership, testing, and code review across teams

  • Apply AI-assisted development to infrastructure problems, and help build the tooling and practices that make the wider team more effective with it

  • Drive cloud cost optimisation at the organisational level - reserved capacity, right-sizing, FinOps practices

Skills & Experience

What you will bring:


  • 5+ years of experience in SRE, Platform Engineering, DevOps or other related roles 

  • Deep understanding of SRE and platform engineering principles, with a track record of setting them as team or organisational standards

  • Extensive experience using, configuring, and setting standards for modern observability tools such as Datadog, Elasticsearch, Prometheus, Grafana

  • Expert-level experience with cloud native and container technology such as Docker, and hands-on experience designing and managing Kubernetes clusters at scale

  • Deep experience defining infrastructure-as-code standards and module libraries using tools such as Terraform

  • Comfortable scripting and developing internal tooling with Bash and at least one programming language (e.g. Python, Go)

  • Fluent with AI-driven development environments like Cursor, Claude Code, or Gemini, with a proven ability to leverage these tools within production engineering workflows

  • Experience working with Linux

  • Strong understanding of networking, distributed systems, and system architecture at scale

  • Proven experience deploying, scaling, and monitoring web applications and databases across multi-region or high-availability environments

  • Expert-level knowledge of AWS and/or GCP platforms, with experience driving cloud cost optimisation and platform decisions at an organisational level

  • Bachelor's degree in Computer Science/Engineering, a postgraduate degree and/or record of academic achievement is also desirable

What you will have:

Ownership & Delivery

  • Takes ambiguous problems and drives them to shipped outcomes - not just code, but results, and takes accountability even without a clear owner

  • Balances speed with quality — knows when to iterate fast and when to invest in durability

  • Manages risk proactively — identifies failure modes and mitigates before they bite

Communication & Influence

  • Creates clarity from ambiguity; documents decisions so others can build on your work

  • Influences through evidence and collaboration, not authority — mentors and unblocks teammates

  • Communicates technical concepts clearly to engineers, product, and business stakeholders

AI-First Mindset

  • Treats AI tools as essential infrastructure, not optional add-ons — continuously experiments with new capabilities

  • Understands LLM strengths and limitations — knows when to prompt and when to build differently

  • Thinks in leverage: automates the repetitive, focuses human attention on judgement calls

  • Helps others adopt AI workflows, shares what works, and raises the floor for the whole team

Why Cover Genius?  

Cover Genius not only cares about being the best in our industry, we care about our team. We’re a business that understands life can be fluid and so we flex to ensure we provide the environment to suit that. What does that mean?  

• Flexible Work Environment - our teams are hybrid. We work from home on Wednesdays and Thursday and attend the office on Monday, Tuesday and Friday with flexibility around start/finish times.

• Global company, with the opportunity to work from any of our offices for 4 weeks a year

• Employee Stock Options - we want our people to share in our success, we reward them with ownership for their contribution in creating a world-class company.

• Work with like-minded people who are passionate about both the work we're doing and giving back. Our CG Gives programs enables us to all become philanthropists through our peer recognition and rewards system.

• Social Initiatives - pictures speak a thousand words!

Sound interesting? If you think you have the best composition of the above, send us your resume and let's chat!

* Cover Genius promotes diversity and inclusivity. We don't tolerate discrimination, demeaning treatment of anyone, or harassment due to race, national origin, gender, gender identity, sexual orientation, protected veteran status, disability, age, or any other legally protected status.

By submitting your application, you acknowledge that we may collect, store and process your personal data for recruitment purposes. To ensure a fair evaluation, we may use AI to assist in sorting applications, but all final decisions are made by our hiring team and no candidate dispositions are automated. We will keep your information on file for three years from the date of your application.  For detailed information about how we handle your data and our use of AI, please review our full Privacy Policy.

Skills Required

  • 5+ years of experience in SRE, Platform Engineering, DevOps or related roles
  • Deep understanding of SRE and platform engineering principles
  • Extensive experience using and configuring observability tools (Datadog, Elasticsearch, Prometheus, Grafana)
  • Expert-level experience with Docker and designing/managing Kubernetes clusters at scale
  • Deep experience defining infrastructure-as-code standards and module libraries using Terraform
  • Comfortable scripting and developing internal tooling with Bash and at least one programming language (Python or Go)
  • Expert-level knowledge of AWS and/or GCP platforms and cloud cost optimisation experience
  • Experience deploying, scaling, and monitoring web applications and databases across multi-region/high-availability environments
  • Experience working with Linux
  • Fluent with AI-driven development environments (Cursor, Claude Code, Gemini) and applying AI to infrastructure workflows
  • Strong understanding of networking, distributed systems, and system architecture at scale
  • Experience defining SLOs, error budgets, incident leadership, and postmortem culture
  • Bachelor's degree in Computer Science/Engineering or postgraduate degree

Cover Genius Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Cover Genius and has not been reviewed or approved by Cover Genius.

  • Flexible Benefits Flexible working conditions and work-from-home options are emphasized across employer materials and role descriptions. These elements are positioned as a core part of the package alongside wellness initiatives.
  • Leave & Time Off Breadth Unlimited PTO and paid parental leave appear prominently, with wellness time also referenced. These policies indicate breadth beyond statutory minimums.
  • Equity Value & Accessibility Employee stock options are presented as standard, with additional peer and founder awards supporting recognition. Equity participation is framed as a way for staff to share in company success.

Cover Genius Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
600 Employees
Year Founded: 2014

What We Do

Cover Genius is the insurtech for embedded protection. Together, we protect the global customers of the world’s largest digital companies including Booking Holdings, owner of Priceline and Booking.com, Intuit, Hopper, Skyscanner, Ryanair, Turkish Airlines, Descartes ShipRush, Zip and SeatGeek. We’re also available at Amazon, Flipkart, eBay, Wayfair and SE Asia’s largest company, Shopee. Cover Genius’ vision is to protect all the customers of the world’s largest online companies through XCover, an award-winning global distribution platform for any line of insurance or warranty, with an API for instant claims payments that holds an industry-leading NPS of +65‡. Cover Genius and its partners co-create solutions that embed protection that’s licensed or authorized in over 60 countries and all 50 US States.

Why Work With Us

We are a vibrant international team that promotes inclusivity and celebrates our differences. We are growing fast, we provide our employees with professional development opportunities and we promote within through our bi-annual performance review cycles. We are bold enough to take chances, to challenge the status quo and inspire each other.

Gallery

Gallery

Similar Jobs

Hybrid
Sydney, New South Wales, AUS
289097 Employees

NBCUniversal Logo NBCUniversal

National Trade Marketing Lead - 18 months contract.

AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Remote or Hybrid
Sydney, New South Wales, AUS

Airwallex Logo Airwallex

Engineering Manager

Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
In-Office
2 Locations
2300 Employees

Airwallex Logo Airwallex

Information Technology Engineer

Artificial Intelligence • Fintech • Payments • Business Intelligence • Financial Services • Generative AI
Hybrid
2 Locations
2300 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account