Reliability & Observability Analyst II

Posted 7 Days Ago
Be an Early Applicant
Sydney, New South Wales, AUS
In-Office
Mid level
Artificial Intelligence • Cloud • Infrastructure as a Service (IaaS) • Renewable Energy
The Role
Performs Level 2 incident analysis across GPU clusters, networks, and facilities; tunes alerts, routing, dashboards, and monitoring; validates AIOps outputs; maintains reliability metrics and ITSM records; implements small automations; supports RCA workflows, shift handoffs, and operational reporting; and coaches junior analysts in a 24/7 IOC/NOC environment.
Summary Generated by Built In

Job Type: Full-time | Location: Sydney |  | Work Location Type: #onsite

IREN is a vertically integrated AI Cloud provider, delivering large-scale data centers and GPU clusters for AI training and inference. IREN’s platform is underpinned by its expansive portfolio of grid-connected land and power in renewable-rich regions across North America, Europe and APAC.
With 100% renewable energy, we build, own and operate our data centers and take pride in being at the forefront of sustainable solutions for the ever-evolving applications of high-performance compute. We believe that human progress is invaluable, but it should be done in the right way – responsibly, sustainably and having a positive impact on the communities we operate in.

We are seeking an IOC Reliability & Observability Analyst II to support our 24/7 HPC Data Center Operations by performing advanced incident triage, improving alert quality and routing, and maintaining high-quality operational telemetry and reporting. This role partners with engineering and operations teams to identify detection gaps, tune monitoring and dashboards, and implement small automations and enrichment to reduce operational toil and improve time-to-action.

This is not a reporting-only role. You will partner closely with IOC, engineering, and operations teams to validate operational signals, tune alerts and dashboards, leverage AIOps outputs during incident response, and ensure telemetry is actionable for real-time triage and escalation.


Job requirements
  • 3–5 years of experience in IOC/NOC/SRE‑adjacent operations, reliability engineering, observability, or production support roles within 24/7 production environments
  • Bachelor’s degree in Computer Science, Data Science, IT, or equivalent hands‑on professional experience 
  • Demonstrated ability to apply reliability engineering principles (e.g., incident lifecycle, MTTD/MTTR, operational risk) to improve detection, response effectiveness, and overall service stability 
  • Strong working knowledge of Linux systems, basic networking, and infrastructure dependencies across compute, network, and facility domains 
  • Practical experience supporting GPU‑based compute environments or high‑density clusters, including analysis of GPU health, performance degradation, and failure patterns to reduce customer impact and improve reliability 
  • Proven experience owning and improving alert quality, including reduction of false positives, missed detections, poor routing, and alert fatigue across complex environments 
  • Hands‑on experience maintaining service health dashboards and operational reliability metrics, including supporting SLI/SLO reporting where defined by engineering or service owners
  • Ability to correlate logs, metrics, and alerts across distributed systems (including GPU, network, and facility telemetry) to accelerate triage and diagnose complex incidents 
  • Experience working with AIOps‑enabled outputs (e.g., anomaly detection, event correlation, automated enrichment), validating accuracy during incident triage and escalating when automated signals do not align with operational conditions
  • Ability to write or modify small automation artifacts (e.g., scripts, templates, configuration‑driven workflows) to standardize triage, enrich alerts or tickets, and reduce manual operational toil
  • Experience ensuring operational data integrity across ticketing systems, incident records, and dashboards to support trend analysis, high‑quality RCAs, and executive reporting 
  • Strong communication skills with the ability to work cross‑functionally with IOC leadership, engineering, and operations teams, including mentoring less‑experienced analysts
  • Strong working experience with IOC/NOC tooling, including ITSM/ticketing systems (e.g., ServiceNow, Jira) and monitoring platforms (e.g., Splunk, Datadog) 
  • Experience producing operational reports, incident summaries, and shift handoff documentation for IOC leadership and stakeholders 
  • Familiarity with RCA workflows, including ensuring incident records, timelines, and artifacts are complete and accurate
Other important requirements
  • This role operates in a 24×7 IOC/NOC environment. The schedule will 5 days a week, 8 hours a day.
  • Pre-employment screening, including background check and substance testing may be required according to company policies

Job responsibilities
  • Perform advanced Level 2 incident analysis by reviewing incident data, system behavior, and operational signals across GPU clusters, networks, and facilities to identify recurring issues, improve triage accuracy, and support faster and more effective escalation
  • Maintain IOC service health dashboards and operational metrics that reflect alert effectiveness, incident response performance (e.g., MTTD/MTTR), and customer impact for use in day‑to‑day operations and leadership reporting
  • Identify alerting and monitoring gaps, under‑monitored systems, and noisy or ineffective alerts; perform day‑to‑day tuning of thresholds, routing, suppression, and enrichment within IOC tooling, and partner with engineering teams when instrumentation changes are required
  • Own operational alert quality outcomes by ensuring sustained reductions in false positives, missed detections, poor routing, and alert fatigue through IOC‑approved standards, validation, and continuous review of alert performance
  • Analyze GPU health and performance signals (errors, degradation, failure indicators) during incidents to support faster triage, improve escalation quality, and reduce customer impact in GPU‑based environments
  • Validate and oversee automated detection and correlation outputs, ensuring alerts, anomalies, and insights are accurate, actionable, and aligned with operational reality 
  • Implement and maintain IOC‑level automation (e.g., alert routing rules, enrichment fields, ticket templates, runbook scripts) to standardize response and reduce manual toil during incidents
  • Ensure ITSM incident and ticket records meet IOC quality standards by validating timelines, categorizations, ownership, and resolution notes; support RCA workflows by providing complete operational inputs and tracking monitoring follow‑ups
  • Provide peer coaching and onboarding support to Analyst I team members on triage patterns, alert interpretation, dashboard usage, and IOC runbooks; contribute to and maintain operational documentation
  • Support IOC shift operations through detailed incident handoffs, queue hygiene, and coordination with on‑call engineering and facilities teams during escalations

Benefits

At IREN, we offer a highly competitive compensation package that includes base salary, annual performance incentives, and opportunities to build long-term wealth through equity programs. These offerings are part of our broader total rewards package, thoughtfully designed to support your health, well-being, and long-term success. 

  • Compensation & Rewards

    • Competitive salary range finalized based on experience and impact

    • Short and long-term incentive programs designed to reward both results and long term company success 

  • Wellbeing & Benefits 

    • Paid vacation to recharge, travel, or simply enjoy more life outside of work

We value diverse perspectives and believe that skills can be developed. If you’re passionate about this role, we want to hear from you — whether you meet every criteria or not. Your unique experiences might be exactly what we need!   

IREN Limited is an equal opportunity employer that is committed to creating an inclusive workplace. We evaluate qualified applicants without regard to race, colour, religion, age, sex, sexual orientation, gender identity, genetic information, national origin, disability, veteran status, and other legally protected characteristics.  
This job will remain posted until filled. While we appreciate all applications we receive, we are only able to contact candidates under consideration. 

By applying for this position and submitting your resume and application materials, you consent to the processing of your personal information in accordance with our Job Applicant Privacy Statement available on our website at www.iren.com.

Skills Required

  • 3-5 years of experience in IOC, NOC, SRE-adjacent operations, reliability engineering, observability, or production support in 24/7 production environments
  • Bachelor's degree in Computer Science, Data Science, IT, or equivalent hands-on professional experience
  • Experience applying reliability engineering principles, including incident lifecycle, MTTD, MTTR, and operational risk
  • Strong knowledge of Linux systems, basic networking, and infrastructure dependencies across compute, network, and facility domains
  • Experience supporting GPU-based compute environments or high-density clusters
  • Experience improving alert quality by reducing false positives, missed detections, poor routing, and alert fatigue
  • Experience maintaining service health dashboards and operational reliability metrics, including SLI/SLO reporting
  • Ability to correlate logs, metrics, and alerts across distributed systems
  • Experience with AIOps outputs such as anomaly detection, event correlation, and automated enrichment
  • Ability to write or modify small automation scripts, templates, or configuration-driven workflows
  • Experience maintaining operational data integrity across ticketing systems, incident records, and dashboards
  • Strong cross-functional communication skills and ability to mentor less-experienced analysts
  • Experience with IOC/NOC tooling, ITSM or ticketing systems such as ServiceNow or Jira, and monitoring platforms such as Splunk or Datadog
  • Experience producing operational reports, incident summaries, and shift handoff documentation
  • Familiarity with root-cause-analysis workflows and maintaining complete incident records, timelines, and artifacts
  • Availability for a 24/7 IOC/NOC schedule, five days per week and eight hours per day
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
644 Employees
Year Founded: 2018

What We Do

IREN is a vertically integrated AI Cloud provider that develops, owns, and operates large-scale data centers and GPU clusters for AI training and inference. Its platform combines direct access to scalable NVIDIA GPUs with grid-connected land and renewable power across North America and other regions. The company emphasizes sustainable, high-performance computing infrastructure for AI builders and enterprise customers across diverse workloads.

Similar Jobs

Dynatrace Logo Dynatrace

Sales Development Coordinator II

Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation
Remote or Hybrid
Sydney, New South Wales, AUS
5600 Employees

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Architect

Artificial Intelligence • Cloud • Information Technology • Consulting
In-Office
3 Locations
85422 Employees

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Architect

Artificial Intelligence • Cloud • Information Technology • Consulting
In-Office
3 Locations
85422 Employees
Hybrid
Sydney, New South Wales, AUS
289097 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account