Operations Engineering Manager (m/f/d)

Posted Yesterday
Be an Early Applicant
London, Greater London, England, GBR
In-Office
Senior level
Information Technology • Business Intelligence • Consulting
The Role
Leads a team responsible for the reliability, performance, and availability of GPU-accelerated HPC infrastructure. Oversees monitoring, incident and problem management, root cause analysis, automation, documentation, change management, and operational metrics. Owns Agile delivery practices, backlog alignment, cross-functional coordination, and continuous improvement. Acts as an escalation point during major incidents while managing team development, resource allocation, stakeholder communication, and third-party support relationships.
Summary Generated by Built In
Job Description

The Operations Engineering Manager will own the operational reliability of the company’s GPUaccelerated HPC infrastructure and lead a team of Operations Engineers. This role combines technical leadership with people management and Agile delivery ownership. 

 

The successful candidate will set the vision for operational excellence, manage and develop the internal team, and work closely with Platform, Network, and Infrastructure teams to build operational excellence. The role will also play a key part in implementing and maturing the company’s scaled Agile Framework across Operations Engineering, ensuring alignment, transparency, and continuous improvement. 

 

YOUR RESPONSIBILITIES

Team Leadership & Management 

  • Lead, coach, and develop an internal team of Operations Engineers. 

  • Set clear goals, priorities, and expectations for the team. 

  • Manage workload and resource allocation across the team. 

  • Support regular 1:1s, performance reviews, and development plans. 

  • Build a collaborative team culture focused on ownership and accountability. 

 

Operational Ownership & Reliability 

  • Own the reliability, performance, and availability of the GPU-accelerated HPC infrastructure from an Operations perspective. 

  • Oversee proactive system monitoring, incident trend analysis, and root cause analysis. 

  • Define, track, and report on key operational metrics. 

  • Ensure strong operational control through effective processes, runbooks, and change management. 

  • Drive proactive improvements in reliability, automation, and performance. 

 

Agile Ways of Working & Framework Oversight 

  • Champion and oversee the scaled Agile Framework within the Operations function. 

  • Collaborate with Product, Platform, and Network teams to align priorities and manage backlogs. 

  • Support Agile ceremonies such as planning, stand-ups, reviews, and retrospectives. 

  • Improve delivery flow, predictability, and cross-team coordination. 

  • Ensure work is prioritised and delivered in line with Agile principles. 

Process, Automation, and Documentation 

  • Own the creation and maintenance of operational documentation, SOPs, and troubleshooting guides. 

  • Drive automation initiatives to reduce manual effort and improve consistency. 

  • Promote best practices in scripting, configuration management, and observability. 

  • Maintain effective knowledge sharing across the team. 

  • Stay informed on relevant trends and assess opportunities for improvement. 

 

People & Stakeholder Communication 

  • Provide regular status updates and reporting to leadership. 

  • Represent Operations in cross-functional planning and strategic discussions. 

  • Act as an escalation point for major incidents and complex technical issues. 

  • Coordinate with Platform, Network, and third-party support teams during critical events. 

  • Communicate clearly to support alignment, risk management, and decision-making. 

 

YOUR QUALIFICATIONS 

Required 

  • 5+ years in infrastructure/operations, with 2+ years managing a technical team. 

  • Advanced Linux administration in production, ideally at scale. 

  • Proven experience running incident/problem management and working with thirdparty or external support teams. 

  • Handson with automation (Ansible or equivalent) and monitoring/observability tools (e.g. Grafana, Prometheus). 

  • Experience with Agile ways of working and exposure to scaled Agile frameworks. 

  • Excellent communication and stakeholder management skills, able to work closely with Platform, Network, and leadership. 

 

Nice to Have 

  • Experience in HPC or GPUaccelerated environments (NVIDIA GPUs, InfiniBand/RDMA, parallel file systems). 

  • Scripting skills in Python and/or Bash for automation and tooling. 

  • Understanding of performance tuning for HPC/GPU systems. 

  • Experience with CI/CD pipelines and modern DevOps tooling. 

  • Background designing or improving oncall rotations, runbooks, and incident readiness. 

WHAT WE OFFER

With us, you will work towards the future of HPC: From new, sustainable building methods for data centers to cooling concepts to software solutions for accelerated compute. 

Your approaches count: In official exchange formats or spontaneously at the coffee machine. At Northern Data, it's the best idea that counts - not the hierarchy. We’re looking forward to getting your inputs!

You make the difference in the company: Unlike in established corporations, at Northern Data you will really help shape things. From implementing new departments, to optimizing processes and culture. 

Best-in-class partners: The best work with Northern Data. This means a knowledge and time advantage from which your career and our customers benefit equally.

Green by heart: Sustainability is at the core of Northern Data. With us, you actively work on the carbon neutrality of datacenters worldwide. Beginning with our infrastructure and continuing with the solutions for our clients, we work towards a green future.

Home Office facts: Work with our international and virtual team flexible from home. And of course, your hardware wishes will be fulfilled to make your ideas for next level HPC come true.

Your wellness matters: At Northern Data we have regular wellbeing initiatives that are designed to promote wellness, diversity, inclusion, and much more, ensuring a supportive and enriching environment for our global team.

Skills Required

  • 5+ years of experience in infrastructure or operations
  • 2+ years managing a technical team
  • Advanced Linux administration in production environments, ideally at scale
  • Experience running incident and problem management
  • Experience working with third-party or external support teams
  • Hands-on automation experience with Ansible or equivalent
  • Experience with monitoring and observability tools such as Grafana and Prometheus
  • Experience with Agile ways of working
  • Exposure to scaled Agile frameworks
  • Excellent communication and stakeholder management skills
  • Experience in HPC or GPU-accelerated environments, including NVIDIA GPUs, InfiniBand/RDMA, or parallel file systems
  • Python and/or Bash scripting skills
  • Understanding of HPC/GPU system performance tuning
  • Experience with CI/CD pipelines and modern DevOps tooling
  • Experience designing or improving on-call rotations, runbooks, and incident readiness
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Frankfurt am Main
124 Employees

What We Do

At Northern Data Group, we believe unlimited High Performance Computing (HPC) will unlock unprecedented opportunities for research and development, business, and ultimately human progress. We power innovation through market-leading HPC infrastructure, operating across our three business divisions: Taiga Cloud, Ardent Data Centers and Peak Mining. Our global organization is rapidly becoming a world leader for GPU-based solutions by designing and operating ultra-efficient green HPC infrastructure. We uniquely combine intelligent and sustainable data centers, cutting-edge hardware and self-developed software for various HPC applications including Generative AI, Machine Learning and Bitcoin Mining. We operate from large-scale custom data centers and proprietary containerized data centers for ultimate site selection flexibility

Similar Jobs

Mastercard Logo Mastercard

Manager, Product Commercialization

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
London, Greater London, England, GBR
38800 Employees

Mastercard Logo Mastercard

Senior Specialist, Product Management

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Hybrid
London, Greater London, England, GBR
38800 Employees

bet365 Logo bet365

Maintenance Engineer, HVAC

Digital Media • Gaming • Software • Esports • Automation
In-Office
Stoke-on-Trent, Staffordshire, England, GBR
10000 Employees

Dscout Logo Dscout

Consultant

Enterprise Web • Mobile • Professional Services • Software
Easy Apply
In-Office
London, Greater London, England, GBR
180 Employees

Similar Companies Hiring

Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account