Platform Engineer

Posted 8 Days Ago
Hiring Remotely in United Kingdom
Remote
Senior level
Artificial Intelligence • Information Technology
The Role
Design, deploy, and operate large-scale GPU-accelerated HPC and AI clusters. Manage scheduling (Slurm), networking (InfiniBand/Ethernet), provisioning, automation, monitoring, incident response, security/RBAC, and vendor engagement to ensure high availability and performance for AI workloads.
Summary Generated by Built In

Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data-centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public-sector organisations


Role Summary:  

We are looking for Platform Engineer (HPC & AI) who can assist in shaping our new Platform team, this role will be customer facing, involve technical troubleshooting, and collaboration with vendor engineering teams to ensure seamless AI platform operations.   

  

Responsibilities:  

  • Designing, deploying, and managing large‑scale HPC and GPU‑accelerated clusters, including NVIDIA based compute environments. 
  • Implementing and administering HPC scheduling and resource‑management systems (e.g., Slurm), including GPU partitioning, workload scheduling, and capacity planning. 
  • Architecting and optimising InfiniBand and Ethernet network topologies. 
  • Ensuring high availability and resilience through failover strategies, planned maintenance coordination, and proactive risk mitigation. 
  • Automating provisioning, configuration, monitoring, and operational workflows across multi‑vendor HPC hardware and software stacks. 
  • Monitoring real‑time performance and leading troubleshooting efforts across compute, storage, interconnect, drivers, and node failures, engaging vendor support for critical issues. 
  • Incident response: node failure management, network issues, driver issues, troubleshooting common issues and then working with vendor support to resolve any critical issues.  
  • Security and access control: Manage user permissions, RBAC, security hardening, data protection.   

 

Required Skills & Experience:  

  • Experience supporting HPE PCAI or other AI/HPC infrastructure and platforms.  
  • System administration experience with OS's like RHEL/CentOS, Ubuntu, tuning Linux kernel. 
  • Proficiency with Ansible, Nvidia and CUDA toolkits, Kubernetes and container orchestration. 
  • Understanding of automation, monitoring and security with GPU as a service. 
  • Extensive experience in system engineering, platform operations or SRE. 
  • Experience with GPU resource allocation (across instances, GPUs count and time).  
  • Advanced networking skills with High performance networking, troubleshooting and fine tuning. 
  • Familiarity with cloud-based platforms, APIs, and distributed systems. 
  • Understanding of AI/ML concepts and tooling (model training, inference, data pipelines basics). 
  • Experience with monitoring/logging tools (e.g., Grafana, Kibana, Splunk).  
  • Excellent communication skills to interface with both customers and internal / vendor teams.  
  • Good understanding of tools requirements for ML engineers and data scientists, and how to optimise the experience. 


Why Join Era4:

You’ll be joining a mission-driven start-up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next-generation company operates at scale. 

 

Diversity & Inclusion:  

Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.  

 

Note: 

We appreciate this is a relatively new skill set and we are open to candidates who may not tick all the boxes but are willing to learn and develop their skillset.  

Skills Required

  • Experience supporting HPE PCAI or other AI/HPC infrastructure and platforms
  • System administration with RHEL/CentOS and Ubuntu, including Linux kernel tuning
  • Proficiency with Ansible for automation
  • Experience with NVIDIA toolkits and CUDA
  • Experience with Kubernetes and container orchestration
  • Implementing and administering HPC schedulers/resource management (e.g., Slurm)
  • GPU resource allocation and management across instances and time
  • Advanced high-performance networking skills, including InfiniBand and Ethernet tuning
  • Experience in system engineering, platform operations or SRE
  • Familiarity with cloud-based platforms, APIs, and distributed systems
  • Understanding of AI/ML concepts and tooling (model training, inference, data pipelines)
  • Experience with monitoring/logging tools such as Grafana, Kibana, or Splunk
  • Strong communication skills to interface with customers, internal teams, and vendors
  • Knowledge of security hardening, RBAC, and data protection for platform environments
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Mereworth
16 Employees

What We Do

Carbon3.ai is building the UK’s sovereign AI platform – secure, sustainable, and designed for real-world impact. AI growth demands are creating new challenges and compute power requirements are outpacing supply. At Carbon3.ai, we’re not just providing infrastructure, we’re building the foundations to overcome these challenges. We are an energy business transforming into the UK’s sovereign choice for AI. Vertically integrated from soil to software transforming legacy industrial sites into renewable powered AI data hubs. Designed, owned, and operated by Carbon3.ai, all infrastructure and data processing are located within the UK and fully subject to UK jurisdiction and regulatory oversight. We generate our own off-grid renewable power, providing low-cost, sustainable energy comparable to Nordic levels, making AI workloads both affordable and sustainable. We own 50+ sites across the UK and are rapidly scaling them into AI data centres, enabling high-density, low-latency, sovereign AI deployment at national scale. Whether you're training models, deploying intelligent agents, or building industry-specific solutions, Carbon3.ai accelerates your journey from concept to production. Backed by strategic partnerships with leading brands and robust investment, we’re building the infrastructure to power the UK’s most ambitious AI innovation – ensuring British enterprises can access world-class AI capabilities securely and sustainably.

Similar Jobs

n8n Logo n8n

Platform Engineer

Artificial Intelligence • Software • Automation
In-Office or Remote
26 Locations
61 Employees

BJAK Logo BJAK

Platform Engineer

Artificial Intelligence • Fintech • Software • Financial Services
Remote or Hybrid
United Kingdom
253 Employees

Monzo Bank Logo Monzo Bank

Platform Engineer

Fintech • Financial Services
In-Office or Remote
2 Locations
2030 Employees
65K-80K Annually

Westpac Logo Westpac

Platform Engineer

Fintech • Financial Services
Remote
Square, Newry Mourne and Down, Northern Ireland, GBR
16000 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account