Era4 develops, owns and operates AI infrastructure across the UK, powered by renewable energy. Converting legacy industrial and energy sites into modern data-centre facilities, Era4 is combining brownfield regeneration opportunities with cleaner, efficient, scalable compute capacity for healthcare, research, finance, enterprise, and public-sector organisations.
Role Summary:
This is a hands-on Technical Lead who can support the build, operation and scale Era4’s sovereign AI/HPC infrastructure. We need someone who can operate at the intersection of HPC platform engineering, GPU infrastructure, Linux systems, Kubernetes/Slurm, automation, observability and production incident response.
You will act as a technical authority, helping shape how GPU clusters are deployed, monitored, automated, supported and improved. You will work closely with SRE, platform engineering, infrastructure, vendors and customer-facing teams to ensure our platform is reliable, observable, scalable and ready for production workloads.
Key Responsibilities:
HPC & GPU Platform Leadership:
- Lead infrastructure, supporting GPU, compute, storage and networking platforms.
- Technical authority across platform engineering, SRE, infrastructure and customer-facing teams.
- Mentor engineers and help define technical standards, best practices and operational excellence.
Platform Reliability & Incident Response:
- Lead technical investigations during major incidents and production outages.
- Improve observability, monitoring and alerting across the platform.
- Drive root-cause analysis and implement long-term reliability improvements.
Automation & Platform Engineering:
- Improve platform scalability, deployment processes and operational efficiency.
- Oversee runbook creation, automation and development of inhouse Agent capabilities to support Operations
- Contribute to the design, build, enhancement of future capabilities
Customer & Technical Engagement:
- Support customer onboarding, complex technical escalations and platform adoption.
- Work with vendors, partners and internal teams to resolve infrastructure issues.
- Translate complex technical challenges into clear, actionable communication.
AI/HPC Infrastructure Evolution:
- Contribute to Technical and Operational Roadmaps
- Contribute to the deployment and optimisation of GPU clusters, Kubernetes environments and next-generation AI infrastructure.
Experience:
You do not need to tick every technology box, but you must bring hands-on experience in production infrastructure and clear depth in HPC, GPU, AI infrastructure, research computing, Neocloud, cloud HPC or high-density compute environments.
- Linux engineering background
- Experience in production of leading teams supporting HPC, GPU infrastructure, AI infrastructure, research computing, cloud HPC, Neocloud, platform engineering or SRE.
- Experience operating or supporting production infrastructure across compute, networking, storage and observability.
- Hands-on experience with at least one workload or orchestration layer such as Slurm, Kubernetes, Run:ai, LSF, PBS or equivalent.
- Experience with Open Source monitoring and troubleshooting using tools such as Prometheus, Grafana, OpenTelemetry, Loki or equivalent.
- Proven involvement in major incidents, on-call, escalation, root-cause analysis or production troubleshooting.
- Ability to mentor engineers, influence technical direction and lead by technical credibility.
- Comfortable working with internal teams, customers, suppliers and vendor engineering teams.
Why Join Era4:
You’ll be joining a mission-driven start-up building critical national infrastructure, where operational excellence directly enables growth. This role offers high visibility with leadership, real autonomy, and the chance to shape how a next-generation company operates at scale.
Diversity & Inclusion:
Era4 is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.
Skills Required
- Proven background within infrastructure operations, HPC, SRE, NOC, managed services, or equivalent mission-critical environment in a management or senior lead role.
- Demonstrated experience across at least two of the three domains: NOC/incident operations, service management, and SRE/platform engineering.
- Working knowledge of observability tooling like Grafana or Prometheus.
- Strong knowledge of Linux, container technology, and infrastructure supporting GPU and HPC workloads.
What We Do
Carbon3.ai is building the UK’s sovereign AI platform – secure, sustainable, and designed for real-world impact. AI growth demands are creating new challenges and compute power requirements are outpacing supply. At Carbon3.ai, we’re not just providing infrastructure, we’re building the foundations to overcome these challenges. We are an energy business transforming into the UK’s sovereign choice for AI. Vertically integrated from soil to software transforming legacy industrial sites into renewable powered AI data hubs. Designed, owned, and operated by Carbon3.ai, all infrastructure and data processing are located within the UK and fully subject to UK jurisdiction and regulatory oversight. We generate our own off-grid renewable power, providing low-cost, sustainable energy comparable to Nordic levels, making AI workloads both affordable and sustainable. We own 50+ sites across the UK and are rapidly scaling them into AI data centres, enabling high-density, low-latency, sovereign AI deployment at national scale. Whether you're training models, deploying intelligent agents, or building industry-specific solutions, Carbon3.ai accelerates your journey from concept to production. Backed by strategic partnerships with leading brands and robust investment, we’re building the infrastructure to power the UK’s most ambitious AI innovation – ensuring British enterprises can access world-class AI capabilities securely and sustainably.






