Senior Site Reliability Engineer

Posted 2 Days Ago
Be an Early Applicant
2 Locations
In-Office
Senior level
Information Technology
The Role
Own reliability, observability, automation, security, and developer tooling for a cloud-based AI platform. Manage Kubernetes infrastructure supporting ML and LLM workloads, including GPU scheduling, model serving, autoscaling, monitoring, incident response, and cost governance. Build reusable Terraform modules, GitHub Actions, GitOps workflows, and platform components across public clouds. Partner cross-functionally to establish technical direction, secure-by-default infrastructure, and reliable software delivery practices while supporting 24x7 operations.
Summary Generated by Built In

At CloudFactory, we are a mission-driven team passionate about unlocking the potential of AI to transform the world. By combining advanced technology with a global network of talented people, we make unusable data usable, driving real-world impact at scale. 

More than just a workplace, we’re a global community founded on strong relationships and the belief that meaningful work transforms lives. Our commitment to earning, learning, and serving fuels everything we do as we strive to connect one million people to meaningful work and build leaders worth following.

Our Culture

At CloudFactory, we believe in building a workplace where everyone feels empowered, valued, and inspired to bring their authentic selves to work. We are:

  • Mission-Driven: We focus on creating economic and social impact.
  • People-Centric: We care deeply about our team’s growth, well-being, and sense of belonging.
  • Innovative: We embrace change and find better ways to do things together.
  • Globally Connected: We foster collaboration between diverse cultures and perspectives.

If you’re passionate about innovation, collaboration, and making a real impact, we’d love to have you on board!

Role Summary

As a Site Reliability Engineer, you will play a key role in keeping all production systems running smoothly. You will work closely with other engineers and operators to fuse engineering principles, operational knowledge, security, and automation to work towards platform/service production excellence from an angle of infrastructure, reliability, and security. 

The SRE team owns the foundation of AI Platform’s Core platform - the services and infrastructure that let us deploy to a multitude of public cloud providers and that powers many ML and LLM powered features. We give every other engineering team a reliable base to build on, and we own the software delivery lifecycle end to end: the tooling, patterns, and automation that reduce friction for the whole org.

This is an exciting opportunity to grow professionally while contributing to a mission-driven organization.

Responsibilities:

What you’ll own

  • Reliability of platform(includes ML and LLM workloads) - model serving and inference infrastructure (GPU-backed endpoints, autoscaling, latency and cost tradeoffs), with SLOs, on-call, and incident response that cover models, not just services
  • Observability(includes ML models) - drift and performance monitoring for ML, plus LLM-specific tracing, evals, and guardrails, wired into the same metrics and logging stacks we run everywhere else
  • Company-wide technical direction: shaping the roadmap and building golden paths that raise the baseline for every team
  • Developer tooling and automation that compounds - reusable GitHub Actions, GitOps workflows, Terraform modules - so every engineer ships faster
  • Reusable components packaging common open-source tools (Grafana, Istio, CloudNative stack, and ML tooling such as model registries and feature stores) for teams to deploy in any environment
  • Secure-by-default infrastructure - baking security, compliance audits, cost governance, and audit trails into the platform in close partnership with our lead/backend/staff engineers.

Requirements

Who you are (must-haves)

  • 5+ years in infrastructure engineering, DevOps, or SRE, operating large-scale, high-availability production systems using Kubernetes
  • Production Operational experience - a live cluster under real load, not a lab. Fluent with Helm, and Terraform or Cloudformation, on at least one major cloud (AWS preferred).
  • Good proficiency in Python or Go or general scripting for automation and tooling(automation with higher language preferred)
  • AI is already in your daily loop - Agentic tooling (Claude Code, Codex, Droid, internal skills) is part of how you ship and not what you are experimenting with. We believe AI tools can be great with human judgement and we want the SRE team to bring the next wave day to day operations. 
  • First-principles reasoning -  Reasoning from constraints and failure modes naming the tradeoff in business terms (reliability vs. velocity, cost vs. blast radius, standardisation vs. one-off) 
  • At least one infrastructure build you owned end to end - with the outcome metric attached (deploy time, MTTR, cost, adoption, availability).
  • Cross-functional strength. Track record working with product, backend/frontend teams to pull through collective initiative. 

ML & AI platform (strongly preferred)

  • Running ML workloads on Kubernetes - GPU scheduling, capacity, and cost management
  • Model serving and inference at production scale (eg KServe, RayServe, Triton, vLLM, or similar) with real latency and cost constraints(preferred RayServe)
  • MLOps pipeline tooling - training pipelines, model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights & Biases, or equivalents)
  • LLMOps in production - inference serving, prompt/version management, and LLM observability (tracing, evals, drift, guardrails, cost per request)
  • Governing ML/LLM workloads as platform capabilities: data-residency and PII controls, and audit trails

Any other General requirements

  • Global Collaboration: Ability to work across global teams and different cultures across various time zones with strong communication skills.
  • Problem Solving: Ability to break down complex problems into simple, actionable solutions.
  • Ownership & Drive: Tendency to go above and beyond to meet deadlines, manage own deliverables, and assist team members.
  • Availability: Willingness to support processes for 24x7 operational support.

Benefits

At CloudFactory, we believe that work should be more than just a job—it should be a platform for growth, impact, and community. Here, you’ll earn with purpose, learn every day, and serve a mission that truly matters. If you're looking for a career where you can develop professionally, contribute meaningfully, and be part of a global movement, we’d love to have you on this journey!

Join us today and be part of our mission to connect people and technology for a better world! Apply now and bring your whole, authentic self to work—we can’t wait to meet you!

Skills Required

  • 5+ years of experience in infrastructure engineering, DevOps, or SRE
  • Experience operating large-scale, high-availability production systems using Kubernetes
  • Production experience operating live Kubernetes clusters under real load
  • Fluency with Helm
  • Experience with Terraform or CloudFormation
  • Experience with at least one major cloud provider, preferably AWS
  • Proficiency in Python, Go, or general scripting for automation and tooling
  • Experience using agentic AI tooling such as Claude Code, Codex, Droid, or equivalent
  • First-principles reasoning and ability to evaluate reliability, velocity, cost, blast-radius, and standardization tradeoffs
  • Ownership of at least one infrastructure build end to end with measurable outcomes
  • Track record of cross-functional collaboration with product, backend, and frontend teams
  • Experience running ML workloads on Kubernetes, including GPU scheduling, capacity, and cost management
  • Production-scale model serving and inference experience with tools such as KServe, Ray Serve, Triton, or vLLM
  • Experience with MLOps pipeline tooling, model registries, feature stores, and lineage
  • Production LLMOps experience, including inference serving, prompt/version management, observability, evaluations, drift monitoring, guardrails, and cost tracking
  • Experience governing ML and LLM workloads with data residency, PII controls, and audit trails
  • Ability to collaborate across global teams, cultures, and time zones
  • Strong problem-solving and communication skills
  • Willingness to support 24x7 operational processes
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Reading
2,969 Employees
Year Founded: 2010

What We Do

CloudFactory is a global leader in combining people and technology to provide workforce solutions for machine learning and business process optimization. Our professionally managed and trained teams work with high accuracy using virtually any tool. We process millions of tasks a day for innovators including Microsoft, GoSpotCheck, Hummingbird Technologies, Ibotta, and Luminar. We exist to create meaningful work for one million talented people in developing nations, so we can earn, learn, and serve our way to become leaders worth following. We’re on four continents, with offices in the U.K., U.S., Nepal, and Kenya. To learn more, visit www.cloudfactory.com.

Similar Jobs

In-Office
Berlin, DEU
3117 Employees

ScorePlay Logo ScorePlay

Senior Platform Engineer

Artificial Intelligence • Digital Media • Software • Sports
In-Office or Remote
27 Locations
70 Employees

Solaris SE Logo Solaris SE

Senior Site Reliability Engineer

Fintech • Payments • Financial Services
In-Office
Berlin, DEU
736 Employees
75K-95K Annually

Scalable Capital Logo Scalable Capital

Site Reliability Engineer

Fintech • Payments • Financial Services
Hybrid
Berlin, DEU
488 Employees

Similar Companies Hiring

Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account