Site Reliability Engineer

Posted 9 Days Ago
Be an Early Applicant
Hiring Remotely in Office, Machaze, Manica, MOZ
Remote or Hybrid
180K-250K Annually
Senior level
Artificial Intelligence • Software • Cybersecurity
The Role
Own and operate a scalable identity platform: build cloud infrastructure (including FedRAMP), implement observability and automation, lead incident response and postmortems, partner with product teams for reliable feature rollout, and drive infrastructure projects with incremental delivery.
Summary Generated by Built In

C1.ai is building a horizontal identity platform
Identity is becoming the control plane for modern companies. Every organization runs on dozens—sometimes hundreds—of systems that need to answer the same basic question: who can do what, when, and why?
The old approach doesn’t hold up. It’s either IT teams duct-taping systems together, or rigid IGA tools that only work across a narrow set of HR and SSO integrations. That’s not going to cut it for the next decade.
We’re building something more foundational: a horizontal identity platform that any company—and increasingly, any AI agent—can plug into to understand and manage access. SaaS apps, infrastructure, internal tools, AI agents—if it has a concept of “who,” it should connect to us.
We’re still small, and we move fast. If you want to help build this from the ground up, you’ll have real ownership here.

What You'll Do
  • Own the reliability and scalability of our platform—design, build, and operate the infrastructure that keeps C1 running for customers who depend on us. You'll work on both our core cloud environment, as well as our FedRAMP environment.

  • Build observability that drives action—create monitoring, alerting, and tooling that helps teams understand system behavior and respond to incidents quickly

  • Drive operational excellence across engineering—partner with product teams to ensure new features are built with reliability in mind from the start

  • Automate relentlessly—if you're doing something twice, build a system to do it for you. We believe in infrastructure as code and eliminating toil.

  • Respond to and learn from incidents—lead incident response, conduct blameless postmortems, and drive systemic improvements

  • Plan and execute infrastructure projects with incremental deliverables—you'll assess technical risks, communicate tradeoffs, and ship iteratively

 
What We're Looking ForYou Likely Have
  • A track record of building and operating production systems at scale—you've kept real systems running for real customers

  • Deep experience with cloud infrastructure (AWS, GCP, or similar) and infrastructure-as-code (Terraform, Pulumi, or similar)

  • Strong programming skills in Go, Python, or similar—you write tools and automation, not just scripts

  • Experience with Kubernetes and container orchestration in production

  • Strong systems thinking—you understand how components interact and where failures cascade

  • High agency—you figure out what needs to be built, not just how to build what you're told. You move fast and unblock yourself.

  • Deep understanding of observability—you know how to instrument systems, build dashboards, and create alerts that actually matter

  • Experience with AI-assisted development (Claude Code, Cursor, Copilot, or similar)—you're already using these tools and excited about what's next

  • Clear, persuasive communication—you can explain complex systems to diverse audiences and drive alignment during incidents

  • Ego in check—you care about getting it right, not being right

You Might Also Have
  • Experience with AI/ML infrastructure or serving LLMs in production

  • Background in security-focused environments or compliance frameworks (SOC 2, FedRAMP, etc.)

  • Experience building developer platforms or internal tooling

  • Familiarity with identity systems and protocols (SCIM, SAML, OAuth, LDAP)

  • Contributions to open source projects or engineering communities

How we work
- Impact matters more than ownership.
- We’d rather ship something real than talk about it.
- People here are trusted to operate independently.
- If something isn’t working, we fix it properly—we don’t paper over it.
- Using AI effectively is expected, not optional.
Compensation & benefits
Salary range: $180k – $250k
- Meaningful equity
- Full medical, dental, and vision coverage
- In-office in Portland
- Solid benefits across the board
Due to the nature of this role and its responsibilities supporting our FedRAMP-authorized environment, candidates must be U.S. citizens. This position may involve work that the U.S. Government has determined can only be performed by U.S. citizens on U.S. soil.

Hiring philosophy
We care about both excellence and grit. You don’t need to check every single box to apply—if this feels like the kind of place you want to build, we’d like to hear from you.
c1.ai is an equal opportunity employer. We consider all qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or veteran status.

Skills Required

  • Proven experience building and operating production systems at scale
  • Deep experience with cloud infrastructure (AWS, GCP, or similar) and infrastructure-as-code (Terraform, Pulumi, or similar)
  • Strong programming skills in Go, Python, or similar (tooling and automation)
  • Experience with Kubernetes and container orchestration in production
  • Strong systems thinking and understanding of failure modes and cascading failures
  • Deep understanding of observability: instrumentation, dashboards, and meaningful alerts
  • Experience leading incident response, blameless postmortems, and driving systemic improvements
  • High agency and ability to operate independently to remove blockers
  • Experience using AI-assisted development tools (Claude Code, Cursor, Copilot, or similar)
  • Clear, persuasive communication and teamwork; humility (ego in check)
  • Experience with AI/ML infrastructure or serving large models in production
  • Background in security-focused environments or compliance frameworks (SOC 2, FedRAMP)
  • Experience building developer platforms or internal tooling
  • Familiarity with identity systems and protocols (SCIM, SAML, OAuth, LDAP)
  • Contributions to open source projects or engineering communities
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
2,100 Employees
Year Founded: 1993

What We Do

C1 is an AI-native identity security platform designed to protect human, non-human, and AI identities. By leveraging platform-level AI and powerful automation, C1 centralizes access visibility, enforces fine-grained controls, enables just-in-time access, and automates user access reviews across all applications, helping enterprises reduce their attack surface and enhance identity governance.

Similar Jobs

Lio (formerly askLio) Logo Lio (formerly askLio)

Site Reliability Engineer

Artificial Intelligence • Software
Remote
Office, Machaze, Manica, MOZ
92 Employees

Illumio Logo Illumio

Site Reliability Engineer

Software • Cybersecurity
Remote
Office, Machaze, Manica, MOZ
552 Employees
141K-162K Annually

Nabla Logo Nabla

Site Reliability Engineer

Artificial Intelligence • Healthtech • Machine Learning
Remote or Hybrid
Office, Machaze, Manica, MOZ
82 Employees
160K-220K Annually

Illumio Logo Illumio

Senior Site Reliability Engineer

Software • Cybersecurity
Remote
Office, Machaze, Manica, MOZ
552 Employees
170K-196K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account