Software Engineer, Infrastructure & Reliability

Reposted One Month Ago
Hiring Remotely in United States
Remote
Mid level
Artificial Intelligence • Software
The Role
Build and operate CrewAI's infrastructure across AWS, Azure, and GCP. Own containerized deployments, CI/CD pipelines, observability, secrets, networking, databases, and on-call reliability. Harden security, automate deployments and installs, partner with runtime and product engineers, and reduce operational toil through tooling and automation.
Summary Generated by Built In
About CrewAI

CrewAI is the leading framework and enterprise platform for building and orchestrating multi-agent AI systems, powering 300M+ agent executions per month across thousands of companies. The Agent Management Platform is our control plane for deploying, monitoring, governing, and scaling agents in production. This role owns the infrastructure foundation that keeps it reliable, secure, and fast.

The Role

You'll build and operate the platform infrastructure behind CrewAI's cloud and enterprise deployments. You'll work across multiple hyperscalers - AWS, Azure, and GCP. You’ll work on containers, CI/CD, deployment automation, observability, secrets, networking, and runtime reliability. Your job is to make the product and runtime teams faster while making customer’s production environments safer.

This is not a pure DevOps support role. You'll write code, improve systems, design deployment paths, harden production, and build the internal platform that lets CrewAI scale and scale our customer deployments.

What You'll Do
  • Own and improve the infrastructure that runs CrewAI's platform: AWS, ECS/ECR, Docker, Kubernetes/Helm, networking, secrets, databases, Redis, and related services.
  • Build and maintain CI/CD pipelines for build, test, image publishing, migrations, environment promotion, rollbacks, and deploy safety.
  • Improve reliability across cloud and enterprise deployments: health checks, alerting, incident response, capacity planning, recovery paths, and operational runbooks - and own the front-line on-call rotation and its SLAs.
  • Partner with runtime engineers on Celery/FastAPI/Redis workloads and with product engineers on Rails/Solid Queue/Postgres production behavior.
  • Manage production observability and telemetry infrastructure: logs, metrics, traces, dashboards, Sentry/OpenTelemetry plumbing, actionable alerts, and telemetry export to customers' own monitoring systems.
  • Harden security and compliance posture across IAM, workload identity, secrets management, vulnerability scanning, dependency/image hygiene, and least-privilege access.
  • Build the tooling and automation that lets field engineers and customers run self-hosted installs themselves - Helm charts, environment config, release artifacts, pre-flight checks, and install runbooks - so engineering does fewer hands-on installs over time.
  • Reduce operational toil by automating recurring workflows and making deployments boring.

RequirementsWhat We're Looking For
  • Strong infrastructure/platform engineering experience in production SaaS environments.
  • Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services.
  • Experience with ECS and/or Kubernetes; Helm experience is a strong plus.
  • Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production.
  • Strong debugging instincts across app, infra, network, deploy, and dependency layers.
  • Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access.
  • Ability to write reliable automation in Python, Ruby, Go, Bash, or similar.
  • Calm, rigorous approach to incidents, rollbacks, migrations, and production change management.
Bonus
  • Experience with AI/agent platforms, workflow runtimes, or high-volume async execution systems.
  • Experience supporting enterprise/self-hosted deployments.
  • Terraform or other IaC experience.
  • SRE background: SLOs, incident review, capacity planning, load testing.
  • Familiarity with Rails, FastAPI, Celery, OpenTelemetry, or multi-service observability.

Skills Required

  • Strong infrastructure/platform engineering experience in production SaaS environments
  • Practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services
  • Experience with ECS and/or Kubernetes
  • Helm experience
  • Experience across multiple hyperscalers (Azure, GCP)
  • Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production
  • Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access
  • Ability to write reliable automation in Python, Ruby, Go, Bash, or similar
  • Strong debugging instincts across app, infra, network, deploy, and dependency layers
  • Calm, rigorous approach to incidents, rollbacks, migrations, and production change management; participate in on-call rotation
  • Experience building and maintaining CI/CD pipelines, deploy safety, and rollback processes
  • Manage observability and telemetry: logs, metrics, traces, dashboards, actionable alerts
  • Experience with Terraform or other IaC
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Middletown, DE
7 Employees

What We Do

Most AI agent frameworks are hard to use. We provide power with simplicity. Automate your most important workflows quickly.

Similar Jobs

LiveKit Logo LiveKit

Senior Software Engineer

Artificial Intelligence • Cloud • Information Technology • Software
In-Office or Remote
30 Locations
83 Employees
135K-300K Annually

LiveKit Logo LiveKit

Senior Software Engineer

Artificial Intelligence • Information Technology • Internet of Things
In-Office or Remote
30 Locations
34 Employees
135K-300K Annually

Micron Technology Logo Micron Technology

Recruiter

Artificial Intelligence • Hardware • Information Technology • Machine Learning
Remote
California, USA
45000 Employees
107K-183K Annually

TransUnion Logo TransUnion

Client Executive - Auto Finance

Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Remote or Hybrid
Chicago, IL, USA
13000 Employees
94K-185K Annually

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account