Senior DevOps Engineer

Posted Yesterday
Be an Early Applicant
2 Locations
In-Office or Remote
Senior level
Artificial Intelligence • Software
The Role
Design, build, and operate AWS and Kubernetes infrastructure powering production AI systems. Own CI/CD, infrastructure as code, GitOps, observability, secrets management, GPU workloads, model serving, reliability, security, scalability, and cost optimization. Troubleshoot production incidents, perform root-cause analysis, and develop automation using Python and Bash while collaborating across Engineering, ML/AI, Security, and Product teams.
Summary Generated by Built In

About Dynamo AI


Dynamo AI helps enterprises deploy AI systems that are reliable, secure, and production-ready. Our market-leading technical controls span AI evaluations, guardrails, agentic risk management, and observability. Backed by world-class talent, we partner with innovative, highly-regulated global organizations — across financial services, government, and beyond — to deploy meaningful AI use cases at scale, securely and compliantly.

About the Role

We are looking for a Senior DevOps Engineer to help build, scale, and operate the infrastructure powering our AI platform. This is a hands-on, high-ownership role. We are looking for someone who can work independently, solve complex infrastructure problems, and thrive in a fast-paced startup environment. You should be comfortable designing systems, automating processes, troubleshooting production issues, and continuously improving reliability, scalability, and cost efficiency.

What You'll Own
  • Design, build, and operate highly available production infrastructure on AWS, with strong expertise in EKS, EC2, VPC, S3, RDS/Aurora, IAM, ECR, ElastiCache, Load Balancers, and other core AWS services.
  • Build and improve CI/CD and release automation using Jenkins, GitHub Actions, Helm, ArgoCD, and GitOps.
  • Manage infrastructure using Terraform and Infrastructure as Code principles.
  • Build and operate Kubernetes platforms at production scale, including cluster management, upgrades, autoscaling, networking, security, and troubleshooting.
  • Run and operate AI/ML workloads in production, with a strong understanding of the infrastructure challenges associated with AI systems.
  • Deploy, scale, monitor, and optimize AI inference and model-serving workloads across Kubernetes and cloud infrastructure.
  • Work with GPU-based workloads, including GPU scheduling, utilization, autoscaling, capacity planning, and optimization.
  • Drive infrastructure efficiency by balancing performance, reliability, scalability, and cost across AI workloads.
  • Own monitoring, logging, and observability using tools such as Prometheus, Grafana, Thanos, and OpenTelemetry.
  • Implement secure secrets management using technologies such as HashiCorp Vault and External Secrets Operator.
  • Develop automation and internal tooling using Python and Bash.
  • Drive improvements around reliability, security, scalability, performance, and infrastructure cost.
  • Participate in production incidents, root-cause analysis, and drive long-term fixes rather than short-term workarounds.
  • Work closely with Engineering, ML/AI, Security, and Product teams to solve infrastructure and platform challenges.
What We're Looking For
  • 5+ years of strong hands-on experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure.
  • Strong production experience with AWS and Kubernetes/EKS.
  • Proven experience running AI/ML workloads or GPU-based workloads in production is highly valuable.
  • Strong understanding of CI/CD, Infrastructure as Code, GitOps, observability, and cloud security.
  • Excellent scripting and automation skills in Python and Bash.
  • Experience operating production systems and troubleshooting complex infrastructure issues independently.
  • Strong understanding of scaling, performance optimization, resource utilization, and cost management, particularly for compute-intensive workloads.
  • Strong ownership mindset with the ability to take a problem from design to production.
  • Experience working in a startup or fast-moving engineering environment is highly valued.
  • Strong communication skills and the ability to work effectively across teams.
Nice to Have
  • Experience with AI inference platforms, model serving, LLM infrastructure, or ML platforms.
  • Experience with GPU infrastructure such as NVIDIA GPUs and Kubernetes GPU scheduling.
  • Experience with multi-region or highly distributed systems.
  • Experience with SOC 2, ISO 27001, or other security/compliance requirements.
  • Experience with PostgreSQL, MongoDB, Redis, Kafka, or similar distributed systems.
The Kind of Engineer We Want

We're looking for someone who builds, automates, and takes ownership — not someone who simply operates existing infrastructure.

You should be comfortable with ambiguity, willing to dive deep into production problems, and constantly looking for ways to make our platform more reliable, secure, scalable, and cost-efficient.

Most importantly, you should understand that AI infrastructure has a different set of operational challenges. We want someone who can help us run AI systems efficiently at scale — making the right trade-offs between GPU utilization, performance, reliability, scalability, and cost.

This is a high-ownership role for someone who wants to make a meaningful impact on the infrastructure behind an AI platform as we scale.


Skills Required

  • 5+ years of hands-on experience in DevOps, SRE, Platform Engineering, or Cloud Infrastructure
  • Strong production experience with AWS and Kubernetes/EKS
  • Experience with CI/CD, Infrastructure as Code, GitOps, observability, and cloud security
  • Strong scripting and automation skills in Python and Bash
  • Experience operating production systems and independently troubleshooting complex infrastructure issues
  • Understanding of scaling, performance optimization, resource utilization, and cost management for compute-intensive workloads
  • Strong ownership mindset and ability to take problems from design through production
  • Strong communication and cross-functional collaboration skills
  • Experience running AI/ML or GPU-based workloads in production
  • Experience working in a startup or fast-moving engineering environment
  • Experience with AI inference platforms, model serving, LLM infrastructure, or ML platforms
  • Experience with NVIDIA GPU infrastructure and Kubernetes GPU scheduling
  • Experience with multi-region or highly distributed systems
  • Experience with SOC 2, ISO 27001, or other security and compliance requirements
  • Experience with PostgreSQL, MongoDB, Redis, Kafka, or similar distributed systems
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
58 Employees
Year Founded: 2021

What We Do

Dynamo AI is pioneering the first end-to-end secure and compliant generative AI infrastructure that runs in any on-premise or cloud environment. With a holistic approach to GenAI compliance, we help accelerate enterprise adoption to deploy secure, reliable, and compliant AI applications at scale. Our platform includes three products: - DynamoEval evaluates GenAI models for security, hallucination, privacy, and compliance risks. - DynamoEnhance remediates identified risks, ensuring more reliable operations. - DynamoGuard offers real-time guardrailing, customizable in natural language and with minimal latency Our client base and partnerships include Fortune 1000 companies across all industries, which underscores our proven success in securing GenAI in highly regulated environments

Similar Jobs

Remote
India
5395 Employees

TechnologyAdvice Logo TechnologyAdvice

Senior Devops Engineer

AdTech • Digital Media • Information Technology • Marketing Tech • News + Entertainment • Social Media • Software
Remote
India
400 Employees
2K-2K Annually

Rapid7 Logo Rapid7

Senior Software Engineer

Artificial Intelligence • Cloud • Information Technology • Sales • Security • Software • Cybersecurity
Remote or Hybrid
Pune, Maharashtra, IND
2400 Employees

Sphera Logo Sphera

Senior Devops Engineer

Cloud • Information Technology • Software
Remote
IN
1300 Employees

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account