DevOps Engineer III

Posted Yesterday
Hiring Remotely in Austin, TX, USA
In-Office or Remote
Senior level
Healthtech • Information Technology • Professional Services
The Role
Own and scale AWS cloud infrastructure, Terraform-based infrastructure as code, Kubernetes/EKS clusters, Docker environments, and GitHub Actions CI/CD pipelines. Automate deployments, maintain observability, backups, disaster recovery, and incident response while supporting HIPAA and SOC 2 compliance. Partner with engineers to operate self-hosted tools, improve reliability and security, document systems, and enable safe self-service. Participate in on-call support and lead infrastructure improvements across multi-account environments.
Summary Generated by Built In
Role Overview

We are looking for a skilled DevOps Engineer to help build, maintain, and scale our cloud infrastructure and deployment pipelines. In this role, you will be the backbone of our engineering delivery, ensuring that our applications are highly available, secure, and deployable at a moment's notice. You will take ownership of our containerized environments using Docker and Kubernetes (EKS), manage our cloud resources in AWS as code with Terraform, and streamline our CI/CD workflows using GitHub Actions.

Our platform supports healthcare operations, so protecting patient data is part of everything we build. You will work independently while partnering closely with our engineers, supporting both the applications we build and a growing set of self-hosted open-source tools. If you are passionate about automating manual processes, eliminating downtime, and building resilient systems that never depend on a single person, this is the role for you.

Key Responsibilities

1. Cloud Infrastructure & Architecture

  • AWS Management: Provision, configure, and maintain scalable cloud infrastructure across various AWS services (e.g., EC2, RDS, S3, VPC, IAM).
  • Infrastructure as Code: Own our Terraform and Terragrunt codebase across AWS, Cloudflare and GitHub. Keep environments reproducible, secure, version-controlled and changed only through reviewed pull requests.
  • Account and environment structure: Plan and carry out changes to our AWS account structure, including moving selected workloads and data between accounts with no data loss and minimal downtime.
  • Security & Compliance: Enforce security best practices across our cloud environments, including access control, network security and encryption. Help keep our infrastructure HIPAA compliant and support SOC 2 work.

2. Containerization & Orchestration

  • Kubernetes Administration: Deploy, manage, upgrade and scale our EKS clusters. Ensure high availability, proper resource allocation and good performance for our services.
  • Dockerization: Work closely with software engineers to containerize applications, optimizing images for size, security and build speed.
  • Self-hosted open-source applications: Run, upgrade and secure the open-source tools we host ourselves (for example, CRM, analytics, workflow automation and internal databases). Help engineering decide when to self-host and when to use a managed service.

3. CI/CD & Automation

  • Pipeline Development: Design, build, and maintain robust Continuous Integration and Continuous Deployment (CI/CD) pipelines using GitHub Actions, including self-hosted runners.
  • Release Engineering: Automate testing, staging, and production deployments to ensure smooth, zero-downtime releases.
  • Process Automation: Identify bottlenecks in the development lifecycle and write scripts to automate repetitive operational tasks.

4. Monitoring, Logging & Reliability

  • Observability: Run our monitoring, alerting, logging and tracing so we can see system health and performance, while keeping PHI out of telemetry.
  • Backups and recovery: Own database backups and disaster recovery, and regularly test that restores actually work.
  • Incident Response: Participate in an on-call rotation to troubleshoot and resolve production issues, conducting blameless root cause analyses (RCAs) to prevent recurrence.

5. Collaboration and shared ownership

  • Partner with engineering: Work independently day to day, in close collaboration with our engineers on architecture, reviews and priorities.
  • Avoid single points of failure: Keep runbooks and documentation current, train at least one backup for each critical system, and make sure there are break-glass access procedures.
  • Enable self-service: Give engineers safe, scoped ways to deploy and operate their own services without waiting on you.

RequirementsQualifications
  • Experience: 5+ years of hands-on experience in a DevOps, Site Reliability (SRE) or Cloud Engineering role, including owning production infrastructure.
  • Infrastructure as Code: Production experience managing infrastructure with Terraform (Terragrunt, Pulumi or CloudFormation experience also counts). This includes state management, imports, refactoring without downtime, and reviewing plans as part of pull requests.
  • Cloud expertise: Strong production experience running infrastructure in AWS, including IAM, networking and managed databases.
  • Container orchestration: Deep understanding of Docker, and production experience running Kubernetes (EKS preferred), including Helm, networking, upgrades and scaling.
  • CI/CD: A track record of building complex, automated pipelines in GitHub Actions, including secure cloud authentication (OIDC).
  • Networking fundamentals: Solid understanding of cloud networking: DNS, load balancing, VPCs, subnets, security groups and private connectivity.
  • Scripting: Strong Python or Bash skills for automating operational work.
  • Independent ownership: A history of being the primary owner of infrastructure while working closely with application engineers. You document your work, share knowledge and design things so no system depends on you alone.
Preferred Skills
  • Healthcare and HIPAA: Experience running infrastructure that handles PHI in a HIPAA-regulated environment. That includes encryption, access logging, retention and working with vendors under BAAs.
  • Self-hosted open source: Experience running third-party open-source apps in production, including version pinning, upgrades with database migrations, and rollbacks.
  • Multi-account AWS: Experience with AWS Organizations, or with moving workloads and data between AWS accounts.
Bonus Skills
  • SOC 2: Experience preparing for or supporting a SOC 2 audit, especially automating evidence collection.
  • Observability tools: Experience with the Grafana stack (Loki, Tempo, Mimir), Prometheus, OpenTelemetry or similar.

Skills Required

  • 5+ years of hands-on experience in DevOps, Site Reliability Engineering, or Cloud Engineering, including ownership of production infrastructure
  • Production experience managing infrastructure with Terraform or comparable infrastructure-as-code tools such as Terragrunt, Pulumi, or CloudFormation
  • Strong production experience running AWS infrastructure, including IAM, networking, and managed databases
  • Deep Docker knowledge and production Kubernetes experience, preferably Amazon EKS, including Helm, networking, upgrades, and scaling
  • Experience building complex automated CI/CD pipelines with GitHub Actions and secure cloud authentication using OIDC
  • Solid understanding of DNS, load balancing, VPCs, subnets, security groups, and private connectivity
  • Strong Python or Bash scripting skills for operational automation
  • History of independently owning infrastructure, documenting systems, sharing knowledge, and eliminating single points of failure
  • Experience running infrastructure handling PHI in a HIPAA-regulated environment
  • Experience running third-party open-source applications in production, including version pinning, database migrations, and rollbacks
  • Experience with AWS Organizations or moving workloads and data between AWS accounts
  • Experience supporting SOC 2 audits, especially automated evidence collection
  • Experience with Grafana, Loki, Tempo, Mimir, Prometheus, OpenTelemetry, or similar observability tools
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Austin, TX
Year Founded: 2017

What We Do

Enable Dental provides portable, at-home dental services, specializing in care for seniors and individuals with special needs, bringing comprehensive dental care directly to patients' homes or communities.

Similar Jobs

TensorWave Logo TensorWave

Devops Engineer

Artificial Intelligence • Cloud • Software
Remote
USA
56 Employees
Remote
United States
174 Employees

Espresso Systems Logo Espresso Systems

Devops Engineer

Blockchain • Software
Remote
USA
32 Employees

Octus Logo Octus

Devops Engineer

Fintech • News + Entertainment • Software • Database • Financial Services
Easy Apply
Remote or Hybrid
United States
808 Employees
165K-190K Annually

Similar Companies Hiring

NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Vitalize Thumbnail
Artificial Intelligence • Healthtech • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account