We’re looking for an exceptional Sr. DevOps Engineer to join our team. As a Lyra Sr. DevOps Engineer, you will be responsible for designing, implementing, and maintaining our cloud infrastructure, as well as supporting our software delivery pipelines. We care deeply about making a difference in people’s lives, and we hope you do too!
You will focus heavily on modernizing our platform tooling, evolving our production Elastic Kubernetes Service (EKS) footprint, and optimizing our managed cloud datastores (AWS RDS, Aurora, ElastiCache). If you love creating frictionless experiences for developers and automating core infrastructure, this role is for you.
This is a remote role however candidates must reside within the United States.
Responsibilities
Design and implement infrastructure using tools such as Terraform/OpenTofu, Python, and Kubernetes onto our AWS cloud platform.
Collaborate closely with the engineering, data science and machine learning teams to continually implement, maintain, and improve our cloud infrastructure and developer experience.
Platform & Tooling Focus: Apply a software product mindset to build, maintain, and abstract internal platform tools, developer templates, and self-service CI/CD pipelines.
Automate infrastructure and operations tasks using scripting languages.
Security & Governance: Operationalize cloud security by enforcing continuous vulnerability remediation, strict IAM policies, and compliance guardrails directly within platform pipelines.
Monitor and analyze systems and network performance, and implement proactive measures to ensure the availability and reliability of our infrastructure.
Enforcing IAM processes and tools that ensure compliance with data privacy and protection regulations.
Working alongside engineers to ensure that security vulnerabilities are remediated.
Qualifications
5+ years in a DevOps, SRE, or Infrastructure role heavily focused on public cloud environments (AWS required).
Experience with Amazon Web Services (AWS) a must, especially many of the following:
EKS, SageMaker, EC2, ECS, VPC, IAM, KMS, RDS, ELB/ALB, MWAA, Identity Center
Strong IaC Practices: Demonstrated experience building modular, reusable Terraform or OpenTofu infrastructure.
Experience building and maintaining templates and packages with Helm.
Experience developing CI/CD pipelines with tools such as GitHub Actions, Jenkins, and/or ArgoCD.
Deep EKS & Container Expertise: Hands-on experience administering, scaling, and troubleshooting multi-tenant Kubernetes (EKS) environments in production.
Strong coding skills in Python, Go, or Bash to write automation scripts, tools, and integrations for internal development teams.
Effective communication and collaboration skills, ability to work independently and as part of a team, and excellent problem-solving and troubleshooting skills.
Preferred Qualifications
Experience with modern cluster autoscaling strategies (e.g., Karpenter) and cloud cost optimization techniques.
Familiarity with developer self-service tooling or internal developer portals
Proficiency with logging analysis, performance monitoring, and performance tuning.
Background supporting ML/Data Science workloads in cloud environments (e.g., SageMaker integrations, feature stores).
Skills Required
- 5+ years in a DevOps, SRE, or Infrastructure role focused on public cloud
- Experience with Amazon Web Services (AWS) including EKS, EC2, ECS, VPC, IAM, KMS, RDS, ELB/ALB, MWAA, Identity Center
- Design and implement infrastructure using Terraform or OpenTofu
- Strong Kubernetes/EKS administration, scaling, and troubleshooting experience in production
- Experience building and maintaining Helm charts and templates
- Develop CI/CD pipelines with GitHub Actions, Jenkins, and/or ArgoCD
- Strong coding/scripting skills in Python, Go, or Bash for automation
- Experience automating infrastructure and operational tasks and operationalizing cloud security (IAM, vulnerability remediation, compliance guardrails)
- Effective communication and collaboration skills; ability to work independently and with teams
- Experience with modern cluster autoscaling (e.g., Karpenter)
- Experience with cloud cost optimization techniques
- Familiarity with developer self-service tooling or internal developer portals
- Proficiency with logging analysis, performance monitoring, and tuning
- Background supporting ML/Data Science workloads in cloud environments (e.g., SageMaker integrations, feature stores)
What We Do
Lyra Health’s mission is to transform mental health care through technology with a human touch — to get more patients the care they need when they need it. If you are an engineer or data scientist who would like to join in this effort, please reach out.
Gallery

.png)






