We are seeking a Remote DevOps Engineer to join our distributed infrastructure team. In this role, you will design, automate, and maintain scalable cloud infrastructure, optimize continuous integration and continuous deployment (CI/CD) pipelines, and ensure high reliability, security, and uptime across our applications.
You will collaborate closely with software engineering, security, and product teams to establish modern infrastructure best practices in a fully remote environment.
Key Responsibilities
- Cloud Infrastructure & Infrastructure as Code (40%): Provision, manage, and scale cloud environments using Infrastructure as Code (IaC) tools like Terraform, CloudFormation, or Pulumi.
- CI/CD Automation (30%): Build, maintain, and optimize robust CI/CD pipelines (e.g., GitHub Actions, GitLab CI, Jenkins) to ensure fast, secure, and reliable software delivery.
- Monitoring, Logging & Incident Management (15%): Implement continuous monitoring, telemetry, and alerting solutions (Datadog, Prometheus, Grafana, ELK Stack) to proactively detect issues and optimize system performance.
- Security & Compliance (15%): Implement security best practices (DevSecOps), zero-trust access controls, secret management, and automated vulnerability scanning across cloud resources.
Qualifications & Requirements
Minimum Requirements:
- Experience: 3+ years of dedicated experience in DevOps, Site Reliability Engineering (SRE), or Cloud Architecture.
- Cloud Platforms: Hands-on expertise with at least one major cloud provider (AWS, GCP, or Azure).
- Infrastructure as Code (IaC): Advanced experience writing and managing modular IaC using Terraform or CloudFormation.
- Containers & Orchestration: Strong proficiency with Docker and Kubernetes (managed services like EKS, GKE, or AKS).
- Scripting: Proficiency in at least one scripting language (Python, Bash, or Go) for system automation.
- Remote Collaboration: Proven ability to work effectively in an asynchronous, distributed remote environment with strong documentation and written communication skills.
Preferred Qualifications (Nice-to-Have):
- AWS, GCP, or Kubernetes (CKA/CKAD) certifications.
- Experience with service mesh frameworks (Istio, Linkerd) and GitOps workflows (ArgoCD, Flux).
- Familiarity with database administration, scaling, and automated backup strategies.
Key Performance Indicators (KPIs)
- System Reliability: Maintaining application uptime (>99.9%) and minimizing Mean Time to Recovery (MTTR).
- Deployment Efficiency: Reducing build/deployment cycle times and improving automated pipeline test coverage.
- Infrastructure Cost Optimization: Efficiently managing cloud resource usage and performance overhead.
What We Offer
- Healthcare & Wellness: Comprehensive medical, dental, and vision insurance options.
- Work-Life Balance: Unlimited/Generous PTO, flexible core hours, and home office setup stipend.
- Growth & Learning: Annual budget for tech certifications, books, and conference attendance.
- Financial Perks: Competitive compensation package, 401(k) matching program, and remote work stipends.
Equal Opportunity Employer
We are an Equal Opportunity Employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All employment decisions are based on business needs, job requirements, and individual qualifications, without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, veteran status, or disability status.
Skills Required
- 3+ years of dedicated experience in DevOps, Site Reliability Engineering (SRE), or Cloud Architecture
- Hands-on expertise with at least one major cloud provider (AWS, GCP, or Azure)
- Advanced experience writing and managing modular Infrastructure as Code (Terraform or CloudFormation)
- Strong proficiency with Docker and Kubernetes (managed services like EKS, GKE, or AKS)
- Proficiency in at least one scripting language for automation (Python, Bash, or Go)
- Proven ability to work effectively in an asynchronous, distributed remote environment with strong documentation and written communication skills
- Experience with CI/CD tools and pipeline optimization (e.g., GitHub Actions, GitLab CI, Jenkins)
- Experience implementing monitoring, logging, and incident management (Datadog, Prometheus, Grafana, ELK Stack)
- AWS, GCP, or Kubernetes (CKA/CKAD) certifications
- Experience with service mesh frameworks (Istio, Linkerd) and GitOps workflows (ArgoCD, Flux)
- Familiarity with database administration, scaling, and automated backup strategies