This is not a traditional operations role.
What you’ll be doing:
Design, deploy, and maintain the reliability, availability, and performance of critical JumpCloud systems and APIs across AWS and GCP.
Operationalize SLIs, SLOs, and error budgets in direct partnership with core application teams.
Build and refine end-to-end observability across microservices and cloud infrastructure using tools like Datadog.
Implement actionable monitoring across Golden Signals (Latency, Traffic, Errors, Saturation) to optimize detection (MTTD) and minimize alert fatigue.
Participate in on-call rotations, incident response, and blameless post-incident reviews to drive continuous systemic improvements.
Manage and operationalize production Kubernetes (EKS) clusters utilizing GitOps delivery workflows (Argo CD, Kargo).
Provision and secure multi-cloud infrastructure using modular Terraform (Infrastructure-as-Code).
Develop and maintain Disaster Recovery (DR) dashboards, runbooks, multi-region failover automation, and validation tests to ensure alignment with defined RTO and RPO targets.
Eliminate operational toil by writing production-grade Python or Go scripts and automation tools.
Leverage AI-assisted development tools (Cursor, Claude Code, GitHub Copilot) to accelerate scripting, runbook generation, and incident triage.
We’re looking for:
5+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical systems.
Python/Go Proficiency: Hands-on capabilities writing code for SRE tools, custom automation, and cloud integrations.
Kubernetes Ecosystem: Production experience with Kubernetes cluster operations, container orchestration, and GitOps pipelines (Argo CD).
Infrastructure as Code: Solid experience writing, maintaining, and modularizing Terraform configurations.
Cloud Architecture: Direct experience operating cloud workloads on AWS (EKS, IAM, VPC networking, Route53, ALB/NLB) or GCP.
FinOps & Cost Visibility: Practical experience setting up cost-allocation tagging, resource right-sizing, and building FinOps dashboards to visualize cloud spend.
Disaster Recovery & Monitoring: Experience building DR dashboards, running failover drills, and configuring monitoring tools to track system health and recovery metrics
Observability & Incident Management: Practical experience with Datadog (or similar), PagerDuty, alerting hygiene, and working within SLI/SLO frameworks.
Solid operational experience configuring and troubleshooting production service meshes (Istio or similar) and managing high-availability proxy solutions (HAProxy, NGINX, or similar).
Problem Solving & Mindset: Strong troubleshooting skills, effective collaboration, and a track record of driving operational efficiency through code.
A strong team player who helps us live by our core values: building connections, thinking big, and getting 1% better every day.
Preferred Qualifications:
Experience with CI/CD tools such as GitHub Actions or GitLab Pipelines.
Basic understanding of chaos engineering principles or testing resilience in staging/production.
Familiarity with secrets management tools (HashiCorp Vault, AWS Secrets Manager, External Secrets Operator).
Basic knowledge of DevSecOps tools and scanning/fixing infrastructure-as-code vulnerabilities.
Skills Required
- 5+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering
- Experience operating 24/7 mission-critical production systems
- Proficiency writing Python or Go for SRE tools, automation, and cloud integrations
- Production Kubernetes cluster operations and container orchestration experience
- Experience with GitOps pipelines, particularly Argo CD
- Experience writing, maintaining, and modularizing Terraform configurations
- Experience operating cloud workloads on AWS or GCP
- Experience with AWS services including EKS, IAM, VPC networking, Route 53, ALB, or NLB
- Experience with FinOps, cost-allocation tagging, resource right-sizing, and cloud-spend dashboards
- Experience building disaster recovery dashboards, conducting failover drills, and monitoring recovery metrics
- Experience with Datadog or similar monitoring tools, PagerDuty, alerting hygiene, and SLI/SLO frameworks
- Production service mesh configuration and troubleshooting, such as Istio
- Experience managing high-availability proxy solutions such as HAProxy or NGINX
- Strong troubleshooting, collaboration, and operational-efficiency skills
- Fluent spoken and written English
- Located in and authorized to work in India
- Experience with CI/CD tools such as GitHub Actions or GitLab Pipelines
- Understanding of chaos engineering or resilience testing
- Familiarity with HashiCorp Vault, AWS Secrets Manager, or External Secrets Operator
- Knowledge of DevSecOps tools and infrastructure-as-code vulnerability remediation
JumpCloud Compensation & Benefits Highlights
-
Healthcare Strength — Multiple medical plan options (including an HSA with employer contribution) plus dental, vision, EAP, disability, and life insurance are explicitly offered. Feedback suggests the overall healthcare package is a core strength.
-
Retirement Support — A 401(k) with company match is clearly provided and repeatedly highlighted in company materials and employee commentary. Feedback suggests retirement benefits contribute meaningfully to total rewards.
-
Leave & Time Off Breadth — Flexible/unlimited PTO, paid sick leave, and parental leave are standard, and feedback suggests the time‑off policy is genuinely usable and encouraged. This breadth of leave is consistently noted alongside a remote‑first setup.
JumpCloud Insights
What We Do
JumpCloud’s mission is to Make Work Happen®, providing simple, secure access to an organization’s technology resources from any device, or any location. The JumpCloud Open Directory Platform gives IT, security operations, and DevOps a single, cloud-based solution to control and manage employee identities and their devices, and apply conditional access controls based on Zero Trust principals. Since launching in 2012, our global user base has grown to more than 150,000 organizations, with more than 5,000 paying customers including Cars.com, GoFundMe, Grab, ClassPass, Uplight and Peloton. JumpCloud has raised over $400M from world-class investors including Sapphire Ventures, General Atlantic, Sands Capital, Atlassian, and CrowdStrike. Our teams are growing fast, too, and we're looking for talent across engineering, sales, customer success, marketing, product management, and more. Join our team of dedicated, passionate, and creative people who are eager to change the IT industry forever. We live by our core values which are: Build Connections Think Big 1% Better Every Day
Why Work With Us
We offer an incredible opportunity to see your impact. Each team member gets an up close personal view and education into building a fast growing startup. We are transparent about what we are doing, how we are doing it, and the decisions that we are making. There is opportunity to progress and flexibility to find unique approaches to our business
Gallery
JumpCloud Offices
Remote Workspace
Employees work remotely.
JumpCloud is committed to being remote-first across the world. We have team members in most U.S. states and in 14 countries.









