What You’ll Be Doing:
Architect, scale, and continuously improve the reliability, availability, and performance of JumpCloud’s multi-region microservices, APIs, and authentication infrastructure (AWS/GCP).
Architect, build, and maintain Disaster Recovery (DR) process, multi-region failover automation, and business continuity strategies to ensure rapid recovery against strict RTO and RPO objectives.
Lead the design and enforcement of SLIs, SLOs, and Error Budget frameworks across multi-disciplinary engineering teams.
Drive end-to-end observability strategy using Datadog, implementing actionable Golden Signals monitoring to drastically reduce MTTD/MTTR and eliminate alert fatigue.
Lead on-call escalation, major incident management, and drive strict adherence to 99.99% availability SLAs.
Facilitate blameless post-incident reviews, executing systemic root-cause remediations to prevent recurring failure modes.
Architect, manage, and scale production Kubernetes (EKS) clusters, implementing advanced GitOps workflows (Argo CD, Kargo) and deployment patterns.
Design and maintain modular, enterprise-grade Infrastructure-as-Code using Terraform across multi-account, multi-region cloud environments.
Design, build, and maintain interactive FinOps and cost-optimization dashboards to provide engineering and leadership teams with actionable insights into multi-cloud spend, unit economics, and resource utilization.
Eliminate complex operational toil by writing production-grade Python or Go tooling, platform automation, and custom integrations.
Champion AI-assisted software development workflows (Cursor, Claude Code, GitHub Copilot) to accelerate automation, runbook creation, and incident triage across the team.
Author operational runbooks, architecture decision records, and mentor mid-level/junior engineers to raise the overall technical bar.
We’re Looking For:
8+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering operating 24/7 mission-critical, highly available distributed systems.
Bachelor's degree in Computer Science, Software Engineering, or equivalent technical discipline.
Strong Python/Go Capabilities: Advanced software engineering skill set for writing internal SRE platforms, tools, and API integrations.
Deep Kubernetes Expertise: Hands-on experience with production EKS/GKE cluster lifecycles, ingress/egress, networking, RBAC, and GitOps tooling (Argo CD).
Advanced IaC & AWS/GCP: Deep Terraform proficiency (module architecture, state management refactoring) across complex multi-account AWS environments (IAM, VPCs, Transit Gateway, ALB/NLB, Route53).
FinOps & Cost Optimization Leadership: Demonstrated experience driving cloud cost-efficiency strategies, resource right-sizing, cost-allocation tagging, workload optimization, and building FinOps dashboards to embed financial accountability into engineering workflows.
Disaster Recovery & High Availability: Proven background in designing and testing multi-region Disaster Recovery architectures, automating failover systems, and monitoring recovery health via DR dashboards.
Observability & Reliability Architecture: Track record of defining SLI/SLOs, managing PagerDuty schedules, and optimizing production observability platforms.
Experience designing and operating enterprise service meshes (Istio, Linkerd, or similar) and production ingress/proxy systems (HAProxy, NGINX, or similar).
Technical Mentorship: Demonstrated ability to lead technical discussions, write architectural design docs/RFCs, and mentor engineering peers.
Strong problem-solving, communication, and collaboration skills with a passion for solving complex distributed systems challenges at scale.
A strong team player who helps us live by our core values: building connections, thinking big, and getting 1% better every day.
Preferred Qualifications:
Basic understanding of chaos engineering principles or testing resilience in staging/production.
Experience with secrets management architectures (Vault, AWS Secrets Manager, External Secrets Operator, Cert-Manager).
Background in DevSecOps practices, service meshes (Istio), and automated vulnerability remediation within cloud infrastructure code.
Background supporting identity services, IAM, enterprise directory platforms, or security-focused SaaS solutions.
Skills Required
- 8+ years of professional software engineering experience in SRE, DevOps, or Platform Engineering
- Bachelor’s degree in Computer Science, Software Engineering, or an equivalent technical discipline
- Advanced Python or Go software engineering capabilities
- Hands-on production Kubernetes expertise, including EKS or GKE cluster lifecycles, networking, RBAC, and GitOps tooling
- Advanced Terraform and AWS/GCP experience across complex multi-account environments
- Experience leading FinOps and cloud cost-optimization strategies
- Experience designing and testing multi-region disaster recovery and high-availability architectures
- Experience defining SLI/SLO frameworks, managing PagerDuty schedules, and optimizing observability platforms
- Experience designing and operating enterprise service meshes and production ingress or proxy systems
- Technical mentorship, architecture documentation, and technical discussion leadership experience
- Strong problem-solving, communication, collaboration, and English-language skills
- Understanding of chaos engineering or resilience testing
- Experience with secrets management architectures
- DevSecOps experience and automated vulnerability remediation
- Experience supporting identity services, IAM, enterprise directories, or security-focused SaaS solutions
JumpCloud Compensation & Benefits Highlights
-
Healthcare Strength — Multiple medical plan options (including an HSA with employer contribution) plus dental, vision, EAP, disability, and life insurance are explicitly offered. Feedback suggests the overall healthcare package is a core strength.
-
Retirement Support — A 401(k) with company match is clearly provided and repeatedly highlighted in company materials and employee commentary. Feedback suggests retirement benefits contribute meaningfully to total rewards.
-
Leave & Time Off Breadth — Flexible/unlimited PTO, paid sick leave, and parental leave are standard, and feedback suggests the time‑off policy is genuinely usable and encouraged. This breadth of leave is consistently noted alongside a remote‑first setup.
JumpCloud Insights
What We Do
JumpCloud’s mission is to Make Work Happen®, providing simple, secure access to an organization’s technology resources from any device, or any location. The JumpCloud Open Directory Platform gives IT, security operations, and DevOps a single, cloud-based solution to control and manage employee identities and their devices, and apply conditional access controls based on Zero Trust principals. Since launching in 2012, our global user base has grown to more than 150,000 organizations, with more than 5,000 paying customers including Cars.com, GoFundMe, Grab, ClassPass, Uplight and Peloton. JumpCloud has raised over $400M from world-class investors including Sapphire Ventures, General Atlantic, Sands Capital, Atlassian, and CrowdStrike. Our teams are growing fast, too, and we're looking for talent across engineering, sales, customer success, marketing, product management, and more. Join our team of dedicated, passionate, and creative people who are eager to change the IT industry forever. We live by our core values which are: Build Connections Think Big 1% Better Every Day
Why Work With Us
We offer an incredible opportunity to see your impact. Each team member gets an up close personal view and education into building a fast growing startup. We are transparent about what we are doing, how we are doing it, and the decisions that we are making. There is opportunity to progress and flexibility to find unique approaches to our business
Gallery
JumpCloud Offices
Remote Workspace
Employees work remotely.
JumpCloud is committed to being remote-first across the world. We have team members in most U.S. states and in 14 countries.









