Senior Site Reliability Engineer (Arlington, VA) - Secret Clearance Required - Relocation Provided

Posted Yesterday
Be an Early Applicant
Hiring Remotely in United States
Remote
180K-220K Annually
Senior level
Software • Defense
Building the future of the military staff.
The Role
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Summary Generated by Built In
Consequential Work. Dedicated People.
About Onebrief

Onebrief builds collaboration and AI-powered workflow software for military planning and operational coordination.

Today, many critical planning workflows still rely on fragmented systems, static documents, and disconnected tools that make collaboration and decision-making unnecessarily difficult. Onebrief brings modern software, AI, and real-time collaboration into those environments, helping teams operate with greater clarity, coordination, and adaptability in situations where decisions carry real-world consequences.

We are a distributed team of builders from military, operational, and technology backgrounds who care deeply about improving how important work gets done. Some team members work remotely, while others work directly alongside customers in operational environments around the world.

Founded in 2019, Onebrief is backed by leading investors including General Catalyst, Battery Ventures, Insight Partners, Sapphire Ventures, and Human Capital. Valued at more than $2 billion, we continue to invest in product innovation, AI capabilities, and team growth.

Security Clearance, Location, and Onsite Notice:

This role requires regularly working on-site at customer locations in Arlington, VA.

If you are not currently within commuting distance, you must be willing to relocate (note that Onebrief will provide relocation assistance).

Active Secret Clearance required.

About The Role

We're hiring a Site Reliability Engineer to join our Infrastructure & Security team. You'll work closely with product engineers, fellow SREs, security, and customer success.

This is an SRE role for someone who's comfortable in application code. Much of the reliability and performance work happens in the codebase (primarily TypeScript), so you'll fix problems at the source rather than working around them in the infrastructure. You'll be a first line of support for our mission-critical deployments across on-prem DoD and AWS environments, and what you learn in the field will feed directly back into the product.

You'll ship code that makes Onebrief more stable, faster, and easier to deploy and operate. The work sits at the seam between engineering and operations, and it's weighted toward engineering.

About You

You treat reliability as a feature, not an afterthought, and you'd rather fix a problem in the code than route around it. You understand the full software development lifecycle (design, review, testing, release) and you know where reliability fits into each step.

You're comfortable reading and writing application code, and you're just as happy dropping into a kubectl shell to triage a production issue. You turn failure modes into guardrails, and you think monitoring, alerting, and clear runbooks are part of building software, not extra credit.

You mentor others and push a culture of blameless postmortems. You work naturally with product and platform teams, helping them move fast without breaking things by giving them the tools, tests, and observability that make quick recovery real.

What You'll Do

You'll help make our production application reliable, scalable, and secure by improving the software itself, not just the systems it runs on. Day to day that looks like:

  • Improving the application: Work directly in the codebase (primarily TypeScript) to fix reliability and performance problems at the source. You'll partner with product engineers on design decisions, review code with reliability and security in mind, and treat "make the app better" as a first-class part of the job rather than something you hand off.

  • Building observability that developers actually use: Design and run our monitoring, logging, and alerting (Prometheus, Loki, Alloy, Grafana). The goal is alerts and dashboards tied to real application behavior, so teams catch issues before users do.

  • Owning reliability targets: Define and measure SLIs and SLOs, wire up alerting that feeds them, and be the person who can say what "reliable" means for our systems and prove it with data.

  • Leading incident response: Act as incident responder, and incident commander when needed. Run blameless post-mortems (AARs) that find the actual root cause and turn it into a code or process fix so it doesn't happen again.

  • Automating away toil: Spot the repetitive operational work and write software to kill it. Share what works with other teams, including those running in air-gapped environments, and help them get production-ready.

What We Look For
  • An active Secret clearance

  • 5+ years in software engineering, SRE, or a related role, with real time spent writing and shipping application code

  • Strong TypeScript (or comparable modern language experience with willingness to work primarily in TypeScript)

  • Solid grasp of the full SDLC: design, code review, testing, release, and how reliability fits into each stage

  • Experience with incident response, root cause analysis, and turning findings into lasting fixes

  • A collaborator who works well across product, platform, and DevOps teams and shares context openly

Technical expertise
  • Application development in TypeScript (Node and/or a modern front-end framework)

  • CI/CD: building and maintaining pipelines (GitHub Actions, GitLab CI/CD, Jenkins)

  • Testing and quality practices as part of the delivery process

  • Comfort with at least one of Python, Go, or Bash for tooling and automation

  • Working knowledge of containers and Kubernetes (enough to debug and deploy, not necessarily to stand up clusters from scratch)

  • Networking fundamentals and secure configuration basics

Bonus points (nice to have)
  • Observability: Grafana stack, ELK, or Datadog

  • Infrastructure as Code (Terraform, Ansible) and cloud experience (AWS or AWS GovCloud)

  • Kubernetes cluster design and operations

  • Designing meaningful SLIs/SLOs with error budgets for distributed systems

  • GitOps practices and toolchains

  • DoD environments and compliance frameworks (RMF, STIGs, ICD 503)

  • Service mesh (Istio, Linkerd)

  • On-prem virtualization (VMware, Proxmox, Nutanix, Hyper-V)

  • Relevant certs (AWS DevOps Engineer, CKA/CKAD)


Notice to Third Party Recruitment Agencies

Please note that Onebrief does not accept unsolicited resumes from recruiters or employment agencies. In the absence of an executed Recruitment Services Agreement, there will be no obligation to any referral compensation or recruiter fee. In the event a recruiter or agency submits a resume or candidate without an agreement Onebrief explicitly reserves the right to pursue and hire those candidate(s) without any financial obligation to the recruiter or agency. Any unsolicited resumes, including those submitted to hiring managers, shall be deemed the property of Onebrief.

Skills Required

  • Active Secret clearance
  • 5+ years in software engineering, SRE, or related role with time writing and shipping application code
  • Strong TypeScript (willingness to work primarily in TypeScript)
  • Application development experience (Node and/or modern front-end frameworks)
  • Experience with incident response, root cause analysis, and implementing lasting fixes
  • Experience designing and running monitoring, logging, and alerting (Prometheus, Loki, Alloy, Grafana)
  • Experience defining and measuring SLIs/SLOs and owning reliability targets
  • CI/CD experience (GitHub Actions, GitLab CI/CD, Jenkins)
  • Testing and quality practices integrated into delivery process
  • Comfort with at least one of Python, Go, or Bash for tooling and automation
  • Working knowledge of containers and Kubernetes and ability to debug/deploy (kubectl)
  • Networking fundamentals and secure configuration basics
  • Ability to work regularly on-site at customer locations in Arlington, VA and willingness to relocate (relocation assistance provided)
  • Collaborative mindset, mentoring, and running blameless postmortems
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
350 Employees
Year Founded: 2019

What We Do

Before Onebrief, military planning and collaboration was slow, inefficient, and resource-intensive. Building slides with no version control as partners collaborated would have staffs spend weeks or months on a single product or document. With Onebrief, these workflows are now simple and collaboration between large commands is efficient. Staff optimization is the key to building a more resilient, more effective military. Today Onebrief users report at least 2x time savings - and growing. Onebrief is a first of its kind software for the military. While many others have tried to build a solution for this problem, Onebrief’s “card” structure for reusing data and enabling real time updates is what makes this possible. Core features and attributes that make this platform powerful include: - Global Collaboration - Real-Time Updates - AI Automation - Interoperability + Integrations - Deployable across Secret and Top Secret Networks Mission Driven Onebrief is composed of professionals from backgrounds of all kinds - spanning veterans across forces and organizations, and technologists from leading-edge software giants. Onebrief is more than just a software platform; it's a mission-driven company dedicated to improving the efficiency and effectiveness of military planning. By joining the team, you'll contribute to solutions that directly support national security and the work of service members. Your work directly addresses critical challenges that military planners and operators face daily. Every line of code and every design decision contributes to real-world outcomes. The software was designed and built by a team of experienced planners - lending a nuanced perspective on the challenges our partners face. Our team embeds alongside users - from the Pentagon to the Indo-Pacific - to build a platform that meets their unique needs. Rapid, Strategic Growth Our users love the platform and growth is scaling, most recently reporting operational usage growth at a 19,600% annualized rate. Stronger utilization is underway and we’re at an exciting period of advancement. As a rapidly growing organization, you'll directly influence its direction and long-term success. Over the past year we’ve seen exciting growth metrics: First, our headcount has grown 150% YoY to keep pace with our product advancement and customer growth. Our funding has skyrocketed, most recently raising our Series C, led by top-tier venture investors who have deep expertise in defense tech.

Why Work With Us

Impactful Transformation At Onebrief, we believe optimizing the military staff is the most impactful thing - on a per-dollar basis - in defense tech right now. This has the potential to save the department of defense billions of dollars and save users countless hours. It’s a longstanding problem that we’re uniquely positioned to solve.

Gallery

Gallery
Gallery
Gallery

Onebrief Offices

Remote Workspace

Employees work remotely.

We’re a fully remote organization - and believe it makes us a more powerful team. We bring together incredible professionals without the constraints of time zones or personal circumstances.

Typical time on-site:
United States

Similar Jobs

Remote
United States
350 Employees
181K-220K Annually
Remote
United States
350 Employees
205K-255K Annually
Remote
United States
350 Employees
240K-290K Annually
Remote
United States
350 Employees
126K-154K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account