Site Reliability Engineer

Posted 3 Hours Ago
Hiring Remotely in San Diego, CA, USA
In-Office or Remote
Senior level
Artificial Intelligence • Machine Learning • Software • Defense
Vannevar builds AI systems to solve America's hardest national security challenges.
The Role
Own platform reliability, observability, incident response, scaling, capacity planning, and deployment automation. Monitor system health, debug and resolve incidents, build logging and monitoring tools, improve CI/CD pipelines, develop self-service automation, and strengthen high-availability delivery systems. The role requires clear incident communication, post-incident learning, and collaboration across engineering teams in secure, high-side environments.
Summary Generated by Built In

Vannevar is a defense technology company building AI to deter our adversaries. In the 21st century, conflict moves at algorithmic speed and foresight equals firepower. Our agentic AI is purpose-built to compete with China—from cross-Strait conflict to gray zone coercion. Trained on the most mission-relevant datasets in defense, our technology models adversary behavior, simulates campaigns, and recommends the best course of action to decision makers. Our AI systems are some of the most trusted in the industry and actively used on the front lines of the Indo-Pacific to keep the peace and save lives.

Exceptional technology starts with exceptional people. Vannevar is a small agile team combining world-class engineers with veteran strategists who bring deep expertise in defense and tradecraft. We’re building a company defined by mission impact, user empathy, and disciplined growth. In just three years, we grew from $3M to $80M in ARR, achieved early profitability, and reached unicorn status—proving that disruption doesn’t require an ego, and staying power doesn’t mean standing still.

About the role

We are looking for an Site Reliability Engineer to own the reliability, health, and deployment automation of the platform at Vannevar Labs. In this role you'll be the person watching the system's pulse — monitoring dashboards, catching health issues before they become incidents, and owning the debugging process from first alert to resolution. Your decisions today will have a large impact on the company's future.
We believe that simple systems are easier to understand, maintain, and scale. You will be making trade-offs as you work to ensure that our systems are prepared to operate reliably in high-side environments at scale. A strong sense of judgment matters here: knowing when to dig deeper into a problem yourself and when to pull in the right people to escalate. Clear, calm communication — during an incident and in day-to-day work — is a must.

What you'll do
  • Monitor dashboards and system telemetry to detect health issues, performance degradation, and reliability risks — often before anyone else notices them.
  • Own the debugging and incident response process end to end, exercising good judgment about when to investigate more deeply and when to escalate.
  • Build logging, monitoring, and observability tooling to visualize the state of the platform and continuously mature our SRE practices.
  • Develop, maintain, and be responsible for overall platform health, scaling, and capacity planning.
  • Understand and help improve the deployment process, and automate build & deployment pipelines.
  • Identify bottlenecks in engineering workflows and drive improvements that make the whole team faster and more reliable.
  • Develop self-service tools and automation to improve engineering efficiency.
  • Play a critical part in implementing a secure, robust, high-availability delivery pipeline.
  • Communicate system status, trade-offs, and post-incident learnings clearly with teammates and stakeholders.
Qualifications
  • 5+ years of experience in SRE, DevOps, or software engineering.
  • Hands-on experience monitoring production systems and responding to incidents — comfortable owning a debugging process and making the call on when to dig in versus escalate.
  • Excellent communication skills, especially the ability to stay clear and organized while troubleshooting live issues.
  • Experience with the PLG stack, Datadog, or other enterprise monitoring/observability tools.
  • Experience participating in an on-call rotation and running or contributing to post-mortems.
  • Knowledge of AWS cloud technologies.
  • Familiarity with infrastructure-as-code technologies such as Terraform and Pulumi.
  • Experience with Python, Bash, or other scripting languages.
  • Experience working in an agile scrum environment, with the ability to work independently.
  • Able to quickly learn new and existing technologies.
  • Strong attention to detail and analytical capabilities.
  • Willingness and ability to work on-site in San Diego, CA.
  • U.S. Citizenship status is required, as this position requires the ability to access U.S.-only data systems and export-controlled data.
  • TS/SCI Clearance required.
Additional Qualifications (Nice to haves)
  • Experience defining and tracking SLOs/SLIs and error budgets.
  • Experience crafting CI/CD processes and automation.
  • Proficient with containerization technologies like Docker.
  • Experience working in AWS GovCloud.
  • Experience with modern web services architectures.
  • Experience with relational database systems, including SQL and relational design.
  • Experience working with Elasticsearch/OpenSearch.
  • Strong collaboration and negotiation skills, with the ability to work on cross-functional projects with internal partner engineering teams.
What we offer
We’re proud to offer competitive benefits that support our employees. Some key highlights of our benefits package include:
  • Health, dental, and vision insurance
  • 100% remote first culture. You can work from anywhere in the US and all full time employees have WeWork access
  • Unlimited PTO including competitive vacation and holiday schedules
  • Lifestyle stipends - Monthly mental health, wellness & fitness stipend, in-home office setup stipend and family planning assistance
  • Salary top-up during military reserve duty
  • Fully paid parental leave
  • Child and pet care reimbursement during travel

Vannevar is an equal opportunity employer, and qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender perception or identity, national origin, age, marital status, protected veteran status, or disability status.
 
We encourage candidates from all backgrounds to apply, even if you don't feel like you're a perfect fit. If you're passionate about contributing to our mission, we'd love to hear from you!
 
IMPORTANT NOTICE
We are committed to protecting the privacy of all applicants. Official emails from the company will come from an @vannevarlabs.com domain. Under no circumstances will a legitimate representative from our company contact you to request passwords, financial information, or other sensitive personal data. Please be vigilant of potential scams.

Skills Required

  • 5+ years of experience in SRE, DevOps, or software engineering
  • Experience monitoring production systems and responding to incidents
  • Excellent communication skills during troubleshooting and live incidents
  • Experience with PLG stack, Datadog, or enterprise monitoring and observability tools
  • Experience participating in an on-call rotation and contributing to post-mortems
  • Knowledge of AWS cloud technologies
  • Familiarity with infrastructure-as-code technologies such as Terraform and Pulumi
  • Experience with Python, Bash, or other scripting languages
  • Experience working in an Agile Scrum environment
  • Ability to work independently and learn new technologies quickly
  • Strong attention to detail and analytical capabilities
  • Willingness and ability to work on-site in San Diego, California
  • U.S. citizenship
  • TS/SCI security clearance
  • Experience defining and tracking SLOs, SLIs, and error budgets
  • Experience crafting CI/CD processes and automation
  • Proficiency with Docker or other containerization technologies
  • Experience working in AWS GovCloud
  • Experience with modern web services architectures
  • Experience with relational databases, SQL, and relational design
  • Experience with Elasticsearch or OpenSearch
  • Collaboration and negotiation skills for cross-functional engineering projects

What the Team is Saying

Kainoa
Vince
Katie
Ann
Scott
Harrison

Vannevar Compensation & Benefits Highlights

  • Healthcare Strength Healthcare is considered comprehensive, with top-tier medical, dental, and vision coverage highlighted alongside mental-health support and wellness stipends. These elements are consistently presented as core pillars of the package.
  • Leave & Time Off Breadth Time off is framed as generous, featuring unlimited PTO, shared downtime aligned to federal holidays, and a company-wide year-end break. The approach signals an expectation that employees actually disconnect and use their time.
  • Parental & Family Support Parental leave is described as fully paid for 12 weeks for all parents, with an additional 4 weeks for new mothers and flexibility to take leave in blocks over a year. Family-planning assistance, monthly caregiving stipends, and travel-related child/pet care reimbursement extend support beyond leave.

Vannevar Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Arlington, Virginia
225 Employees
Year Founded: 2019

What We Do

Vannevar combines Silicon Valley innovation with former special operators to build AI for real-world national security outcomes. Since our founding in 2019, Vannevar has deployed its technology across 125 missions, safeguarding our service members and improving our country's deterrence. We believe our military service members and intelligence officers deserve access to the best technology American innovation can offer, and we are mobilizing frontier technology for this mission. Our founders have 30 years of combined experience across AI labs, national security strategy, intelligence agencies and special operations. We're supported by a number of leading VCs, including Felicis, DFJ Growth, Point72, Costanoa, and General Catalyst. Vannevar is one of the most capital efficient defense tech companies, reaching early profitability and becoming a defense tech unicorn in 2024.

Why Work With Us

We pride ourselves on being a team of passionate, low-ego professionals who share a commitment to engineering excellence and a genuine concern for the people and AI agents we deliver. Vannevar is a company that truly cares about its people and one of our core principles is "put your wellbeing first". ​

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery

Vannevar Offices

Remote Workspace

Employees work remotely.

We are distributed across the map, but united by the mission. We go where the mission takes us and deliver where others can’t. Honolulu, Seattle, San Francisco, Washington D.C., New York. You decide what's next.

Typical time on-site:
Company Office Image
Arlington, Virginia
Company Office Image
Honolulu, Hawaii
Company Office Image
New York, NY
Company Office Image
San Diego, California
Company Office Image
San Francisco, California
Company Office Image
Seattle, Washington
Learn more

Similar Jobs

Vannevar Logo Vannevar

Back-end Engineer

Artificial Intelligence • Machine Learning • Software • Defense
Remote
USA
225 Employees
150K-215K Annually

Vannevar Logo Vannevar

Technical Recruiter

Artificial Intelligence • Machine Learning • Software • Defense
Remote
USA
225 Employees

Vannevar Logo Vannevar

Government Contracts Manager

Artificial Intelligence • Machine Learning • Software • Defense
Remote
USA
225 Employees

Vannevar Logo Vannevar

Application Security Engineer

Artificial Intelligence • Machine Learning • Software • Defense
Remote
USA
225 Employees
160K-210K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account