Senior Site Reliability / DevOps Engineer– AI Products

Reposted 26 Days Ago
Easy Apply
Be an Early Applicant
Tel Aviv, ISR
Hybrid
Senior level
Fintech • Information Technology • Payments • Productivity • Software • Travel • Automation
Travel & expense made easy.
The Role
Build and operate reliable production platforms and AI-integrated services. Improve observability, automate toil, troubleshoot AI provider/API issues, implement IaC, and support on-call incident response tied to SLOs.
Summary Generated by Built In

At Navan, "It's all about the user. All of them." We're passionate about providing a seamless one-stop experience for business travelers, no matter how they travel, where they stay, or where they're going. Navan is building cutting-edge solutions at the intersection of travel, expense, payments, and AI. As a leader in the AI for Travel domain, we are using intelligent, practical AI experiences to make business travel simpler, faster, and more reliable for travelers, travel managers, finance teams, and support teams.

We are seeking a Senior Site Reliability / DevOps Engineer to ensure the scalability, performance, and reliability of our user-facing generative AI features.
In this role, you will bridge the gap between traditional infrastructure and cutting-edge machine learning.
You will build and maintain the high-throughput, low-latency systems required to serve AI models directly to millions of users.
This position is based out of our new Tel Aviv office.


What You'll Do:

  •  Infrastructure Ownership: Design, build, and scale the infrastructure hosting our user-facing AI applications and inference engines.
  • Performance Optimization: Optimize system latency, specifically targeting Time-to-First-Token (TTFT) and total round-trip time for user requests.
  • GPU & Resource Orchestration: Manage and scale GPU clusters within Kubernetes to maximize utilization and minimize operational costs.
  • Resiliency & Fallbacks: Build robust fallback systems, circuit breakers, and rate-limiting infrastructure to handle upstream LLM API failures and traffic spikes.
  • Monitoring & Observability: Implement deep observability for AI workloads, tracking custom metrics like token usage, model drift, and GPU memory saturation.

What We're Looking For:

  • SRE Fundamentals: 4+ years of experience in SRE, DevOps, or Production Engineering roles supporting high-traffic, user-facing applications.
  • LLMOps / AI Infrastructure Expertise: Experience working with AI workloads (such as serving models using vLLM, Server TGI or working with cloud providers like Bedrock, OpenAI, etc.).
  • Container Orchestration: Strong expertise in Kubernetes (EKS, GKE, or AKS) and infrastructure-as-code (Terraform).
  • AI/ML Ecosystem: Hands-on experience with inference servers (e.g., vLLM, TGI) and vector databases (e.g., Pinecone, Milvus, Qdrant).
  • Programming: Proficiency in Python and Go for automation, tooling, and backend optimization.
  • Cloud Architecture: Deep experience managing cloud compute resources, specifically specialized GPU instances
  • Models AI and Code:
    Ability to build automation processes that not only update code versions, but also support testing and safe deployment of new models (Shadow Deployments, Canary releases for models) Product thinking and user orientation (User-Facing)
  • Advanced Observability:
    Mastery of tools like OpenTelemetry, Prometheus, Datadog or Grafana, with the ability to trace agent-based systems and complex model calls.
  • Cost & Capacity Optimization:
    Ability to manage the high costs of GPU/Inference in a productive architecture without compromising availability or performance.
  • Empathy for the end-user experience:
    Understanding that every millisecond of latency or error in the stream directly impacts customer retention.
  • Preferred Qualifications
    Experience building semantic caching layers to reduce LLM API costs.
    Active contributor to open-source LLMOps or MLOps projects. (edited) 

Navan uses AI-assisted Automated Employment Decision Tool (Metaview) to assist with evaluating resumes against job qualifications for this role. All final decisions are made by human recruiters and hiring managers. 

Human oversight: Metaview does not automatically reject candidates or make final hiring decisions. Our recruiters and hiring managers review all outputs and make the final hiring decision regarding every application. 

  • Your rights: If you prefer to have your application reviewed without AI assistance, you may request a human evaluation by entering your email here. Your decision to do so will not affect how your candidacy is evaluated. 

Please refer to our Candidate Privacy Notice for more information about our processing of personal data, and your rights.

Skills Required

  • 5+ years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer.
  • 3+ years operating production, 24x7 customer-facing systems.
  • Hands-on experience delivering production infrastructure, platform tooling, and automation used by engineering teams.
  • Strong software engineering skills in Python, Go, Java, or a similar language, with production-quality code, tests, and documentation.
  • Experience with cloud infrastructure, container orchestration, Linux systems, networking, CI/CD, and IaC such as Terraform or CloudFormation.
  • Experience building, tuning, and automating observability systems (Grafana, Prometheus, New Relic, Datadog, Splunk, or similar).
  • Familiarity with SLOs, incident response, on-call practices, root cause analysis, and blameless postmortems.
  • Practical experience with or strong interest in AI solutions, providers, agents, and AI APIs.
  • Ability to troubleshoot AI tools and provider/API issues including rate limits, auth, quotas, latency, and degraded responses.
  • Excellent communication skills and ability to work with stakeholders across teams.

What the Team is Saying

Brian Guimond
Adamas Victória Cavalcante Robitz
Bastian Martino
Charlotte Delafosse
Daniella Schuh
Alice Rao-Wyckoff
Mily O Loughlin
Anna
Roshni
Henry Statfeld
Jose Soares

Navan Compensation & Benefits Highlights

  • Healthcare Strength Medical, dental, and vision coverage for employees and dependents are highlighted, with mental health resources such as Headspace included. This breadth of core health benefits signals strong baseline coverage.
  • Leave & Time Off Breadth Flexible vacation in the U.S., a company-wide year-end quiet week, and a U.K. policy listing five weeks of PTO point to generous time-away options. These elements indicate meaningful support for rest and recharge across regions.
  • Parental & Family Support Paid parental leave is specified as 16 weeks for the birthing parent and 10 weeks for the non-birthing parent. Clear, above-basic leave durations suggest solid family support.

Navan Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Palo Alto, CA
3,300 Employees
Year Founded: 2015

What We Do

Navan (Nasdaq: NAVN) is the leading all-in-one business travel, payments, and expense management platform that makes travel easy for frequent travelers. From finding flights and hotels to automating expense reconciliation, with 24/7 support along the way, Navan delivers an intuitive experience travelers love and finance teams rely on. See how Navan customers benefit and learn more at navan.com.

Why Work With Us

At Navan, we’re never satisfied with the status quo, and we know breakthrough ideas come from diverse perspectives. We are committed to cultivating a workplace that reflects the diversity of the customers we serve while fostering leadership and innovation.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

Navan Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

In-person connections is the foundation of Navan, the connections forged through face-to-face interactions improve company culture and what we can achieve together. We operate on a hybrid working model, which we define as four days a week in-office.

Typical time on-site: 4 days a week
HQPalo Alto, CA
Austin, TX
Bengaluru, IN
Berlin, DE
Boston, MA
Dallas, TX
Gurugram, IN
Lisbon, PT
London, GB
New Delhi, Delhi
New York, NY
Paris, FR
San Francisco, CA
Singapore
Sydney, AU
Tel Aviv-Yafo, IL
Learn more

Similar Jobs

Navan Logo Navan

Software Engineer

Fintech • Information Technology • Payments • Productivity • Software • Travel • Automation
Easy Apply
Hybrid
Tel Aviv, ISR
3300 Employees

Navan Logo Navan

Director, Engineering

Fintech • Information Technology • Payments • Productivity • Software • Travel • Automation
Easy Apply
Hybrid
Tel Aviv, ISR
3300 Employees

Navan Logo Navan

Product Manager

Fintech • Information Technology • Payments • Productivity • Software • Travel • Automation
Easy Apply
Hybrid
Tel Aviv, ISR
3300 Employees

Navan Logo Navan

Travel Experience Agent

Fintech • Information Technology • Payments • Productivity • Software • Travel • Automation
Easy Apply
Hybrid
Tel Aviv, ISR
3300 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account