At Navan, "It's all about the user. All of them." We're passionate about providing a seamless one-stop experience for business travelers, no matter how they travel, where they stay, or where they're going. Navan is building cutting-edge solutions at the intersection of travel, expense, payments, and AI. As a leader in the AI for Travel domain, we are using intelligent, practical AI experiences to make business travel simpler, faster, and more reliable for travelers, travel managers, finance teams, and support teams.
We are seeking a Senior Site Reliability / DevOps Engineer to ensure the scalability, performance, and reliability of our user-facing generative AI features.
In this role, you will bridge the gap between traditional infrastructure and cutting-edge machine learning.
You will build and maintain the high-throughput, low-latency systems required to serve AI models directly to millions of users.
This position is based out of our new Tel Aviv office.
What You'll Do:
- Infrastructure Ownership: Design, build, and scale the infrastructure hosting our user-facing AI applications and inference engines.
- Performance Optimization: Optimize system latency, specifically targeting Time-to-First-Token (TTFT) and total round-trip time for user requests.
- GPU & Resource Orchestration: Manage and scale GPU clusters within Kubernetes to maximize utilization and minimize operational costs.
- Resiliency & Fallbacks: Build robust fallback systems, circuit breakers, and rate-limiting infrastructure to handle upstream LLM API failures and traffic spikes.
- Monitoring & Observability: Implement deep observability for AI workloads, tracking custom metrics like token usage, model drift, and GPU memory saturation.
What We're Looking For:
- SRE Fundamentals: 4+ years of experience in SRE, DevOps, or Production Engineering roles supporting high-traffic, user-facing applications.
- LLMOps / AI Infrastructure Expertise: Experience working with AI workloads (such as serving models using vLLM, Server TGI or working with cloud providers like Bedrock, OpenAI, etc.).
- Container Orchestration: Strong expertise in Kubernetes (EKS, GKE, or AKS) and infrastructure-as-code (Terraform).
- AI/ML Ecosystem: Hands-on experience with inference servers (e.g., vLLM, TGI) and vector databases (e.g., Pinecone, Milvus, Qdrant).
- Programming: Proficiency in Python and Go for automation, tooling, and backend optimization.
- Cloud Architecture: Deep experience managing cloud compute resources, specifically specialized GPU instances
- Models AI and Code:
Ability to build automation processes that not only update code versions, but also support testing and safe deployment of new models (Shadow Deployments, Canary releases for models) Product thinking and user orientation (User-Facing) - Advanced Observability:
Mastery of tools like OpenTelemetry, Prometheus, Datadog or Grafana, with the ability to trace agent-based systems and complex model calls. - Cost & Capacity Optimization:
Ability to manage the high costs of GPU/Inference in a productive architecture without compromising availability or performance. - Empathy for the end-user experience:
Understanding that every millisecond of latency or error in the stream directly impacts customer retention. - Preferred Qualifications
Experience building semantic caching layers to reduce LLM API costs.
Active contributor to open-source LLMOps or MLOps projects. (edited)
Navan uses AI-assisted Automated Employment Decision Tool (Metaview) to assist with evaluating resumes against job qualifications for this role. All final decisions are made by human recruiters and hiring managers.
Human oversight: Metaview does not automatically reject candidates or make final hiring decisions. Our recruiters and hiring managers review all outputs and make the final hiring decision regarding every application.
Your rights: If you prefer to have your application reviewed without AI assistance, you may request a human evaluation by entering your email here. Your decision to do so will not affect how your candidacy is evaluated.
Please refer to our Candidate Privacy Notice for more information about our processing of personal data, and your rights.
Skills Required
- 5+ years of experience as a Senior SRE, Infrastructure Software Engineer, Production Engineer, or DevOps Engineer.
- 3+ years operating production, 24x7 customer-facing systems.
- Hands-on experience delivering production infrastructure, platform tooling, and automation used by engineering teams.
- Strong software engineering skills in Python, Go, Java, or a similar language, with production-quality code, tests, and documentation.
- Experience with cloud infrastructure, container orchestration, Linux systems, networking, CI/CD, and IaC such as Terraform or CloudFormation.
- Experience building, tuning, and automating observability systems (Grafana, Prometheus, New Relic, Datadog, Splunk, or similar).
- Familiarity with SLOs, incident response, on-call practices, root cause analysis, and blameless postmortems.
- Practical experience with or strong interest in AI solutions, providers, agents, and AI APIs.
- Ability to troubleshoot AI tools and provider/API issues including rate limits, auth, quotas, latency, and degraded responses.
- Excellent communication skills and ability to work with stakeholders across teams.
Navan Compensation & Benefits Highlights
-
Healthcare Strength — Medical, dental, and vision coverage for employees and dependents are highlighted, with mental health resources such as Headspace included. This breadth of core health benefits signals strong baseline coverage.
-
Leave & Time Off Breadth — Flexible vacation in the U.S., a company-wide year-end quiet week, and a U.K. policy listing five weeks of PTO point to generous time-away options. These elements indicate meaningful support for rest and recharge across regions.
-
Parental & Family Support — Paid parental leave is specified as 16 weeks for the birthing parent and 10 weeks for the non-birthing parent. Clear, above-basic leave durations suggest solid family support.
Navan Insights
What We Do
Navan (Nasdaq: NAVN) is the leading all-in-one business travel, payments, and expense management platform that makes travel easy for frequent travelers. From finding flights and hotels to automating expense reconciliation, with 24/7 support along the way, Navan delivers an intuitive experience travelers love and finance teams rely on. See how Navan customers benefit and learn more at navan.com.
Why Work With Us
At Navan, we’re never satisfied with the status quo, and we know breakthrough ideas come from diverse perspectives. We are committed to cultivating a workplace that reflects the diversity of the customers we serve while fostering leadership and innovation.
Gallery
Navan Offices
Hybrid Workspace
Employees engage in a combination of remote and on-site work.
In-person connections is the foundation of Navan, the connections forged through face-to-face interactions improve company culture and what we can achieve together. We operate on a hybrid working model, which we define as four days a week in-office.






















.png)