Site Reliability Engineer (SRE) On-Prem

Posted 15 Days Ago
Be an Early Applicant
Tel Aviv, ISR
In-Office
Mid level
Artificial Intelligence • Security • Cybersecurity
The Role
Own end-to-end reliability, automation, and deployment of Dream's platform at customer sites. Lead on-prem and hybrid Kubernetes deployments, manage Helm charts, build Ansible automation, operate Linux infrastructure, troubleshoot networking and storage, support GPU-enabled nodes and air-gapped installs, use Git and AWS (EC2/S3), and travel to customer sites for staging and deployments.
Summary Generated by Built In
Description

Every nation has data. Few can protect it. Fewer still can act on it.

Dream is the sovereign AI and national cyber-defense company for governments.

We help nations secure their most critical systems, connect fragmented information at a national scale, and turn their most sensitive data into decisions, all fully sovereign.

This is more than a job. It's a Dream job, where you'll work at a global scale alongside some of the best AI researchers, cyber operators, and government experts in the world.

The mission only works if the company behind it does. This role keeps Dream running at the scale our work demands. And our work demands a uniquely global scale.

The Dream Job

We are on an expedition to find an On-Premise Site Reliability Engineer (SRE)- someone who is passionate about building rock-solid, high-performance infrastructure and bringing order to complex environments. In this role, you will own the end-to-end reliability, automation, and deployment of DREAM’s platform across customer sites, working hands-on with cutting-edge AI, bare-metal, and hybrid cloud architectures.

You’ll collaborate closely with Product, R&D, and Architecture teams while serving as the ultimate technical authority for our customer deployments. From designing automated Ansible workflows and mastering Kubernetes to troubleshooting complex network topologies, you will eliminate toil, streamline cluster operations, and ensure every deployment is scalable, seamless, and mission-ready.

The Dream-Maker Responsibilities
  • Lead End-to-End On-Prem & Hybrid Deployments: Own the technical delivery and reliability of DREAM’s platform in close collaboration with Product, R&D, and customer technical teams. 
  • Architect, Execute & Improve K8s Deployments: Take a definitive hands-on role in deploying, configuring, operating, and continuously improving our platform using advanced, enterprise-grade Kubernetes architectures. 
  • Helm Chart Management: Design, modify, and manage Helm charts to package, version, and streamline complex application deployments across different environments. 
  • Drive Automation & Simplification: Design, implement, and maintain robust deployment automation using Ansible. You must have a passion for turning complex manual tasks into reliable, repeatable, single-click operations. 
  • Manage Infrastructure as Code: Utilize Git as the single source of truth to manage configurations, manifests, and automation playbooks, enforcing modern engineering best practices. 
  • Bridge On-Prem and Cloud: Leverage AWS resources (specifically EC2 and S3) for hybrid components, staging environments, or cloud-to-on-prem data flows. 
  • Technical Tier-3 Escalation: Serve as the ultimate technical authority for deployment, Linux networking, and Kubernetes orchestration issues. 
  • Continuous Improvement: Constantly refine our delivery pipelines, optimize bootstrap processes, and create rock-solid technical documentation. 
The Dream Skill Set
  • SRE / Delivery Mindset: 3–5 years of hands-on experience in enterprise infrastructure deployment, systems engineering, or an on-prem operational reliability role. 
  • Kubernetes & Helm Expert: Deep, production-grade experience with Kubernetes architecture, deployment, advanced troubleshooting, and CNI networking. Proven working experience creating, maintaining, and deploying applications using Helm charts. 
  • Ansible Mastery: Proven experience writing clean, scalable Ansible roles and playbooks for configuration management, automation, and infrastructure provisioning. 
  • Modern Workflows (Git & AWS): Solid working experience using Git for version control and collaborating on code/infrastructure. Practical experience provisioning and managing AWS resources (EC2 and S3). 
  • Core Systems & Linux: Strong Linux background (Ubuntu) with a deep understanding of system internals, containerized runtimes, and troubleshooting distributed applications. 
  • Solid Networking Knowledge: Hands-on experience with routing, firewalls, and switching topology (mainly Cisco) 
  • Storage Foundations: Working knowledge of storage protocols (iSCSI, SAN, local NVMe) and enterprise storage arrays (like DELL) interacting with Kubernetes Persistent Volumes. 
  • GPU & Accelerated Compute: Working knowledge of managing GPU-enabled Kubernetes nodes, including NVIDIA drivers/runtime and basic troubleshooting. 
  • Air-Gapped Deployments: Experience deploying and maintaining software in air-gapped or offline environments, including registry mirroring and artifact staging. 
  • Problem-Solver: Strong debugging and problem-solving skills in complex, distributed environments with an intense ownership and accountability mindset. 
  • Willingness to Travel: Ready to travel to customer sites for physical staging and deployments—at least 30%. 
  • Language: High-level English proficiency. 

Nice to have 

  • EU or any other additional citizenship. 
  • Experience with enterprise data components like MongoDB, PostgreSQL, Neo4j, or RabbitMQ. 
  • Working experience with project management tools such as Jira and Monday. 
  • Valid Israeli security clearance. 

Never Stop Dreaming...

If you think this role doesn't fully match your skills but are eager to grow and break glass ceilings, we’d love to hear from you! 

Skills Required

  • 3-5 years hands-on enterprise infrastructure deployment or on-prem SRE/delivery experience
  • Production-grade Kubernetes architecture, CNI networking, and advanced troubleshooting
  • Creating, maintaining, and deploying applications using Helm charts
  • Proven experience writing clean, scalable Ansible roles and playbooks
  • Git as single source of truth for configurations and automation
  • Practical experience provisioning and managing AWS resources (EC2 and S3)
  • Strong Linux (Ubuntu) systems internals and container runtime knowledge
  • Hands-on networking experience with routing, firewalls, and Cisco switching/topologies
  • Working knowledge of storage protocols (iSCSI, SAN, local NVMe) and enterprise storage arrays (e.g., DELL)
  • Experience managing GPU-enabled Kubernetes nodes, NVIDIA drivers/runtime, and basic GPU troubleshooting
  • Experience with air-gapped/offline deployments, registry mirroring and artifact staging
  • Strong debugging/problem-solving skills and ownership mindset in distributed environments
  • Willingness to travel to customer sites for physical staging and deployments (at least 30%)
  • High-level English proficiency
  • EU or additional citizenship
  • Experience with MongoDB, PostgreSQL, Neo4j, or RabbitMQ
  • Experience with Jira or Monday project management tools
  • Valid Israeli security clearance

Dream (dreamgroup.com) Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Dream (dreamgroup.com) and has not been reviewed or approved by Dream (dreamgroup.com).

  • Equity Value & Accessibility Funding and growth stage in a competitive AI/cyber market are framed as enabling competitive cash-and-equity offers, though the company has not disclosed specifics. Candidates are advised to confirm equity structure and eligibility directly given the lack of public detail.

Dream (dreamgroup.com) Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
328 Employees

What We Do

Dream is a pioneering AI cybersecurity company delivering revolutionary defense through artificial intelligence. Our proprietary AI platform creates a unified security system safeguarding assets against existing and emerging generative cyber threats. Dream's advanced AI automates discovery, calculates risks, performs real-time threat detection, and plans an automated response. With a core focus on the "unknowns," our AI transforms data into clear threat narratives and actionable defense strategies. Dream's AI cybersecurity platform represents a paradigm shift in cyber defense, employing a novel, multi-layered approach across all organizational networks in real-time. At the core of our solution is Dream's proprietary Cyber Language Model, a groundbreaking innovation that provides real-time, contextualized intelligence for comprehensive, actionable insights into any cyber-related query or threat scenario.

Similar Jobs

Remitly Logo Remitly

Field Operations Representative - Thai Speaker

eCommerce • Fintech • Payments • Software • Financial Services
In-Office
Tel Aviv, ISR
2800 Employees

Riskified Logo Riskified

Product Manager

Big Data • eCommerce • Fintech • Machine Learning • Payments • Software
Hybrid
Tel Aviv, ISR
680 Employees

Taboola Logo Taboola

Account Manager

AdTech • Big Data • Digital Media • Marketing Tech
Hybrid
Tel Aviv, ISR
1900 Employees

Taboola Logo Taboola

Senior HRBP - Global Growth Sales & Internal Ops

AdTech • Big Data • Digital Media • Marketing Tech
Hybrid
Tel Aviv, ISR
1900 Employees

Similar Companies Hiring

Legora Thumbnail
Artificial Intelligence • Legal Tech • Software
New York, New York
700 Employees
Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account