What will your job look like?
- Build and maintain infrastructure for large‑scale AI and HPC workloads across on‑prem and cloud environments
- Operate and enhance our multi‑cloud, multi‑cluster scheduling platform
- Troubleshoot complex issues across the stack: from kernel-level tuning and drivers to networking, storage, and distributed system bottlenecks.
- Ensure the reliability of critical platform services: queuing systems, time-series databases, and logging pipelines
- Develop deeply integrated automation and tooling
- Collaborate with ML engineers and IT engineers to optimize hardware utilization for data intensive workloads
- Drive best practices in system design, observability, and infrastructure-as-code
All you need is:
- 10+ years of hands‑on experience in SRE, Linux Administration, or Systems Engineering
- Expert-level Linux knowledge: Deep understanding of system internals, debugging, performance tuning, and the ability to solve failures where hardware meets software.
- Kubernetes Expertise: Proven experience managing K8s at scale (both managed EKS and bare-metal deployments)
- Distributed Systems Mastery: Hands-on experience debugging and maintaining:Queuing Systems: RabbitMQ or similar
- Metrics/Observability Stacks: Prometheus, Thanos, and Grafana, or similar
- Logging: Elasticsearch or similar
- Relational Databases: PostgreSQL, or similar
- Infrastructure-as-Code: Proficiency with Terraform, Helm, and configuration management
- Networking & Scripting: Strong fundamentals in networking and proficiency in Bash
- Familiarity with GPU/Accelerator scheduling, AI/ML pipelines
- Experience with multi cloud architectures and hybrid environments
- Experience with workflow orchestration tools (e.g., Argo Workflows)
Skills Required
- 10+ years hands-on experience in SRE, Linux Administration, or Systems Engineering
- Expert-level Linux knowledge including system internals, debugging, performance tuning, kernel and driver troubleshooting
- Kubernetes expertise managing K8s at scale (managed EKS and bare-metal deployments)
- Experience with queuing systems such as RabbitMQ or similar
- Metrics/observability stacks such as Prometheus, Thanos, and Grafana or similar
- Logging systems such as Elasticsearch or similar
- Relational database experience, e.g., PostgreSQL or similar
- Infrastructure-as-code proficiency with Terraform, Helm, and configuration management
- Strong networking fundamentals and proficiency in Bash scripting
- Familiarity with GPU/accelerator scheduling and AI/ML pipelines
- Experience with multi-cloud architectures and hybrid environments
- Experience with workflow orchestration tools (e.g., Argo Workflows)
What We Do
Mobileye is leading the mobility revolution with its autonomous-driving and driver-assistance technologies, harnessing world-renowned expertise in computer vision, machine learning, mapping, and data analysis. Founded in 1999, Mobileye has pioneered such groundbreaking technologies as REM™ crowdsourced mapping, True Redundancy™ sensing, and the RSS™ safety model. These technologies are driving the ADAS and AV fields towards the future of mobility – enabling self-driving vehicles and mobility solutions, powering industry-leading advanced driver-assistance systems and delivering valuable intelligence to optimize mobility infrastructure. Mobileye technology is used in over 170 million vehicles worldwide. In 2022, Mobileye became an independent company while still being majority-owned by Intel. Mobileye’s headquarters and R&D center are based in Jerusalem, with additional offices across Israel and around the world.
Why Work With Us
Our technology enables self-driving vehicles and mobility solutions, powers industry-leading advanced driver assistance systems, and delivers valuable intelligence to optimize mobility infrastructure.
Gallery







