DevOps Tech Lead (BigBrain)

Reposted 7 Hours Ago
Be an Early Applicant
Tel Aviv, ISR
Hybrid
Mid level
Artificial Intelligence • Productivity • Sales • Software
Power the most ambitious teams.
The Role
Build, operate, and scale AI and data platform infrastructure including multi-region EKS, streaming pipelines, ML inference, and CI/CD/GitOps. Drive automation (AIOps), ensure security and observability, enable developer self-service, and evolve data systems to support billions of daily events.
Summary Generated by Built In

About monday.com
monday.com is the AI work platform powering the most ambitious teams. 250,000+ customers across departments use us to bring people, workflows, and AI agents together on one flexible platform where AI doesn't just assist, it executes. We move fast, build things that matter, and foster an ownership-driven culture where you're empowered to shape how organizations work and outpace their competition.

About the team


We are looking for a Senior DevOps Engineer / Tech Lead to join the BigBrain group.

The team that builds and operates monday.com's Data Platform and leads the company's internal AI innovation.

We own the infrastructure behind billions of daily events - from streaming pipelines and data orchestration to the AI Gateway and ML inference platform that power monday.com's intelligent features. We manage some of the most sensitive data in the company, operate across multiple global regions, and are responsible for keeping it all secure and running at scale.

The BigBrain group consists of diverse teams, including Data Scientists, Full-Stack Engineers, Data Engineers, and BI Engineers. As a Senior DevOps Engineer / Tech Lead on this team, you'll own the technical direction for critical infrastructure domains, mentor other engineers, and partner closely with product, security, data, and platform teams to build infrastructure that's secure, scalable, and increasingly autonomous.

This is a high-ownership role with significant influence over the technical roadmap of our data and AI infrastructure.

This is an exciting time to join: we're building an AI Gateway to govern LLM usage across the company, scaling our ML inference platform, designing AIOps agents that automate infrastructure operations, and evolving our data platform with modern technologies - all while keeping the foundation rock-solid across multiple regions.

You'll join our DevOps team based in our headquarters in Tel Aviv, Israel.

Team podcasts - https://www.startupforstartup.com/110-on-the-operations-behind-the-client-facing-teams-growth/ https://pod.link/1595260676/episode/fa4d99bbf8e536e7a139a1ba679692f0
BigBrain: https://www.startupforstartup.com/on-the-bigbrain-that-makes-kpis-accessible-to-every-employee/ https://www.youtube.com/watch?v=x-m0ag0cty0 https://engineering.monday.com/ai-brain-ai-charged-tool-for-internal-usage/

About The Role

  • Lead technical direction — Own the architecture and technical roadmap for critical infrastructure domains, and make high-impact trade-offs across teams.

  • Own and scale platform infrastructure — Manage multi-region Kubernetes clusters (EKS), streaming pipelines (Kafka/MSK, Debezium CDC), and data orchestration (Airflow) that handles billions of daily events.

  • Build and operate AI infrastructure — Deploy and maintain the AI Gateway (governance, observability, guardrails for LLM usage), ML inference platform, and the tooling that enables AI adoption across the company.

  • Drive infrastructure automation with AIOps — Design and build autonomous agents and intelligent tooling (n8n, LangChain, Claude Code) that automate infrastructure operations, analyze workflows, and reduce manual toil.

  • Drive cross-team initiatives — Lead complex, multi-team infrastructure projects end-to-end, from design through rollout.

  • Mentor and grow engineers — Guide and mentor other DevOps/infra engineers, review designs, and raise the technical bar across the team.

  • Build and maintain CI/CD & GitOps pipelines — Design and operate deployment pipelines using GitHub Actions, ArgoCD, Helm, and Terraform (CDKTF), enabling fast, safe, and reliable releases for product teams.

  • Ensure data security — Protect the company's most sensitive data through access control, data governance, and security-first infrastructure design.

  • Evolve the data platform — Work with modern data technologies (Snowflake, ClickHouse, Apache Iceberg, EMR) and contribute to the next generation of our data infrastructure.

  • Enable developer self-service — Improve our internal platform so engineering teams across the company can deploy, configure, and operate services independently.

  • Provide observability and reliability — Build monitoring, alerting, and debugging tools (Datadog, OpenTelemetry, ClickHouse) across all data and AI processes, and lead incident response for critical, high-scale systems.

Our Stack — AWS, Kubernetes (EKS), Kafka/MSK, Debezium, Airflow, Snowflake, ClickHouse, Apache Iceberg, EMR, ArgoCD, Terraform/CDKTF, Helm, GitHub Actions, Docker, Datadog, OpenTelemetry, MLFlow, n8n, API Gateway, Redis, MySQL, Teleport, TypeScript, Node.js, Python.

Your Experience & Skills

  • 6-8+ years of experience as a DevOps / Infrastructure / Platform Engineer, including experience leading technical initiatives, owning architecture decisions, or mentoring other engineers.

  • Deep, hands-on experience with Kubernetes — cluster management, networking, scaling, and troubleshooting at production scale.

  • Strong experience with Infrastructure as Code (Terraform/CDKTF, Helm) and GitOps workflows (ArgoCD or similar).

  • Proven ownership of CI/CD pipelines — designing, maintaining, and optimizing the full release cycle for multiple teams.

  • Deep experience with cloud infrastructure (AWS preferred) — networking, IAM, security, cost optimization.

  • Security-first mindset — strong understanding of application security, access control, and data protection best practices.

  • Experience with data infrastructure — Kafka, Airflow, EMR, Apache Iceberg, Snowflake, ClickHouse, or similar technologies.

  • Fluent in Linux environments, scripting, and at least one programming language (Python, TypeScript, Go).

  • Proven ability to design and own system architecture at scale, and to make and defend complex technical trade-offs across multiple teams and domains.

  • Experience owning production reliability for critical, high-scale systems — including incident leadership, postmortems, and driving systemic fixes.

  • Strong communicator and cross-functional collaborator, comfortable influencing engineering standards and best practices across an organization.

  • Understanding of products and a passion for building software that impacts millions of users.

  • Comfortable operating with high autonomy and ambiguity.

Big advantage:

  • Experience with AI/ML infrastructure — model serving, LLM deployment, AI gateways, GPU workloads, or AI observability tools (MLFlow).

  • Hands-on experience with AI agents and automation — LangChain, LangGraph, n8n, Claude Code, or similar agentic frameworks.

  • Experience building internal developer platforms or self-service tooling.

  • Passion for pushing the boundaries of what AI agents and autonomous

Skills Required

  • 3+ years of experience as a DevOps Engineer
  • Strong technical skills and understanding of systems and infrastructure
  • Experience building the full application release cycle (CI/CD)
  • Familiarity with how modern web applications work and scale
  • Networking, firewall rules management, and application security knowledge
  • Familiar with Linux environment, scripting, and programming
  • Ability to perform system architecture planning and see the bigger picture
  • Product-minded with passion for building software impacting millions of users
  • Team player with strong communication skills and empathy
  • Big Data technologies such as Kafka/Flink
  • Experience supporting AI-native products or LLM-based systems
  • Familiarity with Cloudflare, PostgreSQL, and modern web app infrastructure
  • Background in scalability, performance optimization, and cloud cost management
  • Experience in startup-like settings, rapid prototyping, or 0-to-1 product teams
  • Track record building internal developer platforms or tools to improve productivity

What the Team is Saying

Ruchita
Nate
Kyle
Brad Wisselman
Brad Wisselman
Bianca Collado

monday.com Compensation & Benefits Highlights

  • Healthcare Strength Core coverage is described as comprehensive medical, dental, and vision insurance with FSAs and notable mental health support, including free counseling sessions and wellness app access. Some locations add on‑site wellness activities such as yoga and meditation, reinforcing a broad health focus.
  • Parental & Family Support Parental leave is presented as fully paid for up to 13 weeks for all caregivers, alongside adoption assistance and childcare benefits, with a return‑to‑work program noted. Company‑sponsored family events further signal emphasis on family support.
  • Retirement Support Retirement offerings include a 401(k) plan with automatic company contributions and are paired with other financial protections like life and disability insurance. An ESPP is also available, complementing longer‑term savings options.

monday.com Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
3,048 Employees
Year Founded: 2012

What We Do

At monday.com, we help teams get more work done. We are the best AI work platform that empowers teams to automate, build, and scale their impact end-to-end with tools that actually execute the work for you. With over $1B in ARR, 250,000+ customers, and a global team, we’re serious about building a product people love to use and giving our employees the same ownership and flexibility to shape the way the world works.

Why Work With Us

At monday.com we believe in transparency, accountability, and impact. Together, those values have lent themselves to create a strong culture of professional and creative autonomy where every team member is encouraged to share ideas and help bring them to life!

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

monday.com Teams

Team
Customer Experience
About our Teams

monday.com Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

monday.com embraces a flexible work environment with our hybrid model.

Typical time on-site: 3 days a week
HQNew York, NY
HQTel Aviv
Denver, CO
London
Melbourne
Munich
Paris, France
Sao Paolo
Singapore
Sydney
Tokyo
Warsaw
Learn more

Similar Jobs

monday.com Logo monday.com

AI Internal Product Lead

Artificial Intelligence • Productivity • Sales • Software
Hybrid
Tel Aviv, ISR
3048 Employees

monday.com Logo monday.com

Workplace Maintenance & Operations

Artificial Intelligence • Productivity • Sales • Software
Hybrid
Tel Aviv, ISR
3048 Employees

monday.com Logo monday.com

Application Security Group Lead

Artificial Intelligence • Productivity • Sales • Software
Hybrid
Tel Aviv, ISR
3048 Employees

monday.com Logo monday.com

Program Manager

Artificial Intelligence • Productivity • Sales • Software
Remote or Hybrid
Tel Aviv, ISR
3048 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account