Observability Tech Lead

Posted Yesterday
Easy Apply
Be an Early Applicant
Somerville, MA, USA
Hybrid
160K-200K Annually
Senior level
Enterprise Web • Hardware • Internet of Things • Software
Tulip is fundamentally and materially transforming how people on the frontline of operations work today.
The Role
Lead observability efforts: mentor teams on SLIs/SLOs and reliability, own incident response and production debugging, maintain triage/remediation processes, and design, build, and operate core observability infrastructure and tooling used across engineering.
Summary Generated by Built In

This role is located in Somerville, MA - We are a hybrid work environment and are in the office 3+ days/per week.

Tulip, the leader in AI-native frontline operations, is helping companies around the world equip their workforce with composable, connected apps, leading to higher quality work, improved efficiency, and end-to-end traceability across operations. Tulip’s cloud-native, no-code platform, powered by embedded AI, is driving the digital transformation of industrial environments through composable, human-centric solutions that go beyond disrupting the Manufacturing Execution System (MES) category.

A spinoff out of MIT, Tulip is headquartered in Somerville, MA, with offices in Germany, Hungary, Singapore, and Israel. Tulip has been recognized as a World Economic Forum Global Innovator, a 2024 Deloitte Technology Fast award winner, one of Energage’s Top Workplaces USA, and one of Built In Boston’s “Best Places to Work” and “Best Midsize Places to Work.”

About You:

  • You can reason about systems at scale: their edge cases, failure modes, and life cycles across 
  • You’re excited about setting the technical agenda and coming up with novel, broad ideas
  • You regularly keep up with the newest AI advancements in the realm of Observability & Monitoring
  • You know what a good SLA looks like, and can teach others how to spot one
  • You can communicate as well as you can code. You understand the value of discussion and work best in a team that champions clear and frequent communication

 What skills do I need? 

  • 5+ years of experience working with open source Observability tools (e.g. Loki, Grafana, Tempo, Mimir stack)
  • Hands-on experience instrumenting distributed systems using OpenTelemetry and managing metrics pipelines with Prometheus at scale
  • Direct experience developing and distributing Claude Skills, Gemini Gems, or any other generic AI processes and are able to iterate on their efficacy
  • Experience working with time-series data, ideally using promQL

 Key Responsibilities:

  • Mentor and evangelize on observability best practices, SLIs/SLOs, and reliability culture across engineering teams. 
  • Contributing to and maintaining Tulip's triage & remediation processes as a player / coach
  • Perform incident response and debug production issues across the entire stack 
  • Design, build, and maintain the core infrastructure & tooling used by all of Tulip’s engineering teams

Tech Stack:

  • TS and Go Services running on Kubernetes
  • MongoDB and PostGres DBs
  • Grafana, Loki, Mimir, Tempo, Alloy, Prometheus & OpenTelemetry Observability tooling

Key Collaborators:

  • Engineering
  • Edge 
  • DevOps 
  • Hardware

Working At Tulip

We know even great candidates experience imposter syndrome. Even if you don’t match every requirement, applying gives you the opportunity to be considered. 

We’re building a strong, diverse team that values hard work, families, and personal well-being. Benefits of working with us include:

  • Direct impact on product and culture
  • Company equity
  • Competitive benefits package including Health, Dental, Vision, Short-term Disability, Long-term Disability, Life Insurance, AD&D Insurance, Flexible Spending Account (FSA), Commuter Benefits, Parental Leave, and 401(K)
  • Flexible work schedule and unlimited vacation policy
  • Virtual company events and happy hours
  • Fitness subsidies

We are an equal opportunity employer. At Tulip, we celebrate all. Qualified applicants will receive consideration for employment without regard to race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. Help us build an inclusive community that will transform frontline operations. 

The compensation information displayed on each job posting reflects the range for new hire pay rates for the position across all US locations. Within the range posted, actual compensation will be determined depending on multiple factors including job-related knowledge & skills, experience, business needs, geographical location, market compensation data, and internal equity. Expected compensation ranges for this role may change over time. The salary range for this position is $160,000 - $200,000 per year.


It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.

Please note that we may use AI-based tools to support parts of our hiring process. All data processing is carried out in compliance with local data protection laws, ensuring all personal candidate information is handled securely and ethically.

Skills Required

  • 5+ years experience with open source observability tools (e.g., Loki, Grafana, Tempo, Mimir)
  • Hands-on experience instrumenting distributed systems using OpenTelemetry and managing metrics pipelines with Prometheus at scale
  • Direct experience developing and distributing Claude Skills, Gemini Gems, or other AI processes and iterating on their efficacy
  • Experience working with time-series data, ideally using promQL
  • Experience mentoring and evangelizing observability best practices, SLIs/SLOs, and reliability culture across teams
  • Proven incident response experience and ability to debug production issues across the full stack
  • Experience designing, building, and maintaining core infrastructure and tooling used by engineering teams
  • Experience building or maintaining TypeScript and Go services running on Kubernetes

What the Team is Saying

Xi
Grant
Pablo
Mark

Tulip Compensation & Benefits Highlights

  • Healthcare Strength Medical, dental, and vision coverage sit alongside disability, life insurance, and HRA/HSA options with multiple medical plan choices described. Employer-paid medical premiums for employees and significant contributions for families underscore the depth of coverage.
  • Affordable Benefits Employer funding of medical premiums meaningfully lowers premium costs for employees and dependents. Additional monetary supports such as HSA/FSA contributions, commuter benefits, and fitness subsidies further enhance affordability.
  • Leave & Time Off Breadth Unlimited PTO with flexible work practices is highlighted, alongside paid holidays/sick and wellness days in public benefit summaries. Parental leave is included and supported by return‑to‑work programs.

Tulip Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Somerville, MA
310 Employees
Year Founded: 2014

What We Do

Tulip, the leader in frontline operations, is helping companies around the world equip their workforce with connected apps, leading to higher quality work, improved efficiency, and end-to-end traceability across operations. Companies of all sizes and across industries have implemented composable solutions with Tulip’s cloud-native, no-code platform to solve some of the most pressing challenges in operations: error-proofing processes and boosting productivity, capturing and analyzing real-time data, and continuous improvement.

Why Work With Us

We’re building a strong, diverse team that values hard work, families, and personal well-being. Help us build an inclusive community that will transform frontline operations.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

Tulip Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

Typical time on-site: 3 days a week
Company Office Image
HQSomerville, MA
Israel
Japan
Company Office Image
Budapest, HU
Company Office Image
Munich, DE
Singapore
Learn more

Similar Jobs

Tulip Logo Tulip

Market Intelligence Lead

Enterprise Web • Hardware • Internet of Things • Software
Easy Apply
Hybrid
Somerville, MA, USA
310 Employees
125K-160K Annually

Tulip Logo Tulip

Operations Intern

Enterprise Web • Hardware • Internet of Things • Software
Easy Apply
Hybrid
Somerville, MA, USA
310 Employees
20-25 Hourly

Tulip Logo Tulip

Senior Site Reliability Engineer

Enterprise Web • Hardware • Internet of Things • Software
Easy Apply
Hybrid
Somerville, MA, USA
310 Employees
160K-200K Annually

Tulip Logo Tulip

Devops Engineer

Enterprise Web • Hardware • Internet of Things • Software
Easy Apply
Hybrid
Somerville, MA, USA
310 Employees
150K-190K Annually

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account