Engineering Manager

Posted 7 Hours Ago
Easy Apply
Be an Early Applicant
Bangalore, Bengaluru Urban, Karnataka, IND
In-Office
Entry level
Cloud • Security • Software • Cybersecurity • Automation
The intelligent orchestration platform for DevSecOps
The Role
Lead a globally distributed observability engineering team responsible for metrics, logging, alerting, capacity planning, and reliability platforms. Hire and develop engineers, set priorities, guide technical decisions, improve SLO-based alerting and telemetry, participate in incident response, and sustain effective on-call operations. Collaborate with Site Reliability Engineering and Product Engineering while applying AI tools to engineering workflows and incident triage.
Summary Generated by Built In

GitLab is the intelligent orchestration platform for DevSecOps. GitLab enables organizations to increase developer productivity, improve operational efficiency, reduce security and compliance risk, and accelerate digital transformation. More than 50 million registered users and more than 50% of the Fortune 100* trust GitLab to ship better, more secure software faster.

The same principles built into our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members expected to incorporate AI into their daily workflows to drive efficiency, innovation, and impact. GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Our high-performance culture is driven by our values and continuous knowledge exchange, enabling our team members to reach their full potential while collaborating with industry leaders to solve complex problems. Co-create the future with us as we build technology that transforms how the world develops software.

*Fortune 500® is a registered trademark of Fortune Media IP Limited, used under license. Claim based on GitLab data. Fortune 100 refers to the top 20% ranked companies in the 2025 Fortune 500 list, published in June 2025. Fortune and Fortune Media IP Limited are not affiliated with, and do not endorse products or services of GitLab.

Engineering Manager, Production Engineering - Observability
 
An overview of this role
You'll lead the globally distributed Observability team. The team builds and operates the metrics, logging, alerting, and capacity planning platforms that GitLab engineers use to understand GitLab.com and GitLab Dedicated. You'll help determine how the team collects, stores, queries, and acts on telemetry, balancing reliable signals with scale and cost.
In your first year, you'll guide improvements to Prometheus-based metrics pipelines, log ingestion and retention, alerting driven by service-level objectives (SLOs), and capacity forecasting. You'll work with Site Reliability Engineering, Product Engineering, and other Infrastructure Platforms teams to make it easier for engineers to observe the services they own. You'll also take part in incident response and help keep the team's on-call work sustainable.
What you’ll do
  • Lead, hire, onboard, and develop a distributed engineering team working asynchronously.
  • Set priorities with Site Reliability Engineering, Product Engineering, and GitLab Dedicated teams, and help the team deliver observability services iteratively.
  • Own the reliability, scalability, and cost of the team's metrics, logging, alerting, and capacity planning platforms.
  • Reduce noisy or missing alerts and telemetry gaps, and use SLOs, error budgets, and self-service instrumentation to help engineers maintain the health of their services.
  • Guide technical decisions about time-series storage, high-cardinality metrics, log pipelines, and distributed tracing.
  • Participate in the Incident Manager On Call (IMOC) rotation, coordinating the response to high-severity incidents affecting GitLab.com.
  • Keep the team's on-call rotation sustainable through coverage across time zones, useful runbooks, better alerts, and follow-through on post-incident actions.
  • Use AI tools and agents to support engineering workflows and incident triage, reviewing their output while engineers retain responsibility for decisions.
What you’ll bring
  • Experience leading an observability, platform engineering, or site reliability engineering team operating at scale, including supporting people in a distributed, asynchronous environment.
  • Technical knowledge of metrics systems such as Prometheus and long-term storage, logging platforms such as Elasticsearch or cloud-native services, and alerting design.
  • Experience using SLOs, error budgets, and capacity forecasts to make reliability and investment decisions.
  • Experience operating a large software-as-a-service platform and investigating production issues such as telemetry gaps, ingestion limits, or noisy and missing alerts.
  • Experience participating in and improving production on-call rotations, including incident coordination and balancing operational load with project work.
  • The ability to explain technical tradeoffs to engineering partners and other stakeholders.
  • Experience using AI tools or agents in engineering or management work; you can describe how you would apply them to operational problems such as incident triage.
We welcome different paths into this role, whether through practical experience, formal study, or transferable skills. If the work interests you, please apply even if you don't meet every qualification.
 
About the team
We're part of Production Engineering within Infrastructure Platforms and work asynchronously across regions. Our tools include Tamland, the capacity forecasting tool. We bring lessons from operating GitLab's production systems back into the platforms we build.
 
 
How GitLab Supports Full-Time Employees
  • Benefits to support your health, finances, and well-being
  • Flexible Paid Time Off 
  • Team Member Resource Groups
  • Equity Compensation & Employee Stock Purchase Plan
  • Growth and Development Fund
  • Parental Leave 

Please note that we welcome interest from candidates with varying levels of experience; many successful candidates do not meet every single requirement. Additionally, studies have shown that people from underrepresented groups are less likely to apply to a job unless they meet every single qualification. If you're excited about this role, please apply and allow our recruiters to assess your application.

Country Hiring Guidelines: GitLab hires new team members in countries around the world. All of our roles are remote, however some roles may carry specific location-based eligibility requirements. Our Talent Acquisition team can help answer any questions about location after starting the recruiting process.  

Privacy Policy: Please review our Recruitment Privacy Policy. Your privacy is important to us.

GitLab is proud to be an equal opportunity workplace and is an affirmative action employer. GitLab’s policies and practices relating to recruitment, employment, career development and advancement, promotion, and retirement are based solely on merit, regardless of race, color, religion, ancestry, sex (including pregnancy, lactation, sexual orientation, gender identity, or gender expression), national origin, age, citizenship, marital status, mental or physical disability, genetic information (including family medical history), discharge status from the military, protected veteran status (which includes disabled veterans, recently separated veterans, active duty wartime or campaign badge veterans, and Armed Forces service medal veterans), or any other basis protected by law. GitLab will not tolerate discrimination or harassment based on any of these characteristics. See also GitLab’s EEO Policy and EEO is the Law. If you have a disability or special need that requires accommodation, please let us know during the recruiting process.

Skills Required

  • Experience leading an observability, platform engineering, or site reliability engineering team at scale
  • Experience supporting people in a distributed, asynchronous environment
  • Technical knowledge of Prometheus metrics systems and long-term storage
  • Technical knowledge of Elasticsearch or cloud-native logging services
  • Knowledge of alerting design
  • Experience using SLOs, error budgets, and capacity forecasts for reliability and investment decisions
  • Experience operating a large software-as-a-service platform
  • Experience investigating production issues such as telemetry gaps, ingestion limits, and noisy or missing alerts
  • Experience participating in and improving production on-call rotations and incident coordination
  • Ability to balance operational load with project work
  • Ability to explain technical tradeoffs to engineering partners and stakeholders
  • Experience using AI tools or agents in engineering or management work

What the Team is Saying

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
San Francisco, CA
2,500 Employees
Year Founded: 2014

What We Do

GitLab is the Intelligent Orchestration Platform where software teams and their AI agents stay in flow to amplify their capacity for innovation. Together, they automate repetitive tasks to plan, build, secure, test, deploy and maintain software. With GitLab, software teams spend less time on coordination overhead and more time on the next big idea. What started in 2011 as an open source project to help one team of programmers collaborate is now the intelligent orchestration platform millions of people use to deliver software faster, more efficiently, while strengthening security and compliance. Since the beginning, we've been firm believers in remote work, open source, DevSecOps, and iteration. We get up and log on in the morning to work alongside the GitLab community to deliver new innovations every month that help teams and their AI agents ship great code faster.

Why Work With Us

GitLab is where careers accelerate, innovation flourishes, and every voice is valued. Co-create the future with us as we build technology that transforms how the world develops software.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

GitLab Teams

Team
Software Engineering
Team
Software Engineering
Team
Sales
Team
Customer Success
About our Teams

GitLab Offices

Remote Workspace

Employees work remotely.

All-remote means that each individual in the organization is empowered to work and live where they are most fulfilled; it makes it clear that every team member is equal. No one, not even the executive team, meets in-person on a daily basis.

Typical time on-site: None
San Francisco, CA

Similar Jobs

GitLab Logo GitLab

Engineering Manager

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
2500 Employees

GitLab Logo GitLab

Senior Engineering Manager

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
2500 Employees

GitLab Logo GitLab

Engineering Manager

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
2500 Employees

GitLab Logo GitLab

Engineering Manager

Cloud • Security • Software • Cybersecurity • Automation
Easy Apply
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
2500 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account