Senior Database Reliability Engineer

Posted Yesterday
Be an Early Applicant
Pune, Mahārāshtra, IND
In-Office
Senior level
Information Technology • Security • Cybersecurity
The Role
Owns the reliability and operations of OpenSearch/Elasticsearch, Kafka, and Redis platforms. Responsibilities include cluster administration, troubleshooting, upgrades, scaling, backup and recovery, incident response, observability, capacity planning, and Python automation. The engineer participates in on-call support, improves runbooks and production-readiness standards, executes infrastructure changes, and collaborates with SRE, infrastructure, cloud, storage, networking, and application teams.
Summary Generated by Built In

Come work at a place where innovation and teamwork come together to support the most exciting missions in the world!

Senior Database Reliability Engineer (DBRE) - Data Platform

OpenSearch Must-Have | Kafka | Redis | Python Automation | Production Reliability

Level

Senior DBRE (Senior Individual Contributor)

Domain

Distributed Data Infrastructure

Primary Focus

Hands-on reliability and operational ownership for assigned OpenSearch/Elasticsearch, Kafka, and Redis services; Python automation and disciplined production execution

Work Expectations

U.S. Eastern Time business hours; on-call participation as required


1. Role Summary

We are seeking an experienced Senior DBRE to improve the reliability, performance, scalability, and automation of mission-critical distributed data platforms. OpenSearch/Elasticsearch is the core must-have technology, complemented by production experience with Kafka, Redis, Linux systems, observability, and Python automation.

This is a hands-on role. The Senior DBRE independently operates assigned clusters and services, troubleshoots production issues, performs upgrades and lifecycle work, builds practical automation, and contributes to platform design and production-readiness standards.


2. Level Scope and Leadership Expectations

The Senior DBRE is a high-impact hands-on engineer focused on independent execution within an assigned platform scope. Senior engineers solve common and moderately complex production problems, improve automation and runbooks, and contribute to designs and standards while escalating broader architectural decisions appropriately.


Expectation at This Level

Scope of Ownership

Owns assigned clusters, services, maintenance activities, incidents, upgrades, backup/restore, capacity reviews, and operational improvements with limited supervision.

Technical Depth

Diagnoses common and moderately complex issues involving shards, indexing, search, Kafka lag, Redis memory behavior, JVM, Linux, storage, and networking.

Automation

Builds and maintains tested Python tools, operational CLIs, API integrations, dashboards, alerts, and runbooks that reduce recurring toil.

Execution

Executes migrations, rolling upgrades, scaling, and DR exercises with change-management discipline and appropriate technical review.

Collaboration

Partners with Lead and Staff engineers, SRE, application, infrastructure, storage, network, and cloud teams; documents findings and communicates risks early.

Architecture Influence

Contributes evidence and operational experience to design reviews; does not independently set enterprise-wide platform architecture.


3. Key ResponsibilitiesA. OpenSearch and Elasticsearch Operations

• Operate, scale, and improve assigned OpenSearch/Elasticsearch clusters and node topologies.

• Troubleshoot shard allocation, indexing throughput, slow search and aggregation workloads, cluster state, JVM heap/GC, disk, memory, and recovery issues.

• Implement and maintain index templates, shard strategies, ISM/ILM policies, rollover, retention, snapshot, restore, and rolling-upgrade procedures.

B. Kafka and Redis Operations

• Operate Kafka topics, partitions, replication, retention, consumer groups, and broker health; diagnose lag, rebalance storms, and capacity bottlenecks.

• Operate Redis Cluster and Sentinel environments; address memory fragmentation, eviction behavior, persistence, replication lag, connection limits, and failover issues.

C. Automation and Infrastructure as Code

• Develop clean, modular Python automation for health checks, maintenance, validation, reporting, and safe remediation.

• Use Terraform, Ansible, GitOps workflows, or Kubernetes operators where applicable to make platform operations repeatable.

D. Reliability and Incident Response

• Instrument actionable metrics and alerts using Prometheus, Grafana, OpenTelemetry, or comparable platforms.

• Participate in on-call, stabilize incidents, document root cause, and complete preventive actions.

• Contribute to capacity planning, production-readiness reviews, runbooks, and design reviews.


4. Required Qualifications and Experience

• 5+ years of relevant DBRE, SRE, database, or distributed-systems operations experience in production environments.

• Strong hands-on OpenSearch and/or Elasticsearch experience, including shards, indexing, search, lifecycle management, JVM behavior, and snapshot/restore.

• Production experience with at least one of Kafka or Redis; experience with both is strongly preferred.

• Strong Linux and systems troubleshooting skills across CPU, memory, storage I/O, networking, and process behavior.

• Ability to write maintainable Python automation beyond simple one-off shell scripts.

• Experience with monitoring, incident response, controlled production changes, upgrades, and operational documentation.

• Clear communication, sound judgment, and willingness to escalate risk early.


5. Preferred Qualifications

• Hands-on experience with both Kafka and Redis.

• Terraform, Ansible, GitOps, Kubernetes, ECK/OpenSearch, Strimzi, or Redis operators.

• Cloud-managed data services in AWS, Azure, or GCP.

• Familiarity with ClickHouse, Pinot, vector databases, Flink, Spark Streaming, or Kafka Connect.


6. Expected Outcomes in the First 6-12 Months

• Establish measurable baselines and deliver material improvements in reliability, alert quality, MTTR, and recurring operational toil for assigned services.

• Automate high-frequency maintenance and validation workflows using maintainable Python tooling.

• Improve runbook coverage, upgrade readiness, backup/restore confidence, and capacity visibility.

• Contribute a well-supported evaluation or operational-readiness recommendation for one relevant emerging technology when business needs require it.


7. What This Role Does Not Include

• This is not a people-management role.

• This is not passive monitoring, ticket routing, or a generic operations queue; hands-on troubleshooting and improvement are expected.

• This role does not independently own enterprise-wide architecture or technical strategy.

• This role does not own application code or product feature delivery, although close partnership with application teams is required.

• The role does not require equal mastery of every listed platform; OpenSearch/Elasticsearch is the anchor expertise, with complementary Kafka and/or Redis depth.

8. Work Expectations and Location

This role requires working during U.S. Eastern Time (ET) business hours and participating in on-call support as required.


Skills Required

  • 5+ years of relevant DBRE, SRE, database, or distributed-systems operations experience in production environments
  • Strong hands-on OpenSearch and/or Elasticsearch experience, including shards, indexing, search, lifecycle management, JVM behavior, and snapshot/restore
  • Production experience with at least one of Kafka or Redis
  • Strong Linux and systems troubleshooting skills across CPU, memory, storage I/O, networking, and process behavior
  • Ability to write maintainable Python automation beyond simple one-off shell scripts
  • Experience with monitoring, incident response, controlled production changes, upgrades, and operational documentation
  • Clear communication, sound judgment, and willingness to escalate risk early
  • Hands-on experience with both Kafka and Redis
  • Experience with Terraform, Ansible, GitOps, Kubernetes, ECK/OpenSearch, Strimzi, or Redis operators
  • Experience with cloud-managed data services in AWS, Azure, or GCP
  • Familiarity with ClickHouse, Pinot, vector databases, Flink, Spark Streaming, or Kafka Connect

Qualys Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Qualys and has not been reviewed or approved by Qualys.

  • Affordable Benefits — Benefits costs are widely viewed as low for employees and dependents, with healthcare often described as almost fully paid for. Feedback suggests this affordability helps offset perceptions of lower base pay in some roles.
  • Healthcare Strength — Healthcare offerings are broad, including multiple medical plan options, dental and vision coverage, mental health support, and disability insurance. Benefits are described as “pretty amazing” or “great,” reinforcing perceived quality and coverage depth.
  • Equity Value & Accessibility — Equity participation is accessible through company stock plans and an employee stock purchase plan. Compensation packages commonly include equity alongside salary and bonus, which some consider a meaningful part of total rewards.

Qualys Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Foster City, CA
2,736 Employees
Year Founded: 1999

What We Do

Qualys, Inc. (NASDAQ: QLYS) is a pioneer and leading provider of disruptive cloud-based security, compliance and IT solutions with more than 10,000 subscription customers worldwide, including a majority of the Forbes Global 100 and Fortune 100. Qualys helps organizations streamline and automate their security and compliance solutions onto a single platform for greater agility, better business outcomes, and substantial cost savings. The Qualys Cloud Platform leverages a single agent to continuously deliver critical security intelligence while enabling enterprises to automate the full spectrum of vulnerability detection, compliance, and protection for IT systems, workloads and web applications across on premises, endpoints, servers, public and private clouds, containers, and mobile devices. Founded in 1999 as one of the first SaaS security companies, Qualys has strategic partnerships and seamlessly integrates its vulnerability management capabilities into security offerings from cloud service providers, including Amazon Web Services, the Google Cloud Platform and Microsoft Azure, along with a number of leading managed service providers and global consulting organizations. For more information, please visit http://www.qualys.com

Similar Jobs

Modulr Logo Modulr

Reliability Engineer

Payments • Software
In-Office
Pune, Mahārāshtra, IND
333 Employees

TransUnion Logo TransUnion

Manager, Web Application Development

Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Hybrid
Pune, Mahārāshtra, IND
13000 Employees

CSC Logo CSC

Accountant

Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Remote or Hybrid
2 Locations
8500 Employees

TransUnion Logo TransUnion

Senior Consultant

Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Hybrid
4 Locations
13000 Employees

Similar Companies Hiring

Milestone Systems Thumbnail
Artificial Intelligence • Security • Software • Analytics • Big Data Analytics
Lake Oswego, OR
1500 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account