Site Reliability Engineer

Posted 14 Days Ago
Be an Early Applicant
5 Locations
In-Office
1M-3M Annually
Senior level
Artificial Intelligence • HR Tech • Professional Services • Software
The Role
Manage, deploy, and optimize Apache Kafka clusters and large-scale streaming platforms. Monitor performance, troubleshoot production incidents, automate operational tasks, implement capacity planning and disaster recovery, ensure security and compliance, create runbooks, and participate in on-call incident response to maintain highly available, scalable platform infrastructure.
Summary Generated by Built In

๐—ง๐—ต๐—ถ๐˜€ ๐—ฟ๐—ผ๐—น๐—ฒ ๐—ถ๐˜€ ๐—ณ๐—ผ๐—ฟ ๐—ผ๐—ป๐—ฒ ๐—ผ๐—ณ ๐˜๐—ต๐—ฒ ๐—ช๐—ฒ๐—ฒ๐—ธ๐—ฑ๐—ฎ๐˜†'๐˜€ ๐—ฐ๐—น๐—ถ๐—ฒ๐—ป๐˜๐˜€

๐—ฆ๐—ฎ๐—น๐—ฎ๐—ฟ๐˜† ๐—ฟ๐—ฎ๐—ป๐—ด๐—ฒ: ๐—ฅ๐˜€ ๐Ÿญ๐Ÿฎ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ - ๐—ฅ๐˜€ ๐Ÿฎ๐Ÿด๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ๐Ÿฌ (๐—ถ๐—ฒ ๐—œ๐—ก๐—ฅ ๐Ÿญ๐Ÿฎ-๐Ÿฎ๐Ÿด ๐—Ÿ๐—ฃ๐—”)

Experience: 6+ yrs

Location: Hyderabad, Bengaluru, Pune, Chennai, Tamil Nadu, India, Mumbai, Maharashtra, India

Job Type: Full-time

We are seeking an experienced Kafka Platform Engineer with strong expertise in distributed systems, large-scale messaging platforms, and production operations. This role is ideal for professionals who are passionate about building, maintaining, and optimizing highly available streaming infrastructure while ensuring reliability, scalability, and operational excellence across enterprise environments.

As a Kafka Platform Engineer, you will be responsible for managing mission-critical messaging platforms, improving platform performance, automating operational processes, and supporting production environments. You will collaborate with infrastructure, application, DevOps, and engineering teams to deliver resilient streaming solutions, troubleshoot complex production issues, and continuously enhance platform reliability through automation, monitoring, and best practices.


RequirementsKey Responsibilities
  • Design, deploy, manage, and optimize Apache Kafka clusters and large-scale messaging or streaming platforms.
  • Monitor platform health, system performance, and resource utilization using modern monitoring, logging, and alerting tools.
  • Troubleshoot production incidents, identify root causes, and implement long-term solutions to improve system stability.
  • Automate operational tasks, deployments, and maintenance activities using scripting and infrastructure automation techniques.
  • Collaborate with application and infrastructure teams to support messaging architecture, integrations, and production workloads.
  • Implement performance tuning, capacity planning, and scalability improvements for distributed systems.
  • Maintain high availability, fault tolerance, and disaster recovery strategies for messaging infrastructure.
  • Ensure platform security, system compliance, and operational best practices across production environments.
  • Develop operational documentation, runbooks, and knowledge-sharing resources to improve support efficiency.
  • Participate in on-call support, incident response, and continuous improvement initiatives to maintain service reliability.
What Makes You a Great Fit
  • 6+ years of experience managing distributed systems, production infrastructure, or platform engineering environments.
  • Strong hands-on experience with Apache Kafka or large-scale messaging and event streaming platforms.
  • Deep understanding of distributed systems architecture, scalability, fault tolerance, and production operations.
  • Experience with monitoring, logging, alerting, and observability tools for enterprise infrastructure.
  • Proficiency in at least one scripting or programming language such as PythonBash, or Java.
  • Strong knowledge of Linux system administration, networking fundamentals, and troubleshooting methodologies.
  • Experience automating operational workflows and improving platform reliability through scripting and infrastructure automation.
  • Excellent analytical, debugging, and problem-solving skills with a proactive operational mindset.
  • Strong communication and collaboration skills with the ability to work effectively across cross-functional engineering teams.
  • Passion for building reliable, secure, and highly available platform infrastructure while continuously improving operational excellence.

Skills Required

  • 6+ years experience managing distributed systems, production infrastructure, or platform engineering
  • Strong hands-on experience with Apache Kafka or large-scale event streaming platforms
  • Deep understanding of distributed systems architecture, scalability, and fault tolerance
  • Experience with monitoring, logging, alerting, and observability tools for enterprise infrastructure
  • Proficiency in at least one scripting or programming language such as Python, Bash, or Java
  • Strong knowledge of Linux system administration, networking fundamentals, and troubleshooting methodologies
  • Experience automating operational workflows and infrastructure automation
  • Excellent analytical, debugging, problem-solving, and cross-functional collaboration skills
  • Familiarity with high availability, disaster recovery, security, and compliance for messaging infrastructure
  • Passion for building reliable, secure, and highly available platform infrastructure
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Year Founded: 2021

What We Do

Weekday is an AI-powered recruitment platform that helps startups hire top-tier engineering and product talent. By leveraging a massive database of white-collar professionals and advanced outreach tools, the company streamlines the hiring process through automated sourcing, AI-driven resume screening, and white-glove contingency services. Their mission is to modernize recruitment by enabling companies to discover and engage passive candidates efficiently, ensuring high-quality hires for critical roles.

Similar Jobs

M&G Logo M&G

Chief Technology Officer

Fintech • Financial Services
In-Office
Pune, Mahārāshtra, IND
2729 Employees

Qualys Logo Qualys

Site Reliability Engineer

Information Technology • Security • Cybersecurity
In-Office
Pune, Mahārāshtra, IND
2736 Employees

Qualys Logo Qualys

Site Reliability Engineer

Information Technology • Security • Cybersecurity
In-Office
Pune, Mahārāshtra, IND
2736 Employees

CrowdStrike Logo CrowdStrike

Site Reliability Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Hybrid
Pune, Mahārāshtra, IND
11000 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account