Sr. DevOps Engineer

Posted 5 Days Ago
Be an Early Applicant
Pune, Maharashtra, IND
In-Office
Senior level
Information Technology • Software
The Role
Own the reliability, scalability, security, and performance of large-scale Linux infrastructure supporting backup and disaster recovery services. Manage Kubernetes, infrastructure as code, CI/CD pipelines, monitoring, observability, incident response, and production troubleshooting. Automate operational tasks with Shell, Python, or Go; implement security controls; and collaborate with engineering and SRE teams to improve deployment processes and platform reliability.
Summary Generated by Built In

About Kaseya

Kaseya is the leading provider of AI-powered IT management and cybersecurity software, serving Managed Service Providers (MSPs) and internal IT organizations worldwide. Our comprehensive platform helps organizations efficiently manage, secure, and automate their IT environments, driving operational efficiency and long-term business success.

Backed by Insight Partners, a leading global software investor, Kaseya has experienced sustained double-digit growth and continues to expand its global footprint. Today, Kaseya supports customers in more than 20 countries and manages over 15 million endpoints worldwide.

Founded in 2000, Kaseya has built a culture centered around innovation, accountability, and results. We are a high-growth, high-performance organization that values individuals who are driven, adaptable, and committed to delivering exceptional outcomes for our customers and teammates alike.

At Kaseya, success comes from embracing challenges, moving with urgency, and continuously raising the bar. 


Senior DevOps Engineer

Experience: 8–12 Years
Location: Pune (Hybrid/Onsite)

About the Role

We are looking for an experienced Senior DevOps Engineer with deep expertise in Linux Administration to join our Backup Platform Engineering team. In this role, you will own the reliability, scalability, and performance of large-scale Linux infrastructure that powers our next-generation backup and disaster recovery platform. You will work closely with software engineering, SRE, and platform teams to automate infrastructure, improve operational excellence, and ensure highly available production environments.


Key Responsibilities

  • Design, build, and manage highly available Linux-based production infrastructure supporting mission-critical backup services.
  • Administer and optimize large-scale Linux environments, including performance tuning, kernel configuration, storage, networking, and system troubleshooting.
  • Manage Kubernetes clusters, ensuring reliability, scalability, security, and efficient resource utilization.
  • Build and maintain Infrastructure as Code (Terraform/Pulumi) following reusable and modular design principles.
  • Design and enhance CI/CD pipelines using GitHub Actions, Jenkins, ArgoCD, or similar tools.
  • Develop automation using Shell scripting, Python, or Go to eliminate manual operational tasks.
  • Implement monitoring, logging, and observability using Prometheus, Grafana, Datadog, or similar platforms.
  • Drive incident response, root cause analysis, postmortems, and continuous operational improvements.
  • Collaborate with development teams to improve deployment processes, platform reliability, and production readiness.
  • Implement infrastructure security best practices including RBAC, secrets management, vulnerability scanning, and audit logging.
  • Troubleshoot complex infrastructure, networking, storage, and Linux operating system issues in production environments.

Required Skills

  • 8–12 years of experience in DevOps, Platform Engineering, Linux Administration, or Site Reliability Engineering (SRE).
  • Strong expertise in Linux system administration, including:
    • Performance tuning
    • Kernel parameters
    • Storage & filesystem management
    • Process management
    • System troubleshooting
    • Networking fundamentals
  • Hands-on experience managing Kubernetes clusters in production.
  • Strong knowledge of Infrastructure as Code using Terraform or Pulumi.
  • Experience designing and maintaining CI/CD pipelines using GitHub Actions, Jenkins, ArgoCD, or equivalent.
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, ELK, etc.
  • Strong scripting skills in Shell, Python, or Go for infrastructure automation.
  • Experience managing production incidents, on-call support, alert tuning, and operational excellence.
  • Strong understanding of:
    • DNS
    • Load Balancers
    • Firewalls  
    • VPCs  
    • Hybrid networking
  • Experience implementing security best practices including Vault, RBAC, secrets management, and audit logging.

Preferred Skills

  • Experience with OpenStack (Nova, Swift, Neutron, Cinder).
  • Experience managing workloads across AWS, Azure, GCP, and private cloud environments.
  • Exposure to large-scale distributed infrastructure (5,000+ nodes).
  • Experience with storage platforms, backup infrastructure, and disaster recovery concepts (RPO/RTO).
  • Knowledge of cost optimization (FinOps), storage tiering, and infrastructure capacity planning.
  • Experience with Chaos Engineering and resilience testing.
  • Familiarity with bare-metal provisioning technologies such as PXE, MaaS, or Ironic.
  • Ability to read and troubleshoot Go-based services and contribute to automation tooling.

Additional information
Kaseya provides equal employment opportunity to all employees and applicants without regard to race, religion, age, ancestry, gender, sex, sexual orientation, national origin, citizenship status, physical or mental disability, veteran status, marital status, or any other characteristic protected by applicable law.

Skills Required

  • 8-12 years of experience in DevOps, Platform Engineering, Linux Administration, or Site Reliability Engineering
  • Strong Linux system administration expertise, including performance tuning, kernel parameters, storage, filesystems, processes, troubleshooting, and networking
  • Production experience managing Kubernetes clusters
  • Infrastructure as Code experience with Terraform or Pulumi
  • Experience designing and maintaining CI/CD pipelines using GitHub Actions, Jenkins, ArgoCD, or equivalent
  • Experience with monitoring and observability tools such as Prometheus, Grafana, Datadog, or ELK
  • Strong scripting skills in Shell, Python, or Go
  • Experience with production incidents, on-call support, alert tuning, and operational excellence
  • Strong understanding of DNS, load balancers, firewalls, VPCs, and hybrid networking
  • Experience implementing security best practices, including Vault, RBAC, secrets management, and audit logging
  • Experience with OpenStack, including Nova, Swift, Neutron, or Cinder
  • Experience managing workloads across AWS, Azure, GCP, and private cloud environments
  • Experience with large-scale distributed infrastructure of 5,000 or more nodes
  • Experience with storage platforms, backup infrastructure, and disaster recovery concepts such as RPO and RTO
  • Knowledge of FinOps, storage tiering, and infrastructure capacity planning
  • Experience with Chaos Engineering and resilience testing
  • Familiarity with PXE, MaaS, or Ironic bare-metal provisioning
  • Ability to read and troubleshoot Go-based services and contribute to automation tooling

Kaseya Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Kaseya and has not been reviewed or approved by Kaseya.

  • Leave & Time Off Breadth PTO is commonly described around 20–21 days per year plus standard holidays. Some indicate they can fully disconnect while on leave.
  • Equity Value & Accessibility Equity or option grants are available to many roles, offering potential upside beyond base pay. This exposure is presented as a meaningful component of total compensation for some roles.
  • Affordable Benefits The high‑deductible medical plan is described as having low or employer‑covered employee‑only premiums in some cases. This can reduce out‑of‑pocket costs for those who select the HDHP.

Kaseya Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Miami, FL
5,000 Employees
Year Founded: 2000

What We Do

Kaseya is a premier provider of unified IT management and security software for managed service providers (MSPs) and small to medium-sized businesses (SMBS). Through its customer-centric approach, Kaseya delivers best-in-breed technologies that allow organizations to efficiently manage, secure and backup IT. Kaseya offers a broad array of IT management solutions, including well-known names: Kaseya, IT Glue, RapidFire Tools, Spanning Cloud Apps, ID Agent, Graphus, RocketCyber, TruMethods and Unitrends. These solutions empower businesses to command all of IT centrally, easily manage remote and distributed environments, simplify backup and disaster recovery, safeguard against cybersecurity attacks, effectively manage compliance and network assets, streamline IT documentation and automate across IT management functions. Headquartered in Miami, Florida, Kaseya is privately held with a presence in over 20 countries.

Gallery

Gallery

Similar Jobs

MetLife Logo MetLife

Senior Devops Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Hybrid
Pune, Maharashtra, IND
43000 Employees

SailPoint Logo SailPoint

Staff Devops Engineer

Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Hybrid
Pune, Maharashtra, IND
2461 Employees

SailPoint Logo SailPoint

Senior Devops Engineer

Artificial Intelligence • Cloud • Sales • Security • Software • Cybersecurity • Data Privacy
Remote or Hybrid
India
2461 Employees
In-Office
3 Locations
3464 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account