Senior DevOps Engineer

Posted 2 Days Ago
Be an Early Applicant
Pune, Mahārāshtra, IND
In-Office
Senior level
Professional Services • Software • Analytics • Business Intelligence
The Role
Lead implementation, automation, and operational support of scalable cloud and container platforms. Build and optimize CI/CD and GitOps pipelines, manage Kubernetes clusters and workloads, implement IaC with Terraform/Ansible, improve observability and security, perform advanced incident troubleshooting and RCA, mentor engineers, and collaborate with architects and stakeholders to maintain production-ready, cost-effective infrastructure.
Summary Generated by Built In
Role Summary
We are seeking an experienced Senior DevOps Engineer who will work closely with solution architects, application architects, development teams, security teams, and project stakeholders to implement, automate, secure, and maintain scalable cloud infrastructure and deployment platforms. The role is strongly hands-on and focuses on translating approved architecture into reliable implementations, optimizing delivery pipelines, operating cloud and container platforms, improving observability, and resolving complex production issues.
The Senior DevOps Engineer will contribute implementation and operational inputs during design discussions. Kubernetes architecture will be defined collaboratively with the architects, while this role will be responsible for implementation, validation, migration, upgrades, optimization, and ongoing operational support.

Key Roles and Responsibilities
 Work closely with architects, development teams, security teams, QA teams, customers, and project stakeholders to understand technical requirements, project KPIs, service expectations, and operational constraints.
 Translate approved infrastructure and application architecture into secure, scalable, maintainable, and production-ready implementations.
 Provide practical DevOps inputs on deployment feasibility, automation, resilience, observability, scalability, security, performance, and cloud cost.
 Build, maintain, and optimize CI/CD pipelines covering build, test, quality checks, security scanning, artifact management, deployment, environment promotion, and rollback.
 Select and implement suitable CI/CD, GitOps, infrastructure automation, monitoring, logging, and security tools in consultation with architects and technical leads.
 Implement and manage GitOps-based deployments using Argo CD and/or FluxCD.
 Implement, configure, administer, migrate, and upgrade Kubernetes clusters and workloads in collaboration with architects.
 Manage Kubernetes manifests, Helm charts, namespaces, ingress, storage, networking, RBAC, autoscaling, requests and limits, secrets, ConfigMaps, and workload health.
 Hands-on experience implementing and supporting Kubernetes Custom Resource Definitions (CRDs) and operators, including troubleshooting reconciliation, permissions, lifecycle, upgrade, and integration issues where required.
 Create, optimize, secure, and troubleshoot Docker images and containers, including multi-stage builds, startup failures, networking, volumes, resource constraints, and health checks.
 Work closely with development teams to resolve build, dependency, configuration, runtime, networking, and deployment issues across all environments.
 Perform advanced troubleshooting and root cause analysis for infrastructure, cloud, network, container, deployment, database-connectivity, and application-runtime incidents.
 Lead incident investigation, coordinate resolution across teams, document findings, and implement corrective and preventive actions.
 Implement Infrastructure as Code and configuration automation using Terraform, Ansible, and cloud-native provisioning tools.
 Develop reusable modules, deployment templates, automation scripts, and operational tooling using Shell/Bash and Python.
 Automate repetitive deployment, maintenance, monitoring, backup, reporting, access-management, and security activities wherever possible.
 Implement preventive maintenance, patching, upgrade planning, backup verification, restore testing, certificate renewal, capacity reviews, and infrastructure health checks.
 Manage secrets, credentials, certificates, access keys, and automated secret-rotation processes using appropriate cloud or platform services.
 Implement infrastructure security controls, vulnerability scanning, access governance, backup policies, disaster-recovery procedures, and compliance requirements.
 Configure application and infrastructure monitoring, centralized logging, dashboards, alerts, notification channels, and escalation workflows.
 Set up and manage custom alerts in AWS and Azure, including CloudWatch, Azure Monitor, Application Insights, and Log Analytics Workspaces.
 Monitor availability, performance, capacity, security posture, deployment health, and cloud cost across environments.
 Manage Linux-based production and non-production servers and troubleshoot system, process, disk, memory, package, service, and network issues.
 Support release planning, production deployments, rollback activities, planned maintenance, and incident response.
 Maintain architecture implementation documents, infrastructure diagrams, deployment guides, runbooks, SOPs, RCA reports, and disaster-recovery procedures.
 Provide periodic progress, risk, incident, capacity, and infrastructure-status reports to management, architects, customers, and project teams.
 Mentor DevOps engineers and review CI/CD pipelines, Dockerfiles, Terraform code, Ansible configurations, Kubernetes manifests, Helm charts, and operational procedures.
 Use AI-assisted tools responsibly for troubleshooting, scripting, automation, documentation, code review, and operational efficiency.
 Communicate technical risks, dependencies, constraints, and recommendations clearly and professionally in English.
 Be willing to work in the EST time zone or provide significant overlap with EST business hours.

Required Technical Skills

Linux and Troubleshooting
 Strong hands-on Linux administration experience, including services, processes, users, permissions, storage, networking, package management, SSH, systemd, cron, and log analysis.
 Excellent troubleshooting and root cause analysis skills across infrastructure and application environments.
 Ability to interpret application logs and runtime behavior for services developed in Python, PHP, Java, Node.js, Ruby, or similar technologies.

Containers and Kubernetes
 Strong hands-on experience with Docker, image optimization, secure image practices, multi-stage builds, Docker Compose, and container troubleshooting.
 Strong Kubernetes administration and implementation experience, including migrations, version upgrades, cluster operations, workload management, ingress, networking, storage, RBAC, autoscaling, and monitoring.
 Strong experience with Helm, along with hands-on Kubernetes CRD and operator experience, including deployment, configuration, troubleshooting, and upgrade support.
 Experience with managed Kubernetes platforms such as Amazon EKS, Azure Kubernetes Service, Google Kubernetes Engine, or Linode Kubernetes Engine.

CI/CD and GitOps
 Strong experience designing and optimizing CI/CD pipelines using GitLab CI/CD, GitHub Actions, Jenkins, Azure DevOps, Bitbucket Pipelines, or AWS CodePipeline.
 Hands-on experience with Argo CD and/or FluxCD.
 Knowledge of branching strategies, release management, approvals, artifact handling, security gates, deployment strategies, and rollback mechanisms.

Infrastructure as Code and Automation
 Strong hands-on experience with Terraform and Ansible.
 Strong Shell/Bash scripting skills and good Python automation skills.
 Experience creating reusable infrastructure modules, templates, and operational automation.
 Exposure to AWS CloudFormation is desirable.

Cloud Platforms
 AWS: EKS, EC2, ECS, ECR, RDS, S3, IAM, VPC, Route 53, ELB, CloudFront, CloudWatch, EventBridge, Systems Manager, Secrets Manager, CodeBuild, CodeDeploy, CodePipeline, and CloudFormation.
 Azure: Azure DevOps, AKS, Container Apps, Container Registry, Virtual Machines, Storage, Key Vault, Azure Monitor, Log Analytics Workspace, Application Insights, networking, Load Balancer, Application Gateway, Front Door, Managed Identities, and RBAC.
 Working knowledge of Google Cloud Platform services.
 Linode experience will be considered an added advantage.

Monitoring, Logging, Security, and Compliance
 Experience with Prometheus, Grafana, CloudWatch, Azure Monitor, Log Analytics Workspace, Application Insights, and centralized logging solutions.
 Experience configuring custom dashboards, log queries, alerts, notification channels, and operational metrics.
 Strong understanding of IAM, RBAC, least privilege, network security, secret management, certificate rotation, vulnerability management, backup, disaster recovery, and audit readiness.
 Exposure to SOC 2, ISO 27001, GDPR, or similar compliance frameworks.
 Experience with DevSecOps tools such as Trivy, SonarQube, OWASP ZAP, Gitleaks, Snyk, or Checkov is desirable.

Kubernetes Responsibility Boundary
The Senior DevOps Engineer is not expected to independently own the complete Kubernetes architecture.
The role will:
 Collaborate with architects to understand and refine the proposed Kubernetes design.
 Provide implementation, operability, security, scalability, and maintainability feedback during design discussions.
 Implement the approved architecture and manage the platform throughout its lifecycle.
 Handle cluster and workload configuration, migrations, upgrades, troubleshooting, documentation, and operational readiness.
 Escalate architectural trade-offs and final design decisions to the relevant solution or application architects.

Technologies We Are Looking For
 Cloud platforms: AWS and Microsoft Azure; exposure to Google Cloud Platform and Linode is advantageous.
 Containers and orchestration: Docker, Docker Compose, Kubernetes, Amazon EKS, Azure Kubernetes Service (AKS), Google Kubernetes Engine (GKE), and Linode Kubernetes Engine (LKE).
 Kubernetes ecosystem: Helm, Ingress controllers, RBAC, autoscaling, persistent storage, networking, Custom Resource Definitions (CRDs), and Kubernetes operators.
 CI/CD and GitOps: GitLab CI/CD, GitHub Actions, Jenkins, Azure DevOps Pipelines, AWS CodeBuild/CodeDeploy/CodePipeline, Argo CD, and FluxCD.
 Infrastructure as Code and configuration management: Terraform, Ansible, and AWS CloudFormation.
 Monitoring and observability: Prometheus, Grafana, AWS CloudWatch, Azure Monitor, Application Insights, Log Analytics Workspace, New Relic, Datadog, ELK/OpenSearch, and Loki.
 Security and DevSecOps: Trivy, SonarQube, OWASP ZAP, Gitleaks, Snyk, Checkov, cloud IAM/RBAC, secrets management, and certificate rotation.
 Source control and collaboration: Git, GitLab, GitHub, Bitbucket, and Jira.
 Web, database, and middleware technologies: Nginx, Apache, PostgreSQL, MySQL, MongoDB, and Redis.
 Scripting and application ecosystems: Shell/Bash, Python, and troubleshooting exposure to PHP, Java, Node.js, and Ruby applications.

Soft Skills and Preferred Qualifications
 Strong verbal and written English communication skills.
 Strong ownership, accountability, analytical thinking, and documentation practices.
 Ability to mentor engineers and collaborate effectively with architects, developers, security teams, management, and customers.
 Experience supporting production systems, customer-facing environments, and distributed teams.
 AWS and/or Microsoft Azure certifications are preferred. Relevant Kubernetes, Terraform, Linux, or DevOps

certifications will be an added advantage.
 Experience with multi-cloud environments, cost optimization, capacity planning, and infrastructure governance is preferred.

Skills Required

  • Strong hands-on Linux administration and troubleshooting experience
  • Excellent troubleshooting and root cause analysis across infrastructure and applications
  • Experience interpreting application logs and runtime behavior for Python, PHP, Java, Node.js, or Ruby
  • Hands-on Docker experience including image optimization and multi-stage builds
  • Strong Kubernetes administration and lifecycle management (migrations, upgrades, cluster ops, workload management)
  • Experience with Helm, Kubernetes manifests, CRDs, and operators
  • Experience with managed Kubernetes platforms (EKS, AKS, GKE)
  • Designing and optimizing CI/CD pipelines using GitLab CI/CD, GitHub Actions, Jenkins, Azure DevOps or similar
  • Hands-on experience with GitOps tools (Argo CD and/or FluxCD)
  • Infrastructure as Code using Terraform and configuration management with Ansible
  • Strong Shell/Bash scripting and Python automation skills
  • Practical experience with AWS services (EKS, EC2, ECS, ECR, RDS, S3, IAM, VPC, Route 53, ELB, CloudFront, CloudWatch, EventBridge, Systems Manager, Secrets Manager, CodeBuild/CodeDeploy/CodePipeline)
  • Practical experience with Microsoft Azure services (AKS, Azure DevOps, Container Registry, Virtual Machines, Storage, Key Vault, Azure Monitor, Log Analytics, Application Insights)
  • Experience configuring monitoring, centralized logging, dashboards, alerts (Prometheus, Grafana, CloudWatch, Azure Monitor, Log Analytics, Application Insights)
  • Strong understanding of IAM/RBAC, secrets management, vulnerability management, backup and disaster recovery practices
  • Experience managing production Linux servers and troubleshooting system/process/network issues
  • Ability to mentor engineers and review CI/CD pipelines, Dockerfiles, Terraform, Ansible, Kubernetes manifests and Helm charts
  • Willingness to work EST time zone or provide significant overlap with EST business hours
  • Exposure to AWS CloudFormation
  • Working knowledge of Google Cloud Platform
  • Linode experience
  • Exposure to SOC 2, ISO 27001, GDPR or similar compliance frameworks
  • Experience with DevSecOps tools (Trivy, SonarQube, OWASP ZAP, Gitleaks, Snyk, Checkov)
  • AWS and/or Microsoft Azure certifications or relevant Kubernetes/Terraform/Linux/DevOps certifications
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
9 Employees
Year Founded: 2009

What We Do

Benchmarkit helps B2B SaaS companies make metrics-informed decisions through industry benchmarks, research, media, and events. Formerly known as RevOps Squared, the company aims to enable B2B SaaS organizations to increase revenue growth efficiency and enterprise value by providing access to timely, contextual benchmarks and industry best practices across the entire customer journey, including customer acquisition, retention, and expansion.

Similar Jobs

Xplor Technologies Logo Xplor Technologies

Senior Devops Engineer

Cloud • Payments • Software
In-Office
Pune, Mahārāshtra, IND
3100 Employees
Hybrid
Pune, Mahārāshtra, IND
1135 Employees

NVIDIA Logo NVIDIA

Senior Devops Engineer

Artificial Intelligence • Computer Vision • Hardware • Robotics • Metaverse
In-Office or Remote
2 Locations
21960 Employees

Nue.io Logo Nue.io

Senior Devops Engineer

Information Technology • Software • Analytics
In-Office or Remote
2 Locations
175 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account