Senior DevOps Engineer

Posted 6 Days Ago
Be an Early Applicant
Gurugram, Haryana, IND
In-Office
Senior level
Cloud • Information Technology • Consulting • Automation
The Role
Own and optimize multiple client cloud environments across AWS, Azure, and/or GCP. Design and operate Kubernetes platforms, infrastructure as code, CI/CD and GitOps workflows, monitoring, security, disaster recovery, and cost optimization practices. Lead production troubleshooting, incident response, architecture decisions, and reliability improvements. Provide technical leadership and mentorship to 3–5 engineers while managing client relationships, technical escalations, and stakeholder communication.
Summary Generated by Built In
AWS / Azure / GCP | Multi-Client Ownership | Technical Leadership

About the Role
We are looking for a Senior DevOps Engineer with 5–8 years of strong hands-on production experience to independently own multiple client environments and provide technical leadership to a team of engineers.
In this role, you will typically manage 2–3 client environments, drive cloud and infrastructure architecture decisions, troubleshoot complex production challenges, and lead initiatives across reliability, security, automation, performance, and cost optimization.
You will also provide technical direction and mentorship to 3–5 engineers, ensuring high-quality delivery, strong engineering practices, and effective execution.
The role requires strong architectural thinking, engineering judgment, production troubleshooting skills, and the ability to communicate effectively with both technical teams and client stakeholders.
Key Responsibilities1. Cloud Infrastructure & Architecture
  • Design, deploy, manage, and optimize production cloud environments across AWS, Azure, and/or GCP.
  • Design and review highly available, scalable, secure, and cost-efficient cloud architectures.
  • Work extensively with core cloud services including compute, storage, databases, networking, IAM, and load balancers.
  • Own production workloads on EKS, AKS, GKE, including cluster design, upgrades, scaling, troubleshooting, and maintenance.
  • Evaluate architectural options and recommend solutions based on client requirements, technical constraints, and business objectives.
  • Identify technical debt, architectural risks, and opportunities to improve reliability, scalability, security, and operational efficiency.
  • Implement and improve backup, disaster recovery, high availability, and business continuity practices.
2. Kubernetes, Infrastructure as Code & Automation
  • Design and operate production Kubernetes environments, addressing challenges related to networking, scheduling, scaling, resource utilization, availability, and application behavior.
  • Establish reusable Kubernetes and infrastructure patterns across client environments.
  • Develop and maintain infrastructure using Terraform and Infrastructure as Code best practices.
  • Build reusable IaC modules and standards to improve consistency, scalability, and operational reliability.
  • Design and implement scalable CI/CD and GitOps workflows using tools such as ArgoCD, Flux, Spinnaker, or similar platforms.
  • Automate operational processes using Bash, Python, and other appropriate scripting or automation tools.
  • Identify and eliminate repetitive operational work through automation and engineering improvements.
3. Reliability, Monitoring & Production Operations
  • Own production reliability and operational excellence across assigned client environments.
  • Lead troubleshooting of complex infrastructure, Kubernetes, networking, and application performance issues.
  • Configure and improve monitoring, logging, alerting, and observability using tools such as Prometheus, Grafana, Coralogix, New Relic, Datadog, CloudWatch, or equivalent platforms.
  • Define appropriate SLIs, SLOs, dashboards, and alerting standards for production environments.
  • Lead production incident response, root cause analysis, and preventive remediation initiatives.
  • Identify systemic reliability risks and drive improvements to prevent recurring incidents.
4. Security & Compliance
  • Implement cloud security best practices across infrastructure and production environments.
  • Apply principles of IAM, RBAC, least privilege, secrets management, vulnerability management, and OS hardening.
  • Identify security gaps and incorporate security considerations into infrastructure and architecture decisions.
  • Work with engineering teams to improve the overall security posture of client environments.
5. Cloud Cost & Performance Optimization
  • Drive cloud cost and resource optimization initiatives across client environments.
  • Identify cloud waste through right-sizing, capacity planning, workload optimization, and appropriate cloud pricing models.
  • Balance cost, performance, reliability, security, and scalability when making technical decisions.
  • Contribute to FinOps practices and help clients achieve sustainable cloud cost efficiency.
6. Client & Stakeholder Leadership
  • Own the technical relationship for approximately 2–3 client environments.
  • Lead technical discussions with client engineering teams, architects, and leadership stakeholders.
  • Translate business requirements into practical, scalable, and maintainable technical solutions.
  • Present architecture recommendations, technical risks, trade-offs, and improvement plans to clients.
  • Handle technical escalations and production incidents with clear ownership and proactive communication.
  • Challenge inefficient or technically risky approaches and recommend better alternatives.
  • Build client trust through technical credibility, effective communication, and predictable delivery.
7. Team Leadership & Mentoring
  • Provide technical leadership to a team of 3–5 engineers across assigned client environments.
  • Plan and delegate work based on technical capability, priorities, client requirements, and business impact.
  • Review technical implementations and ensure adherence to engineering standards and best practices.
  • Mentor engineers and actively contribute to their technical development.
  • Identify capability gaps and create opportunities for knowledge sharing and skill development.
  • Provide technical guidance during complex incidents, design discussions, and implementation challenges.
  • Maintain accountability for technical quality and delivery standards.
Engineering Expectations
  • Make independent and well-reasoned technical decisions in complex or ambiguous production situations.
  • Evaluate trade-offs across reliability, scalability, security, performance, cost, and complexity.
  • Demonstrate strong ownership of production systems and follow issues through to resolution.
  • Establish and continuously improve engineering and operational standards.
  • Proactively identify technical debt, operational risks, and improvement opportunities.
  • Approach problems with a structured, analytical, and solution-oriented mindset.
Required Skills & ExperienceMust Have
  • 5–8 years of hands-on experience in DevOps, Cloud Engineering, SRE, or a similar role.
  • Strong production experience with AWS, Azure, and/or GCP.
  • Strong hands-on experience with Kubernetes, preferably EKS, AKS, or GKE.
  • Strong experience with Terraform and Infrastructure as Code.
  • Strong understanding of CI/CD and GitOps practices.
  • Strong Linux and networking fundamentals.
  • Proven experience troubleshooting complex production environments and handling incidents.
  • Strong scripting and automation skills using Bash, Python, or equivalent.
  • Good understanding of cloud security fundamentals.
  • Experience designing and reviewing production architectures.
  • Demonstrated ownership of production environments and technical initiatives.
  • Experience mentoring or providing technical leadership to engineers.
  • Strong written and verbal communication skills.
  • Experience working with multiple clients, products, or production environments simultaneously.
What Success Looks Like
Within the role, you will be expected to:
  • Independently own 2–3 client environments and their technical outcomes.
  • Lead and mentor 3–5 engineers while maintaining high engineering standards.
  • Resolve complex production incidents and technical escalations effectively.
  • Review and continuously improve client architectures.
  • Make sound technical decisions with minimal supervision.
  • Drive measurable improvements in reliability, security, automation, performance, and cloud cost.
  • Lead technical conversations confidently with client stakeholders.
  • Identify risks and improvement opportunities proactively rather than responding only to incidents.
  • Raise the technical maturity and engineering capabilities of the team.

Skills Required

  • 5–8 years of hands-on experience in DevOps, Cloud Engineering, SRE, or a similar role
  • Strong production experience with AWS, Azure, and/or GCP
  • Strong hands-on experience with Kubernetes, preferably EKS, AKS, or GKE
  • Strong experience with Terraform and Infrastructure as Code
  • Strong understanding of CI/CD and GitOps practices
  • Strong Linux and networking fundamentals
  • Experience troubleshooting complex production environments and handling incidents
  • Strong scripting and automation skills using Bash, Python, or equivalent
  • Good understanding of cloud security fundamentals
  • Experience designing and reviewing production architectures
  • Demonstrated ownership of production environments and technical initiatives
  • Experience mentoring or providing technical leadership to engineers
  • Strong written and verbal communication skills
  • Experience working with multiple clients, products, or production environments simultaneously
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
39 Employees

What We Do

Infra360 is a fast-growing cloud, DevOps, and infrastructure services company headquartered in Gurugram (Gurgaon), India, that helps startups and enterprises modernize, secure, and scale their platforms across AWS, Azure, and GCP. Founded in 2022 by Deepak Agrawal, it provides customer-centric cloud consulting for technology-driven organizations, focusing on cloud excellence and business resilience, and is recognized as a DPIIT-recognized startup.

Similar Jobs

Citi Logo Citi

Senior Devops Engineer

Fintech • Financial Services
In-Office
DLF Cybercity, Gurugram, Haryana, IND
223850 Employees

Ericsson Logo Ericsson

Devops Engineer

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
5 Locations
88000 Employees

PAR Technology Logo PAR Technology

Senior Devops Engineer

Food • Software • Hospitality
Hybrid
Gurugram, Haryana, IND
2000 Employees
In-Office
Gurugram, Haryana, IND
215 Employees

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account