The Role
Design, implement, and maintain scalable AWS infrastructure; automate with Terraform/CloudFormation/CDK; build CI/CD and GitOps pipelines; manage Kubernetes/Docker clusters; implement monitoring with Prometheus/Grafana/OpenTelemetry; troubleshoot production incidents, perform RCA, and improve reliability; script automation in Python/Go/Shell; support HA, DR, capacity planning, and maintain runbooks.
Summary Generated by Built In
This role is for one of the Weekday's clients
Min Experience: 6+ years
Location: Chennai
JobType: full-time
RequirementsKey Responsibilities
- Design, implement, and maintain scalable AWS cloud infrastructure.
- Manage and automate infrastructure using Terraform, CloudFormation, or AWS CDK.
- Build and maintain CI/CD pipelines using GitHub Actions, Jenkins, and ArgoCD.
- Manage Kubernetes and Docker environments and troubleshoot deployments.
- Implement monitoring and observability using Prometheus, Grafana, and OpenTelemetry.
- Troubleshoot production issues, perform RCA, and improve system reliability and performance.
- Automate operational processes using Python, Go, or Shell scripting.
- Work with engineering teams to improve deployment, scalability, availability, and operational efficiency.
- Support high-availability, disaster recovery, and capacity planning initiatives.
- Maintain infrastructure documentation and operational runbooks.
- Collaborate with development, QA, and architecture teams to ensure smooth and reliable releases.
- 6–9 years of hands-on experience in DevOps, Cloud, or SRE.
- Strong hands-on experience with AWS, Terraform, Kubernetes, and Docker.
- Experience with ArgoCD / GitOps and CI/CD tools such as GitHub Actions or Jenkins.
- Strong knowledge of Linux, networking, cloud infrastructure, and microservices.
- Experience with Prometheus, Grafana, and OpenTelemetry.
- Proficiency in Python, Go, or Shell scripting
- Strong troubleshooting, incident management, RCA, and production support experience.
- Good understanding of scalability, high availability, reliability, and distributed systems.
- Working knowledge of PostgreSQL, MongoDB, or other RDBMS/NoSQL databases.
- Good understanding of REST, gRPC, API Gateway, and Load Balancing.
DevOps, AWS, Argocd
Good-to-have skillsContainers, CI/CD, Scripting
Skills Required
- 6-9 years of hands-on experience in DevOps, Cloud, or SRE
- Strong hands-on experience with AWS, Terraform, Kubernetes, and Docker
- Experience with ArgoCD / GitOps and CI/CD tools such as GitHub Actions or Jenkins
- Experience managing and automating infrastructure using CloudFormation or AWS CDK
- Experience implementing monitoring and observability using Prometheus, Grafana, and OpenTelemetry
- Proficiency in Python, Go, or Shell scripting
- Strong knowledge of Linux, networking, cloud infrastructure, and microservices
- Strong troubleshooting, incident management, RCA, and production support experience
- Good understanding of scalability, high availability, reliability, and distributed systems
- Working knowledge of PostgreSQL, MongoDB, or other RDBMS/NoSQL databases
- Good understanding of REST, gRPC, API Gateway, and Load Balancing
- Must-have skills: DevOps, AWS, ArgoCD
- Good-to-have skills: Containers, CI/CD, Scripting
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Weekday is an AI-powered recruitment platform that helps startups hire top-tier engineering and product talent. By leveraging a massive database of white-collar professionals and advanced outreach tools, the company streamlines the hiring process through automated sourcing, AI-driven resume screening, and white-glove contingency services. Their mission is to modernize recruitment by enabling companies to discover and engage passive candidates efficiently, ensuring high-quality hires for critical roles.

_0.png)







