Senior DevOps/Cloud Platform Engineer (AWS| Kubernetes|AI Infrastructure)

Posted Yesterday
Be an Early Applicant
Indore, Madhya Pradesh, IND
In-Office
Senior level
Agency • Artificial Intelligence • Information Technology • Professional Services
The Role
Design, deploy, and manage secure, highly available AWS and Amazon EKS platforms across multiple environments. Build CI/CD and infrastructure-as-code automation, optimize Kubernetes and database performance, support GPU-based LLM inference, and implement monitoring, disaster recovery, security, compliance, and cloud cost optimization. Collaborate with software, AI/ML, architecture, and security teams to operate resilient enterprise-scale platforms.
Summary Generated by Built In
About the Role

We are looking for a highly skilled Senior DevOps / Cloud Platform Engineer with 5–8 years of experience in designing, deploying, and managing secure, scalable, and highly available cloud infrastructure on AWS.

The ideal candidate will have extensive hands-on experience with AWS, Amazon EKS, Kubernetes, CI/CD automation, infrastructure as code, and cloud security. This role also requires experience supporting AI workloads, including deploying and optimizing Large Language Models (LLMs) on CPU and GPU infrastructure.

You will work closely with software engineers, AI/ML engineers, architects, and security teams to build and maintain cloud platforms that are secure, resilient, cost-efficient, and compliant with SOC 2 and HITRUST requirements.


Key Responsibilities
Cloud Infrastructure
  • Design, deploy, and manage scalable, highly available, and secure AWS infrastructure.
  • Architect cloud environments capable of supporting enterprise-scale applications and AI workloads.
  • Optimize infrastructure for performance, reliability, scalability, and cost.
  • Implement high availability, disaster recovery, backup, and failover strategies.
  • Design multi-environment infrastructure (Development, QA, UAT, Production).

Kubernetes & Container Platform
  • Design, deploy, and manage production-grade Kubernetes clusters using Amazon EKS.
  • Optimize Kubernetes workloads for high availability and resource utilization.
  • Configure namespaces, RBAC, network policies, autoscaling, ingress controllers, and service meshes where applicable.
  • Troubleshoot Kubernetes networking, scheduling, storage, and performance issues.
  • Manage rolling deployments, blue-green deployments, and canary releases.

AWS Services

Strong hands-on experience with:

  • Amazon EKS
  • Amazon EC2
  • Auto Scaling Groups
  • Elastic Load Balancer (ALB/NLB)
  • Amazon S3
  • Amazon RDS
  • AWS Lambda
  • Amazon ECR
  • Amazon CloudWatch
  • IAM
  • Route 53
  • VPC
  • NAT Gateway
  • Security Groups
  • AWS WAF
  • AWS Secrets Manager
  • Systems Manager (SSM)
  • CloudFront
  • EventBridge
  • SNS
  • SQS

CI/CD & DevOps Automation
  • Design and implement end-to-end CI/CD pipelines.
  • Automate application deployments across multiple environments.
  • Implement infrastructure automation and GitOps practices.
  • Build deployment strategies with minimal downtime.
  • Integrate automated testing, security scanning, and quality gates into CI/CD pipelines.

Experience with:

  • GitHub Actions
  • Jenkins
  • GitLab CI
  • ArgoCD

Infrastructure as Code

Develop and manage infrastructure using:

  • Terraform
  • AWS CloudFormation
  • Kubernetes YAML

AI & LLM Infrastructure
  • Deploy and manage Small Language Models (SLMs) and Large Language Models (LLMs) in production environments.
  • Build scalable inference infrastructure for AI workloads.
  • Configure GPU-enabled Kubernetes nodes for model serving.
  • Optimize CPU and GPU utilization for AI inference.
  • Manage model deployments, scaling, versioning, and monitoring.
  • Support vector databases and AI inference services.
  • Work closely with AI/ML engineers to optimize model performance and infrastructure costs.

Database Infrastructure & Performance
  • Deploy and manage Amazon RDS databases.
  • Monitor and optimize database performance.
  • Implement backup, recovery, and replication strategies.
  • Tune database configurations for high-throughput applications.
  • Monitor slow queries, indexing strategies, and connection pooling.
  • Collaborate with engineering teams on database performance optimization.

Monitoring & Observability

Implement monitoring and observability using:

  • CloudWatch
  • Prometheus
  • Grafana
  • ELK / OpenSearch
  • Loki

Responsibilities include:

  • Infrastructure monitoring
  • Application monitoring
  • Log aggregation
  • Alerting
  • Capacity planning
  • Incident response

Security & Compliance
  • Implement AWS security best practices.
  • Design secure IAM policies and access controls.
  • Manage secrets and encryption.
  • Perform infrastructure hardening.
  • Ensure compliance with:
    • SOC 2
    • HITRUST
    • HIPAA
  • Participate in security audits and vulnerability remediation.
  • Maintain audit logs and infrastructure documentation.

Cost Optimization
  • Continuously optimize AWS infrastructure costs.
  • Right-size EC2 instances and EKS node groups.
  • Optimize storage and networking costs.
  • Implement Savings Plans and Reserved Instances where appropriate.
  • Optimize GPU utilization for AI workloads.
  • Monitor cloud spending and recommend cost-saving initiatives.


RequirementsRequired Qualifications
  • Bachelor's or Master's degree in Computer Science, Information Technology, or a related field.
  • 5–8 years of hands-on experience in DevOps, Cloud Engineering, or Platform Engineering.
  • Strong experience designing and managing production AWS environments.
  • Extensive experience with Kubernetes and Amazon EKS.
  • Experience managing enterprise-scale cloud infrastructure.
  • Proven experience automating deployments and infrastructure management.
Required Technical Skills
Cloud Platforms
  • Amazon Web Services (AWS)
AWS Services
  • Amazon EC2
  • Amazon EKS
  • Amazon ECS
  • Amazon RDS
  • Amazon S3
  • Lambda
  • ECR
  • CloudFront
  • IAM
  • Route 53
  • VPC
  • CloudWatch
  • Systems Manager
  • WAF
  • Secrets Manager
  • SNS
  • SQS
  • EventBridge
Containers & Orchestration
  • Docker
  • Kubernetes
  • Amazon EKS
  • Helm
  • Kubernetes Networking
  • Ingress Controllers
  • Horizontal & Vertical Pod Autoscaling
Infrastructure as Code
  • Terraform
  • CloudFormation
  • Helm
  • Customize
CI/CD
  • GitHub Actions
  • Jenkins
  • GitLab CI
  • ArgoCD
Databases
  • Amazon RDS
  • PostgreSQL
  • MySQL
  • Redis

Experience with:

  • Performance tuning
  • Replication
  • Backup & recovery
  • Connection pooling
  • Query optimization
AI Infrastructure

Experience deploying and managing:

  • LLMs and SLMs
  • GPU-based inference workloads
  • NVIDIA GPU infrastructure
  • CUDA-enabled environments (preferred)
  • Hugging Face models
  • vLLM, Ollama, or similar inference frameworks
  • Model serving and autoscaling
Monitoring & Logging
  • Prometheus
  • Grafana
  • CloudWatch
  • ELK/OpenSearch
  • Loki
Security & Compliance

Strong understanding of:

  • SOC 2
  • HITRUST
  • HIPAA
  • IAM
  • RBAC
  • Network Security
  • Encryption
  • Secrets Management
  • Vulnerability Management
Preferred Qualifications
  • AWS Certified Solutions Architect – Professional or Associate.
  • AWS Certified DevOps Engineer – Professional.
  • Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD).
  • Experience with AI platforms, MLOps, or GPU infrastructure.
  • Experience deploying high-availability, multi-tenant SaaS applications.
  • Familiarity with service mesh technologies (Istio or Linkerd) is a plus.


Skills Required

  • Bachelor's or Master's degree in Computer Science, Information Technology, or a related field
  • 5-8 years of hands-on experience in DevOps, Cloud Engineering, or Platform Engineering
  • Strong experience designing and managing production AWS environments
  • Extensive experience with Kubernetes and Amazon EKS
  • Experience managing enterprise-scale cloud infrastructure
  • Proven experience automating deployments and infrastructure management
  • Experience with AWS, Amazon EC2, EKS, ECS, RDS, S3, Lambda, ECR, CloudFront, IAM, Route 53, VPC, CloudWatch, Systems Manager, WAF, Secrets Manager, SNS, SQS, and EventBridge
  • Experience with Docker, Kubernetes networking, ingress controllers, horizontal and vertical pod autoscaling, Helm, Terraform, CloudFormation, Kustomize, GitHub Actions, Jenkins, GitLab CI, and ArgoCD
  • Experience with PostgreSQL, MySQL, Redis, database performance tuning, replication, backup and recovery, connection pooling, and query optimization
  • Experience deploying and managing LLMs, SLMs, GPU-based inference workloads, NVIDIA GPU infrastructure, model serving, and autoscaling
  • Experience with CUDA-enabled environments
  • Experience with Hugging Face models and vLLM, Ollama, or similar inference frameworks
  • Experience with Prometheus, Grafana, CloudWatch, ELK or OpenSearch, and Loki
  • Strong understanding of SOC 2, HITRUST, HIPAA, IAM, RBAC, network security, encryption, secrets management, and vulnerability management
  • AWS Certified Solutions Architect certification
  • AWS Certified DevOps Engineer Professional certification
  • Certified Kubernetes Administrator or Certified Kubernetes Application Developer certification
  • Experience with AI platforms, MLOps, or GPU infrastructure
  • Experience deploying high-availability, multi-tenant SaaS applications
  • Familiarity with Istio or Linkerd service mesh technologies
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
1,220 Employees
Year Founded: 2016

What We Do

Annova Solutions is an India-based outsourcing and offshoring services company providing healthcare services, AI/ML operations, and end-to-end human-resources support. Founded in 2016 and headquartered in Indore, it delivers business-process services through dedicated, industry-specific teams. Its recruiting listing identifies the employer's operating field as BPO/ITES, while its broader service profile combines technology-enabled operations with specialized healthcare and workforce-support capabilities for client organizations.

Similar Jobs

MongoDB Logo MongoDB

Technical Director, Builder Relations

Big Data • Cloud • Software • Database
Easy Apply
Remote or Hybrid
India
5550 Employees

Capco Logo Capco

Product Manager

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

Capco Logo Capco

Visualisation Analyst (Liquidity & Treasury)

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

CSC Logo CSC

Security Engineer

Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Remote or Hybrid
2 Locations
8500 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account