As a Junior SRE/DevOps Engineer, you will support the setup and maintenance of Infra, CI/CD and help sustain different key projects used throughout globally, working under the guidance of senior engineers.
As a member of our geographically distributed development team your communication and analytical skills are essential to the role.
Key Responsibilities:
· Assist in designing and maintaining cloud infrastructure that is secure, scalable, and highly available on AWS/Azure
· Work collaboratively with software engineering to define infrastructure and deployment requirements
· Provision, configure and maintain AWS cloud infrastructure defined as code.
· Containerization using Docker and Kubernetes
· Troubleshoot problems across a wide array of services and functional areas
· Build and maintain operational tools for deployment, monitoring, and analysis of AWS infrastructure and systems
· Assist with infrastructure cost analysis and support optimization efforts.
· Support the development of self-healing and automated remediation mechanisms using AI/ML techniques
· Assist in integrating AI/LLM capabilities into DevOps workflows (e.g., log analysis, automated RCA, deployment insights)
· Support monitoring strategy enhancements by leveraging intelligent alerting, noise reduction, and pattern-based anomaly detection across logs, metrics, and traces.
· Assist in building and maintaining MLOps pipelines for model training, deployment, and continuous improvement.
· Collaborate with a global team of engineers in a highly agile DevOps environment, focused on efficient operation of daily activities, developer productivity and continuous improvement of the framework.
· Support the development, implementation, and maintenance of CI/CD frameworks, and contribute to tools development for hybrid environments (Cloud, On premise) with a vision to achieve “CI/CD” objectives for large-scale integration of systems in order to reduce manual build and deploy efforts.
· Work with geographically dispersed teams including multi-vendor into Scrum teams to meet “CI/CD”
Required Knowledge & Skills:
· 2-4 years of experience building and maintaining AWS infrastructure (VPC, EC2, Security Groups, IAM, ECS/EKS, CloudFront, S3, RDS, SQS, SNS, Lambda Function, Batch jobs, AWS Glue)
· Working understanding of how to secure AWS environments and meet compliance requirements
· Working knowledge of deploying and managing infrastructure with Terraform.
· Exposure to or working knowledge of LLMs (OpenAI, Azure OpenAI, Claude etc.)
· Basic awareness of LLMOps concepts (prompt management, model evaluation, versioning, fine-tuning lifecycle)
· Familiarity with MLOps tools such as MLflow, SageMaker, Kubeflow or equivalent.
· Familiarity with AIOps platforms/tools for intelligent monitoring and incident management.
· Ability to apply AI techniques to improve deployment speed, reliability, and monitoring effectiveness.
· Working experience on windows & Linux based environments.
· Experience with Docker, GitHub, Jenkins, Azure DevOps, ELK and deploying applications on AWS.
· Good command on scripting languages like Python, Bash/Shell, Powershell etc
· Knowledge in log analytics tools like Elastic search and Kibana.
· Basic knowledge of Cloud Migration/Disaster Recovery/Blue Green Deployment implementation.
· Good understanding about monitoring the services and alerting using Cloudwatch, Datadog, Prometheus or Azure monitor.
· Good to hire a candidate with certification
Personal Attributes:
· Very good communication skills.
· Ability to easily fit into a distributed development team.
· Customer service oriented.
· Enthusiastic/High initiative.
· Ability to manage timelines of multiple initiatives.
· Very good attention to detail and the ability to always follow up.
Skills Required
- 2-4 years building and maintaining AWS infrastructure (VPC, EC2, Security Groups, IAM, ECS/EKS, CloudFront, S3, RDS, SQS, SNS, Lambda, Batch, AWS Glue)
- Working understanding of securing AWS environments and meeting compliance requirements
- Working knowledge of deploying and managing infrastructure with Terraform
- Exposure to or working knowledge of LLMs (OpenAI, Azure OpenAI, Claude)
- Basic awareness of LLMOps concepts (prompt management, model evaluation, versioning, fine-tuning lifecycle)
- Familiarity with MLOps tools such as MLflow, SageMaker, Kubeflow or equivalent
- Familiarity with AIOps platforms/tools for intelligent monitoring and incident management
- Experience with Docker, Kubernetes and deploying applications on AWS
- Experience with CI/CD tools: GitHub, Jenkins, Azure DevOps and building/maintaining CI/CD frameworks
- Working experience on Windows and Linux environments
- Proficient scripting with Python, Bash/Shell, PowerShell
- Knowledge of log analytics tools (Elasticsearch, Kibana) and ELK stack
- Basic knowledge of Cloud Migration, Disaster Recovery, and Blue-Green Deployment implementation
- Monitoring and alerting knowledge using CloudWatch, Datadog, Prometheus or Azure Monitor
- Certification (cloud/SRE/DevOps)
What We Do
AlgoLeap specializes in AI-powered software solutions, digital product engineering, and IT consulting services, focusing on digital transformation and AI-driven innovation.






