The Role
Build and scale reliable AWS cloud platforms through infrastructure design, Infrastructure-as-Code, CI/CD automation, observability, incident response, and production debugging. The role improves system resilience, scalability, security, cost efficiency, and operational maturity using scripting and reusable tooling. Responsibilities include developing deployment and recovery workflows, supporting monitoring and alerting, collaborating with engineering teams, and modernizing existing systems while preserving production stability.
Summary Generated by Built In
We are looking for a SRE / DevOps Engineer to build and scale enterprise-grade cloud platforms. This is a balanced role (70% engineering, 30% operations) focused on:
- building reliable and scalable cloud infrastructure,
- driving automation and platform engineering,
- improving observability and operational maturity,
- enabling resilient production systems.
This role is ideal for engineers who are curious about how systems behave in production, enjoy debugging and automation, and want to grow into strong Site Reliability Engineers over time.
Work closely with senior engineers and platform teams to improve reliability, scalability, deployment workflows, and production operations across enterprise-grade cloud platforms.
Key Responsibilities
- Design and implement scalable AWS infrastructure for production systems
- Build Infrastructure-as-Code modules for consistent and reproducible environments
- Develop and maintain CI/CD pipelines for deployment, testing, and validation
Build automation for: deployment workflows, system health verification, smoke testing, recovery validation and operational efficiency
- Contribute to monitoring, alerting, logging, and observability systems
- Participate in production issue debugging, incident response, and system stability improvements
- Collaborate across engineering teams to improve platform capabilities and operational maturity
- Contribute to scalable and cost-aware cloud infrastructure design alongside senior engineers
- Improve reliability and reduce operational toil through scripting, automation, and reusable tooling
- Work on secure, resilient, and highly available cloud environments
- Support modernization and improvement of existing systems without disrupting production stability
Requirements
- 2–4 years of experience in DevOps / SRE / Cloud Engineering roles
Strong hands-on experience with:
AWS production environments
Infrastructure-as-Code (Terraform or CloudFormation)
CI/CD pipelines (Jenkins, GitHub Actions, or similar)
- Strong scripting/programming skills in Python or Bash (must)
- Proven experience with debugging production issues, improving system stability, automating infrastructure or operational workflows, and working with cloud-native systems.
- Good understanding of distributed systems, cloud architecture, observability, scalability, and cost-aware infrastructure practices.
- Familiarity with containerized environments such as Docker, Kubernetes, or ECS
- Strong problem-solving mindset with willingness to learn and take ownership
- Good communication and collaboration skills
Tech Stack
- AWS (RDS, Lambda, EventBridge, ECS/Kubernetes, CloudWatch, IAM, VPC)
- Terraform / CloudFormation (IaC)
- CI/CD: Jenkins, GitHub Actions
- Observability: CloudWatch, Prometheus, Grafana
- Scripting/Development: Python, Bash (Node.js a plus)
- Chaos Engineering tools (AWS FIS, Gremlin, etc.) are good to have, not mandatory
Good to Have
- Exposure to production incident handling or on-call support
- Experience with Kubernetes or ECS
- Exposure to monitoring, alerting, and observability tooling
- Basic understanding of reliability engineering concepts
- Exposure to database operations, backup/recovery, or disaster recovery concepts
- Background in backend engineering before moving to DevOps/SRE
- Curiosity toward automation, reliability engineering, and platform scalability
Benefits
- Opportunity to work on large-scale cloud platforms and mission-critical systems
- Work closely with experienced SRE and platform engineering teams
- Exposure to advanced areas such as Digital Twin, AI/ML systems, and cloud-native architectures
- Opportunity to grow into reliability engineering and platform ownership roles
- Work with a collaborative and engineering-focused team culture
- Be part of a company passionate about solving real engineering problems through technology
Skills Required
- 2-4 years of experience in DevOps, SRE, or Cloud Engineering roles
- Strong hands-on experience with AWS production environments
- Strong hands-on experience with Infrastructure-as-Code using Terraform or CloudFormation
- Strong hands-on experience with CI/CD pipelines such as Jenkins or GitHub Actions
- Strong scripting or programming skills in Python or Bash
- Experience debugging production issues and improving system stability
- Experience automating infrastructure or operational workflows
- Experience working with cloud-native systems
- Understanding of distributed systems, cloud architecture, observability, scalability, and cost-aware infrastructure practices
- Familiarity with containerized environments such as Docker, Kubernetes, or ECS
- Strong problem-solving mindset and willingness to learn and take ownership
- Good communication and collaboration skills
- Exposure to production incident handling or on-call support
- Exposure to monitoring, alerting, and observability tooling
- Exposure to database operations, backup and recovery, or disaster recovery concepts
- Background in backend engineering
- Experience with Kubernetes or ECS
- Experience with chaos engineering tools such as AWS FIS or Gremlin
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
CCTech, the Centre for Computational Technologies, is a digital transformation company developing CAD, CFD, artificial intelligence, machine learning, 3D web, augmented reality, digital twin, and enterprise applications. Its product division includes simulationHub, a cloud-based CFD platform, while its consulting division helps engineering organizations with computational engineering, CAD/CFD software development, digital transformation, and related technology services.







