Culture at CloudifyOps :
About the Role :
We’re looking for someone who genuinely wants to understand why systems fail, not just respond to alerts. This role sits at the crossroads of cloud infrastructure and production reliability. You’ll own monitoring, handle on-call, and be the person who digs in when things go wrong. At the same time, we’re building an AI-powered pipeline monitoring tool and need someone curious enough to contribute to shaping it, not just watching over it.
What you’ll do:
Handle the on-call rotation and own incidents end-to-end triage, mitigation, escalation where needed, and clean resolution. You don’t pass the baton and disappear.
Write clear, structured RCAs after every significant incident what happened, when, why, and what changes going forward. These go to clients, so they need to work for both an engineer and a non-technical reader.
Maintain and improve the monitoring stack across environments dashboards, alerting rules, log pipelines, and distributed traces. Treat noisy alerts as a problem to fix, not something to mute.
Provision and manage cloud infrastructure on AWS using Terraform. This is a hands-on role not just reviewing what others set up.
Work with Kubernetes across multiple environments, debugging pod and node issues.
Monitor CI/CD pipeline health via Jenkins and support teams using Rancher for workload and cluster management.
Track application performance using APM tooling and JVM metrics: spot anomalies, investigate degradation, and flag systemic issues before they become incidents.
Contribute to the AI monitoring tool initiative: prototype, test, iterate. This is early-stage work and needs someone willing to figure things out, not just execute a finished design.
Tech Stack:
Cloud & Infrastructure : AWS · Kubernetes (K8s) · Terraform · Linux
Observability & Metrics : Prometheus · Grafana · APM (Datadog / New Relic / Kfuse) · JVM Metrics & GC Analysis · ELK / EFK Stack · Distributed Tracing
CI/CD & Platform : Jenkins · ArgoCD · Rancher · Git · Docker
Good to Have(Not Mandatory) : Python / Bash scripting · OpenTelemetry · Zenduty / OpsGenie · ML / AI basics
Expectations:
On-call here is real.Incidents happen outside business hours and when they do, it’s your responsibility to pick them up and drive them forward. That’s not unusual for this type of role but we want to be direct about it upfront.
Client expectations are high. You’ll produce RCAs, incident timelines, and status communications that clients read closely. Your writing needs to be clear, structured, and free of vagueness. “We investigated and fixed the issue” isn’t good enough. What was the issue, why did it happen, what was the business impact, and what prevents recurrence.
We expect precision regarding your own work. After a change, an incident, or a deployment, you should be able to clearly explain what you did and why without being prompted. Ownership doesn’t end when the alert clears.
Who we’re looking for:
The ideal candidate should have 2.5 years to 5 years of work experience.
Strong fundamentals. You understand how distributed systems actually behave under load, not just that a dashboard went red. You can read logs, metrics, and traces together.
Ownership without prompting. If you find a gap in monitoring coverage, you close it. If an RCA feels incomplete, you go back and make it precise. You don’t wait to be asked.
Writes clearly under pressure. During an incident, your updates should help not add noise. After one, your documentation should be good enough that anyone picking it up six months later understands what happened.
Curious about what comes next. The AI tooling initiative needs someone interested in figuring it out, not just waiting for a ticket. Some comfort with experimentation and ambiguity goes a long way here.
CloudifyOps is proud to be an equal opportunity employer with a global culture that embraces diversity. We are committed to providing an environment free of unfair discrimination and harassment. We do not discriminate based on age, race, color, sex, religion, national origin, disability, pregnancy, marital status, sexual orientation, gender reassignment, veteran status, or other protected category.
Skills Required
- 2.5 to 5 years of professional work experience
- Strong understanding of distributed systems under load
- Ability to interpret logs, metrics, and distributed traces together
- Experience with cloud infrastructure, AWS, Terraform, and Kubernetes
- Ability to participate in an on-call rotation and manage incidents end-to-end
- Strong monitoring, observability, and application performance troubleshooting skills
- Clear technical writing for RCAs, incident timelines, and client communications
- Ownership, initiative, and comfort working with ambiguity
- Python or Bash scripting experience
- OpenTelemetry experience
- Zenduty or OpsGenie experience
- Machine learning or artificial intelligence fundamentals
What We Do
CloudifyOps is a Cloud technology firm enabling businesses to maximise profitability, become more agile and innovative through our comprehensive portfolio of Cloud transformation, DevOps consulting and Managed IT Services. We are a proud Advanced Consulting Partner of Amazon Web Services and have deep expertise in Microsoft Azure and Google Cloud Platform solutions. We are driven by passion to deliver extraordinary value for our customers. We assure to deliver reduced TCO, greater IT-business alignment and higher SLAs. Our Service Offerings: Cloud Services - We make your transformation to Cloud effortless and efficient : CloudifyOps delivers engineering services to support Cloud Infrastructure as a Service (IaaS). Once you identify your business requirements, we bring in our comprehensive expertise in Analysis, Design, Deployment and End-to-End Support of the Cloud solution. DevOps Services - You build your application and we help you deliver to the market FASTER : At CloudifyOps, we do a thorough strategy assessment, create a pilot framework and a complete tool stack construction aligned to your application/product. In simple words - End-to-End DevOps Implementation. Site Reliability Engineering (SRE) - We help you keep your revenue-critical systems up and running and achieve scalability and reliability through continuous improvement and automation : The CloudifyOps team will establish the current state, set up an SRE backlog for proactive improvement, cross-pollinate learnings and help you scale up to reliable and scalable processes that are fully automated and data driven. Managed IT Services - We handle your IT while you continue to grow your business : We make sure your business can keep generating revenue even when an emergency strikes (High Availability & Disaster Recovery Management). Data Capability Platforms, Accelerators and Frameworks







