Platform Engineer

Reposted 2 Days Ago
Be an Early Applicant
3 Locations
In-Office
Senior level
Artificial Intelligence • Big Data • Machine Learning
The Role
Maintain platform stability through Tier-1 incident resolution, system monitoring, alerting, runbook execution, SLA adherence, reporting, and escalation. Build and maintain scalable Kubernetes and container infrastructure, CI/CD pipelines, deployment strategies, health checks, and observability. Develop production-grade Python services using Flask and Gunicorn, integrating databases and streaming systems such as Redis and Kafka. Support Linux environments, video-processing workloads, scheduled maintenance, and rotational operations shifts.
Summary Generated by Built In

While technology is the heart of our business, a global and diverse culture is the heart of our success. We love our people and we take pride in catering them to a culture built on transparency, diversity, integrity, learning and growth.
If working in an environment that encourages you to innovate and excel, not just in professional but personal life, interests you- you would enjoy your career with Quantiphi!

Role: Platform Engineer (Junior/Senior)

Experience Level: 1 to 6 Years 

Work location: Remote (Base Location - Mumbai, Bangalore & Trivandrum) Please apply if you are comfortable with rotational shifts and can join immediately.

Role & Responsibilities: 

  • Work with a highly responsive, proactive team to provide support for maintaining system stability with minimal issues

  • Thrive to consistently achieve service levels through focus, metrics, and capacity management

  • Utilize Best practices/Tools for SLA reporting 

  • Minimize people dependency and move the projects to autopilot mode

  • Have consistent and complete communication and collaboration with all partners

Daily Responsibilities:

Resolve Tier-1 level of incidents (as mentioned below) -

  • Common L1 fixes may include System Restarts, Network Connectivity Checks, Model Resets, Data Flow Checks, Re-Processing Alerts

  • Daily monitoring and report alerting metrics 

  • RCA and Resolution request review

  • Runbook fixes for each incident 

  • Provide immediate resolution, monitoring and daily reporting

  • Regularly verifying the the server environment 

  • Provide daily/weekly reports on system performance, incidents, and key metrics

  • Schedule  regular downtime for system updates or maintenance, ensuring minimal disruption to operations

  • If the issue cannot be resolved via the run book, escalate to L2/L3 support for model or hardware security issues

  • Adhering Service Level Agreements (SLAs) for response and resolution, prioritizing issues based on severity

  • Check-in via email when the Park opens

  • Check Real Time Overlay

  • Check Security Desk

  • Constantly monitor pagerduty channel and take necessary actions according to the run book

  • Redis db cron job check

Must Have Skills:

Kubernetes Developer:

  • Certification that best reflects below skills: CKAD

  • High Level Skills:

  • Create container infrastructure platform

  • Create and maintain highly scalable Kubernetes architecture

  • Enhance delivery systems with Continuous Integration

  • Improve monitoring and alerting

  • Troubleshoot issues and resolve problems

Granular Skills:

  • Expertise in Linux

  • Maintenance of Docker container and Kubernetes Clusters

  • Implement design patterns for distributed systems

  • Simple to complex k8s resources like Deployments, Pods, Jobs, StatefulSets, ConfigMaps, Service, NodePort, Ingress, Volumes

  • Templating Kubernetes deployments using Helm Charts

  • Multi container pod design using sidecar, init containers etc.

  • Knowledge of hybrid cloud environments - (bonus)

  • Strong expertise in DevOps and CI/CD implementation

  • Strong knowledge of deployment strategies (Blue/Green, Canary etc.) and rolling updates.

  • Implement probes and health checks

  • Utilizing container logs for debugging

  • Custom Resource Definitions


Backend Developer:

High Level Skills:

  • Strong grasp of Core Python concepts

  • Sound knowledge of web frameworks

  • Strong understanding of Object Oriented concepts and Software Design patterns

  • Experience with writing production grade code (NOT "SCRIPTS")

  • Good grasp of database integration and streaming pipelines (redis, kafka etc)

  • Multiprocess architecture

  • Strong Git and GitHub knowledge

  • Good knowledge of Video Streaming and image processing in python preferable

  • Good knowledge of Cloud Services

Granular Skills (extensive hands on experience required):

  • Core Python concepts (must have): Iterators, Exception Handling, File Handling, Data types and structures, OOP

  • Web frameworks (must have): Flask, Gunicorn

  • Video (extremely useful and highly rated): Gstreamer, ffmpeg, OpenCV

  • Containerization: Docker, Kubernetes core concepts

  • Strong knowledge of Linux shell


Shift Schedule: 

Work in rotational shifts:

Shift A - 6 AM IST to 2 PM IST

Shift B - 2 PM IST to 10 PM IST
Shift C - 10 PM IST to 6 AM IST 

Two consecutive days off in a week

.

If you like wild growth and working with happy, enthusiastic over-achievers, you'll enjoy your career with us!

Skills Required

  • 1 to 6 years of experience
  • CKAD certification or certification reflecting Kubernetes development skills
  • Expertise in Linux and Linux shell
  • Hands-on maintenance of Docker containers and Kubernetes clusters
  • Experience creating scalable Kubernetes architectures and container infrastructure
  • Knowledge of Kubernetes Deployments, Pods, Jobs, StatefulSets, ConfigMaps, Services, NodePort, Ingress, and Volumes
  • Experience templating Kubernetes deployments with Helm Charts
  • Knowledge of multi-container pods, including sidecar and init containers
  • Strong DevOps and CI/CD implementation experience
  • Knowledge of blue-green, canary, rolling update, and other deployment strategies
  • Experience implementing probes and health checks
  • Experience troubleshooting Kubernetes and container logs
  • Knowledge of Custom Resource Definitions
  • Strong grasp of core Python concepts, including iterators, exception handling, file handling, data types, data structures, and OOP
  • Experience with Flask and Gunicorn
  • Production-grade software development experience
  • Knowledge of web frameworks, software design patterns, and object-oriented programming
  • Experience integrating databases and streaming pipelines such as Redis and Kafka
  • Experience with multiprocess architecture
  • Strong Git and GitHub knowledge
  • Knowledge of cloud services
  • Comfort working rotational shifts and joining immediately
  • Knowledge of hybrid cloud environments
  • Knowledge of GStreamer, FFmpeg, and OpenCV
  • Knowledge of video streaming and image processing in Python

Quantiphi Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Quantiphi and has not been reviewed or approved by Quantiphi.

  • Wellbeing & Lifestyle Benefits Wellbeing initiatives such as monthly meeting-free AMA-Zen Days, health check-ups, and wellness counseling are designed to reduce burnout and support day-to-day balance. Broader wellness programs reinforce both physical and mental health.
  • Flexible Benefits Remote/hybrid options with flexible working hours provide meaningful autonomy over where and when work gets done. Flexible leave constructs, including sabbaticals and special day leaves, add practical adaptability to the package.
  • Parental & Family Support Paid parental leave in the U.S., alongside maternity and childcare support, signals solid backing for families. These family-oriented policies integrate with a wider health and wellness focus.

Quantiphi Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Marlborough, MA
3,494 Employees
Year Founded: 2013

What We Do

Quantiphi is an award-winning AI-first digital engineering company driven by the desire to solve transformational problems at the heart of business. Quantiphi solves the toughest and complex business problems by combining deep industry experience, disciplined cloud, and data-engineering practices, and cutting-edge artificial intelligence research to achieve quantifiable business impact at unprecedented speed.

Similar Jobs

MetLife Logo MetLife

Platform Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Hybrid
Pune, Maharashtra, IND
43000 Employees

MetLife Logo MetLife

Platform Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Hybrid
Pune, Maharashtra, IND
43000 Employees

MetLife Logo MetLife

Platform Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Hybrid
Pune, Maharashtra, IND
43000 Employees

MetLife Logo MetLife

Platform Engineer

Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Hybrid
Pune, Maharashtra, IND
43000 Employees

Similar Companies Hiring

Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account