SDE - 2/3 | Backend & Agentic AI | Granary

Posted 8 Days Ago
Be an Early Applicant
Mumbai, Maharashtra, IND
In-Office
Entry level
eCommerce
The Role
Build and operate scalable backend services, distributed event-processing pipelines, resilient recovery systems, incident analysis tools, and autonomous AI agents. Responsibilities include performance optimization, observability, on-call incident response, root-cause analysis, tool-enabled LLM workflows, and production reliability. The role requires strong Python and distributed-systems expertise, hands-on disaster recovery experience, and the ability to evaluate AI systems for correctness, latency, cost, and failure behavior.
Summary Generated by Built In
Fynd is a frontier technology company. We started at the intersection of technology and retail because that is where technology was the least available. Over the years, we became one of India’s largest retail technology platforms. But retail was the entry point, not the boundary.

Today, Fynd builds intelligent software that runs business operations. Not tools that help people work faster, but systems that absorb entire functions: manufacturing, marketing, logistics, commerce, quality control. We sit inside our customers’ businesses, harvest deep domain context, and build AI systems that operate autonomously. We are expanding from retail into manufacturing, generative media, physical AI, and healthcare.

About the role
Retail OS powers live store operations across fulfilment, delivery, complaints, and operational incidents. We are hiring an SDE - 2 or 3 who has built and operated products at scale and can bring that engineering discipline to AI agents, root-cause analysis, attribution, and autonomous workflows.
You will own services and features from design through production, combining performance, availability, and recovery engineering with applied AI. You will partner with senior engineers, operations teams, and upstream platform owners.

What you will do
  • Build scalable backend services and event-processing pipelines with measurable throughput, latency, and availability objectives.
  • Improve performance through profiling, efficient data access, concurrency management, capacity planning, and load testing.
  • Engineer resilience through replication, failover, backpressure, graceful degradation, and recovery procedures; validate backup restoration and disaster recovery.
  • Build incident correlation, root-cause analysis, and attribution systems that reconstruct operational timelines and link conclusions to evidence.
  • Develop agents that investigate incidents using operational data, metrics, and runbooks, then execute permitted actions with deduplication, audit trails, escalation, and verified closure.
  • Own production health through observability, on-call participation, incident response, postmortems, and preventive improvements.
What you must bring
  • Experience operating products at scale: Direct ownership of production services with significant traffic, event volumes, or concurrency. You can explain their scale, bottlenecks, performance targets, and availability outcomes.
  • Strong backend and distributed-systems fundamentals: Python, APIs, asynchronous processing, data modelling, consistency, idempotency, duplicate and late events, checkpoints, and replay.
  • Hands-on availability and recovery experience: Replication, failover, backup and restore validation, and disaster-recovery exercises, including recovery time and recovery point objectives (RTO/RPO).
  • Production debugging and performance depth: Experience diagnosing application, database, and infrastructure failures using logs, metrics, traces, profiling, and query analysis.
  • Applied AI engineering: Experience shipping LLM applications or agents with tool calling, structured outputs, retrieval, and evaluations of correctness, latency, cost, and failure behaviour.
  • AI-native development practices: Effective use of coding agents while independently reviewing, testing, and taking ownership of the resulting software.
  • Sound operational judgment: Clear reasoning about evidence, uncertainty, permissions, rollback, and when human intervention is required.

Useful additional experience
Commerce, fulfilment, logistics, payments, observability, or workflow automation; Kubernetes/GCP; MongoDB; React; and operational attribution systems.

Our environment
Python, FastAPI, MongoDB, GKE, Databricks, React, Prometheus, and Grafana.

What success looks like
Your services meet agreed performance and availability objectives, recover predictably during failures, and have tested recovery procedures. Your AI workflows reduce investigation effort, improve attribution quality, and complete permitted actions with observable, verifiable outcomes.

What do we offer?
Growth
At Fynd, growth is limitless. We nurture a culture that encourages innovation, embraces challenges, and supports continuous learning. As we expand into new product lines and global markets, we’re seeking talented individuals eager to grow with us.
We believe in empowering our people to take ownership, lead with confidence, and shape their careers.
  • Learning Wallet: Enrol in external courses or certifications to upskill—we’ll reimburse the costs to support your development.
Culture
We believe in building strong teams and lasting connections.
  • Regular community engagement and team-building activities
  • Biannual events to celebrate achievements, foster collaboration, and strengthen our workplace culture
Wellness
Your well-being is our priority. Comprehensive Mediclaim policy for you, your spouse, children, and parents

Work Environment
We thrive on collaboration and creativity. Our teams work from the office five days a week to encourage open communication, teamwork, and innovation.

Join us to be part of a dynamic environment where your ideas make an impact!


Skills Required

  • Direct experience owning and operating production products at scale, including significant traffic, event volumes, or concurrency
  • Strong backend and distributed-systems fundamentals, including Python, APIs, asynchronous processing, data modeling, consistency, idempotency, duplicate and late events, checkpoints, and replay
  • Hands-on experience with replication, failover, backup and restore validation, disaster-recovery exercises, RTO, and RPO
  • Production debugging and performance optimization using logs, metrics, traces, profiling, and query analysis
  • Experience shipping LLM applications or agents with tool calling, structured outputs, retrieval, and evaluations
  • Effective use of coding agents while independently reviewing, testing, and owning resulting software
  • Sound operational judgment concerning evidence, uncertainty, permissions, rollback, and human intervention
  • Experience with commerce, fulfillment, logistics, payments, observability, or workflow automation
  • Experience with Kubernetes, GCP, MongoDB, React, and operational attribution systems
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Mumbai
997 Employees
Year Founded: 2012

What We Do

Fynd is India's largest omnichannel ecosystem and multi-platform tech company. Headquartered in Mumbai and founded by Farooq Adam, Harsh Shah, and Sreeraman MG in 2012. We have modernized retail strategies for more than 1000 brands & created a rich suite of tech products. Rooted in technology & innovation, we have products in applied machine learning, big data, gaming+crypto, image editing, and learning space. Our constant innovation and expertise in technology has been noticed worldwide. Fynd made it to Fast Company's list of Top 10 most innovative Asia-Pacific companies of 2022. We are a fast growing team of 1000+ fun, skilled and ambitious people. We explore the unexplored, innovate unafraid, and have the time of our life while we do. Be a part of the new. Join us. For more information about our products, visit us at www.omnifynd.com

Similar Jobs

Remote or Hybrid
3 Locations
1100 Employees

CrowdStrike Logo CrowdStrike

Full-stack Engineer

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
India
11000 Employees

WHOOP Logo WHOOP

Marketplaces Lead, India

Fitness • Hardware • Healthtech • Sports • Wearables
Hybrid
Mumbai, Maharashtra, IND
500 Employees
100K-150K Annually

Cencora Logo Cencora

Senior Engineer

Healthtech • Logistics • Pharmaceutical
In-Office
Pune, Maharashtra, IND
51000 Employees

Similar Companies Hiring

Munchkin, Inc. Thumbnail
Consumer Web • eCommerce • Food • Kids + Family • Design • Manufacturing
Milton, Ontario
325 Employees
Scotch Thumbnail
Artificial Intelligence • eCommerce • Fintech • Payments • Retail • Software • Analytics
US
35 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account