- Build scalable backend services and event-processing pipelines with measurable throughput, latency, and availability objectives.
- Improve performance through profiling, efficient data access, concurrency management, capacity planning, and load testing.
- Engineer resilience through replication, failover, backpressure, graceful degradation, and recovery procedures; validate backup restoration and disaster recovery.
- Build incident correlation, root-cause analysis, and attribution systems that reconstruct operational timelines and link conclusions to evidence.
- Develop agents that investigate incidents using operational data, metrics, and runbooks, then execute permitted actions with deduplication, audit trails, escalation, and verified closure.
- Own production health through observability, on-call participation, incident response, postmortems, and preventive improvements.
- Experience operating products at scale: Direct ownership of production services with significant traffic, event volumes, or concurrency. You can explain their scale, bottlenecks, performance targets, and availability outcomes.
- Strong backend and distributed-systems fundamentals: Python, APIs, asynchronous processing, data modelling, consistency, idempotency, duplicate and late events, checkpoints, and replay.
- Hands-on availability and recovery experience: Replication, failover, backup and restore validation, and disaster-recovery exercises, including recovery time and recovery point objectives (RTO/RPO).
- Production debugging and performance depth: Experience diagnosing application, database, and infrastructure failures using logs, metrics, traces, profiling, and query analysis.
- Applied AI engineering: Experience shipping LLM applications or agents with tool calling, structured outputs, retrieval, and evaluations of correctness, latency, cost, and failure behaviour.
- AI-native development practices: Effective use of coding agents while independently reviewing, testing, and taking ownership of the resulting software.
- Sound operational judgment: Clear reasoning about evidence, uncertainty, permissions, rollback, and when human intervention is required.
- Learning Wallet: Enrol in external courses or certifications to upskill—we’ll reimburse the costs to support your development.
- Regular community engagement and team-building activities
- Biannual events to celebrate achievements, foster collaboration, and strengthen our workplace culture
Skills Required
- Direct experience owning and operating production products at scale, including significant traffic, event volumes, or concurrency
- Strong backend and distributed-systems fundamentals, including Python, APIs, asynchronous processing, data modeling, consistency, idempotency, duplicate and late events, checkpoints, and replay
- Hands-on experience with replication, failover, backup and restore validation, disaster-recovery exercises, RTO, and RPO
- Production debugging and performance optimization using logs, metrics, traces, profiling, and query analysis
- Experience shipping LLM applications or agents with tool calling, structured outputs, retrieval, and evaluations
- Effective use of coding agents while independently reviewing, testing, and owning resulting software
- Sound operational judgment concerning evidence, uncertainty, permissions, rollback, and human intervention
- Experience with commerce, fulfillment, logistics, payments, observability, or workflow automation
- Experience with Kubernetes, GCP, MongoDB, React, and operational attribution systems
What We Do
Fynd is India's largest omnichannel ecosystem and multi-platform tech company. Headquartered in Mumbai and founded by Farooq Adam, Harsh Shah, and Sreeraman MG in 2012. We have modernized retail strategies for more than 1000 brands & created a rich suite of tech products. Rooted in technology & innovation, we have products in applied machine learning, big data, gaming+crypto, image editing, and learning space. Our constant innovation and expertise in technology has been noticed worldwide. Fynd made it to Fast Company's list of Top 10 most innovative Asia-Pacific companies of 2022. We are a fast growing team of 1000+ fun, skilled and ambitious people. We explore the unexplored, innovate unafraid, and have the time of our life while we do. Be a part of the new. Join us. For more information about our products, visit us at www.omnifynd.com









