Beacon is a cloud-native, cross-asset risk analytics and quantitative development platform used by top-tier asset managers and banks. It provides a transparent, extensible environment for building, deploying and scaling valuation models, pricing libraries and risk analytics across asset classes — including derivatives, structured products and fixed income. It combines pre-built financial applications with a flexible developer infrastructure that lets quants and model developers write custom pricing and valuation logic, run scenario analysis, and integrate models directly into front-office and risk workflows.
The role- You will own the reliability of Clearwater's risk and analytics applications in engineering terms — designing and building the automation, tooling and operational architecture that keeps the overnight and intraday pipelines correct and on time for asset managers and insurers who need to see their portfolios each morning.
- This is a senior engineering role on the application layer. The distinction that matters at this level: you are expected to retire classes of failure, not instances of them. When the same problem appears for the third time, we want the engineer who stops patching it and builds the thing that makes it impossible to fail. That work competes for time with the model and feature changes the quant and risk teams are asking for, so you will also need to argue for it in their terms — what the recurring failure is costing them now, and what fixing it properly buys them. Cloud fluency matters because the applications run on AWS and you need to operate them competently there, but you will not be provisioning the infrastructure beneath them.
- You will be based in Mumbai alongside the market data team — a group with front-office training who partner on the genuinely hard analytics questions, not only on data plumbing. That adjacency is the point of the placement: you will sit with the quant, risk and market data teams and be expected to learn what the systems actually compute. When a curve build fails or an end-of-day risk run finishes late, we want engineers who can reason about which clients are affected and what the number means.
- The team is part of a globally distributed platform organisation spanning Mumbai, Tokyo and the US, with US-based coverage for the London close, so every major market cutoff falls inside a staffed window. You will join and help shape a follow-the-sun on-call rotation covering the end-of-day cycles of multiple regions. We are explicit about this because the three-region model is what keeps it sustainable — there is no permanent night shift.
- You will have real ownership of the services you support and direct influence over how reliability is engineered into them. You will report to the support lead in Mumbai.
- Own the diagnostic tooling strategy — including AI-assisted tooling — so that establishing whether an issue is infrastructure or market data, and which clients it touches, takes seconds rather than a manual trawl through logs and jobs. Tooling, not documentation alone, is how we intend to get response times down.
- Write and review production-quality Python that eliminates toil: runbook automation, health checks, self-service diagnostics and remediation used across the team.
- Design automated recovery and self-healing for analytics jobs, so a single failed task does not become a missed client deliverable.
- Set the observability standard — design monitoring and alerting around business outcomes (late valuations, missing risk measures, stale prices) rather than host and container metrics, and drive alert noise down so pages are actionable and rare.
- Own incident response for client-impacting issues: command the incident, mitigate, communicate, and run blameless postmortems that end in fixes rather than findings.
- Investigate at the application level — logs, stack traces, job histories and the code itself — to establish root cause, then fix it or hand the development team a diagnosis precise enough to act on.
- Track service commitments and recurring pain: report on timeliness and failure trends, and push through the fixes that retire repeat issues.
- Set the boundary for first-line infrastructure work — stale hosts, memory pressure, connectivity failures, host loss — so routine cases are remediated locally and deeper platform work escalates cleanly.
- Own and harden large-scale batch and EOD processing: job orchestration, dependency management, retries, idempotency, backfills and safe reruns.
- Keep the overnight window on schedule — find where time is going across stages and drive the fixes with the development teams.
- Support the reference data, market data and timeseries stores behind pricing and risk — schema and index design, replication, retention and query performance in NoSQL and timeseries systems.
- Operate data quality and reconciliation controls with the market data team: completeness checks, stale-data and anomaly alerting, vendor feed failure handling.
- Debug data-shaped production problems — a missing fixing, a bad identifier mapping, a holiday calendar difference, a failed download job, a partial vendor file — traced through to the affected downstream analytics.
- Act as a single escalation path with the market data team on front-office-facing issues, sharing triage and root cause rather than passing tickets between functions.
- Partner with quant developers and risk engineers to make reliability a design input rather than a post-launch concern.
- Mentor engineers on incident triage, investigation technique and on-call practice, and set the standards for automation and operational readiness.
- Contribute to production readiness reviews for new services and client onboardings.
- Support client escalations on risk and analytics output, and be the person client-facing teams want in the room during an incident.
- Document architecture, runbooks and failure modes.
- 7+ years in software engineering for a business-critical platform, with real ownership of services.
- Strong Python, plus shell. You must be comfortable reading and debugging application code to find root cause — this role is not runbook execution.
- Demonstrated troubleshooting depth — working from a vague client symptom through logs, job histories, data and code to a precise diagnosis.
- Hands-on experience operating batch and scheduled workloads at scale, including workflow orchestration (Airflow, Prefect, Step Functions or equivalent).
- Solid grounding in distributed systems — you can reason about partial failure, retries, idempotency, consistency, and queue and backpressure behaviour under load.
- Working knowledge of AWS — enough to operate applications confidently (compute, S3, IAM, CloudWatch, basic networking) and to tell an infrastructure problem from an application one.
- Practical NoSQL experience (MongoDB, DynamoDB, Cassandra or similar): data modelling, indexing, replication and performance troubleshooting.
- Observability in practice — Prometheus, Grafana, Datadog, OpenTelemetry or equivalent — and sensible alert design.
- On-call experience with structured escalation and postmortem writing, and the judgement to lead an incident rather than only participate in one.
- A general understanding of financial systems and genuine appetite to learn the business. You should be able to hold a conversation about what a valuation or a risk number is for, and be willing to go deeper.
- Clear written and verbal communication, and the ability to work effectively in a globally distributed team.
- Degree (B.Tech / M.Tech / MCA / MSc) in computer science, engineering or a related field.
- Experience with timeseries data at scale — storage and partitioning strategies, retention, and query patterns for large historical datasets.
- Experience supporting trading systems, order management systems, risk management systems or market data platforms.
- Familiarity with financial data concepts — instrument reference data, curves and fixings, corporate actions, EOD valuation cycles, VaR and sensitivities.
- Kubernetes and containerised workloads in production, from an operations perspective.
- Infrastructure as code and CI/CD — Terraform, Docker, GitHub Actions or GitLab CI — enough to make and review changes when the work calls for it.
- Streaming and event-driven infrastructure (Kafka, Kinesis, MSK) and distributed compute (Spark, EMR, Dask, Ray).
- MongoDB Atlas, Redis / ElastiCache, or managed timeseries and analytical stores (Timestream, ClickHouse, Redshift).
- Working in a regulated financial environment — change management, access controls, audit evidence and client-facing SLAs.
- Fluency with AI-assisted development and operations tooling — using LLM-based agents and assistants to build diagnostics, automate investigation and accelerate delivery.
- Experience mentoring engineers and setting operational standards across a team.
- A track record of turning a fragile, hand-held process into a stable service with measurably less manual intervention.
- Experience designing automated recovery, failover and monitoring specifically for analytics or compute-heavy batch jobs.
- Tooling you built that measurably shortened time-to-diagnosis — and a view on where AI-assisted tooling genuinely helps in an operations context and where it does not.
- The ability to translate a technical failure into client impact, and to be the person client-facing teams want in the room during an incident.
- Curiosity about the domain: engineers who ask what the model is doing and why the number matters build far better reliability into these systems.
Skills Required
- 7-10 years of experience in software engineering, site reliability engineering, DevOps, or platform engineering.
- Strong programming skills in Python and ability to write production-quality code with tests.
- Hands-on experience with at least one major cloud provider (AWS or Azure) including networking (VPCs/VNets, subnets, security groups, load balancers, VPN), IAM/RBAC, storage, and compute.
- Working knowledge of infrastructure-as-code, ideally Terraform, and managing multiple environments from shared modules and per-environment configuration.
- Solid Linux fundamentals: reading logs, tracing processes, debugging services, and automating fixes.
- Experience extending provisioning and deployment pipelines (Terraform, configuration generation) to speed onboarding and deployment.
- A collaborative, service-oriented mindset working directly with onboarding, support, and client success teams.
- An automation reflex: build tools to eliminate repeated manual operational work.
- Experience operating multi-tenant or fleet-style environments.
- Observability stack experience (metrics, log aggregation, alerting, dashboards).
- Formal incident management experience (on-call, postmortems, blameless RCA).
- Exposure to financial services, fintech, or other regulated environments.
Clearwater Analytics (CWAN) Compensation & Benefits Highlights
-
Healthcare Strength — Employer-provided medical, dental, and vision coverage is consistently listed in current job postings, with disability insurance also referenced. Feedback suggests these core health benefits are solid even if plan richness is not portrayed as top-tier across sources.
-
Retirement Support — A 401(k) plan with employer matching is consistently cited in company and employer-verified materials. This reliable match supports long-term savings and is presented as a standard component of total rewards.
-
Leave & Time Off Breadth — Immediate eligibility for paid time off, holidays, and volunteer time, along with parental leave, appears across job postings. This breadth of leave options provides practical flexibility from day one.
Clearwater Analytics (CWAN) Insights
What We Do
CWAN was founded on a simple belief: investment professionals deserve modern technology that actually works for them. Not legacy systems that slow them down. Not fragmented data that creates confusion. But one comprehensive platform that gives you complete visibility and crystal-clear insights. The result? Investment management that works as seamlessly as your investment strategy. Since our founding in 2004, CWAN has been the trusted technology partner powering the world’s leading institutional investors — from insurance companies, asset managers, and hedge funds to asset owners like corporations, endowments, and pension funds managing over $10 trillion in assets.
Why Work With Us
We continue to grow, fueled by a strong foundation, an ambitious vision, and a commitment to delivering exceptional value to our clients, partners, and team members around the world. What started as a bold idea in Boise, Idaho has rapidly transformed into a global presence. We’ve expanded our footprint significantly—now operating out of 24 offices
Gallery
Clearwater Analytics (CWAN) Offices
Hybrid Workspace
Employees engage in a combination of remote and on-site work.


_1.jpg)








_1.jpg)





