Principal Engineer - SRE

Posted Yesterday
Be an Early Applicant
2 Locations
In-Office
Expert/Leader
Cloud • Fintech • Information Technology • Software • Financial Services
The Role
Principal SRE individual contributor responsible for the reliability, scalability and operational maturity of a Kubernetes-native AWS platform. Own HA/DR, observability, CI/CD and IaC, troubleshoot production incidents, drive automation and tooling, lead on-call rotations, and produce architecture, runbooks, and RFCs to improve system stability and resilience.
Summary Generated by Built In

Company Overview

Arcesium is a global financial technology firm that solves complex data-driven challenges faced by some of the world’s most sophisticated financial institutions. We constantly innovate our platform and capabilities to meet tomorrow’s challenges, anticipate the risks our clients encounter, and design advanced solutions to help our clients achieve transformational business outcomes.   

Financial technology is a high-growth industry as change and innovation continue to disrupt the status-quo and prompt major transformation. Arcesium is at a particularly interesting time in our own growth as we look to leverage our successfully established market position and expand operations in pursuit of strategic new business opportunities. We value intellectual curiosity, proactive ownership, and collaboration with colleagues, and we empower you to meaningfully contribute from day one and accelerate your professional development.

About the Role and the Team 

Arcesium seeks an exceptional engineer to join our Infrastructure team and provide expert-level technical leadership for our front-office Portfolio and Order Management System recently integrated into the Arcesium product suite.

This is a hands-on individual contributor role that owns the architectural direction for our most complex site reliability challenges and will be responsible for:

  1. Reliability, scalability, and operational maturity of a Kubernetes-native, AWS-hosted platform.
  2. High availability, disaster recovery, observability across components and client pods
  3. Troubleshooting live production issues with a deep focus on rapid incident resolution, creating RCA and building tools that enhance system stability and resilience.
  4. Bringing strong opinions on how infrastructure should be run - backed by experience running large-scale production systems.
What You Will Do
  • Own the reliability and availability of core platform infrastructure including Kubernetes clusters, ArgoCD deployments, PostgreSQL databases, and supporting AWS services
  • Plan and execute infrastructure maintenance - upgrades, patching, and validation of platform components with zero-to-minimal client impact
  • Improve CI/CD pipelines and developer tooling, including migration of locally executed Terraform apply processes to automated, auditable workflows
  • Build and enhance infrastructure monitoring, alerting, and observability to reduce mean time to detection and resolution
  • Evaluate and adopt new infrastructure tooling and capabilities (service mesh, cost optimization, capacity planning) that improve reliability and scalability.
  • Contribute to the convergence of Limina infrastructure with Arcesium's broader infrastructure ecosystem, supporting standardization of tooling, networking, security, and operational best practices
  • Establish and participate in on-call rotations, building sustainable operational coverage across global time zones as the client base grows
  • Troubleshoot and resolve complex cross-system issues spanning application, infrastructure, and cloud layers
  • Write and review infrastructure-as-code, automation scripts, and internal tooling
  • Lead by example with exemplary code, design documents, RFC's, runbooks, system and architecture design documents and documenting operational procedures to build institutional knowledge
What You Will Need
  • A bachelor’s or master’s degree in computer science, Engineering, or a related field with 9+ years of professional engineering experience, including significant time in a principal-level or equivalent individual contributor role.
  • Hands-on experience in site reliability engineering, infrastructure engineering, DevOps, or platform engineering roles
  • Strong track record of owning and operating production Kubernetes environments (cluster lifecycle, networking, storage, RBAC, troubleshooting)
  • Hands-on experience running critical workloads on AWS (EC2, EKS, RDS, IAM, VPC, S3, CloudWatch, and related services)
  • Experience with observability platforms (Datadog, Prometheus, Grafana, ELK), modern alerting design, infrastructure-as-code (Terraform, CloudFormation), and CI/CD pipelines (GitLab CI, Jenkins).
  • Strong programming and scripting ability in Python; familiarity with Bash, Go, or Java is a plus
  • Experience with GitOps-based deployment workflows (ArgoCD, FluxCD, or similar)
  • Familiarity with PostgreSQL or similar relational databases in a production context - backup, recovery, upgrades, and performance tuning
  • Demonstrated experience with on-call rotations, incident management, and post-incident review processes
  • A proven track record designing and delivering large-scale reliability initiatives (HA/DR, observability, automation platforms) with measurable outcomes.

Arcesium's Personal Data Privacy Notice for Candidates is linked here.


Recruiting Security
Emails from genuine Arcesium recruiters who are employees of the company will always come from the @arcesium.com domain. In some cases, you may also be contacted by independent search firms engaged to recruit on our behalf; emails from their employees should always come from their firm's applicable domain. We'll never ask for your banking information or any payment as part of the recruiting process. If something seems off or you're contacted by an unexpected third party, please reach out to us at [email protected] (US/UK), [email protected] (India) or [email protected] (Portugal/Sweden). 

Arcesium is an equal opportunity employer.

Skills Required

  • Bachelor's or Master's degree in Computer Science, Engineering, or related field with 9+ years professional engineering experience including principal-level experience
  • Hands-on experience in site reliability engineering, infrastructure engineering, DevOps, or platform engineering roles
  • Proven track record owning and operating production Kubernetes environments (cluster lifecycle, networking, storage, RBAC, troubleshooting)
  • Hands-on experience running critical workloads on AWS (EC2, EKS, RDS, IAM, VPC, S3, CloudWatch)
  • Experience with observability platforms and modern alerting design (Datadog, Prometheus, Grafana, ELK)
  • Experience with infrastructure-as-code (Terraform, CloudFormation) and improving CI/CD pipelines (GitLab CI, Jenkins)
  • Strong programming and scripting ability in Python
  • Experience with GitOps-based deployment workflows (ArgoCD, FluxCD, or similar)
  • Familiarity with PostgreSQL or similar relational databases in production (backup, recovery, upgrades, performance tuning)
  • Demonstrated experience with on-call rotations, incident management, and post-incident review processes
  • Proven track record designing and delivering large-scale reliability initiatives (HA/DR, observability, automation) with measurable outcomes
  • Familiarity with Bash, Go, or Java
  • Experience evaluating/adopting new infrastructure tooling (service mesh, cost optimization, capacity planning)

Arcesium Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Arcesium and has not been reviewed or approved by Arcesium.

  • Healthcare Strength Healthcare coverage is described as comprehensive, including medical, dental, vision, mental‑health programs, and protections for life, accident, and disability. The scope indicates strong support for physical and mental well‑being.
  • Retirement Support Retirement offerings include a 401(k) and a pension alongside financial protections such as life and disability insurance. These elements point to meaningful long‑term financial security within total rewards.
  • Parental & Family Support Parental leave is characterized as generous, with extended maternity and supportive paternity leave in some locations, plus family and adoption assistance. This family focus complements broader time‑off and caregiving supports.

Arcesium Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
1,500 Employees
Year Founded: 2015

What We Do

Arcesium is a global financial technology and professional services firm, delivering post-investment and enterprise data management solutions to some of the world's most sophisticated financial institutions, including hedge funds, banks, institutional asset managers, and private equity firms. Expertly designed to achieve a single source of truth throughout a client's ecosystem, Arcesium's cloud-native technology is built to systematize the most complex workflows and help clients achieve scale. Building on a platform developed and tested by investment and technology development firm, the D. E. Shaw group, Arcesium was launched as a joint venture with Blackstone Alternative Asset Management. J.P. Morgan, another large client, later joined as our third partner. Today, Arcesium services over $679 billion in global client AUM with a staff of over 1,500 software engineering, accounting, operations, and treasury professionals.

Similar Jobs

In-Office
Hyderabad, Telangana, IND
3062 Employees
Easy Apply
In-Office
Hyderabad, Telangana, IND
900 Employees
Hybrid
3 Locations
289097 Employees

MassMutual India Logo MassMutual India

Architect

Big Data • Fintech • Information Technology • Insurance • Financial Services
In-Office
Hyderabad, Telangana, IND

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account