Lead Cloud Platform Engineer (Data & Execution Platform)

Posted 9 Days Ago
Be an Early Applicant
2 Locations
In-Office
Senior level
Fintech
The Role
Lead a cloud-native data execution platform: implement, scale, and operate Spark and Airflow-based distributed systems on AWS/Kubernetes. Own performance tuning, observability, CI/CD, SLI/SLOs, incident response, and mentor engineers to ensure reliable, secure, and performant production pipelines.
Summary Generated by Built In
Job Description:

Note: Fidelity is not providing immigration sponsorship for this position.  

The Role

We are seeking a hands-on Lead Cloud Platform Engineer to implement, scale, and operate cloud-native infrastructure and services that power large-scale data processing systems. This role focuses on translating defined architectures into production-grade platforms that are reliable, observable, secure, and performant. You will lead the implementation and operation of a modern execution platform built on Apache Spark for distributed compute and an Airflow orchestration layer and DAG execution environment. The ideal candidate brings deep production experience in Spark and Airflow, and excels at troubleshooting, tuning, and operationalizing distributed systems in AWS environments, while leveraging modern developer productivity tools such as AI-assisted coding and LLM-based workflows.

The Expertise and Skills You Bring

  • Implement and operate cloud-native platform services for distributed data systems

  • Scale fault-tolerant, high-throughput systems aligned with architectural patterns

  • Own Spark data pipelines and Airflow orchestration layer and DAG execution

  • Tune Spark workloads (partitioning, memory, execution plans, shuffle optimization)

  • Troubleshoot Spark jobs and Airflow DAGs across performance and failures

  • Operate and optimize Kubernetes-based execution environments, including node group scaling, workload placement, and resource utilization

  • Troubleshoot Kubernetes infrastructure and workload issues, including scheduling, networking, and runtime performance

  • Leverage developer productivity tools (e.g., GitHub Copilot, LLMs) to accelerate development, debugging, and operational workflows.

  • Drive operational excellence including monitoring, incident response, and RCA

  • Implement observability (metrics, logging, tracing, dashboards, alerting)

  • Define and manage SLIs/SLOs for platform reliability

  • Deploy solutions using AWS services (EKS, EC2, S3, Lambda, RDS, etc.) (Implement secure networking (VPCs, IAM, subnets, load balancing)

  • Maintain CI/CD pipelines and deployment automation

  • Lead execution across planning, delivery, and cross-team coordination

  • Mentor engineers and promote reliability and scalability best practices

  • Strong understanding of distributed systems (fault tolerance, scalability, consistency

  • Expertise in Apache Spark (tuning, debugging, optimization)

  • Expertise in Apache Airflow (DAG execution, orchestration, troubleshooting)

  • Strong experience operating Kubernetes (EKS preferred) including cluster scaling and lifecycle management

  • Hands-on management of node groups, autoscaling, and capacity planning

  • Deep understanding of Kubernetes networking and security (security groups, network policies, ingress/egress)

  • Experience with Kubernetes resources (Deployments, StatefulSets, Jobs, CronJobs)

  • Familiarity with Custom Resources (CRDs) and advanced configuration via annotations and labels

  • Experience monitoring Kubernetes clusters (metrics, logs, events) and integrating with observability tools

  • Troubleshooting Kubernetes workloads (scheduling failures, resource contention, networking issues)

  • Experience with AWS services and cloud-native design patterns

  • Proficiency in Python, Java, or Go

  • Experience with Docker and Kubernetes

  • Hands-on observability (metrics, logging, tracing)

  • Experience with SLI/SLO-based reliability models

  • Practical experience using AI-assisted development tools (e.g., GitHub Copilot, LLMs) to improve code quality, debugging, and productivity

  • Networking fundamentals (DNS, TCP/IP, TLS, VPC design)

  • Strong troubleshooting and performance tuning skills

  • Strong communication and leadership skills

  • Bachelor’s or Master’s degree in Computer Science or related field (or equivalent experience)

  • 8 plus years in software, platform, or cloud engineering roles

  • Experience operating large-scale distributed systems in production

  • Strong experience with AWS cloud platforms

  • Mandatory hands-on experience with Apache Spark and Apache Airflow in production

  • Experience supporting ETL, data platforms, or workflow execution systems at scale

Fidelity’s Onsite Working Model
Fidelity is transitioning to a full-time onsite working model through a phased rollout across regions and roles. Currently, some roles and locations require 100% onsite presence, while others require less. Onsite expectations are likely to evolve as the rollout continues. This transition does not apply to fully remote roles.

Certifications:

Category:Information Technology

Please be advised that Fidelity’s business is governed by the provisions of the Securities Exchange Act of 1934, the Investment Advisers Act of 1940, the Investment Company Act of 1940, ERISA, numerous state laws governing securities, investment and retirement-related financial activities and the rules and regulations of numerous self-regulatory organizations, including FINRA, among others. Those laws and regulations may restrict Fidelity from hiring and/or associating with individuals with certain Criminal Histories.

Skills Required

  • Mandatory hands-on experience with Apache Spark in production (tuning, debugging, optimization)
  • Mandatory hands-on experience with Apache Airflow in production (DAG execution, orchestration, troubleshooting)
  • 8+ years in software, platform, or cloud engineering roles
  • Experience operating large-scale distributed systems in production
  • Strong experience with AWS services and cloud-native design patterns (EKS, EC2, S3, Lambda, RDS, etc.)
  • Operate and optimize Kubernetes-based execution environments, cluster scaling, node group management, autoscaling
  • Deep understanding of Kubernetes networking and security (network policies, ingress/egress, security groups)
  • Proficiency in Python, Java, or Go
  • Experience with Docker and containerized workloads
  • Implement observability: metrics, logging, tracing, dashboards, alerting
  • Define and manage SLIs/SLOs for platform reliability
  • Maintain CI/CD pipelines and deployment automation
  • Strong troubleshooting and performance tuning skills for Spark, Airflow, and Kubernetes workloads
  • Networking fundamentals (DNS, TCP/IP, TLS, VPC design)
  • Drive operational excellence including monitoring, incident response, and RCA
  • Experience supporting ETL, data platforms, or workflow execution systems at scale
  • Bachelor's or Master's degree in Computer Science or related field, or equivalent experience
  • EKS experience (preferred)
  • Familiarity with Custom Resources (CRDs) and advanced configuration via annotations and labels
  • Practical experience using AI-assisted development tools (e.g., GitHub Copilot, LLMs) to improve productivity

Fidelity Investments Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Fidelity Investments and has not been reviewed or approved by Fidelity Investments.

  • Strong & Reliable Incentives Bonuses, commissions, and profit-sharing are presented as generous and meaningful components of total compensation, with certain roles achieving high total earnings through multiple pay streams. Variable pay is consistently framed as a positive contributor beyond base salary.
  • Retirement Support A 401(k) match up to 7% alongside additional profit-sharing up to 10% materially enhances long-term compensation. These retirement features are highlighted as standout strengths of the overall package.
  • Parental & Family Support Generous paid parental leave (16 weeks maternity, 12 weeks parental), backup dependent care, and adoption assistance provide robust family support. Hybrid work and caregiving resources further ease family responsibilities.

Fidelity Investments Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Boston, MA
58,848 Employees
Year Founded: 1946

What We Do

At Fidelity, our goal is to make financial expertise broadly accessible and effective in helping people live the lives they want. We do this by focusing on a diverse set of customers: - from 23 million people investing their life savings, to 20,000 businesses managing their employee benefits to 10,000 advisors needing innovative technology to invest their clients’ money. We offer investment management, retirement planning, portfolio guidance, brokerage, and many other financial products. Privately held for nearly 70 years, we’ve always believed by providing investors with access to the information and expertise, we can help them achieve better results. That’s been our approach- innovative yet personal, compassionate yet responsible, grounded by a tireless work ethic—it is the heart of the Fidelity way.

Similar Jobs

Ericsson Logo Ericsson

Warehouse Operator

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office
Lewisville, TX, USA
88000 Employees

Ericsson Logo Ericsson

Systems Engineer

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
Hybrid
2 Locations
88000 Employees

Optum Logo Optum

Investigator

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Houston, TX, USA
160000 Employees
50K-89K Annually

Optum Logo Optum

VP Managed Care Performance & Analytics - Kelsey Seybold Clinics, Pearland, TX.

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Pearland, TX, USA
160000 Employees
159K-273K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account