Principal Site Reliability Engineer

Posted 17 Hours Ago
Be an Early Applicant
Durham, NC, USA
In-Office
Senior level
Fintech
The Role
Leads enterprise reliability strategy and architects highly available, scalable systems in AWS. Designs performance, load, stress, and chaos testing frameworks; implements observability, SLOs, SLIs, error budgets, monitoring, and incident response. Automates infrastructure and operational workflows using Python, Shell, Kubernetes, and CI/CD technologies. Analyzes performance data, improves resilience, mentors engineers, advises leadership, and develops recommendations for system optimization across financial services trading platforms.
Summary Generated by Built In
Job Description:

Note: Fidelity will not provide immigration sponsorship for this position.

Position Description:

Deploys and supports distributed, multi-tiered systems at scale while ensuring high availability and fault tolerance across multiple environments. Builds and operates resilient platforms in Amazon Web Services (AWS) using Elastic Compute Cloud (EC2), Simple Storage Service (S3), and Auto Scaling Groups for dynamic resource management. Designs, develops, and executes performance tests using Java-based frameworks, Apache JMeter, k6, and Rush-hour to validate system behavior under day-to-day traffic patterns. Defines and implements observability practices to monitor system health, latency, and error rates through metrics, logs, and distributed tracing using Datadog, Grafana, Splunk, and the Elasticsearch, Logstash, and Kibana (ELK) stack. Automates operational workflows with Python and Shell scripting to enhance efficiency and reduce manual tasks. Supports consistent build, deployment, and orchestration processes using cloud computing and DevOps technologies -- Continuous Integration and Continuous Delivery (CI/CD) pipelines and Kubernetes. Supports Site Reliability Engineering (SRE) functions by establishing Service Level Objectives (SLOs), Service Level Indicators (SLIs), and error budgets, and implementing proactive monitoring and incident response strategies. Builds and refines methodologies for performance, load, stress, and chaos testing and develops analytics and reports aligned with business needs to improve system resilience and optimization.

Primary Responsibilities: 

  • Defines and leads enterprise-level reliability strategies.
  • Architects resilient systems and infrastructure.
  • Creates and publishes performance test results report with recommendations on quality improvement.
  • Maintains scalability and resiliency of complex environment.
  • Implements advanced observability practices and techniques at scale.
  • Manages and interprets large datasets using query languages and visualization tools.
  • Advises senior leadership on reliability engineering best practices.
  • Mentors junior engineers.
  • Performs independent and complex technical and functional analysis for multiple divisional initiatives.
  • Develops innovative solutions to improve system availability, scalability, and performance.
  • Designs, implements, and maintains performance test frameworks.

Education and Experience:

Bachelor’s degree in Computer Science, Engineering, Information Technology Management, Information Systems Security, Business Administration, or a closely related field (or foreign education equivalent) and five (5) years of experience as a Principal Site Reliability Engineer (or closely related occupation) implementing highly available trading systems in a financial services environment.

Or, alternatively, Master’s degree in Computer Science, Engineering, Information Technology Management, Information Systems Security, Business Administration, or a closely related field (or foreign education equivalent) and three (3) years of experience as a Principal Site Reliability Engineer (or closely related occupation) implementing highly available trading systems in a financial services environment.

Skills and Knowledge:

Candidate must also possess:

  • Demonstrated Expertise (“DE”) performing software performance benchmarking and engineering for online financial web applications, Application Programming Interfaces (APIs), and mobile transactions according to DevOps practices, using performance benchmarking tools Rushhour, Locust, K6, and JMeter; and configuring CI/CD and test automation, using Jenkins, Sonar, Ant, Maven, Artifactory, and Terraform in AWS.
  • DE solutioning, designing, architecting, and building scalable and resilient enterprise-grade software platforms using cloud-based architecture and AWS services (EC2, Elastic Container Service (ECS), Lambda, Elastic MapReduce (EMR), and CloudFormation); developing microservices on Elastic Kubernetes Service (EKS), implementing CI/CD pipelines using DevOps tools (Bitbucket, GitHub, Artifactory, Sonar, Veracode, and Helm), and adhering to DevOps practices along with leveraging Java, Python, Spring Boot, Docker, EKS, and AWS.
  • DE analyzing and monitoring system and application performance across Apache, NGINX, Java, and Node.js platforms, and Linux and Windows environments, using Splunk, Datadog, Kibana, Grafana, and AWS CloudWatch; diagnosing performance bottlenecks, recommending tuning strategies, reducing Mean Time to Detect (MTTD) and Mean Time to Repair (MTTR), using Application Performance Monitoring (APM) tools -- Dynatrace, New Relic, Splunk, and Datadog; and performing capacity planning to optimize Central Processing Unit (CPU), memory, and process configurations.
  • DE instrumenting advanced observability practices at scale across cloud-native and hybrid environments; defining and tracking SLOs and SLIs to ensure reliability and performance metrics, using Python automation, Infrastructure as Code (IaC) methodologies, and observability tools (Datadog, Splunk, Dynatrace, Grafana, and the ELK stack; developing custom dashboards, alerting rules, and automated incident response workflows to proactively detect and resolve performance degradations, using Datadog, Catchpoint, Grafana, ELK stack, and Cloudwatch; and enabling actionable insights through trace-level correlation of end-to-end (E2E) user journeys and system behaviors, using Dynatrace, Splunk, Draw.io, and Miro.

#PE1M2

#LI-DNI

Fidelity’s Onsite Working Model
Fidelity is transitioning to a full-time onsite working model through a phased rollout across regions and roles. Currently, some roles and locations require 100% onsite presence, while others require less. Onsite expectations are likely to evolve as the rollout continues. This transition does not apply to fully remote roles.

Certifications:

Category:Information Technology

Please be advised that Fidelity’s business is governed by the provisions of the Securities Exchange Act of 1934, the Investment Advisers Act of 1940, the Investment Company Act of 1940, ERISA, numerous state laws governing securities, investment and retirement-related financial activities and the rules and regulations of numerous self-regulatory organizations, including FINRA, among others. Those laws and regulations may restrict Fidelity from hiring and/or associating with individuals with certain Criminal Histories.

Skills Required

  • Bachelor's degree in Computer Science, Engineering, Information Technology Management, Information Systems Security, Business Administration, or a closely related field, plus five years of experience as a Principal Site Reliability Engineer or closely related role implementing highly available trading systems in financial services
  • Alternatively, master's degree in Computer Science, Engineering, Information Technology Management, Information Systems Security, Business Administration, or a closely related field, plus three years of relevant experience implementing highly available trading systems in financial services
  • Expertise in software performance benchmarking and engineering for financial web applications, APIs, and mobile transactions
  • Experience with Rushhour, Locust, k6, and JMeter
  • Experience configuring CI/CD and test automation using Jenkins, Sonar, Ant, Maven, Artifactory, and Terraform in AWS
  • Experience designing scalable, resilient enterprise software platforms using AWS cloud architecture and services
  • Experience developing microservices on EKS and implementing CI/CD with Bitbucket, GitHub, Artifactory, Sonar, Veracode, and Helm
  • Experience with Java, Python, Spring Boot, Docker, EKS, and AWS
  • Experience analyzing and monitoring Apache, NGINX, Java, and Node.js platforms across Linux and Windows environments
  • Experience with Splunk, Datadog, Kibana, Grafana, AWS CloudWatch, Dynatrace, New Relic, and APM tools
  • Experience with observability practices, SLOs, SLIs, Python automation, Infrastructure as Code, dashboards, alerting, and automated incident response

Fidelity Investments Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Fidelity Investments and has not been reviewed or approved by Fidelity Investments.

  • Strong & Reliable Incentives Bonuses, commissions, and profit-sharing are presented as generous and meaningful components of total compensation, with certain roles achieving high total earnings through multiple pay streams. Variable pay is consistently framed as a positive contributor beyond base salary.
  • Retirement Support A 401(k) match up to 7% alongside additional profit-sharing up to 10% materially enhances long-term compensation. These retirement features are highlighted as standout strengths of the overall package.
  • Parental & Family Support Generous paid parental leave (16 weeks maternity, 12 weeks parental), backup dependent care, and adoption assistance provide robust family support. Hybrid work and caregiving resources further ease family responsibilities.

Fidelity Investments Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Boston, MA
58,848 Employees
Year Founded: 1946

What We Do

At Fidelity, our goal is to make financial expertise broadly accessible and effective in helping people live the lives they want. We do this by focusing on a diverse set of customers: - from 23 million people investing their life savings, to 20,000 businesses managing their employee benefits to 10,000 advisors needing innovative technology to invest their clients’ money. We offer investment management, retirement planning, portfolio guidance, brokerage, and many other financial products. Privately held for nearly 70 years, we’ve always believed by providing investors with access to the information and expertise, we can help them achieve better results. That’s been our approach- innovative yet personal, compassionate yet responsible, grounded by a tireless work ethic—it is the heart of the Fidelity way.

Similar Jobs

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Inside Sales Representative

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
10 Locations
40000 Employees
45K-85K Annually

Samsara Logo Samsara

Integration Engineer

Artificial Intelligence • Cloud • Computer Vision • Hardware • Internet of Things • Software
Easy Apply
Remote or Hybrid
United States
4000 Employees
87K-118K Annually

Zeta Global Logo Zeta Global

Senior Customer Success Manager

AdTech • Artificial Intelligence • Marketing Tech • Software • Analytics
Easy Apply
Remote or Hybrid
United States
2429 Employees
80K-100K Annually

Apryse Logo Apryse

Commercial Account Executive

Productivity • Software • App development • Automation
In-Office or Remote
10 Locations
665 Employees
150K-170K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Artificial Intelligence • Fintech • Software
New York, New York
9 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account