Site Reliability Engineer/L3 Support

Posted 14 Days Ago
3 Locations
In-Office or Remote
110K-130K Annually
Mid level
Fintech • Software
The Role
Owner of operational health and availability for a FedRAMP High cloud platform. Act as L3 escalation, monitor observability, lead incident response, perform root cause analysis, automate toil, support deployments, maintain runbooks, and ensure compliance while partnering with engineering to improve reliability and resilience.
Summary Generated by Built In

As a leading financial services and healthcare technology company based on revenue, SS&C is headquartered in Windsor, Connecticut, and has 27,000+ employees in 35 countries. Some 20,000 financial services and healthcare organizations, from the world's largest companies to small and mid-market firms, rely on SS&C for expertise, scale, and technology.

Job Description

Job Title: Site Reliability Engineer (SRE) / L3 Support Engineer REMOTE

Getting to know us:

As a leading financial services and healthcare technology company based on revenue, SS&C is headquartered in Windsor, Connecticut, and has 27,000+ employees in 35 countries. Some 20,000 financial services and healthcare organizations, from the world's largest companies to small and mid-market firms, rely on SS&C for expertise, scale, and technology.

Kick off your software engineering career on our Quality & Automation team. You will learn modern test engineering practices while contributing real code, automated tests, and quality improvements. We welcome candidates new to finance

Why You Will Love It Here! 

  • Flexibility: Hybrid Work Model & a Business Casual Dress Code, including jeans

  • Your Future: 401k Matching Program, Professional Development Reimbursement

  • Work/Life Balance: Flexible Personal/Vacation Time Off, Sick Leave, Paid Holidays

  • Your Wellbeing: Medical, Dental, Vision, Employee Assistance Program, Parental Leave

  • Wide Ranging Perspectives: Committed to Celebrating the Variety of Backgrounds, Talents and Experiences of Our Employees 

  • Training: Hands-On, Team-Customized, including SS&C University

  • Extra Perks: Discounts on fitness clubs, travel and more!

What You Will Get To Do:

  • We are looking for a Site Reliability Engineer (SRE) to join our Platform Engineering team and take ownership of the operational health, reliability, and availability of our FedRAMP High cloud platform.

  • This role combines modern Site Reliability Engineering practices with advanced production support responsibilities. You will act as the highest level of operational support (L3), proactively identifying and resolving issues before they impact customers, driving continuous improvement, and working closely with engineering teams to improve the reliability and operability of the platform.

  • This is not a traditional operations role. You will use automation, observability, and engineering best practices to reduce operational toil while helping development teams build resilient, secure services.

Due to the nature of the environment, this position requires the successful candidate to be a U.S. Citizen and eligible to work on systems supporting FedRAMP High workloads.

What you will get to do:

  • Monitor the health, availability, performance, and security of production services.

  • Proactively identify emerging issues using telemetry, logs, metrics, and distributed tracing.

  • Investigate, troubleshoot, and resolve complex production incidents across application and infrastructure layers.

  • Act as the L3 escalation point for operational issues that cannot be resolved by L1 or L2 support.

  • Participate in an on-call rotation for critical production incidents.

  • Lead incident response activities, including coordination, communication, and post-incident reviews.

  • Perform root cause analysis and ensure corrective actions are implemented to prevent recurrence.

  • Develop and maintain operational runbooks, dashboards, alerts, and standard operating procedures.

  • Improve platform observability by enhancing monitoring, alerting, dashboards, and service-level indicators.

  • Work closely with software engineering teams to improve service reliability, scalability, and resilience.

  • Identify opportunities to automate operational tasks and eliminate repetitive manual work.

  • Support production deployments, infrastructure changes, and maintenance activities.

  • Assist with disaster recovery exercises, resilience testing, and operational readiness reviews.

  • Ensure operational activities comply with FedRAMP High security and compliance requirements.

  • Contribute to continuous improvement initiatives across reliability, performance, and operational excellence.

What you will Bring:

  • U.S. Citizenship (required).

  • 3–6 years of experience in Site Reliability Engineering, Production Engineering, DevOps, Platform Engineering, or a senior production support role.

  • Experience supporting mission-critical cloud-based production systems.

  • Strong understanding of Linux operating systems and networking fundamentals.

  • Experience troubleshooting distributed applications running in Kubernetes.

  • Experience with public cloud platforms, preferably AWS.

  • Experience with infrastructure as code and configuration management.

  • Strong scripting or programming skills (e.g. Python, Bash, PowerShell, Go, or similar).

  • Experience using monitoring and observability platforms such as Prometheus, Grafana, CloudWatch, Datadog, Splunk, or OpenTelemetry.

  • Experience analysing application logs, metrics, and traces to diagnose production issues.

  • Understanding of incident management, problem management, and root cause analysis.

  • Strong analytical and troubleshooting skills.

  • Excellent written and verbal communication skills.

Preferred Qualifications

  • Experience supporting systems operating under FedRAMP High, DoD IL5/IL6, or similar regulated environments.

  • Experience with Kubernetes in production.

  • Experience with AWS services including EKS, RDS, IAM, CloudWatch, Route 53, VPC networking, and AWS Backup.

  • Experience with CI/CD pipelines and deployment automation.

  • Knowledge of service mesh technologies such as Istio.

  • Familiarity with security best practices including IAM, least privilege, vulnerability management, and compliance monitoring.

  • Experience with PagerDuty, Jira Service Management, or similar incident management platforms.

  • AWS certification (Associate or Professional) is desirable.

What Success Looks Like

Within your first year you will:

  • Maintain high platform availability and service reliability.

  • Detect and resolve issues before customers experience impact.

  • Reduce mean time to detect (MTTD) and mean time to recover (MTTR).

  • Improve monitoring coverage and reduce unnecessary alert noise.

  • Increase operational automation and reduce manual support effort.

  • Produce high-quality incident reviews with actionable improvements.

  • Partner effectively with engineering teams to continuously improve platform resilience.

Why Join Us?

You'll work on a modern cloud-native SaaS platform running in a highly secure FedRAMP High environment, where reliability, automation, and engineering excellence are fundamental. You'll collaborate with software engineers, platform engineers, and security specialists to build and operate systems that customers depend upon every day.



Unless explicitly requested or approached by SS&C Technologies, Inc. or any of its affiliated companies, the company will not accept unsolicited resumes from headhunters, recruitment agencies, or fee-based recruitment services.



SS&C Technologies offers a comprehensive total rewards package designed to support your wellbeing, growth, and future. Our benefits include medical, dental, and vision coverage; a 401(k) plan with company match; paid time off, holidays, and parental leave; and professional development reimbursement opportunity.

Actual base salary will vary based on several factors, including but not limited to relevant skills, prior experience, education, demonstrated performance, and geographic location.

New York: The expected base salary for the position is between 110000 USD to 130000 USD.

In addition, employees in this role may be eligible for consideration on an annual basis for a discretionary bonus and/or equity awards, such as restricted stock units or stock options, based upon individual and business performance at the company’s discretion. 


Applications will be accepted on an ongoing basis until the position is filled.



SS&C Technologies is an Equal Employment Opportunity employer and does not discriminate against any applicant for employment or employee on the basis of race, color, religious creed, gender, age, marital status, sexual orientation, national origin, disability, veteran status or any other classification protected by applicable discrimination laws.

Skills Required

  • U.S. Citizenship
  • 3-6 years experience in SRE, Production Engineering, DevOps, Platform Engineering, or senior production support
  • Experience supporting mission-critical cloud-based production systems
  • Strong understanding of Linux operating systems and networking fundamentals
  • Experience troubleshooting distributed applications running in Kubernetes
  • Experience with public cloud platforms, preferably AWS
  • Experience with infrastructure as code and configuration management
  • Strong scripting or programming skills (e.g., Python, Bash, PowerShell, Go)
  • Experience using monitoring and observability platforms (Prometheus, Grafana, CloudWatch, Datadog, Splunk, OpenTelemetry)
  • Experience analyzing application logs, metrics, and traces to diagnose production issues
  • Understanding of incident management, problem management, and root cause analysis
  • Strong analytical and troubleshooting skills
  • Excellent written and verbal communication skills
  • Experience supporting systems under FedRAMP High, DoD IL5/IL6, or similar regulated environments
  • Experience with Kubernetes in production (preferred)
  • Experience with AWS services including EKS, RDS, IAM, CloudWatch, Route 53, VPC networking, and AWS Backup
  • Experience with CI/CD pipelines and deployment automation
  • Knowledge of service mesh technologies such as Istio
  • Familiarity with security best practices including IAM, least privilege, vulnerability management, and compliance monitoring
  • Experience with PagerDuty, Jira Service Management, or similar incident management platforms
  • AWS certification (Associate or Professional)

SS&C Technologies Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about SS&C Technologies and has not been reviewed or approved by SS&C Technologies.

  • Leave & Time Off Breadth Leave policies are described as generous, including flexible or unlimited vacation and broadly positive views of PTO as a meaningful part of the overall package.
  • Retirement Support Retirement benefits are positioned as a notable strength, with repeated references to a 401(k) plan with company matching as a valued component of rewards.
  • Equity Value & Accessibility Equity and stock-related incentives are highlighted as a bright spot, with stock incentives described as excellent in some roles and contributing positively to perceived total rewards.

SS&C Technologies Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Birmingham, AL
22,000 Employees
Year Founded: 1986

What We Do

SS&C is a global provider of investment and financial services and software for the financial services and healthcare industries. Named to Fortune 1000 list as top U.S. company based on revenue, SS&C is headquartered in Windsor, Connecticut and has 22,000+ employees in over 150 offices in 35 countries. Some 18,000 financial services and healthcare organizations, from the world's largest institutions to local firms, manage and account for their investments using SS&C's products and services.

Similar Jobs

In-Office or Remote
4 Locations
22000 Employees
110K-120K Annually

Dropbox Logo Dropbox

Director of Media Strategy & Activation (Paid Media)

Artificial Intelligence • Cloud • Consumer Web • Productivity • Software • App development • Data Privacy
Remote
United States
2500 Employees
206K-278K Annually

Globe Life Logo Globe Life

Customer Service Representative

Insurance • Financial Services
Remote
USA
3000 Employees

Coinbase Logo Coinbase

Senior Product Marketing Manager

Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Easy Apply
Remote
USA
4700 Employees
171K-201K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account