Senior Site Reliability Engineer

Posted 18 Hours Ago
Hiring Remotely in US
Remote
104K-140K Annually
Senior level
Software • Financial Services
The Role
Own day-to-day AWS and database operations for a serverless production platform, manage backups and disaster recovery, monitor and debug production, lead incident response and on-call, maintain infrastructure-as-code (SST/Pulumi), optimize costs, and mentor the team on operational best practices.
Summary Generated by Built In

As a Senior Site Reliability Engineer on our cloud engineering team, you'll keep our production environment healthy, secure, and running smoothly. This is an operations-focused role: you'll own the day-to-day administration of our AWS accounts and databases, backup posture across our data stores, and production monitoring and debugging for a fully serverless platform. Your work will span the operational side of the software development life cycle — from deployment to maintenance and updates — always striving for continuous improvement. You'll keep our infrastructure clean, easily deployable, and scalable, creating a stable operating environment for the whole team.

Responsibilities

  • Own day-to-day administration across AWS services, accounts, and access, as well as database administration across PostgreSQL and our other data stores.

  • Own backup posture across databases, S3 buckets, and queues; verify restores regularly and maintain a tested disaster recovery plan.

  • Proactively monitor production — CloudWatch dashboards, metric alarms, log-based metrics, and Slack alerting — addressing operational issues before they impact users.

  • Lead production debugging and incident response: build and maintain runbooks, participate in the on-call rotation, and resolve queue and dead-letter-queue failures through retry, redrive, and recovery.

  • Continuously refine our infrastructure to ensure it is easily deployable and scalable: keep infrastructure as code (SST/Pulumi) accurate, retire unused infrastructure, and keep cost visible and justified.

  • Share your knowledge of production operations with the team, fostering a culture of learning and growth.

Qualifications: Knowledge, Skills, & Abilities

  • Bachelor's degree and 4-6 years of related experience or equivalent work experience.

  • 5+ years of experience in DevOps, site reliability, or platform operations, with significant responsibility for production systems.

  • 3+ years of hands-on experience with AWS, with an emphasis on serverless services (Lambda, SQS, EventBridge, CloudWatch, S3).

  • Strong database administration experience: PostgreSQL operations, backup and recovery, and query performance; comfort administering other data stores.

  • Proficiency in scripting languages such as TypeScript, Python, and bash for production automation and operational tooling.

  • Strong understanding of Linux, DNS, TLS, Docker, GitHub Actions, and infrastructure as code (SST, Pulumi, or Terraform).

  • Experience with production monitoring and alerting, incident response, and on-call ownership.

Skills Required

  • Bachelor's degree or equivalent experience (4-6 years) or equivalent work experience
  • 5+ years in DevOps, site reliability, or platform operations with production responsibility
  • 3+ years hands-on AWS experience, emphasis on serverless services (Lambda, SQS, EventBridge, CloudWatch, S3)
  • Strong PostgreSQL database administration experience (backup/recovery, query performance)
  • Proficiency in scripting for automation and tooling (TypeScript, Python, bash)
  • Experience with Linux, DNS, TLS, and Docker
  • Experience with CI/CD and automation (GitHub Actions)
  • Experience with infrastructure-as-code (SST, Pulumi, or Terraform)
  • Experience with production monitoring, alerting, incident response, runbooks, and on-call ownership
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Costa Mesa, CA
522 Employees
Year Founded: 1998

What We Do

Pioneering Technologies for Your Financial Institution Since 1998, we have been creating innovative technologies that transform the way financial institutions operate by solving complex problems with streamlined, user-friendly solutions. Our robust and secure technologies empower lenders and consumers to get reliable, accurate information every time, at any time. As well-established industry leaders, we continue to set the industry standard for web-based credit reporting and lending for financial institutions of every size.

Similar Jobs

NBCUniversal Logo NBCUniversal

Senior Site Reliability Engineer

AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Remote or Hybrid
Centennial, CO, USA
130K-160K Annually

Circle Logo Circle

Senior Site Reliability Engineer

Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
In-Office or Remote
San Francisco, CA, USA
1050 Employees
153K-205K Annually
Remote or Hybrid
San Francisco, CA, USA
1100 Employees
147K-278K Annually
Remote
United States
350 Employees
180K-220K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account