Senior Cloud Infrastructure Engineer

Posted 2 Days Ago
Be an Early Applicant
San Francisco, CA, USA
Hybrid
200K-230K Annually
Senior level
Fintech • Software
The Role
Own and evolve Collective’s multi-cloud infrastructure across AWS and GCP using Terraform. Lead CI/CD, observability, identity, security, disaster recovery, automation, incident response, and platform reliability. Partner with product engineering and AI Platform teams to deliver scalable infrastructure, reduce operational toil, improve system performance, and mentor engineers. Participate in on-call rotations and drive post-incident improvements.
Summary Generated by Built In

About Collective:

Collective is on a mission to redefine the way businesses-of-one work. Our technology and team of trusted advisors help members achieve financial independence by taking care of everything from business incorporation to accounting, bookkeeping, tax services, and access to a thriving community, all in one integrated platform. We believe in empowering self-employed people to enjoy the same tax savings that big companies get, so they can focus on their passion, not paperwork.

Featured in Forbes, Business Insider, Yahoo, Bloomberg, Financial Times, TechCrunch, and more. We are backed by General Catalyst, Sound Ventures, QED Investors, Google’s Gradient Ventures, Expa, and other investors who have financed iconic companies like YouTube, Substack, Twitch, Box, Hims, Instacart, and Lyft.

About the role:

You'll join Collective's Infrastructure team, a small group that owns the platform every engineer here builds on: AWS and GCP foundations, CI/CD, IaC, observability, secrets management, and identity. We partner with product engineering to unblock shipping, with the security team to keep the platform safe, and with the AI Platform team on the infra behind Collective's growing AI investments. As a Senior Cloud Infrastructure Engineer on this team, you'll pave new paths, test recovery plans, and manage quiet, stable systems. You'll be the person justifying design choices with production experience and the one others turn to when a system needs rethinking.

What you'll do: 

  • Use Infrastructure as Code (IaC) with Terraform to provision, deploy, and manage cloud resources on AWS and GCP as well as other SaaS vendors.

  • Embed security best practices into the infrastructure by enforcing zero-trust architecture principles like least privilege and identity-based access to protect systems and data.

  • Build scalable, reliable, and cost-effective systems that hold up as Collective grows.

  • Develop and test disaster recovery plans.

  • Own the CI/CD system engineering teams ship on. Set standards, drive reliability and speed improvements, and mentor teams on best practices.

  • Reduce operational toil across the platform through automation, leveraging AI tooling where it accelerates safe, high-quality work.

  • Work closely with product engineering teams to understand application needs and translate them into scalable infrastructure solutions.

  • Own the observability stack (monitoring, logging, alerting) and use it to proactively identify and remediate performance bottlenecks.

  • Participate in the on-call rotation to respond to outages, recover systems, own incident response and post-mortem.

  • Stay current with emerging technologies and best practices in Cloud Infrastructure, DevOps, and Platform Engineering.

What you'll bring:

  • At least 5 years of hands-on experience as a Cloud Infrastructure Engineer, DevOps, or SRE with a proven track record of operating production cloud environments at scale.

  • You operate effectively in ambiguous, fast-changing environments. You can pick up a half-defined problem, define the path forward, and drive it to production without waiting for a playbook.

  • Cloud Platforms: Proficiency in multi-cloud operations. AWS is highly preferred; GCP is a plus.

  • Experience implementing infrastructure and security policy as code.

  • Strong software development skills, preferably in Python or another high level language

  • Strong written and verbal communication skills for driving cross-team alignment. You must be able to clearly and persuasively communicate complex concepts and risks in an engineering-driven environment.

  • Experience mentoring engineers, leading post-incident reviews, or driving cross-team infra initiatives to completion. You're comfortable being the person other engineers ask when something breaks.

  • Ability to write clean, maintainable code for automation and tooling. Experience building internal tools or services to eliminate manual work is a plus.

  • Familiarity with foundational networking concepts and protocols (TCP/IP, DNS, load balancing, VPCs, firewalls) and their application in cloud and hybrid environments.

  • Strong hands-on skills with Linux and command-line tools; you are comfortable using terminals and utilities to manage and debug systems efficiently.

  • Comfort using AI tooling as leverage for infra automation, tooling, and debugging. Bonus if you've built or contributed to AI-assisted DevOps workflows.

Our stack:

  • Cloud: AWS (EC2, IAM, VPC, ECS, Fargate, Lambda, RDS, Elasticache, Opensearch), GCP (BigQuery)

  • Monitoring/Observability: Datadog, Sentry, Amplitude

  • Github, GHA

  • IAC: Terraform, HCP

  • Security tooling

  • Codebase: Python/Typescript/React

What we offer:

  • Hybrid Work Model: Based in San Francisco with a balance of in-office and remote flexibility.

  • Fresh Lunch: Provided on in-office days.

  • Commuter Support: $150 monthly reimbursement for transit expenses.

  • Health & Wellness: $200 quarterly reimbursement to support your well-being.

  • Time Off: Flexible PTO plus 14 company holidays.

  • Comprehensive Coverage: 100% medical, dental, and vision for employees; 75% coverage for dependents.

  • Parental Leave: 16 weeks fully paid.

  • Retirement & Ownership: 401k plan plus an equity package.

  • Team Connection: Quarterly virtual events and an annual in-person summit.

Skills Required

  • At least 5 years of hands-on experience as a Cloud Infrastructure Engineer, DevOps Engineer, or SRE
  • Experience operating production cloud environments at scale
  • Proficiency in multi-cloud operations, including AWS; GCP experience is a plus
  • Experience implementing infrastructure and security policy as code
  • Strong software development skills in Python or another high-level programming language
  • Strong written and verbal communication skills
  • Experience mentoring engineers, leading post-incident reviews, or driving cross-team infrastructure initiatives
  • Ability to write clean, maintainable automation and tooling code
  • Experience building internal tools or services to eliminate manual work
  • Familiarity with TCP/IP, DNS, load balancing, VPCs, firewalls, and cloud or hybrid networking
  • Strong hands-on Linux and command-line skills
  • Comfort using AI tooling for infrastructure automation, tooling, and debugging
  • Experience building or contributing to AI-assisted DevOps workflows
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
229 Employees
Year Founded: 2020

What We Do

Collective is on a mission to redefine the way Businesses-of-One work. Collective is the first online concierge financial platform designed to give self-employed people the technology and team they need so they can focus on their passion, not their paperwork. Collective handles company formation, taxes, accounting, bookkeeping, and more.

Similar Jobs

Kepler Logo Kepler

Infrastructure Engineer

Artificial Intelligence • Hardware • Semiconductor
In-Office
San Jose, CA, USA
55 Employees
190K-225K Annually
In-Office
Sunnyvale, CA, USA
3411 Employees
190K-270K Annually

Handshake Logo Handshake

Senior Cloud Engineer

Edtech • Enterprise Web • HR Tech • Software
In-Office
San Francisco, CA, USA
700 Employees
176K-220K Annually

Trener Robotics Logo Trener Robotics

Infrastructure Engineer

Artificial Intelligence • Robotics • Automation • Manufacturing
In-Office
San Jose, CA, USA
45 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account