Staff Software Engineer (L4)

Posted 2 Days Ago
Be an Early Applicant
Hiring Remotely in Ireland, IRL
Remote
Senior level
Productivity • Software • Conversational AI
The Role
Staff Site Reliability Engineer responsible for the reliability, availability, performance, capacity, and recovery of production services. Designs resilient distributed systems, defines SLIs and SLOs, builds reliability automation, leads incident response, conducts post-mortems, reduces operational toil, and coordinates complex cross-team projects. The role requires hands-on coding, architecture, observability, production operations, mentorship, and ownership of significant platform areas.
Summary Generated by Built In

Who we are 

At Twilio, we’re shaping the future of communications, all from the comfort of our homes. We deliver innovative solutions to hundreds of thousands of businesses and empower millions of developers worldwide to craft personalized customer experiences.

Our dedication to remote-first work, and strong culture of connection and global inclusion means that no matter your location, you’re part of a vibrant team with diverse experiences making a global impact each day. As we continue to revolutionize how the world interacts, we’re acquiring new skills and experiences that make work feel truly rewarding. Your career at Twilio is in your hands.
.

Hiring and how we work

We use Artificial Intelligence (AI) to help make our hiring process efficient. That said, every hiring decision is made by real Twilions! 

Also, while we are a remote-first company, you may be asked to report in person on an ad-hoc basis for team gatherings, functional off-sites or customer meetings. 
.

See yourself at Twilio

Join the team as Twilio’s next Staff SRE, Platform Engineering

About the job

Twilio is looking for a Staff Site Reliability Engineer to join our Platform Engineering organization. SRE owns production health and resiliency at Twilio — we are accountable for whether our services stay available, performant, and recoverable for the customers who build their businesses on us. SREs here are software engineers who focus on reliability: you will architect systems with reliability designed in from the outset, define the SLIs and SLOs that determine whether we are meeting customer expectations, and build the automation that keeps our infrastructure efficiently ahead of capacity and performance demand.

At the Staff SRE you work without day-to-day guidance, applying deep subject-matter knowledge and industry-leading practice to improve the products, processes, and services that Twilio runs on. You will own the reliability posture of significant parts of our production estate. Your impact will be felt across multiple teams rather than within one.

This is a hands-on engineering role. You will write and deploy code that improves service reliability, orchestrate complex changes across systems, lead the response when production is degraded, and raise the quality bar for the engineers around you.

Responsibilities

In this role, you’ll:

  • Own the reliability posture of production services in your area — availability, latency, capacity, efficiency, performance, and the monitoring and alerting that makes them visible
  • Define, instrument, and operate against SLIs and SLOs, and use error budgets to drive engineering priorities
  • Identify trends and problem areas that threaten stability, and provide a path forward to mitigate risk before it reaches customers
  • Drive down repair items and prevent classes of incidents rather than resolving them one at a time
  • Improve detection, response, and recovery — reducing time to acknowledge, engage, mitigate, and restore, with fewer people pulled in
  • Design for failure: strengthen failure domains, validate recovery paths, and make production changes safer to ship and safer to roll back
  • Participate in on-call for the services you support, and lead the response when production is degraded
  • Write post-mortems that identify true root causes, and drive the follow-up work to completion
  • Oversee efforts to identify, diagnose, report, and document production problems across all reliability dimensions
  • Write, configure, and deploy code that measurably improves service reliability — maintainable, reviewed, documented, and well tested
  • Orchestrate complex changes across systems and services, documenting design changes, technical decisions, migration plans, and upgrades
  • Lead debugging, troubleshooting, and analysis of service architecture and design
  • Use code review to drive up the quality of your coworkers' code
  • Reduce the operational overhead required to run infrastructure and services
  • Drive projects from conception to completion for efforts spanning the concerns of your team
  • Coordinate across programs, collaborating with others to estimate and communicate delivery timelines
  • Break projects into milestones and tasks, track progress, and communicate updates to stakeholders
  • Identify and communicate changes that may impact stability

Qualifications 

Twilio values diverse experiences from all kinds of industries, and we encourage everyone who meets the required qualifications to apply. If your career is just starting or hasn't followed a traditional path, don't let that stop you from considering Twilio. We are always looking for people who will bring something new to the table!

*Required:

  • 8+ years of related engineering experience, with a substantial portion focused on reliability, infrastructure, or platform engineering
  • Demonstrated accountability for production systems — you have carried a pager for services that mattered and owned the outcome when they failed
  • Strong software engineering fundamentals: you build and ship production code, not only configure tooling
  • Experience defining and operating against SLIs and SLOs, and using error budgets to inform engineering priorities
  • Depth in production operations: incident command, post-mortem analysis, capacity planning, and observability
  • A record of preventing recurrence — reducing incident classes and operational toil, not just closing tickets
  • Experience driving changes that span multiple teams, and the communication skills to build alignment without formal authority
  • A track record of improving the engineers around you through code review, design feedback, and mentorship
  • Experience with large-scale distributed systems in a cloud environment

Desired:

  • Familiarity with infrastructure-as-code, container orchestration, and GitOps-style delivery
  • Experience with multi-region architecture, failure-domain design, or regional expansion work
  • Background in chaos engineering, game days, or other proactive resilience validation

Location

This role will be remote, and based in Ireland.

Travel 

We prioritize connection and opportunities to build relationships with our customers and each other. For this role, you may be required to travel occasionally to participate in project or team in-person meetings.

What We Offer

Working at Twilio offers many benefits, including competitive pay, generous time off, ample parental and wellness leave, healthcare, a retirement savings program, and much more. Offerings vary by location.

Twilio thinks big. Do you?

We like to solve problems, take initiative, pitch in when needed, and are always up for trying new things. That's why we seek out colleagues who embody our values — something we call Twilio Magic. Additionally, we empower employees to build positive change in their communities by supporting their volunteering and donation efforts.

So, if you're ready to unleash your full potential, do your best work, and be the best version of yourself, apply now! If this role isn't what you're looking for, please consider other open positions.

.

Stay alert to recruitment fraud

We care about your safety. Scammers sometimes impersonate Twilio recruiters through fake job postings, emails, websites, or messages. Please ensure you are engaging with an official @twilio.com email address. We will never ask for payment, gift cards, cryptocurrency, or banking information during the recruiting process. We do not make job offers without a formal interview process or conduct interviews exclusively through text-based messaging apps.
.

Twilio is proud to be an equal opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or other applicable legally protected characteristics. We also consider qualified applicants with criminal histories, consistent with applicable federal, state and local law. Qualified applicants with arrest or conviction records will be considered for employment in accordance with the Los Angeles County Fair Chance Ordinance for Employers and the California Fair Chance Act. Additionally, Twilio participates in the E-Verify program in certain locations, as required by law.

Skills Required

  • 8+ years of related engineering experience, including substantial experience in reliability, infrastructure, or platform engineering
  • Demonstrated accountability for production systems and experience carrying a pager for critical services
  • Strong software engineering fundamentals and experience building and shipping production code
  • Experience defining and operating against SLIs and SLOs and using error budgets
  • Experience with incident command, post-mortem analysis, capacity planning, and observability
  • Record of preventing incident recurrence and reducing operational toil
  • Experience driving changes across multiple teams and building alignment without formal authority
  • Experience improving engineers through code review, design feedback, and mentorship
  • Experience with large-scale distributed systems in a cloud environment
  • Familiarity with infrastructure as code, container orchestration, and GitOps-style delivery
  • Experience with multi-region architecture, failure-domain design, or regional expansion
  • Background in chaos engineering, game days, or proactive resilience validation

Twilio Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Twilio and has not been reviewed or approved by Twilio.

  • Healthcare Strength — Healthcare coverage spans medical, dental, vision, mental-health programs, and core protections like life and disability insurance. Wellness options and stipends are also emphasized to support ongoing wellbeing.
  • Parental & Family Support — Paid leave for maternity, paternity, and adoption is described as robust. Financial assistance for adoption and fertility needs further extends family support.
  • Leave & Time Off Breadth — Generous PTO, paid holidays, sick days, and dedicated volunteer time are highlighted, with some teams operating an unlimited time-off approach. Company-wide breaks reinforce the ability to take meaningful time away from work.

Twilio Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
6,355 Employees
Year Founded: 2008

What We Do

Millions of developers around the world have used Twilio to unlock the magic of communications to improve any human experience. Twilio has democratized communications channels like voice, text, chat, video, and email by virtualizing the world’s communications infrastructure through APIs that are simple enough for any developer to use, yet robust enough to power the world’s most demanding applications. By making communications a part of every software developer’s toolkit, Twilio is enabling innovators across every industry — from emerging leaders to the world’s largest organizations — to reinvent how companies engage with their customers.

Similar Jobs

Circle (circle.so) Logo Circle (circle.so)

Designer

Artificial Intelligence • Consumer Web • Digital Media • Information Technology • Social Impact • Software
In-Office or Remote
2 Locations
250 Employees
100K-120K Annually

Circle (circle.so) Logo Circle (circle.so)

Senior Quality Engineer

Artificial Intelligence • Consumer Web • Digital Media • Information Technology • Social Impact • Software
In-Office or Remote
2 Locations
250 Employees
120K-130K Annually

Circle (circle.so) Logo Circle (circle.so)

Lead Engineer, AI Quality

Artificial Intelligence • Consumer Web • Digital Media • Information Technology • Social Impact • Software
In-Office or Remote
2 Locations
250 Employees
170K-170K Annually

Mastercard Logo Mastercard

Software Engineer

Blockchain • Fintech • Payments • Consulting • Cryptocurrency • Cybersecurity • Quantum Computing
Remote or Hybrid
Dublin, IRL
38800 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account