Senior Site Reliability Engineer

Posted 12 Days Ago
Be an Early Applicant
2 Locations
In-Office or Remote
Senior level
Cloud • Security • Software • Cybersecurity
Akamai powers and protects life online by solving the toughest challenges and turning the impossible into the possible.
The Role
Designs, develops, and manages reliable applications and infrastructure for large-scale compute services. Responsibilities include building automation and tooling, optimizing deployments and monitoring, supporting Kubernetes and containerized systems, participating in on-call incident response, improving virtualization performance, and contributing to capacity planning and autoscaling for AI compute infrastructure. The role also mentors engineers and promotes reliability, scalability, and operational excellence.
Summary Generated by Built In

Are you passionate about cutting edge technology?

Do solving some of the Internet's most difficult content delivery challenges interest you?

Join our highly skilled Site Reliability team

Our team designs, develops, and manages applications and infrastructure that support Akamai's Compute products and services. We do this while maintaining Akamai's mission at the forefront of what we do. Make life better for billions of people, billions of times a day.

Partner with the best

The Senior Engineer creates solutions to improve automation and efficiency for systems and teams. Responsibilities include optimizing workflows, infrastructure, and applications. Expertise in Linux administration, configuration management, and performance tuning is essential. Collaborate on deployment, monitoring, and resolving incidents. Focus on reliability, scalability, and efficiency through automation and resource optimization. Promote continuous improvement and operational excellence across all systems.

As a Senior Site Reliability Engineer, you will be:

  • Providing support and mentorship for other engineers within the department

  • Developing and maintaining automated tools and scripts to enhance system reliability, deployment processes, and incident response efficiency.

  • Improving our system monitoring to speed error detection and remediation, enhancing performance and reliability of virtualization platform

  • Participating in on-call rotations, guiding restoration and repair of service-impacting issues

  • Writing automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response

  • Contributing to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure

Do what you love

To be successful in this role you will:

  • Possess expert level experience in a SysAdmin (Linux/Unix Administration), DevOps or SRE role, working with large scale distributed systems

  • Demonstrate expertise in Kubernetes and large-scale containerization systems.

  • Possess at least one programming language (Python/Golang) and configuration management with Terraform/SaltStack/Ansible

  • Define SLOs and work with observability tools like Prometheus, Grafana, and distributed tracing to enhance system monitoring.

  • Have experience with architecting software and infrastructure at scale

  • Demonstrate accountability for reliability, develop automation and monitoring, and collaborate effectively with an engineering team unfamiliar with SRE practices.

Build your career at Akamai


Our ability to shape digital life today relies on developing exceptional people like you. The kind that can turn impossible into possible. We’re doing everything we can to make Akamai a great place to work. A place where you can learn, grow and have a meaningful impact.

With our company moving so fast, it’s important that you’re able to build new skills, explore new roles, and try out different opportunities. There are so many different ways to build your career at Akamai, and we want to support you as much as possible. We have all kinds of development opportunities available, from programs such as GROW and Mentoring, to internal events like the APEX Expo and tools such as Linkedin Learning, all to help you expand your knowledge and experience here.

Learn more

Not sure if this job is the right match for you or want to learn more about the job before you apply? Schedule a 15-minute exploratory call with the Recruiter and they would be happy to share more details.


Skills Required

  • Expert-level experience in Linux or Unix administration, DevOps, or Site Reliability Engineering
  • Experience working with large-scale distributed systems
  • Expertise in Kubernetes and large-scale containerization systems
  • Proficiency in at least one programming language, including Python or Golang
  • Experience with configuration management using Terraform, SaltStack, or Ansible
  • Experience defining service-level objectives and using observability tools such as Prometheus, Grafana, and distributed tracing
  • Experience architecting software and infrastructure at scale
  • Experience developing automation and monitoring and collaborating with engineering teams

Akamai Technologies Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Akamai Technologies and has not been reviewed or approved by Akamai Technologies.

  • Leave & Time Off Breadth — Unlimited PTO in the U.S., wellness days, paid volunteering time, and the FlexBase model contribute to broad time‑off flexibility. These practices are described as supporting strong work‑life balance.
  • Parental & Family Support — Generous paid parental leaves, paid family care leave, subsidized backup childcare, and inclusive family‑building benefits via Carrot form a comprehensive family support suite. Coverage spans fertility support through caregiving resources.
  • Retirement Support — A 401(k) with a substantial company match and immediate vesting, plus Roth and Mega Backdoor Roth options, underpin long‑term savings. These features provide reliable long‑term savings opportunities.

Akamai Technologies Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Cambridge, MA
10,285 Employees
Year Founded: 1998

What We Do

At Akamai, we make life better for billions of people, billions of times a day. Every moment, billions of people, all over the world, are using the internet to shop, play games, look after finances, learn remotely, share videos, connect across the world, and so much more. These life-shaping digital experiences wouldn’t be possible without Akamai. We power and protect life online. It’s an extraordinary mission, and our global teams achieve it by solving the toughest challenges, and turning the impossible into the possible. With the world’s most distributed compute platform — from cloud to edge — we make it easy for businesses to develop and run applications, while we keep experiences closer to users and threats farther away. That’s why innovative companies worldwide choose Akamai to build, deliver, and secure their digital experiences. Thanks to our world’s most distributed platform for cloud computing, security, and content delivery. Akamai keeps applications and experiences closer and threats farther away. Devoted, determined problem-solvers who share a passion for technology, we’re always pushing ground-breaking ideas and driving innovation. Do you want to power and protect life online, by solving the toughest challenges with us? Be part of an amazing team!

Why Work With Us

Our people are devoted, determined problem-solvers who share a passion for technology. We find solutions through persistence, and push ground-breaking ideas forward with urgency and courage. This creates an environment where we harness creative energy and drive innovation. We are inclusive and diverse, and we love to collaborate across the world.

Gallery

Gallery

Similar Jobs

Elastic Logo Elastic

Senior Site Reliability Engineer

Cloud • Security • Software • Generative AI
Remote
Poland
3222 Employees
359K-465K Annually

Pragmatike Logo Pragmatike

Senior Site Reliability Engineer

Information Technology • Software
In-Office or Remote
8 Locations
11 Employees

Pragmatike Logo Pragmatike

Senior Site Reliability Engineer

Information Technology • Software
In-Office or Remote
9 Locations
11 Employees

Pragmatike Logo Pragmatike

Senior Site Reliability Engineer

Information Technology • Software
In-Office or Remote
8 Locations
11 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account