Site Reliability Engineer

Posted Yesterday
Be an Early Applicant
3 Locations
Remote
Senior level
Big Data • Cloud • Digital Media • Machine Learning • Mobile • Software • Industrial
Autodesk is a global leader in design and make technology that helps innovators everywhere solve today's challenges.
The Role
Designs, builds, and operates reliable, scalable, secure cloud infrastructure for SaaS applications. Responsibilities include Infrastructure as Code, AWS and container management, automation, observability, security hardening, SLO and SLI management, incident response, on-call support, capacity planning, and post-incident improvement. The role collaborates with engineering and cross-functional teams to improve system performance, availability, and operational efficiency.
Summary Generated by Built In

Job Requisition ID #

26WD98218

Position Overview

An exciting opportunity is available for a Site Reliability Engineer to join Autodesk's Product Design and Manufacturing Solutions (PDMS) Platform Site Reliability Engineering (SRE) team. In this role, you will wear multiple hats, including first responder, performance analyst, system architect, capacity planner, and monitoring expert.

You will bring strong technical and communication skills, a passion for learning new technologies, and a problem-solving mindset. You will help build and operate reliable, scalable, secure, and high-performing cloud infrastructure that supports Autodesk products and customers.

Responsibilities

  • Architect and implement hosting solutions for highly dynamic Software as a Service (SaaS) web applications, ensuring reliability, scalability, and performance
  • Design, implement, and maintain Infrastructure as Code (IaC) solutions to support scalable, reliable, and secure global environments
  • Develop and maintain well-documented engineering standards, processes, and best practices
  • Implement infrastructure and application security best practices, including system hardening and the principle of least privilege
  • Use modern infrastructure management tools such as Docker, Terraform, Amazon Web Services (AWS) CloudFormation, and AWS Cloud Development Kit (CDK) to manage and deploy containers and virtual machines
  • Collaborate with Development, Quality Assurance, and Documentation teams throughout the product development lifecycle to ensure quality and reliability
  • Automate operational processes and integrate new technologies to improve efficiency, reliability, and scalability
  • Define and monitor Service Level Objectives (SLOs) and Service Level Indicators (SLIs) and manage error budgets to ensure reliability goals are achieved
  • Partner with stakeholders to align technical strategies with business requirements
  • Participate in on-call support and incident management, ensuring timely resolution and clear stakeholder communication
  • Conduct blameless post-incident reviews to identify root causes, document learnings, and drive continuous improvement
  • Take ownership of initiatives and contribute to a culture of continuous learning, operational excellence, and continuous improvement

Minimum Qualifications

  • 5+ years of experience in DevOps, Site Reliability Engineering (SRE), or a related role supporting cloud-based applications
  • Bachelor's degree in Computer Science or a related technical field
  • Advanced hands-on experience with Linux administration, including monitoring, troubleshooting, reliability, performance, and security
  • Experience managing large-scale cloud infrastructure, preferably on Amazon Web Services (AWS)
  • Strong scripting skills using languages such as Bash, Python, or Perl
  • Expert-level knowledge of AWS services, including Amazon Elastic Compute Cloud (EC2), Elastic Container Service (ECS), Elastic Kubernetes Service (EKS), AWS Lambda, Elastic Load Balancing (ELB), Amazon Simple Storage Service (S3), Identity and Access Management (IAM), Virtual Private Cloud (VPC), Amazon DynamoDB, and Amazon Relational Database Service (RDS)
  • Hands-on experience with Docker, Kubernetes, and container technologies
  • Proficiency with Infrastructure as Code (IaC) tools such as Terraform and AWS CloudFormation
  • Experience with Continuous Integration and Continuous Deployment (CI/CD) tools and technologies such as Jenkins, JFrog Artifactory, and Git
  • Experience with logging, monitoring, and observability tools such as Amazon CloudWatch, Splunk, Dynatrace, New Relic, and Grafana
  • Experience with relational database technologies such as MySQL, PostgreSQL, and Microsoft SQL Server, along with Structured Query Language (SQL)
  • Excellent analytical and problem-solving skills with the ability to work independently
  • Excellent written and verbal communication skills

Preferred Qualifications

  • Experience using Artificial Intelligence (AI)-assisted engineering tools and development practices
  • Experience designing and operating highly available, distributed cloud-native systems
  • Experience with Site Reliability Engineering practices, including observability, capacity planning, incident management, and error budget management
  • Experience automating infrastructure and operational processes at scale
  • Knowledge of cloud security, infrastructure hardening, and compliance best practices
  • Experience working in globally distributed engineering teams

The Ideal Candidate

The ideal candidate is a technically strong and highly motivated Site Reliability Engineer who combines cloud infrastructure expertise, automation skills, and a reliability-first mindset to build and operate secure, scalable, and resilient systems

  • Demonstrates strong technical expertise in cloud infrastructure, Linux administration, containers, automation, and observability
  • Applies Site Reliability Engineering principles to improve system availability, performance, scalability, and operational efficiency
  • Uses analytical thinking and structured problem-solving to troubleshoot complex production issues and identify root causes
  • Takes ownership of services and infrastructure while proactively identifying and addressing reliability risks
  • Automates repetitive operational activities to improve engineering efficiency and reduce manual intervention
  • Responds effectively to incidents while maintaining clear communication and driving timely resolution
  • Learns from incidents and contributes to a blameless culture focused on continuous improvement
  • Collaborates effectively with Software Engineering, Quality Assurance, and cross-functional teams to improve product and platform reliability
  • Demonstrates curiosity and continuously develops expertise in emerging cloud, automation, observability, and Artificial Intelligence technologies
  • Communicates technical concepts clearly and contributes to a collaborative culture focused on reliability, accountability, and engineering excellence

#LI-KJ2

Learn More

About Autodesk

Welcome to Autodesk! Amazing things are created every day with our software – from the greenest buildings and cleanest cars to the smartest factories and biggest hit movies. We help innovators turn their ideas into reality, transforming not only how things are made, but what can be made.

We take great pride in our culture here at Autodesk – it’s at the core of everything we do. Our culture guides the way we work and treat each other, informs how we connect with customers and partners, and defines how we show up in the world.

When you’re an Autodesker, you can do meaningful work that helps build a better world designed and made for all. Ready to shape the world and your future? Join us!

Salary transparency

Salary is one part of Autodesk’s competitive compensation package. Offers are based on the candidate’s experience and geographic location. In addition to base salaries, our compensation package may include annual cash bonuses, commissions for sales roles, stock grants, and a comprehensive benefits package.

Belonging
We take pride in cultivating a culture of belonging where everyone can thrive. Learn more here: https://www.autodesk.com/company/global-belonging


In-Person Onboarding and Identity Verification

This role may require in-person onboarding and/or in-person ID verification.

Are you an existing contractor or consultant with Autodesk?

Please search for open jobs and apply internally (not on this external site).

Skills Required

  • 5+ years of experience in DevOps, Site Reliability Engineering, or a related cloud application support role
  • Bachelor's degree in Computer Science or a related technical field
  • Advanced hands-on Linux administration experience, including monitoring, troubleshooting, reliability, performance, and security
  • Experience managing large-scale cloud infrastructure, preferably AWS
  • Strong scripting skills with Bash, Python, or Perl
  • Expert-level knowledge of AWS services including EC2, ECS, EKS, Lambda, ELB, S3, IAM, VPC, DynamoDB, and RDS
  • Hands-on experience with Docker, Kubernetes, and container technologies
  • Proficiency with Infrastructure as Code tools such as Terraform and AWS CloudFormation
  • Experience with CI/CD tools such as Jenkins, JFrog Artifactory, and Git
  • Experience with logging, monitoring, and observability tools such as CloudWatch, Splunk, Dynatrace, New Relic, and Grafana
  • Experience with MySQL, PostgreSQL, Microsoft SQL Server, and SQL
  • Excellent analytical, problem-solving, written communication, and verbal communication skills
  • Experience designing and operating highly available, distributed cloud-native systems
  • Experience with SRE practices including observability, capacity planning, incident management, and error budget management
  • Experience using AI-assisted engineering tools and development practices
  • Experience automating infrastructure and operational processes at scale
  • Knowledge of cloud security, infrastructure hardening, and compliance best practices
  • Experience working in globally distributed engineering teams

Autodesk Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Autodesk and has not been reviewed or approved by Autodesk.

  • Leave & Time Off Breadth Time away is considered expansive, combining discretionary time off for salaried roles, company holidays/Autodays, and a periodic paid sabbatical. These options provide flexibility beyond standard accrual-based PTO.
  • Equity Value & Accessibility Total rewards prominently include RSUs and an employee stock purchase plan with a discount and lookback, alongside annual bonus or commission programs. These elements are available to eligible employees and can materially augment base pay.
  • Parental & Family Support Family-building support includes reimbursement for adoption, surrogacy, IVF/co‑maternity, and fertility benefits, plus dedicated coaching and Cleo resources for parenting and caregiving. These services extend support before, during, and after leave.

Autodesk Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Francisco, CA
13,285 Employees
Year Founded: 1982

What We Do

Autodesk makes software for people who make things. If you’ve ever driven a high-performance car, admired a towering skyscraper, used a smartphone, or watched a great film, chances are you’ve experienced what millions of Autodesk customers are doing with our software. Autodesk gives you the power to make anything. Over 100 million people use Autodesk software like AutoCAD, Revit, Maya, 3ds Max, Fusion 360, SketchBook, and more to unlock their creativity and solve important design, business and environmental challenges. Our software runs on both personal computers and mobile devices and taps the infinite computing power of the cloud to help teams around the world collaborate, design, simulate and fabricate their ideas in 3D. We provide exceptional compensation/benefit packages and we’d love for you to join us. We’re proud to be an equal opportunity employer and we consider all qualified applicants without regard to race, gender, disability, veteran status or other protected category. To see our culture in action, check out #AutodeskLife. We are headquartered in the San Francisco Bay Area and have more than 10,000 employees worldwide.

Why Work With Us

Our work is impactful. Our people are innovative. And our culture is inclusive. As our software shapes new solutions to the world’s biggest challenges, you shape your career path. With us, you lead the way in achieving sustainability, resilient communities, and an equitable workforce. Discover #AutodeskLife. 

Gallery

Gallery

Similar Jobs

Binance Logo Binance

Site Reliability Engineer

Blockchain • Fintech • Software • Cryptocurrency • Metaverse
In-Office or Remote
17 Locations
7696 Employees

Hyphen Connect Limited Logo Hyphen Connect Limited

Site Reliability Engineer

Agency • Artificial Intelligence • Blockchain • Web3
Remote
3 Locations
7 Employees

DBS Bank Ltd Logo DBS Bank Ltd

Full-stack Engineer

Fintech • Information Technology • Software • Financial Services
In-Office or Remote
17 Locations
41000 Employees

QuEra Computing Logo QuEra Computing

Site Reliability Engineer

Hardware • Quantum Computing
Remote
Tsukuba, Ibaraki, JPN
58 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account