Lead Site Reliability Engineer (AI & Cloud Operations)

Posted Yesterday
2 Locations
In-Office or Remote
111K-145K Annually
Senior level
Information Technology
The Role
Leads SRE and cloud operations for highly available, scalable platforms and AI-powered solutions. Designs AWS infrastructure, Kubernetes environments, CI/CD pipelines, Infrastructure as Code, observability, automation, and incident response processes. Develops AI-driven operational capabilities, supports MLOps and model lifecycle management, establishes reliability metrics and SLOs, and mentors engineering teams. Partners with software, data, machine learning, security, and product teams to improve platform performance, resilience, and operational excellence.
Summary Generated by Built In

We are seeking a highly skilled Lead Site Reliability Engineer (AI & Cloud Operations) to drive the reliability, scalability, automation, and operational excellence of our cloud-native platforms and AI-powered solutions. This role will serve as a technical leader responsible for building resilient infrastructure, implementing modern DevOps and SRE practices, and enabling enterprise AI capabilities through automation, observability, and operational intelligence.

The ideal candidate combines deep expertise in AWS cloud technologies, Kubernetes, infrastructure automation, CI/CD, and incident management with hands-on experience supporting AI/ML and Generative AI platforms. This individual will partner closely with software engineering, data engineering, machine learning, security, and product teams to establish highly available systems, streamline deployments, optimize platform performance, and accelerate innovation through AI-driven operations.

Responsibilities
  • Lead the design, implementation, and continuous improvement of Site Reliability Engineering (SRE) practices to ensure highly available, scalable, and resilient cloud platforms.
  • Architect, deploy, and support AWS-based infrastructure and services, including containerized and serverless environments.
  • Build, maintain, and optimize CI/CD pipelines and Infrastructure as Code (IaC) solutions to accelerate and standardize deployments.
  • Develop automation solutions, operational tooling, and self-healing capabilities using Python, Shell, and modern DevOps technologies.
  • Manage Kubernetes and container platforms, ensuring performance, scalability, and operational stability.
  • Establish and enhance observability through monitoring, logging, alerting, and performance management tools to proactively identify and resolve issues.
  • Lead incident response, root cause analysis, problem management, and service reliability improvement initiatives.
  • Partner with engineering, data, AI/ML, and security teams to support enterprise applications, analytics platforms, and cloud-native solutions.
  • Design and implement AI-driven operational capabilities, including intelligent monitoring, automated remediation, predictive analytics, and chatbot-enabled support workflows.
  • Support MLOps and AI platform operations, including model deployment, monitoring, governance, and lifecycle management.
  • Define and track reliability metrics, service-level objectives (SLOs), and operational KPIs to drive continuous improvement.
  • Mentor and provide technical leadership to engineering teams while promoting best practices in reliability, automation, cloud operations, and AI-enabled innovation.


 

Qualifications
  • 5+ years of experience in Site Reliability Engineering (SRE), DevOps, or Cloud Operations.
  • Strong hands-on experience with AWS services including EC2, EKS, ECS, Lambda, S3, RDS, IAM, CloudWatch, and VPC.
  • Experience building and managing CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI/CD, Azure DevOps, or similar platforms.
  • Strong scripting and automation skills using Python, Shell, or similar languages.
  • Experience with Infrastructure as Code (IaC) tools such as Terraform or CloudFormation.
  • Expertise in containerization and orchestration technologies (Docker, Kubernetes).
  • Experience with observability tools such as Prometheus, Grafana, Datadog, Splunk, or ELK Stack.
  • Understanding of analytics platforms, data pipelines, and operational data analysis.
  • Strong troubleshooting, problem-solving, and incident management skills.
  • AI & Automation Experience
  • Experience implementing AI/ML or Generative AI solutions within enterprise environments.
  • Familiarity with AI platforms such as Azure OpenAI, AWS Bedrock, Amazon SageMaker, OpenAI APIs, LangChain, or NVIDIA AI ecosystem.
  • Experience building AI-assisted operational workflows, chatbots, intelligent monitoring, predictive analytics, or automated remediation solutions.
  • Understanding of MLOps concepts, model deployment, monitoring, and governance.

WHAT WE BELIEVE

At Perficient, we promise to challenge, champion, and celebrate our people. You will experience a unique and collaborative culture that values every voice. Join our team, and you’ll become part of something truly special.

 

We believe in developing a workforce that is as diverse and inclusive as the clients we work with. We’re committed to actively listening, learning, and acting to further advance our organization, our communities, and our future leaders… and we’re not done yet.

 

Perficient, Inc. proudly provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, gender, sexual orientation, national origin, age, disability, genetic information, marital status, amnesty, or status as a protected veteran in accordance with applicable federal, state and local laws. Perficient, Inc. complies with applicable state and local laws governing non-discrimination in employment in every location in which the company has facilities. This policy applies to all terms and conditions of employment, including, but not limited to, hiring, placement, promotion, termination, layoff, recall, transfer, leaves of absence, compensation, and training. Perficient, Inc. expressly prohibits any form of unlawful employee harassment based on race, color, religion, gender, sexual orientation, national origin, age, genetic information, disability, or covered veterans. Improper interference with the ability of Perficient, Inc. employees to perform their expected job duties is absolutely not tolerated.

 

Disability Accommodations:

 

Perficient is committed to providing a barrier-free employment process with reasonable accommodations for qualified individuals with disabilities and disabled veterans in our job application procedures. If you need assistance or accommodation due to a disability, please contact us.

 

Applications will be accepted until the position is filled or the posting is removed.

 

The salary range for this position takes into consideration a variety of factors, including but not limited to skill sets, level of experience, applicable office location, training, licensure and certifications, and other business and organizational needs. The new hire salary range displays the minimum and maximum salary targets for this position across all US locations, and the range has not been adjusted for any specific state differentials. It is not typical for a candidate to be hired at or near the top of the range for their role, and compensation decisions are dependent on the unique facts and circumstances regarding each candidate. A reasonable estimate of the current salary range for this position is $111,300 to $144,600. Please note that the salary range posted reflects the base salary only and does not include benefits or any potential variable compensation programs. Information regarding the benefits available for this position are in our benefits overview.

 

Disclaimer:  The above statements are not intended to be a complete statement of job content, rather to act as a guide to the essential functions performed by the employee assigned to this classification.  Management retains the discretion to add or change the duties of the position at any time. 

#LI-RS1

 

About UsPerficient is the global AI and technology consulting firm disrupting the traditional consulting model. Powered by our 7,000+ advisors, engineers, and designers, Perficient implements AI-first solutions that break conventions and deliver outcomes that matter. Proudly serving clients that represent the world’s most innovative brands, and in collaboration with our powerful technology partner ecosystem, we bring deep industry expertise and data-driven design to redefine how businesses run and succeed. Perficient is different. For real. Learn more at perficient.com.

Skills Required

  • 5+ years of experience in Site Reliability Engineering, DevOps, or Cloud Operations
  • Hands-on experience with AWS services including EC2, EKS, ECS, Lambda, S3, RDS, IAM, CloudWatch, and VPC
  • Experience building and managing CI/CD pipelines using Jenkins, GitHub Actions, GitLab CI/CD, Azure DevOps, or similar platforms
  • Strong scripting and automation skills using Python, Shell, or similar languages
  • Experience with Infrastructure as Code tools such as Terraform or CloudFormation
  • Expertise in containerization and orchestration technologies including Docker and Kubernetes
  • Experience with observability tools such as Prometheus, Grafana, Datadog, Splunk, or ELK Stack
  • Understanding of analytics platforms, data pipelines, and operational data analysis
  • Strong troubleshooting, problem-solving, and incident management skills
  • Experience implementing AI/ML or Generative AI solutions within enterprise environments
  • Familiarity with AI platforms such as Azure OpenAI, AWS Bedrock, Amazon SageMaker, OpenAI APIs, LangChain, or the NVIDIA AI ecosystem
  • Experience building AI-assisted operational workflows, chatbots, intelligent monitoring, predictive analytics, or automated remediation solutions
  • Understanding of MLOps concepts, model deployment, monitoring, and governance

Perficient Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Perficient and has not been reviewed or approved by Perficient.

  • Retirement Support — Retirement offerings are positioned as robust, including a 401(k) with company match, an Employee Stock Purchase Plan, and an option for after-tax contributions (mega backdoor Roth). Eligibility details are described as clear for core benefits, supporting confidence in plan access timing.
  • Parental & Family Support — Parental benefits are described as structured, with paid maternity recovery time and paid parental leave for all new parents. Company-paid disability coverage is also highlighted, strengthening the overall family support posture.
  • Fair & Transparent Compensation — Compensation is characterized as generally market-aligned for a portion of roles, with examples of pay being viewed as fair or decent in certain contexts (such as remote or region-specific situations). Variable pay potential appears stronger in some tracks, improving perceived competitiveness for those roles.

Perficient Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Saint Louis, MO
3,295 Employees
Year Founded: 1997

What We Do

Perficient is a leading global digital consultancy. We imagine, create, engineer, and run digital transformation solutions that help our clients exceed customers’ expectations, outpace competition, and grow their business. With unparalleled strategy, creative, and technology capabilities, we bring big thinking and innovative ideas, along with a practical approach to help the world’s largest enterprises and biggest brands succeed.

Similar Jobs

Circle Logo Circle

Staff Software Engineer

Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3
In-Office or Remote
12 Locations
1050 Employees
195K-258K Annually

Ericsson Logo Ericsson

Architect

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office or Remote
2 Locations
88000 Employees

Toast Logo Toast

Senior Manager, Customer Training

Cloud • Fintech • Food • Information Technology • Software • Hospitality
Remote
United States
5000 Employees
149K-238K Annually

Wipfli Logo Wipfli

Tax Senior Manager - Real Estate

Cloud • Fintech • Software • Business Intelligence • Consulting • Financial Services
Remote or Hybrid
Minneapolis, MN, USA
2900 Employees
145K-195K Annually

Similar Companies Hiring

Axle Health Thumbnail
Artificial Intelligence • Healthtech • Information Technology • Logistics
Santa Monica, CA
25 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account