Senior Operations Engineer

Posted Yesterday
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
600K-3M Annually
Senior level
Artificial Intelligence • HR Tech • Professional Services • Software
The Role
Maintain and optimize production Generative AI platforms: automate infrastructure with Python, support GenAI deployments, monitor incidents, implement RAG workflows, and ensure high availability on cloud environments (AWS/GCP).
Summary Generated by Built In

This role is for one of the Weekday's clients

Salary range: Rs 600000 - Rs 2500000 (ie INR 6 - 25 LPA)

Min Experience: 4+ years

Location: Bengaluru
JobType: full-time

We are seeking a highly skilled Senior Operations Engineer to support and optimize modern AI-powered production platforms. The ideal candidate will have strong expertise in Python, Site Reliability Engineering (SRE), DevOps, and GenAI application support. You will be responsible for ensuring the reliability, scalability, and operational excellence of AI workloads while collaborating with engineering teams to deploy and maintain production-grade Generative AI solutions.

This role is ideal for engineers who enjoy solving complex operational challenges, automating infrastructure, and supporting next-generation AI applications in cloud environments.


RequirementsKey Responsibilities
  • Manage, monitor, and optimize production environments supporting Generative AI applications and services.
  • Develop automation tools and operational utilities using Python to improve system reliability and operational efficiency.
  • Support deployment, maintenance, and troubleshooting of GenAI applications built using modern AI frameworks.
  • Monitor production systems, investigate incidents, perform root cause analysis, and implement preventive measures.
  • Collaborate with software engineering, platform, and infrastructure teams to ensure high availability and performance.
  • Deploy, maintain, and optimize AI workloads across AWS, Google Cloud Platform (GCP), or similar cloud environments.
  • Support Retrieval-Augmented Generation (RAG) pipelines and AI application workflows.
  • Implement monitoring, logging, alerting, and incident response processes for production systems.
  • Optimize Linux-based infrastructure, networking, and system performance.
  • Contribute to CI/CD automation, infrastructure improvements, and operational best practices.
  • Maintain operational documentation, runbooks, and incident management procedures.
RequirementsMust-Have Skills
  • Strong programming expertise in Python.
  • KARAT assessment completion (mandatory, if applicable to the hiring process).
  • Hands-on experience developing or supporting Generative AI (GenAI) applications.
  • Experience building or supporting applications using LangChain, LangGraph, and RAG (Retrieval-Augmented Generation) pipelines.
  • Experience working with AWS, Google Cloud Platform (GCP), or OpenAI platforms.
  • Strong experience in Site Reliability Engineering (SRE), DevOps, Production Engineering, or Production Support.
  • Good understanding of Linux systems administration and networking concepts.
  • Experience with monitoring, troubleshooting, incident management, and production operations.
  • Strong understanding of cloud-native infrastructure and operational best practices.
Good-to-Have Skills
  • Experience with LangChain.
  • Experience with LangGraph.
  • Knowledge of Kubernetes, Docker, or container orchestration platforms.
  • Familiarity with CI/CD pipelines and Infrastructure as Code (IaC).
  • Experience supporting Machine Learning or AI production workloads.
  • Exposure to observability tools, monitoring platforms, and log management solutions.
  • Knowledge of automation frameworks and scripting for infrastructure management.
Preferred Qualifications
  • Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related field.
  • 4–10 years of experience in SRE, DevOps, Production Engineering, or Platform Operations.
  • At least 1 year of experience supporting Generative AI or Machine Learning production workloads is preferred.
  • Experience working in cloud-native, distributed, or AI-driven environments.
Soft Skills
  • Strong analytical and troubleshooting skills.
  • Excellent problem-solving and critical thinking abilities.
  • Effective communication and cross-functional collaboration skills.
  • Ability to manage production incidents under pressure.
  • Strong ownership mindset with a focus on operational excellence.
  • Self-motivated, adaptable, and eager to learn emerging AI technologies.
Success Metrics

Success in this role will be measured by maintaining high availability, reliability, and performance of production AI platforms while minimizing system downtime and incident resolution times. Performance will also be evaluated based on the successful deployment and operational stability of GenAI applications, the effectiveness of automation initiatives, and improvements in system monitoring and observability.

Additionally, success will be reflected in proactive incident prevention, continuous optimization of cloud infrastructure, efficient support for AI workloads, and strong collaboration with engineering teams to ensure secure, scalable, and reliable production environments.

Skills Required

  • Strong programming expertise in Python
  • KARAT assessment completion (mandatory, if applicable)
  • Hands-on experience developing or supporting Generative AI (GenAI) applications
  • Experience building or supporting applications using LangChain, LangGraph, and RAG pipelines
  • Experience working with AWS, Google Cloud Platform (GCP), or OpenAI platforms
  • Strong experience in Site Reliability Engineering (SRE), DevOps, Production Engineering, or Production Support
  • Good understanding of Linux systems administration and networking concepts
  • Experience with monitoring, troubleshooting, incident management, and production operations
  • Strong understanding of cloud-native infrastructure and operational best practices
  • Knowledge of Kubernetes, Docker, or container orchestration platforms
  • Familiarity with CI/CD pipelines and Infrastructure as Code (IaC)
  • Experience supporting Machine Learning or AI production workloads
  • Exposure to observability tools, monitoring platforms, and log management solutions
  • Knowledge of automation frameworks and scripting for infrastructure management
  • Bachelor's or Master's degree in Computer Science, IT, Engineering, or related field
  • 4-10 years of experience in SRE, DevOps, Production Engineering, or Platform Operations
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Year Founded: 2021

What We Do

Weekday is an AI-powered recruitment platform that helps startups hire top-tier engineering and product talent. By leveraging a massive database of white-collar professionals and advanced outreach tools, the company streamlines the hiring process through automated sourcing, AI-driven resume screening, and white-glove contingency services. Their mission is to modernize recruitment by enabling companies to discover and engage passive candidates efficiently, ensuring high-quality hires for critical roles.

Similar Jobs

EchoStar Logo EchoStar

Senior Engineer

Aerospace • Cloud • Digital Media • Information Technology • Mobile • News + Entertainment • Generative AI
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
14500 Employees

EchoStar Logo EchoStar

Senior Engineer - NOC Core Advanced Operations

Aerospace • Cloud • Digital Media • Information Technology • Mobile • News + Entertainment • Generative AI
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
14500 Employees

EchoStar Logo EchoStar

Senior NOC Core Advanced Operations Engineer

Aerospace • Cloud • Digital Media • Information Technology • Mobile • News + Entertainment • Generative AI
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
14500 Employees
Hybrid
Bengaluru, Bengaluru Urban, Karnataka, IND
205000 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account