Senior SDET – Performance Engineering

Posted 5 Days Ago
Be an Early Applicant
Pune, Mahārāshtra, IND
Hybrid
Senior level
Cloud • Information Technology
The Role
Lead performance engineering for cloud-native microservices on Azure. Design automated performance frameworks, build simulators and traffic generators, define KPIs, integrate tests into CI/CD, execute load/stress/chaos tests, analyze bottlenecks across app/database/infrastructure layers, tune AKS, caching and databases, implement observability, and drive RCA, capacity planning and resilience improvements.
Summary Generated by Built In
Company Description

BETSOL is a cloud-first digital transformation and data management company offering products and IT services to enterprises in over 40 countries. BETSOL team holds several engineering patents, is recognized with industry awards, and BETSOL maintains a net promoter score that is 2x the industry average.

BETSOL’s open source backup and recovery product line, Zmanda (Zmanda.com), delivers up to 50% savings in total cost of ownership (TCO) and best-in-class performance.

BETSOL Global IT Services (BETSOL.com) builds and supports end-to-end enterprise solutions, reducing time-to-market for its customers.

BETSOL offices are set against the vibrant backdrops of Broomfield, Colorado and Bangalore, India.

We take pride in being an employee-centric organization, offering comprehensive health insurance, competitive salaries, 401K, volunteer programs, and scholarship opportunities. Office amenities include a fitness center, cafe, and recreational facilities.

Learn more at betsol.com

Job Description

Role Overview
We are looking for a Senior SDET specializing in Performance Engineering in a cloud-native Azure environment. This role focuses on driving scalability, reliability, and performance validation across distributed microservices systems. The candidate will design automated performance frameworks, build simulators and mocks, define KPIs, and partner with engineering, architecture, and SRE teams to ensure production-grade resilience.
Key Responsibilities

  • Design and implement end-to-end performance and load testing strategies for microservices-based systems
  • Build custom simulators, traffic generators, and mocks for complex system dependencies
  • Define, measure, and track performance KPIs (latency, throughput, error rate, saturation, scalability limits)
  • Develop fully automated performance test frameworks integrated into CI/CD pipelines (Jenkins, GitHub Actions, GitLab)
  • Execute load, stress, spike, endurance, and chaos testing
  • Collaborate with architects, developers, product owners, and SRE teams to optimize system performance
  • Analyze bottlenecks across application, database, and infrastructure layers
  • Work closely with Azure services (AKS, compute, storage, networking) for performance tuning
  • Implement observability using Prometheus, Grafana, and APM tools
  • Optimize Redis caching, database queries (MariaDB, MySQL, etc), and messaging systems
  • Support resilience engineering and chaos testing (Chaos Monkey or equivalent)
  • Drive RCA for performance issues and production incidents
  • Contribute to capacity planning and scalability strategy

Qualifications

Required Skills

  • Strong experience in performance testing tools (K6, JMeter, Gatling, and creating custom frameworks)
  • Proficiency in scripting (Python, C#, Java, or similar)
  • Deep understanding of distributed systems and microservices architecture
  • Hands-on experience with Kubernetes (AKS preferred)
  • Strong knowledge of Azure cloud ecosystem
  • Experience with CI/CD and DevOps practices
  • Understanding of SRE principles (SLI/SLO, error budgets)
  • Experience with observability and monitoring tools
  • Strong database performance tuning expertise

Preferred Skills

  • Experience in contact center / SaaS platforms
  • Exposure to Kafka, RabbitMQ
  • Knowledge of AIOps, AI-driven testing and anomaly detection
  • Experience building custom performance tools or simulators

Qualifications

  • Bachelor’s/Master’s in Computer Science or related field
  • 10+ years experience in QA, development and automation, with strong focus on performance engineering

Additional Information

AI-Driven Performance Engineering (GenAI & AIOps)

  • Leverage Generative AI (GenAI) to auto-generate performance test scenarios, workloads, and synthetic datasets
  • Implement AI-driven anomaly detection for identifying performance regressions and system bottlenecks
  • Use machine learning models for predictive capacity planning and workload forecasting
  • Integrate AIOps tools for intelligent alerting, noise reduction, and automated root cause analysis (RCA)
  • Apply AI techniques for log analysis, pattern recognition, and failure prediction
  • Build self-healing test systems with automated remediation triggers
  • Enhance observability platforms (Prometheus, Grafana) with AI-based insights
  • Utilize AI for dynamic test optimization based on real-time system behavior
  • Collaborate with data science teams to implement advanced analytics for performance insights

AI & Observability Tooling (Real-World Examples)

  • Azure Monitor + Application Insights (with AI capabilities): Smart detection, failure anomaly detection, and auto-root cause insights
  • Azure OpenAI / GenAI integrations: Generate performance scenarios, synthetic workloads, and intelligent test data
  • Dynatrace (Davis AI): Automatic dependency mapping, causal AI for root cause analysis, and real-time anomaly detection
  • Datadog AI / Watchdog: Automated anomaly detection, performance regression identification, and alert correlation
  • New Relic AI: Predictive alerting and performance intelligence across distributed systems
  • Prometheus + Grafana (with ML plugins): Advanced metric analysis and anomaly detection extensions
  • Elastic Stack (ELK) with ML: Log anomaly detection, pattern recognition, and predictive insights
  • Chaos Engineering tools (Gremlin, Chaos Monkey): Integrated with observability platforms for resilience validation
  • k6 + AI-based extensions: Intelligent load modeling and performance insights
  • Custom AI/ML pipelines: Python-based models for predictive scaling, workload modeling, and anomaly detection

Skills Required

  • Strong experience in performance testing tools (k6, JMeter, Gatling, custom frameworks)
  • Proficiency in scripting (Python, C#, Java, or similar)
  • Deep understanding of distributed systems and microservices architecture
  • Hands-on experience with Kubernetes (AKS preferred)
  • Strong knowledge of Azure cloud ecosystem
  • Experience with CI/CD and DevOps practices (Jenkins, GitHub Actions, GitLab)
  • Understanding of SRE principles (SLI/SLO, error budgets)
  • Experience with observability and monitoring tools (Prometheus, Grafana, APM tools)
  • Strong database performance tuning expertise (MariaDB, MySQL, Redis)
  • Experience executing load, stress, spike, endurance, and chaos testing
  • Experience building custom simulators, traffic generators, and mocks
  • Exposure to Kafka or RabbitMQ
  • Knowledge of AIOps, AI-driven testing, anomaly detection, or GenAI integrations
  • Experience with Chaos Engineering tools (Chaos Monkey, Gremlin)
  • Bachelor's or Master's in Computer Science or related field
  • 10+ years experience in QA, development and automation with strong focus on performance engineering
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Denver, CO
775 Employees
Year Founded: 2002

What We Do

BETSOL is a cloud-first digital transformation and data management company offering products and IT services to enterprises in over 40 countries. Our team holds several engineering patents, is recognized with industry awards, and maintains a net promoter score that is 2x the industry average. BETSOL’s open source-based backup and recovery product line, Zmanda (Zmanda.com), delivers 80% savings in total cost of ownership (TCO) and best-in-class performance. BETSOL Global IT Services (BETSOL.com) builds and supports end-to-end enterprise solutions, reducing time-to-market for our customers. Our work locations are set against the vibrant backdrops of Broomfield, Colorado and Bangalore, India. We take pride in being an employee-centric organization, offering comprehensive health insurance, competitive salaries, 401K, volunteer programs, and scholarship opportunities. Our office amenities include a fitness center, cafe, and recreational facilities. Learn more at betsol.com

Similar Jobs

Optum Logo Optum

Lead Full-stack Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Pune, Mahārāshtra, IND
160000 Employees

Optum Logo Optum

Lead Full-stack Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Pune, Mahārāshtra, IND
160000 Employees

Cloudflare Logo Cloudflare

Customer Engineer, India (Based in Mumbai)

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
India
4400 Employees

Mondelēz International Logo Mondelēz International

Data Scientist

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
India
90000 Employees

Similar Companies Hiring

Amplify Platform Thumbnail
Fintech • Financial Services • Consulting • Cloud • Business Intelligence • Big Data Analytics
Scottsdale, AZ
62 Employees
Standard Template Labs Thumbnail
Artificial Intelligence • Information Technology • Software
New York, NY
25 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account