Senior Site Reliability Engineer (KR)

Reposted One Month Ago
Be an Early Applicant
Yeoksam-dong, Gangnam-gu, Seoul, KOR
Hybrid
Senior level
Artificial Intelligence
The Role
Own platform reliability for internal cloud and customer-managed environments: maintain Kubernetes/EKS clusters, monitoring, alerting, incident first-response, automation, capacity planning, CI/CD, and drive continuous improvement to keep Panoptes available, performant, and scalable.
Summary Generated by Built In
Gauss Labs is an industrial AI company on a mission to revolutionize manufacturing with AI, starting with the semiconductor sector. Panoptes is an AI-based virtual metrology solution deployed in high-volume manufacturing fabs, helping customers improve yield, reduce costs, and accelerate production. Our software runs in our customers' own managed environments, and we're seeking a Site Reliability Engineer to own the reliability of the infrastructure and platform that Panoptes runs on. You will keep the platform available, performant, and scalable; own monitoring, alerting, incident first-response, and the on-call rotation; and build the automation and observability that let engineering teams operate their services safely.

Responsibilities

  • Platform reliability and operations: Own platform-layer reliability across both environments. In our internal cloud environment: full ownership — cluster health, resource management (CPU/memory/OOM), scheduling, autoscaling, Kubernetes/EKS lifecycle. In the customer environment: operate directly at the application-namespace level and for the customer-controlled cluster/node layer, diagnose and clearly communicate what's needed, and operate the platform within their setup, decisions, and constraints.
  • Monitoring and Alerting: Build and maintain robust monitoring and alerting for the infrastructure and platform layer to proactively identify and resolve issues before they impact the platform.
  • Incident Response: Own incident first-response for the platform layer and participate in the on-call rotation to minimize downtime and restore service quickly.
  • Automation: Develop automation tools and scripts to streamline operations, reduce manual effort, and enable engineering teams to operate their own services safely.
  • Capacity Planning: Forecast resource needs, optimize resource utilization, and ensure the platform infrastructure can handle increasing workloads.
  • Deployment infrastructure: Build and maintain CI/CD pipelines and deployment infrastructure for the platform.
  • Continuous Improvement: Drive a culture of continuous improvement by identifying opportunities to enhance platform reliability, performance, and efficiency.

Basic Qualifications

  • Bachelor's degree in computer science, engineering, or a related discipline
  • 5+ years of industry experience as a Site Reliability Engineer or in platform/infrastructure engineering
  • Hands-on experience operating Kubernetes in production (EKS preferred): cluster lifecycle, scheduling, autoscaling, resource management
  • Experience with cloud platforms (AWS preferred) and containerization technologies (Docker, Kubernetes)
  • Experience with observability and alerting tools (Prometheus, Grafana, ElasticSearch, Jaeger)
  • Experience with scripting languages (Python, Bash)
  • Working knowledge of GitHub, GitHub Actions, and CI/CD concepts
  • Strong problem-solving and troubleshooting skills
  • Working proficiency in English for internal documentation and technical coordination

Preferred Qualifications

  • Knowledge of AI/ML infrastructure and workloads.
  • Knowledge of database technologies (MongoDB, PostgreSQL)
  • Experience operating software in customer-managed (on-prem or customer-cloud) environments
  • Exposure to manufacturing, semiconductor, or enterprise B2B customer environments

[Interview process]
Application reivew - Phone interview - Virtual onsite interview - VP interview/Core Value interview - CEO interview

Skills Required

  • Bachelor's degree in computer science, engineering, or related discipline
  • 5+ years industry experience as a Site Reliability Engineer or platform/infrastructure engineer
  • Hands-on experience operating Kubernetes in production (cluster lifecycle, scheduling, autoscaling, resource management)
  • Experience with EKS
  • Experience with cloud platforms
  • Experience with AWS
  • Containerization technologies (Docker, Kubernetes)
  • Observability and alerting tools (Prometheus, Grafana, ElasticSearch, Jaeger)
  • Scripting languages (Python, Bash)
  • Working knowledge of GitHub, GitHub Actions, and CI/CD concepts
  • Strong problem-solving and troubleshooting skills
  • Working proficiency in English for internal documentation and technical coordination
  • Knowledge of AI/ML infrastructure and workloads
  • Knowledge of database technologies (MongoDB, PostgreSQL)
  • Experience operating software in customer-managed (on-prem or customer-cloud) environments
  • Exposure to manufacturing, semiconductor, or enterprise B2B customer environments
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Jose, CA
51 Employees
Year Founded: 2020

What We Do

We normalize AI. Gauss Labs aims to revolutionize manufacturing by building industrial AI systems beyond human capabilities. Founded in August 2020 with two international locations in San Jose, CA, and Seoul, Korea, Gauss Labs is home to Gaussians who are enthusiastic about pursuing this goal under balanced and inspiring leadership.

Similar Jobs

Datadog Logo Datadog

Manager 2, Technical Enablement Management

Artificial Intelligence • Cloud • Security • Software • Cybersecurity
Easy Apply
Hybrid
3 Locations
6500 Employees

Ericsson Logo Ericsson

Head of MA Asia EAS - Ericsson Antenna System

Cloud • Information Technology • Internet of Things • Machine Learning • Software • Cybersecurity • Infrastructure as a Service (IaaS)
In-Office or Remote
19 Locations
88000 Employees

Coursera + Udemy  Logo Coursera + Udemy

Enterprise Account Executive

Artificial Intelligence • Consumer Web • Edtech • Enterprise Web • HR Tech • Social Impact • Generative AI
Remote or Hybrid
South Korea
1500 Employees

UL Solutions Logo UL Solutions

Engineer

Automotive • Professional Services • Software • Consulting • Energy • Chemical • Renewable Energy
Remote or Hybrid
Korea (Republic of)
15000 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account