Founding Infra Engineer

Posted 2 Days Ago
Be an Early Applicant
Gurugram, Haryana, IND
In-Office
Mid level
Artificial Intelligence • HR Tech • Professional Services • Software
The Role
Design, build, and operate scalable, highly available infrastructure and GPU compute for ML/AI workloads. Manage Kubernetes clusters, optimize GPU utilization, implement IaC and CI/CD, establish monitoring and incident response, and collaborate with ML and application teams to ensure production readiness and cost-efficient performance.
Summary Generated by Built In

This role is for one of Weekday’s clients

Min Experience: 4+ years
Location: Gurugram, Haryana, India, Gurgaon, Haryana, India
JobType: full-time

We are looking for a highly skilled and entrepreneurial Founding Infrastructure Engineer with 4–10 years of experience to build, scale, and own the infrastructure powering our next generation of products and AI workloads. This is an early engineering role with significant ownership, where you will work closely with the founding team to design infrastructure from the ground up and establish systems that are reliable, scalable, secure, and cost-efficient.

The ideal candidate has strong hands-on expertise in infrastructure, GPU computing, and Kubernetes, with a deep understanding of distributed systems and cloud-native technologies. You should be comfortable operating in an ambiguous, fast-paced environment and taking projects from architecture and design through implementation and production operations.


RequirementsKey Responsibilities
  • Design, build, and operate highly available and scalable infrastructure for production and compute-intensive workloads.
  • Architect and manage GPU infrastructure for machine learning, AI, and other high-performance computing workloads.
  • Design, deploy, and maintain Kubernetes clusters across cloud and/or on-premise environments.
  • Optimize GPU utilization, scheduling, networking, storage, and compute resources to maximize performance and cost efficiency.
  • Build infrastructure automation using Infrastructure as Code and modern DevOps practices.
  • Establish reliable deployment, monitoring, observability, alerting, and incident-response systems.
  • Develop scalable solutions for container orchestration, workload scheduling, resource allocation, and service discovery.
  • Manage infrastructure lifecycle, including provisioning, upgrades, capacity planning, performance tuning, and disaster recovery.
  • Work closely with application and ML engineers to provide reliable infrastructure for model training, inference, experimentation, and production services.
  • Identify infrastructure bottlenecks and proactively improve system reliability, performance, scalability, and security.
  • Define engineering best practices around infrastructure architecture, Kubernetes operations, CI/CD, and production readiness.
  • Participate in technical strategy and help shape the infrastructure roadmap as an early member of the engineering team.
Must-Have Skills
  • 4–10 years of hands-on experience in infrastructure engineering, platform engineering, DevOps, or SRE.
  • Strong expertise in infrastructure architecture and operations across production environments.
  • Deep hands-on experience with Kubernetes, including cluster architecture, deployments, networking, storage, scheduling, and troubleshooting.
  • Strong understanding of GPU infrastructure, GPU provisioning, utilization, scheduling, and performance optimization.
  • Experience working with containerization technologies such as Docker and Kubernetes-based workloads.
  • Strong understanding of Linux systems, networking, compute, storage, and distributed systems.
  • Experience with cloud infrastructure and services, preferably AWS, GCP, or Azure.
  • Proficiency with Infrastructure as Code tools such as Terraform and configuration-management/automation tools.
  • Experience building CI/CD pipelines and automated infrastructure workflows.
  • Strong debugging and problem-solving skills across complex production environments.
Good-to-Have Skills
  • Experience with NVIDIA GPUs, CUDA, GPU operators, or GPU orchestration.
  • Experience managing large-scale GPU clusters or AI/ML infrastructure.
  • Knowledge of Kubernetes operators, Helm, service meshes, and cluster autoscaling.
  • Experience with high-performance networking, distributed storage, and workload schedulers.
  • Exposure to ML platforms, model serving, inference infrastructure, or large-scale training systems.
  • Experience with observability tools such as Prometheus, Grafana, OpenTelemetry, or similar technologies.
  • Experience building infrastructure at an early-stage startup or as an early engineering hire.

Skills Required

  • 4-10 years of experience in infrastructure engineering, platform engineering, DevOps, or SRE
  • Hands-on experience with Kubernetes (cluster architecture, networking, storage, scheduling, troubleshooting)
  • Experience architecting and managing GPU infrastructure for ML/AI and HPC workloads
  • Experience with containerization (Docker) and Kubernetes-based workloads
  • Strong understanding of Linux systems, networking, compute, storage, and distributed systems
  • Experience with cloud infrastructure and services (AWS, GCP, or Azure)
  • Proficiency with Infrastructure as Code tools such as Terraform and configuration-management/automation tools
  • Experience building CI/CD pipelines and automated infrastructure workflows
  • Strong debugging and problem-solving skills in complex production environments
  • Experience with NVIDIA GPUs, CUDA, GPU operators, or GPU orchestration
  • Knowledge of Kubernetes operators, Helm, service meshes, and cluster autoscaling
  • Experience with observability tools such as Prometheus, Grafana, OpenTelemetry
  • Experience building infrastructure at an early-stage startup or as an early engineering hire
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
Year Founded: 2021

What We Do

Weekday is an AI-powered recruitment platform that helps startups hire top-tier engineering and product talent. By leveraging a massive database of white-collar professionals and advanced outreach tools, the company streamlines the hiring process through automated sourcing, AI-driven resume screening, and white-glove contingency services. Their mission is to modernize recruitment by enabling companies to discover and engage passive candidates efficiently, ensuring high-quality hires for critical roles.

Similar Jobs

Capco Logo Capco

GCB 5-PMO

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

Cloudflare Logo Cloudflare

Senior Manager, Customer Engineering, India

Cloud • Information Technology • Security • Software • Cybersecurity
Remote or Hybrid
India
4400 Employees

Mondelēz International Logo Mondelēz International

Controller

Big Data • Food • Hardware • Machine Learning • Retail • Automation • Manufacturing
Remote or Hybrid
India
90000 Employees

BlackRock Logo BlackRock

Product Manager

Fintech • Information Technology • Financial Services
In-Office
Gurugram, Haryana, IND
25000 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account