SDE II - AI Platform

Posted 11 Days Ago
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
Mid level
Cybersecurity
The Role
Build and operate Safe’s AI platform, including GPU clusters, Kubernetes infrastructure, inference serving, model lifecycle management, evaluation systems, observability, and secure multi-tenant services. Optimize GPU utilization, latency, throughput, and cost while supporting production LLM workloads. Develop APIs, SDKs, CI/CD workflows, and Terraform-based infrastructure on cloud platforms. Own reliability, security, disaster recovery, incident response, code reviews, mentoring, and end-to-end delivery across platform engineering initiatives.
Summary Generated by Built In

Most boards and executives are currently flying blind when it comes to cyber risk. They are guessing. At Safe, we’ve built an AI-driven engine that finally gives the C-Suite a clear, quantified, and real-time view of their security posture. We don’t just provide data; we provide certainty.

We are a $170M Series C-funded category leader. We don’t play in the mid-market; we operate at the highest levels of global enterprise. Today, we are proud to serve 10% of the Fortune 500, protecting global icons such as Apple, Netflix, AT&T, Verizon, and Victoria’s Secret.

As we scale toward our next chapter, we are looking for high-performers who want to do the best work of their careers at the intersection of AI and Cybersecurity.

The Culture Memo: Our Operating System

Safe is not a typical corporate environment. We are a high-intensity, mission-driven team. We value builders who want to define a category and work alongside people who are equally committed to excellence.

  • Extreme Ownership: We don’t do "not my job." We hire people who see a gap and own the solution from start to finish.

  • The Elite Standard: We serve the most sophisticated companies on the planet. Our work must be bulletproof. Whether it’s a line of code or a sales deck, we aim for Tier-1 quality every time.

  • Methodology & Rigor: We don’t wing it. From Force Management and MEDDICC in sales to data-driven sprints in engineering, we rely on proven frameworks to stay disciplined and predictable.

  • Radical Candor: We move too fast for politics or sugar-coating. We value direct, honest feedback that helps us find the right answer quickly.

  • The Series C Hustle: We have the stability of a well-funded leader but the heart of a startup.

The Perks & Ownership:

We want our team to feel like owners because they are owners. We trust our people to manage their results and their time.

  • Meaningful Equity: Every "Safestar" is a shareholder. You aren’t just an employee; you are a partner in our success.

  • Unlimited Leaves: We don’t believe in clock-watching. We offer unlimited leave because we trust you to take the time you need to recharge while staying committed to the mission.

  • Comprehensive Benefits: We provide top-tier medical insurance and wellness benefits to ensure you and your family are well cared for.

  • Career Trajectory: We are growing aggressively. For high-performers, the path for advancement moves at the speed of your ambition.


AI is not a side project at Safe - it is the engine behind how we quantify cyber risk for the world's largest enterprises. Agentic services, LLM analytics, and inference workloads run in production, on real customer data, under enterprise security constraints.
 
You will own the platform underneath all of it: GPU clusters, inference serving, model lifecycle, and the developer-facing abstractions on top, so every engineering and AI/ML team at Safe ships AI workloads without rebuilding infrastructure each time.
 
This is a platform and infrastructure engineering role, not an ML research role. We want an engineer who has operated GPUs in production, not only consumed them.

AI platform is your primary charter. As an SDE II on the platform team, you will also contribute to core platform engineering, multi-tenant microservices, APIs, and cloud architecture, and carry the same review, mentoring, and delivery ownership as every SDE II at Safe.

What You'll Do:

  • Build the AI platform: APIs, SDKs, and self-service workflows so AI/ML engineers deploy, version, evaluate, and monitor workloads without touching raw infrastructure.

  • Own GPU infrastructure: Cluster design, provisioning, scheduling, isolation, and utilisation optimisation across shared multi-team demand.

  • Run inference in production: vLLM, Triton, KServe, or Ray — continuous batching, autoscaling, model routing, and multi-model serving against real latency and throughput SLOs.

  • Optimize relentlessly: Drive tokens/sec, p99 latency, GPU utilization, and cost per token through quantization, batching strategy, and capacity planning.

  • Engineer GPU-aware Kubernetes: GPU Operator, device plugins, GPU-aware scheduling, MIG partitioning, and distributed workloads over NCCL.

  • Debug the hard layer: GPU OOMs, driver and CUDA runtime mismatches, interconnect bottlenecks, throughput regressions.

  • Build the MLOps backbone: Model registry and versioning, CI/CD for AI workloads, and safe rollout/rollback.

  • Own the evaluation layer: Offline and online eval harnesses, golden datasets, LLM-as-a-judge scoring, accuracy and hallucination metrics, and automated regression gates so no model, prompt, or quantisation change ships without a measured quality verdict.

  • Make it observable: Response quality and accuracy drift, latency, throughput, GPU utilisation, and per-tenant cost — with SLOs, alerting, and incident response.

  • Secure and multi-tenant by default: AuthN/AuthZ, secrets, data protection, tenant isolation, and governed access to models and GPUs.

  • Codify and lead: Terraform for everything, HA and DR designed in, plus architecture direction and mentoring across teams.

  • Contribute beyond AI: Design and build secure, multi-tenant microservices and APIs on AWS, run thorough code reviews, and own feature delivery end-to-end with Product and Design.

What We're Looking For:

  • 2-4 years building and operating production software and infrastructure, with senior ownership of systems end to end.

  • Bachelor's or Master's in Computer Science, Engineering, or equivalent practical experience.

  • Strong Python and/or Go — you write and review production services, not just scripts and manifests.

  • Deep Kubernetes and Docker: scheduling, resource management, operators, networking, debugging under load.

  • Strong Linux, networking, storage, and distributed systems fundamentals.

  • Production AWS/Azure/GCP experience — Lambda, API Gateway, EC2, S3, RDS and equivalents — with Terraform as your default way of working.

  • Working knowledge of SQL and NoSQL databases, including schema design and performance tuning.

  • Experience building and operating backend services and APIs in a multi-tenant SaaS product.

  • Leadership, code review, and mentoring skills, with end-to-end ownership of delivery in an agile environment.

  • Hands-on with a managed ML/AI platform — AWS SageMaker, Bedrock, GCP Vertex AI, or Azure ML — including where it fits and where self-managed infrastructure wins on cost or control.

  • Track record on HA, scalability, and DR in multi-tenant environments.

  • Solid observability and CI/CD practice — metrics, traces, SLOs, automated delivery.

  • Experience building or operating AI evaluation systems — accuracy and quality measurement, LLM-as-a-judge pipelines, eval datasets, and regression testing for models and prompts.

  • Background in platform engineering, infrastructure, SRE, distributed systems, AI infrastructure, or MLOps/LLMOps.

GPU & AI Experience (Required)

  • Hands-on GPU cluster operations: provisioning, capacity planning, and day-2 ownership.

  • NVIDIA ecosystem and CUDA — drivers, container runtime, toolkit compatibility, and their failure modes.

  • GPU scheduling, allocation, isolation, and utilisation optimisation across competing workloads.

  • Kubernetes GPU workloads: GPU Operator, device plugins, GPU-aware scheduling.

  • GPU troubleshooting and tuning: memory/OOM, driver faults, interconnect and throughput bottlenecks.

  • LLM inference and model serving in production (vLLM, Triton, KServe, Ray or equivalent) with demonstrated cost and performance gains.

  • Managed AI platform experience — SageMaker (training jobs, endpoints, inference components) or equivalent — alongside self-managed GPU serving.

Nice to Have:

    SageMaker HyperPod, Bedrock, or Vertex AI at scale · DCGM and Prometheus/Grafana for GPU fleets · MIG and GPU partitioning in multi-tenant setups · NCCL and distributed GPU workloads · quantization and inference optimization (FP8/INT8, AWQ/GPTQ, speculative decoding, KV-cache tuning) · LLM observability and tracing (Langfuse, OpenTelemetry for LLMs) · eval frameworks and human-in-the-loop labeling · PyTorch profiling · GPU cluster schedulers and queueing · vector databases, RAG, LLM gateways, agentic systems · TypeScript/Express API development · React or other modern frontends · security or compliance-bound environments.

What Success Looks Like:

  • AI teams deploy and roll back models through self-service workflows, with no manual infrastructure work per deployment.

  • GPU utilisation is measured, forecast, and consistently optimised across the shared fleet.

  • Cost per token falls quarter over quarter, with the numbers on a dashboard.

  • Inference SLOs hold under peak enterprise load, and GPU incidents drop in frequency and time-to-resolve.

  • No model, prompt, or optimisation change reaches production without passing automated accuracy and quality evaluation.

If you’re passionate about cyber risk, thrive in a fast-paced environment, and want to be part of a team that’s redefining security, we want to hear from you! 🚀

Skills Required

  • 2-4 years building and operating production software and infrastructure
  • Bachelor's or Master's degree in Computer Science, Engineering, or equivalent practical experience
  • Strong Python and/or Go production development experience
  • Deep Kubernetes and Docker experience, including scheduling, resource management, operators, networking, and debugging
  • Strong Linux, networking, storage, and distributed systems fundamentals
  • Production experience with AWS, Azure, or GCP
  • Terraform infrastructure-as-code experience
  • SQL and NoSQL database experience, including schema design and performance tuning
  • Experience building and operating backend services and APIs in multi-tenant SaaS products
  • Leadership, code review, mentoring, and end-to-end delivery ownership
  • Experience with a managed ML/AI platform such as SageMaker, Bedrock, Vertex AI, or Azure ML
  • Experience designing high availability, scalability, and disaster recovery in multi-tenant environments
  • Observability and CI/CD experience, including metrics, traces, SLOs, and automated delivery
  • Experience building or operating AI evaluation systems, including quality measurement, evaluation datasets, LLM-as-a-judge pipelines, and regression testing
  • Background in platform engineering, infrastructure, SRE, distributed systems, AI infrastructure, or MLOps/LLMOps
  • Hands-on GPU cluster operations, including provisioning, capacity planning, and day-two ownership
  • NVIDIA ecosystem and CUDA experience, including drivers, container runtime, toolkit compatibility, and troubleshooting
  • GPU scheduling, allocation, isolation, and utilization optimization
  • Kubernetes GPU workload experience with GPU Operator, device plugins, and GPU-aware scheduling
  • GPU troubleshooting and tuning, including memory, OOM, driver faults, interconnect, and throughput bottlenecks
  • Production LLM inference and model serving experience with vLLM, Triton, KServe, Ray, or equivalent
  • Managed AI platform experience alongside self-managed GPU serving
  • Experience with SageMaker HyperPod, Bedrock, or Vertex AI at scale
  • Experience with DCGM and Prometheus/Grafana for GPU fleets
  • MIG and GPU partitioning experience in multi-tenant environments
  • NCCL and distributed GPU workload experience
  • Quantization and inference optimization experience, including FP8/INT8, AWQ/GPTQ, speculative decoding, or KV-cache tuning
  • LLM observability and tracing experience
  • Evaluation frameworks and human-in-the-loop labeling experience
  • PyTorch profiling experience
  • GPU cluster schedulers and queueing experience
  • Vector databases, RAG, LLM gateways, or agentic systems experience
  • TypeScript/Express API development experience
  • React or other modern frontend experience
  • Experience in security or compliance-bound environments
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Palo Alto, CA
403 Employees
Year Founded: 2012

What We Do

Safe Security is a pioneer in the “Cybersecurity and Digital Business Risk Quantification” (CRQ) space. It helps organizations measure and mitigate enterprise-wide cyber risk in real-time using it’s ML Enabled API-First SAFE Platform by aggregating automated signals across people, process and technology, both for 1st & 3rd Party to dynamically predict the breach likelihood (SAFE Score) & $$ Value at Risk of an organization Headquartered in Palo Alto, Safe Security has over 200 customers worldwide including multiple Fortune 500 companies averaging an NPS of 73 in 2020. Backed by John Chambers and senior executives from Softbank, Sequoia, PayPal, SAP, and McKinsey & Co., it was also one of the Top Contributors to the National Vulnerability Database(NVD) of the U.S. Government in 2019 and the ATT&CK MITRE Contributor in 2020. The company, since 2018, has also been working with MIT in joint research for the development of their SAFE Scoring Algorithm. Safe Security has received several awards including the Morgan Stanley CTO Innovation Award.

Similar Jobs

Capco Logo Capco

Accounts BA

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

Capco Logo Capco

BA - Advisory (Wealth) GCB4

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Software Engineer

Artificial Intelligence • Cloud • Information Technology • Consulting
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
85422 Employees

Hewlett Packard Enterprise Logo Hewlett Packard Enterprise

Senior Engineer

Artificial Intelligence • Cloud • Information Technology • Consulting
In-Office
Bengaluru, Bengaluru Urban, Karnataka, IND
85422 Employees

Similar Companies Hiring

Copia Automation Thumbnail
Cybersecurity • Industrial
New York, New York
50 Employees
SEON Thumbnail
Artificial Intelligence • Cybersecurity
Budapest, Budapest
415 Employees
NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account