Forward Deployed Inference Engineer

Posted 3 Days Ago
Be an Early Applicant
Munich, Bayern, DEU
In-Office
Entry level
Artificial Intelligence • Software • Generative AI
The Role
Own customer model deployment and optimization on Tensordyne’s AI inference platform. Profile and benchmark workloads, enable models, validate quality, diagnose performance across software and hardware, support customer PoCs and deployments, and translate recurring technical learnings into tooling, documentation, benchmarks, and product improvements. Collaborate across compiler, runtime, kernel, systems, Product, Sales, and customer teams.
Summary Generated by Built In
About Tensordyne

Tensordyne is building a new class of AI inference system designed for high-performance, power-efficient deployment of the world’s most demanding generative AI workloads.

Our platform combines purpose-built silicon, new AI math, optimized scale-up networking, and memory architecture into a tightly integrated system purpose built for large-scale AI inference. We work with hyperscalers, Neoclouds, frontier model developers, enterprises, and infrastructure partners operating at the leading edge of AI.

As Tensordyne moves from system development into silicon bring-up, customer validation, beta deployments, and production rollout, we are building the technical customer organization that will sit directly between our engineering teams and the companies deploying the platform.

Role summary

We are looking for a Forward Deployed Inference Engineer who combines deep AI systems expertise with strong customer instincts. This person will own the path from a customer workload or model request to a technical result and, where needed, to an optimized model running successfully on Tensordyne hardware and software.

The role sits at the intersection of model architecture, inference performance, systems optimization, developer tooling, and customer deployment. You will work hands-on with engineering while also acting as a technical bridge to Product, BizDev, Sales, and customers.

What you will do
  • Turn customer workloads into fast, credible performance answers through profiling & benchmarking. Define relevant KPIs, compare against competitive baselines, and keep our evaluation methodology current with external benchmarks.
  • Model enablement & optimization: convert and bring up customer models on the Tensordyne stack, validate numerical quality, identify performance bottlenecks, and work with compiler, runtime, kernel, and system teams to improve results.
  • Deployment / forward engineering: work directly with customers and partners on technical PoCs, integration, deployment, and debugging; translate requirements into measurable acceptance criteria for quality, latency, throughput, and other relevant KPIs.
  • Track profiling-to-hardware accuracy by continuously comparing profiling/simulation results with actual hardware deployments, explain material gaps, and flag missing capabilities in the compiler, SDK, inference server, KV-cache management, or adjacent systems to the owning teams.
  • Turn repeated customer-specific learnings into reusable tooling, documentation, benchmarks, or product improvements.
Core qualifications
  • Strong hands-on experience with AI models and inference systems, especially dense and MoE LLMs (Llama, DeepSeek, Qwen, GPT-OSS, Kimi, GLM), and VLM, speech and diffusion models.
  • Strong Python and PyTorch skills and the ability to understand and modify model code.
  • Experience profiling, benchmarking, or optimizing model inference and reasoning about latency, throughput, memory, and utilization.
  • Strong problem-solving and communication skills, with the ability to drive ambiguous technical problems across team boundaries.
  • Proficiency in using AI-powered developer tools (e.g., Claude Code, Cursor).
Strong pluses
  • Experience with LLM serving and deployment stacks such as vLLM, SGLang, or similar systems.
  • Experience working directly with customers or external technical partners.
  • Experience bringing models up on new accelerators or non-standard hardware, including performance debugging across framework/runtime/hardware boundaries.
  • Practical experience with production inference techniques or environments such as quantization, distributed inference, or Kubernetes.
  • Experience navigating and contributing to Rust codebases.
Tensordyne's culture was built on the following values
  • Put people first. We only succeed when our people succeed.
  • Ethics and integrity always; Being open, honest, and respectful of everyone.
  • Think Big. Be ambitious and have audacious goals.
  • Aim for excellence. Quality and excellence count in everything we do.
  • Own it and get it done. Results matter!
  • Make each person better together, than they would be as an individual.
  • Embrace each others’ differences, and embrace that there will be differences.

Tensordyne is an equal opportunity employer. We believe diverse teams are better equipped to solve complex problems and build exceptional technology. All qualified applicants will receive consideration for employment without regard to age, color, gender identity or expression, marital status, national origin, disability, protected veteran status, race, religion, pregnancy, sexual orientation, or any other characteristic protected by applicable laws, regulations, and ordinances.

A note to recruitment agencies: Please do not contact Tensordyne employees or leaders regarding this role. We do not accept unsolicited agency resumes and are not responsible for fees associated with unsolicited submissions.

Skills Required

  • Strong hands-on experience with AI models and inference systems, especially dense and MoE LLMs, VLMs, speech models, and diffusion models
  • Strong Python and PyTorch skills, including the ability to understand and modify model code
  • Experience profiling, benchmarking, or optimizing model inference, including latency, throughput, memory, and utilization analysis
  • Strong problem-solving and communication skills, with the ability to drive ambiguous technical problems across team boundaries
  • Proficiency using AI-powered developer tools such as Claude Code or Cursor
  • Experience with LLM serving and deployment stacks such as vLLM or SGLang
  • Experience working directly with customers or external technical partners
  • Experience bringing models up on new accelerators or non-standard hardware and debugging across framework, runtime, and hardware boundaries
  • Practical experience with quantization, distributed inference, or Kubernetes
  • Experience navigating and contributing to Rust codebases
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: San Jose, CA
113 Employees

What We Do

Every leap in AI has followed the same pattern: first we made models bigger, then we tailored them, and now we let them think longer. Each step (scaling law) truly adds intelligence. But in the end we land on the same runway: inference, and demand is exploding while power supply lags. We asked: what if the next step isn’t stacking another law on top, but a zeroth law beneath them all. A law that changes AI math. Because after all, AI is math, trillions of multiplies, and multiplication burns watts. Tensordyne uses logarithmic compute to turn multiplies into adds, cutting power at the root. We’ve cast our proprietary logarithmic math into custom silicon, hardware, interconnect, and system software. The result: one integrated system for multimodal GenAI inference designed for Hyperscaler and Neo Cloud data centers. What this means for our customers: With Tensordyne they can run the world’s largest multimodal models for thousands of users, with fewer racks, less power, and lower cost. We’re well-funded and fast-moving, with co-headquarters in Sunnyvale, California and Munich, Germany, and a distributed team across North America and Europe. Join us to change how the world runs Gen AI.

Similar Jobs

Zello Logo Zello

Enterprise Account Executive

Logistics • Mobile • Productivity • Software • Transportation
Hybrid
Munich, Bayern, DEU
80 Employees

Magna International Logo Magna International

OT Developer - OT and AMR (m/f/x)

Automotive • Hardware • Robotics • Software • Transportation • Manufacturing
Hybrid
4 Locations
171000 Employees

Magna International Logo Magna International

OT Specialist - robot and collaborative robot (m/f/x)

Automotive • Hardware • Robotics • Software • Transportation • Manufacturing
Hybrid
4 Locations
171000 Employees

Magna International Logo Magna International

OT Developer - OT and Camera / Vision system (m/f/x)

Automotive • Hardware • Robotics • Software • Transportation • Manufacturing
Hybrid
4 Locations
171000 Employees

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
70 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account