Staff Machine Learning Engineer

Reposted 4 Days Ago
Be an Early Applicant
Bengaluru, Bengaluru Urban, Karnataka, IND
In-Office
Senior level
Artificial Intelligence • Cloud • Machine Learning • Retail • Software
The Role
Build and operate production ML infrastructure to convert prototype models (LLM, NLP, recommendation, forecasting, tabular) into low-latency, reliable services. Own CI/CD, containers, orchestration, inference microservices, feature pipelines, model/feature lifecycle, observability, safety/guardrails, and runtime reliability while collaborating with Applied Sciences and Product.
Summary Generated by Built In
About Tekion:

Positively disrupting an industry that has not seen any innovation in over 50 years, Tekion has challenged the paradigm with the first and fastest cloud-native automotive platform that includes the revolutionary Automotive Retail Cloud (ARC) for retailers, Automotive Enterprise Cloud (AEC) for manufacturers and other large automotive enterprises and Automotive Partner Cloud (APC) for technology and industry partners. Tekion connects the entire spectrum of the automotive retail ecosystem through one seamless platform. The transformative platform uses cutting-edge technology, big data, machine learning, and AI to seamlessly bring together OEMs, retailers/dealers and consumers. With its highly configurable integration and greater customer engagement capabilities, Tekion is enabling the best automotive retail experiences ever. Tekion employs close to 3,000 people across North America, Asia and Europe.

Build and operate the production backbone that takes models from Applied Sciences (AS) and delivers reliable, low-latency ML services across Tekion’s DMS, CRM, Digital Retail, Service, Payments, and enterprise products. You’ll own pipelines, microservices, CI/CD, observability, and runtime reliability—working hand-in-hand with Applied Sciences and Product to turn ideas into measurable dealer and consumer impact. 

Why this Role Matters 

  • Accelerate the rollout of LLM-powered and agent-driven features across Tekion products. 
  • Enable agentic workflows that automate, reason, and interact on behalf of users and internal stakeholders. 
  • Operationalize secure, compliant, and explainable LLM and agentic services at scale. 
  • Convert Applied Sciences models into scalable, compliant, cost‑efficient production services. 
  • Standardize how models are trained, validated, deployed, and monitored across Tekion products. 
  • Power real-time, context-aware experiences by integrating batch/stream features, graph context, and online inference. 

What You’ll Do 

  • Turn Applied Sciences prototype models (tabular, NLP/LLM, recommendation, forecasting) into fast, reliable services with well-defined API contracts. 
  • Integrate with the LLM Gateway/MCP, prompt/config versioning. 
  • Build and orchestrate CI/CD pipelines. 
  • Review data science models; refactor and optimize code; containerize; deploy; version; and monitor for quality. 
  • Collaborate with data scientists, data engineers, product managers, and architects to design enterprise systems. 
  • Monitor, detect, and mitigate risks unique to LLMs and agentic systems. 
  • Implement prompt management: versioning, A/B testing, guardrails, and dynamic orchestration based on feedback and metrics. 
  • Design batch/stream pipelines (Airflow/Kubeflow, Spark/Flink, Kafka) and online features linked to our domain graph. 
  • Build inference microservices (REST/gRPC) with schema versioning, structured outputs, and stringent p95 latency targets. 
  • Manage the model/feature lifecycle: feature store strategy, model/agent registry, versioning, and lineage. 
  • Instrument deep observability: traces/logs/metrics, data/feature drift, model performance, safety signals, and cost tracking. 
  • Ensure real-time reliability: autoscaling, caching, circuit breakers, retries/fallbacks, and graceful degradation. 
  • Develop templates/SDKs/CLIs, sandbox datasets, and documentation that make shipping ML the default path. 

Desired Skills and Experience 

  • 8 - 11+ years in ML engineering/MLOps or backend/platform engineering with production ML. 
  • Experience with LLMs, retrieval systems, vector stores, and graph/knowledge stores. 
  • Strong software engineering fundamentals: Python plus one of Java/Go/Scala; API design; concurrency; testing. 
  • Hands-on with orchestration frameworks and libraries (LangChain, LlamaIndex, OpenAI Function Calling, AgentKit, etc.). 
  • Knowledge of agent architectures (reactive, planning, retrieval-augmented agents), and safe execution patterns. 
  • Pipelines and data: Airflow/Kubeflow or similar; Spark/Flink; Kafka/Kinesis; strong data quality practices. 
  • Microservices and runtime: Docker/Kubernetes, service meshes, REST/gRPC; performance and reliability engineering. 
  • Model ops: experiment tracking, registries (e.g., MLflow), feature stores, A/B and shadow testing, drift detection. 
  • Observability: OpenTelemetry/Prometheus/Grafana; debugging latency, tail behavior, and memory/CPU hotspots. 
  • Cloud: AWS preferred (IAM, ECS/EKS, S3, RDS/DynamoDB, Step Functions/Lambda), with cost optimization experience. 
  • Security/compliance: secrets management, RBAC/ABAC, PII handling, auditability. 

Preferred Mindset 

  • Product-oriented: You measure success by dealer and consumer outcomes, not just technical metrics. 
  • Reliability- and safety-first: You move fast with guardrails, rollbacks, and clear SLOs. 
  • Systems thinker: You design for multi-tenant scale, portability, and cost efficiency. 
  • Collaborative: You translate between Applied Sciences, Product, and the Data & AI Platform; you document and teach. 
  • Pragmatic: You automate the 80% and leave room for rapid experimentation. 

Tekion is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, gender (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, victim of violence or having a family member who is a victim of violence, the intersectionality of two or more protected categories, or other applicable legally protected characteristics. 

For more information on our privacy practices, please refer to our Applicant Privacy Notice here.

Skills Required

  • 8-11+ years in ML engineering, MLOps, or backend/platform engineering with production ML
  • Experience with LLMs, retrieval systems, vector stores, and graph/knowledge stores
  • Proficient in Python plus one of Java, Go, or Scala; strong API design, concurrency, and testing skills
  • Hands-on with agent/LLM orchestration libraries (LangChain, LlamaIndex, OpenAI Function Calling, AgentKit, etc.)
  • Knowledge of agent architectures (reactive, planning, RAG) and safe execution patterns
  • Experience designing pipelines and data systems (Airflow/Kubeflow, Spark/Flink, Kafka/Kinesis) and strong data quality practices
  • Experience building inference microservices (Docker/Kubernetes), REST/gRPC, service meshes, and meeting strict latency SLOs
  • Model ops experience: experiment tracking, model/agent registries (e.g., MLflow), feature stores, A/B and shadow testing, drift detection
  • Observability and performance debugging (OpenTelemetry, Prometheus, Grafana); monitor traces/logs/metrics and tail behavior
  • Cloud experience, AWS preferred (IAM, ECS/EKS, S3, RDS/DynamoDB, Step Functions, Lambda) and cost optimization
  • Security and compliance practices: secrets management, RBAC/ABAC, PII handling, and auditability
  • Build and orchestrate CI/CD pipelines, containerize and deploy models, and implement prompt/version management and A/B testing

Tekion Compensation & Benefits Highlights

The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about Tekion and has not been reviewed or approved by Tekion.

  • Healthcare Strength Healthcare coverage includes 100% employer-paid medical, dental, and vision for many U.S. roles, often extending to families. Additional options like HSA/FSA and pet insurance are cited as part of a robust health offering.
  • Fair & Transparent Compensation Pay is considered competitive overall, with total rewards (pay, stock, equity, and benefits) frequently viewed favorably. Many accounts characterize compensation as strong alongside comprehensive benefits.
  • Wellbeing & Lifestyle Benefits Perks such as free meals/snacks, wellness resources, and on-site amenities are highlighted. Flexible or remote options are available for some roles, adding convenience and flexibility.

Tekion Insights

Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Pleasanton, CA
1,858 Employees
Year Founded: 2016

What We Do

At Tekion, we believe that business applications don’t have to be boring. In fact, they should be simple, fun and cool! They should be as delightful to use as your favorite social or consumer application, yet powerful enough to seamlessly and efficiently run global businesses that provide unparalleled consumer experience without compromise. Founded by visionary entrepreneur and innovator Jay Vijayan, we are building the world’s best business applications on the cloud starting with the automotive retail industry. We inherently use cutting-edge technologies like big data, machine learning/AI, and human computer interaction (voice, touch, vision, sensors and IoT). We are inventing new technology along the way to overcome barriers and solve big problems, all while having a blast doing it! Our flagship product offering, Automotive Retail Cloud ™- an industry-first cloud-native retail platform, including all functionalities of a Dealer Management System (DMS) launched recently.

Similar Jobs

Walmart Global Tech Logo Walmart Global Tech

Software Engineer

Big Data • Cloud • Logistics • Machine Learning • Retail
Hybrid
Bangalore, Bengaluru Urban, Karnataka, IND
578950 Employees

Conga Logo Conga

Machine Learning Engineer

Cloud • eCommerce • Information Technology • Payments • Software
Hybrid
Bangalore, Bengaluru Urban, Karnataka, IND
1468 Employees

Infrrd Logo Infrrd

Staff Software Engineer

Artificial Intelligence • Software
In-Office
Bangalore, Bengaluru Urban, Karnataka, IND
227 Employees
Remote or Hybrid
2 Locations
3661 Employees

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account