AI/ML Architect

Posted 4 Days Ago
Be an Early Applicant
Chennai, Tamil Nadu, IND
In-Office
Senior level
Artificial Intelligence • Information Technology • Consulting • Automation
The Role
Architect scalable, low-latency AI/ML systems and real-time inference pipelines for voice, vision, and NLP applications. Lead model productionization, infrastructure and framework selection, architecture reviews, optimization, observability, and MLOps best practices. Collaborate with product, data science, and engineering teams while mentoring engineers. The role requires expertise in distributed systems, LLM inference optimization, vector databases, RAG, containerization, cloud platforms, CI/CD, edge deployment, and software architecture.
Summary Generated by Built In
Job Summary

We are seeking an experienced AI/ML Architect to lead the design and development of scalable, real-time AI systems. You will work closely with product, data, and engineering teams to architect end-to-end solutions — from model development and deployment to system integration and production monitoring.

Key Responsibilities

·         Design and architect AI/ML systems that are scalable, low-latency, and production-ready

·         Lead development of real-time inference pipelines for use cases like voice, vision, or NLP

·         Select and integrate appropriate tools, frameworks, and infrastructure (e.g., Kubernetes, Kafka, TensorFlow, PyTorch, ONNX, Triton, VLLM etc.)

·         Collaborate with data scientists and ML engineers to productionize models

·         Ensure reliability, observability, and performance of deployed systems

·         Conduct architecture reviews, POCs, and system optimizations

·         Mentor engineers and help set best practices for ML lifecycle (MLOps)



Requirements

·         6+ years of experience building and deploying ML systems in production

·         Proven expertise in real-time, low-latency system design (e.g., streaming inference, event-driven pipelines)

·         Strong understanding of scalable architectures — microservices, message queues, distributed training/inference

·         Proficient in Python and popular ML/DL frameworks (scikit-learn, TensorFlow, PyTorch)

·         Hands-on experience with LLM inference optimization using frameworks like vLLM, TensorRT-LLM, and SGLang

·         Familiarity with vector databases, embedding-based retrieval, and RAG pipelines

·         Experience with containerized environments (Docker, Kubernetes) and managing multi-container applications

·         Working knowledge of cloud platforms (AWS, GCP, or Azure) and CI/CD practices for ML workflows

·         Exposure to edge deployments and model compression/optimization techniques

·         Strong foundation in software engineering principles and system design

 

Nice to Haves

·         Experience in Linux (Ubuntu)

·         Terminal/Bash Scripting

 



Skills Required

  • 6+ years of experience building and deploying machine learning systems in production
  • Expertise in real-time, low-latency system design, including streaming inference and event-driven pipelines
  • Strong understanding of scalable architectures, microservices, message queues, and distributed training or inference
  • Proficiency in Python and machine learning/deep learning frameworks such as scikit-learn, TensorFlow, and PyTorch
  • Hands-on experience optimizing LLM inference with vLLM, TensorRT-LLM, and SGLang
  • Familiarity with vector databases, embedding-based retrieval, and RAG pipelines
  • Experience with containerized environments, Docker, Kubernetes, and multi-container applications
  • Working knowledge of AWS, GCP, or Azure and CI/CD practices for machine learning workflows
  • Exposure to edge deployments and model compression or optimization techniques
  • Strong foundation in software engineering principles and system design
  • Experience with Linux, particularly Ubuntu
  • Terminal or Bash scripting experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
189 Employees
Year Founded: 2018

What We Do

Recode Solutions is a technology solutions provider helping enterprise businesses accelerate digital transformation. It delivers technology consulting and Agentic AI-enabled digital workers for enterprise process transformation, including data and analytics, robotic process automation, integration, software development and operations, quality assurance, digital commerce, and computer-vision-based industrial automation. Its solutions combine custom programming with selected enterprise platforms to optimize operations and improve intelligent, scalable processing.

Similar Jobs

Sutherland Logo Sutherland

Architect

Artificial Intelligence • Analytics
In-Office
Chennai, Tamil Nadu, IND
39547 Employees

Indium Logo Indium

Architect

Artificial Intelligence • Big Data • Cloud • Analytics • Business Intelligence • Generative AI • Big Data Analytics
In-Office
CORP Colony, Tondiarpet, Chennai, Tamil Nadu, IND
5000 Employees

TransUnion Logo TransUnion

Developer, Java, Spark and GCP

Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Hybrid
2 Locations
13000 Employees

TransUnion Logo TransUnion

Rep I

Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
Hybrid
Chennai, Tamil Nadu, IND
13000 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account