We are seeking an experienced AI/ML Architect to lead the design
and development of scalable, real-time AI systems. You will work closely with
product, data, and engineering teams to architect end-to-end solutions — from
model development and deployment to system integration and production monitoring.
· Design and architect AI/ML systems that are scalable, low-latency, and
production-ready
· Lead development of real-time inference pipelines for use cases like
voice, vision, or NLP
· Select and integrate appropriate tools, frameworks, and infrastructure
(e.g., Kubernetes, Kafka, TensorFlow, PyTorch, ONNX, Triton, VLLM etc.)
· Collaborate with data scientists and ML engineers to productionize
models
· Ensure reliability, observability, and performance of deployed systems
· Conduct architecture reviews, POCs, and system optimizations
· Mentor engineers and help set best practices for ML lifecycle (MLOps)
Requirements
· 6+ years of experience building
and deploying ML systems in production
· Proven expertise in real-time,
low-latency system design (e.g., streaming inference, event-driven
pipelines)
· Strong understanding of scalable
architectures — microservices, message queues, distributed
training/inference
· Proficient in Python and
popular ML/DL frameworks (scikit-learn, TensorFlow, PyTorch)
· Hands-on experience with LLM
inference optimization using frameworks like vLLM, TensorRT-LLM, and SGLang
· Familiarity with vector
databases, embedding-based retrieval, and RAG pipelines
· Experience with containerized
environments (Docker, Kubernetes) and managing multi-container applications
· Working knowledge of cloud
platforms (AWS, GCP, or Azure) and CI/CD practices for ML workflows
· Exposure to edge deployments and model compression/optimization techniques
· Strong foundation in software
engineering principles and system design
· Experience in Linux (Ubuntu)
· Terminal/Bash Scripting
Skills Required
- 6+ years of experience building and deploying machine learning systems in production
- Expertise in real-time, low-latency system design, including streaming inference and event-driven pipelines
- Strong understanding of scalable architectures, microservices, message queues, and distributed training or inference
- Proficiency in Python and machine learning/deep learning frameworks such as scikit-learn, TensorFlow, and PyTorch
- Hands-on experience optimizing LLM inference with vLLM, TensorRT-LLM, and SGLang
- Familiarity with vector databases, embedding-based retrieval, and RAG pipelines
- Experience with containerized environments, Docker, Kubernetes, and multi-container applications
- Working knowledge of AWS, GCP, or Azure and CI/CD practices for machine learning workflows
- Exposure to edge deployments and model compression or optimization techniques
- Strong foundation in software engineering principles and system design
- Experience with Linux, particularly Ubuntu
- Terminal or Bash scripting experience
What We Do
Recode Solutions is a technology solutions provider helping enterprise businesses accelerate digital transformation. It delivers technology consulting and Agentic AI-enabled digital workers for enterprise process transformation, including data and analytics, robotic process automation, integration, software development and operations, quality assurance, digital commerce, and computer-vision-based industrial automation. Its solutions combine custom programming with selected enterprise platforms to optimize operations and improve intelligent, scalable processing.








