The Role
Build, integrate, deploy, troubleshoot, and optimize production generative AI, RAG, agentic AI, and voice AI solutions. Develop Python backend services, APIs, RAG pipelines, conversational systems, and enterprise integrations. Work directly with customers to translate business needs into technical solutions, taking applications from proof of concept to scalable production deployments across cloud, on-premises, and customer environments.
Summary Generated by Built In
SCRY AI is an innovative AI-driven technology company focused on building intelligent, scalable, and high-performance solutions. We work with modern technologies across Software Engineering, AI/ML, and Data to solve real-world business problems and deliver impactful digital products. We are seeking a skilled Forward Deployed AI/ML Software Engineer with strong hands-on experience in Generative AI, RAG, LLMs, backend engineering, and Voice AI. The candidate should combine strong software engineering skills with practical AI/ML expertise and be comfortable working directly with customers to build, integrate, deploy, troubleshoot, and optimize production AI solutions across varied environments.
Key Responsibilities
- Design, develop, and deploy scalable GenAI and RAG applications using LLMs such as OpenAI, Claude, Llama, Qwen, and other open-source models.
- Build complete RAG pipelines including document ingestion, chunking, embeddings, vector databases, retrieval, reranking, prompt engineering, and response generation.
- Develop and integrate Voice Bot / Conversational AI solutions connecting ASR, LLM/SLM, TTS, telephony, APIs, and business workflows.
- Build robust Python backend services and APIs using frameworks such as FastAPI/Flask.
- Develop Agentic AI and workflow automation solutions using frameworks such as LangChain/LlamaIndex, including tool calling and external system integrations.
- Integrate AI applications with SQL/NoSQL databases, APIs, enterprise systems, and customer-specific data sources.
- Work directly with customers and stakeholders to understand requirements, troubleshoot issues, demonstrate solutions, and translate business needs into technical implementations.
- Take AI/ML solutions from POC to production, focusing on scalability, reliability, latency, monitoring, and maintainability.
- Deploy and troubleshoot AI applications using Docker across cloud, on-premises, and customer environments.
- Optimize AI applications for latency, throughput, cost, model performance, and resource utilization.
- Follow strong software engineering practices including Git, CI/CD, testing, logging, monitoring, code reviews, and documentation.
- Collaborate with Data Scientists, ML Engineers, Product, Infrastructure, and customer teams to deliver and support production-ready AI solutions.
Key Qualifications
- 3+ years of experience in Software Engineering, AI/ML Engineering, Data Science, or GenAI, with strong hands-on development experience.
- Strong proficiency in Python, backend development, and software engineering fundamentals.
- Strong hands-on experience building and deploying RAG-based applications and LLM solutions.
- Experience with LangChain, LlamaIndex, Hugging Face, vector databases, embeddings, reranking, or equivalent technologies.
- Good understanding of LLMs, prompt engineering, Agentic AI, tool calling, MCP, and workflow automation.
- Experience with Voice AI / Voice Bots, preferably including ASR, TTS, telephony, and real-time conversational pipelines.
- Experience with PostgreSQL/SQL, MongoDB, Redis, or similar databases, along with REST APIs and backend integrations.
- Strong experience with FastAPI/Flask, microservices, Docker, and backend architecture; Kubernetes is a plus.
- Experience with AWS/Azure/GCP and cloud or on-premises deployments.
- Knowledge of CI/CD, automated testing, Git, logging, monitoring, production debugging, and AI/ML inference optimization.
- Strong problem-solving, communication, customer-facing, and collaboration skills, with the ability to work independently in fast-paced environments.
Good to Have
- Experience with Kubernetes, cloud-native architectures, CI/CD, and production observability.
- Exposure to MCP, Agentic AI frameworks, AI evaluation, model optimization, and LLMOps/MLOps.
- Experience working with customer deployments, POCs, enterprise integrations, and troubleshooting in cloud/on-premises environments.
Skills Required
- 3+ years of experience in software engineering, AI/ML engineering, data science, or generative AI
- Strong proficiency in Python, backend development, and software engineering fundamentals
- Hands-on experience building and deploying RAG-based applications and LLM solutions
- Experience with LangChain, LlamaIndex, Hugging Face, vector databases, embeddings, reranking, or equivalent technologies
- Understanding of LLMs, prompt engineering, agentic AI, tool calling, MCP, and workflow automation
- Experience with Voice AI or voice bots, including ASR, TTS, telephony, and real-time conversational pipelines
- Experience with PostgreSQL or SQL, MongoDB, Redis, similar databases, REST APIs, and backend integrations
- Strong experience with FastAPI or Flask, microservices, Docker, and backend architecture
- Experience with AWS, Azure, GCP, and cloud or on-premises deployments
- Knowledge of CI/CD, automated testing, Git, logging, monitoring, production debugging, and AI/ML inference optimization
- Strong problem-solving, communication, customer-facing, collaboration, and independent working skills
- Experience with Kubernetes, cloud-native architectures, CI/CD, and production observability
- Exposure to MCP, agentic AI frameworks, AI evaluation, model optimization, and LLMOps/MLOps
- Experience with customer deployments, proof-of-concepts, enterprise integrations, and cloud or on-premises troubleshooting
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Scry AI is a research-led enterprise AI company that develops intelligent platforms for businesses in banking, financial services, insurance, logistics, and industrial sectors. Its suite includes Auriga for conversational AI, Collatio for document intelligence, and Concentio for cognitive IoT and operational intelligence. The platforms process fragmented data, automate document and workflow operations, support compliance, and generate actionable insights to improve enterprise efficiency.






