Design and extend production-grade LLM applications and agentic workflows using
NestJS, XState v5, and the OpenAI SDK — flows include RAG, intent detection,
clarification, fulfillment, escalation, tool-use, and human-in-the-loop state machines
- Build and maintain the conversation-machine substrate: guard/action registries, flow
validation (ajv), DB-driven flow configs, and design-time tooling in Epicenter admin
- Build and evolve the AI systems behind Epic Support Assistant (ESA), the
player-facing support chatbot, and Agent Support Assistant, the AI copilot used by
customer support agents
- Integrate with MCP servers (Model Context Protocol) for tool-use and agentic behaviors
- Evaluate, benchmark, and tune models across providers including OpenAI, Gemini,
Anthropic, and future providers; own model selection decisions balancing quality,
latency, throughput, reliability, and cost
- Troubleshoot production LLM issues including hallucinations, retrieval failures, prompt
regressions, model drift, token inefficiencies, latency bottlenecks, and provider outages
- Build resilience mechanisms: retries, fallback routing, caching, streaming, rate limiting,
and provider routing
- Instrument and tune model quality using Langfuse (tracing, evals, prompt
management), evaluation datasets, A/B testing, prompt versioning, and production
telemetry
- Manage async workloads via BullMQ and caching with Redis; PostgreSQL persistence
via Kysely
Requirements
Must-Have
- Proven experience building and operating production LLM-powered systems
similar in scope to chatbots, AI assistants, agent copilots, RAG systems, or LLM
orchestration platforms
- Strong TypeScript/Node.js engineering; TypeScript strict-mode fluency
- Production AI experience: prompt engineering, RAG pipelines, agent design, tool
calling, model evaluation, observability, and failure-mode analysis — you've shipped AI
features, not just prototyped them
- Fullstack depth: comfortable moving between NestJS APIs, React UIs, databases,
infrastructure, and production operations; you don't artificially limit yourself to one layer
- Ability to evaluate tradeoffs between model quality, latency, reliability, throughput,
and cost
- Ability to troubleshoot AI systems across prompts, retrieval pipelines, model
configuration, infrastructure, and application code
- State machine thinking — you naturally model complex async workflows; XState or
similar experience is a strong signal
- Solid understanding of REST API design, async patterns (queues, events), and caching
strategies
- Strong testing culture: unit, integration, and contract tests are first-class deliverables, not
afterthoughts
- Experience working in a monorepo with multiple interconnected services
Strong Plus
- Hands-on experience with MCP (Model Context Protocol) or building tool-use agentic
workflows
- Familiarity with Langfuse or other LLM observability/evaluation platforms
- Experience operating AI workloads at scale
- Experience evaluating multiple foundation models and providers
- Experience building AI copilots, assistants, or conversational products
- Experience with semantic search and retrieval architectures
- Experience with AI gateways such as Portkey or similar platforms
- Experience with NestJS specifically: modules, providers, guards, interceptors, DI
patterns
- Background in customer support or player support platforms — you understand the
stakes of getting AI-generated responses wrong
- Experience shipping under low-latency constraints (chatbot response time budgets,
streaming)
- Previous work in gaming or high-volume consumer products
Skills Required
- Proven experience building and operating production LLM-powered systems such as chatbots, AI assistants, agent copilots, RAG systems, or orchestration platforms
- Strong TypeScript and Node.js engineering skills, including TypeScript strict-mode fluency
- Production AI experience with prompt engineering, RAG pipelines, agent design, tool calling, model evaluation, observability, and failure-mode analysis
- Full-stack experience across NestJS APIs, React UIs, databases, infrastructure, and production operations
- Ability to evaluate tradeoffs among model quality, latency, reliability, throughput, and cost
- Ability to troubleshoot AI systems across prompts, retrieval pipelines, model configuration, infrastructure, and application code
- State-machine design experience; XState or similar experience is strongly preferred
- Understanding of REST API design, asynchronous patterns such as queues and events, and caching strategies
- Strong testing practices covering unit, integration, and contract tests
- Experience working in a monorepo with multiple interconnected services
- Hands-on MCP experience or experience building tool-use agentic workflows
- Familiarity with Langfuse or other LLM observability and evaluation platforms
- Experience operating AI workloads at scale
- Experience evaluating multiple foundation models and providers
- Experience building AI copilots, assistants, or conversational products
- Experience with semantic search and retrieval architectures
- Experience with AI gateways such as Portkey or similar platforms
- Experience with NestJS modules, providers, guards, interceptors, and dependency injection patterns
- Background in customer support or player support platforms
- Experience shipping under low-latency constraints, including chatbot response budgets and streaming
- Previous work in gaming or high-volume consumer products
What We Do
Codurance is a global software consultancy that helps businesses build a better sustainable technical capability to support growth via Software Modernisation, Product Development, Feature Delivery and Platform Engineering. We believe that productive partnerships, collaboration, fast feedback, and small iterations are the best way to deliver successful software projects, using Agile methodologies and Extreme Programming practices, like Test-Driven Development, Simple Design, Pair-Programming and Continuous Integration, in all our projects. We are software craftspeople, passionate about our profession, collaborating with our clients, to help them move into the next stage of growth. Remember to follow us on Twitter (https://twitter.com/codurance) and subscribe to our YouTube channel (https://youtube.com/c/codurance)






