Remote · at least 4 hours overlap with US Pacific · Contractor · Initial 3-month engagement
The Role
We are prototyping a new game where AI-driven conversations with NPCs are a core part of the player experience.
We're looking for a mid-level Conversational AI Engineer to help us design and build the AI technology stack behind those interactions.
The core problem: a player speaks to an NPC, and the NPC answers out loud, in character, fast enough to feel like a conversation - target, e.g. under 1 second. If you have built a loop like that and measured where the time goes, we want to hear from you. If you haven't, this role is not a fit.
This is a hands-on engineering role. You will choose the technologies, design the conversational architecture, connect the pieces and build the system in Unity.
What You'll Do
You will own the technical implementation of the prototype's conversational AI system, including:
- Design and build the end-to-end voice conversation pipeline between players and NPCs.
- Evaluate and integrate speech-to-text (STT), LLM, text-to-speech (TTS), and speech-to-speech technologies.
- Evaluate different approaches to conversational memory, context, personality, and NPC behavior.
- Prototype real-time conversational interactions and optimize for latency, quality, reliability, and cost.
- Determine which components should run on-device versus in the cloud.
- Build and integrate the resulting systems into Unity/C#.
- Establish an architecture that can evolve from a prototype into a production system.
You'll work closely with the game design and engineering team to determine what the technology needs to accomplish and how best to make it feel natural and responsive to players.
Our Stack
We have no fixed stack. You will choose it, and you will justify each choice with numbers you measured: response time, cost per conversation minute and how often it fails.
What We're Looking For
Must have (all three):
- 5+ years of professional software engineering.
- At least one shipped project built in Unity with C# (a released game, app or production tool). You will be asked to link it.
- A working voice conversation loop you built yourself, end to end: speech in, language model, speech out. You will be asked to link or demo it.
Nice to have:
- Game AI, NPC systems or interactive entertainment
- Conversational memory and agent architectures
- On-device speech or language models
- Multilingual or language-focused AI systems
Not A Fit If
- Your AI experience is mostly prompt writing, or chat apps built on a single API.
- You have not shipped anything in Unity.
- You want to research or train models rather than build a playable prototype.
Who You'll Be
You'll be deciding what sits between the player's voice and the NPC's reply, not implementing an architecture someone hands you.
You like getting from "let's try this" to a working build fast, with little process and a lot of ownership.
Engagement
Remote · at least 4 hours overlap with US Pacific · Contractor · Initial 3-month engagement
The initial engagement will focus on designing and building a working prototype of the conversational AI system in Unity.
For the right person, there may be an opportunity to continue working with the team as the project moves beyond the prototype stage.
To Apply
We look at the work before the resume. Applications without all three items below are not reviewed.
- Your resume and cover letter.
- A link to a voice or conversational system you built (repo, video, live demo or write-up). Tell us which parts you built yourself.
- Short answers, a few sentences each, in your own words:
- In that system, what was the time from the player finishing speaking to the reply starting, and where did that time go?
- What did you run on-device and what in the cloud, and why?
- What would you try first to make an NPC remember a player between sessions?
Shortlisted people will walk us through their demo on a 30-minute call.
Skills Required
- 5+ years of professional software engineering experience
- At least one shipped project built in Unity with C#, such as a released game, app, or production tool
- A working end-to-end voice conversation loop built independently, including speech input, a language model, and speech output
- Experience with game AI, NPC systems, or interactive entertainment
- Experience with conversational memory and agent architectures
- Experience with on-device speech or language models
- Experience with multilingual or language-focused AI systems
What We Do
At Endless Studios, we envision a digitally empowered future for every student. We're a youth-centric game-making studio where budding creators collaborate with professionals to develop games. We are designed to mirror the authentic environment of a professional game studio to train technical skills like coding and design, alongside soft skills such as leadership and problem-solving. By blending game development with an enriching learning experience, we aim to nurture a community of young creators ready to innovate in a digital future. Endless Studios' parent company, E-Line Media, was founded over a decade ago to develop games that catalyze curiosity, explore meaningful themes, and provide gateways to new perspectives and interests. Our consumer games with meaningful themes (e.g. Never Alone, Beyond Blue) and our game-based learning platforms (e.g. Gamestar Mechanic, MinecraftEDU) have reached a global audience of over 10 million players across channels and been used by over 15,000 schools and after-school programs. Endless Network, our other parent organization, is a global network of companies, foundations, non-profits, technologists, and advocates who want to unlock human potential through technology. We are dedicated to empowering the next generation to succeed in the digital economy. We work to transform the lives of kids and young adults by expanding access to technology and information; and by helping them develop 21st century skills.









