The Role
Build production AI systems combining voice interfaces, browser agents, computer vision, and large language models. Responsibilities include developing speech-to-text and text-to-speech systems, browser automation, prompt and LLM workflows, multimodal pipelines, evaluation frameworks, real-time infrastructure, observability, and production performance improvements.
Summary Generated by Built In
The Opportunity
Core Responsibilities
Technical Requirements
Nice to Have
Why Karumi
Compensation
Join our AI engineering team in the US to build the core intelligence behind our platform. You'll work at the intersection of voice AI, browser automation, and large language models - creating agents that can listen, speak, navigate interfaces, and interact naturally with users in real-time.
This role combines cutting-edge AI with practical systems work. You'll design voice experiences, build browser agents that understand and control web applications, and optimize LLM behavior for production reliability. We ship working AI features that solve real problems, balancing innovation with pragmatic constraints.
We sponsor visas for qualified candidates.
- Build and optimize voice AI systems using speech-to-text and text-to-speech models
- Design browser agents that navigate, understand, and interact with web applications
- Implement browser automation with computer vision and DOM understanding
- Engineer prompt systems and LLM workflows for consistent, intelligent behavior
- Create evaluation frameworks to measure voice quality, agent accuracy, and user experience
- Integrate multimodal AI - combining voice, vision, and language understanding
- Build real-time AI pipelines where latency and reliability are critical
- Manage the AI Infrastructure and take care of it
- Monitor and improve AI system performance in production environments
- Production experience with LLMs (OpenAI, Anthropic, or open-source models)
- Hands-on work with speech AI (STT/TTS systems like Deepgram, ElevenLabs, Whisper)
- Experience with browser automation (Playwright, Puppeteer, Selenium) or computer vision
- Strong Python skills with async programming and real-time systems
- Understanding of prompt engineering, retrieval systems, and agent frameworks
- Ability to debug complex AI behaviors and build observability tools
- Software engineering fundamentals for production AI systems
- Experience building autonomous agents or multi-step AI workflows
- Knowledge of computer vision for UI understanding and visual grounding
- Fine-tuning or training language models for specialized tasks
- Real-time audio processing and streaming architectures
- Background in NLP, machine learning research, or AI systems
- Meaningful equity stake in a backed, fast-growing company\
- Work on cutting-edge voice AI and browser agents in production\
- Shape how AI systems interact with users and software interfaces\
- Small team with direct impact on core product capabilities\
- Gym
- Visa sponsorship available
The base pay range for this role is €50,000 – €125,000 per year.
Skills Required
- Production experience with large language models, including OpenAI, Anthropic, or open-source models
- Hands-on experience with speech AI, including speech-to-text and text-to-speech systems
- Experience with browser automation tools such as Playwright, Puppeteer, or Selenium, or with computer vision
- Strong Python skills, including asynchronous programming and real-time systems
- Understanding of prompt engineering, retrieval systems, and agent frameworks
- Ability to debug complex AI behaviors and build observability tools
- Software engineering fundamentals for production AI systems
- Experience building autonomous agents or multi-step AI workflows
- Knowledge of computer vision for user-interface understanding and visual grounding
- Experience fine-tuning or training language models for specialized tasks
- Experience with real-time audio processing and streaming architectures
- Background in NLP, machine learning research, or AI systems
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Karumi is a San Francisco-based YC F25 startup building AI agents for personalized, real-time software demonstrations. Its agent joins live video calls with prospects, navigates a company’s product, adapts the demo to each user, and operates 24/7 in any language. The platform serves SaaS go-to-market teams by explaining products and creating faster buyer “aha moments” across demo, onboarding, and support.









