Location: Remote (Hybrid opportunity if live in Denver, CO)
Type: Full-Time
Clearance: Must pass FBI fingerprint and background check in multiple states (U.S. Citizenship is strongly preferred)
GovWorx is helping public safety rise to today's greatest challenge: the loss of experience. Our AI-powered platform, CommsCoach, supports 9-1-1 and emergency communications centers across the country by automating quality assurance, training, and real-time call evaluation—allowing agencies to strengthen their teams and better serve their communities. or the one you already have.
Position OverviewWe're looking for an experienced AI Engineer to help build and improve the next generation of AI systems used by public safety agencies across the country. This role sits at the intersection of AI engineering, prompt engineering, and data science.
You'll own the evaluation and continuous improvement of production AI systems, developing automated evaluation pipelines, designing prompt experiments, analyzing model performance, and building tooling that enables rapid iteration. You'll work closely with data scientists, data engineers, and product managers to ensure our AI systems remain accurate, reliable, and trustworthy in real-world public safety environments.
Key ResponsibilitiesDesign, build, and maintain automated AI evaluation pipelines for production LLM applications
Develop prompt engineering strategies and iterate on prompts and compare LLMs using quantitative evaluation methods
Build offline evaluation datasets and regression testing frameworks to measure AI performance over time
Analyze production AI behavior using Python, SQL, and statistical techniques to identify opportunities for improvement
Design experiments, A/B tests, and benchmarking methodologies for evaluating prompt and model changes
Develop dashboards and reporting that communicate AI quality, reliability, and performance metrics
Partner with engineering and product teams to safely deploy and monitor improvements to production AI systems
Investigate model failures through detailed error analysis and recommend improvements to prompts, evaluation datasets, and workflows
Help establish best practices for Responsible AI, evaluation methodologies, and continuous model improvement
Must pass FBI fingerprint and background check in multiple states (U.S. Citizenship is strongly preferred)
3+ years of experience in software engineering, machine learning, data science, or a related technical field
Experience designing evaluation metrics and interpreting AI model performance
Understanding of statistical methods including hypothesis testing and experiment design
Strong Python development experience
Strong SQL skills with experience analyzing large datasets
Experience building or supporting production LLM or Generative AI applications
Experience with prompt engineering and systematic prompt evaluation
Experience using AI evaluation or observability platforms such as Langfuse, LangSmith, MLflow, or Label Studio
Experience with AWS services such as Bedrock, Lambda, S3, Glue, or SageMaker
Experience building dashboards using Tableau, Sisense, Power BI, or similar tools
Knowledge of Responsible AI principles and evaluation methodologies
Help build AI systems that directly support first responders and emergency communications professionals
Own AI quality, evaluation, and continuous improvement for production applications
Work on cutting-edge LLM technologies and help shape the future of Responsible AI
Collaborate with a high-performing team across AI, engineering, product, and data science
Solve technically challenging problems with real-world impact on public safety
Influence AI strategy and evaluation practices across a growing technology company
Skills Required
- Pass FBI fingerprint and background checks in multiple states
- U.S. citizenship
- 3+ years of experience in software engineering, machine learning, data science, or a related technical field
- Experience designing evaluation metrics and interpreting AI model performance
- Understanding of statistical methods, including hypothesis testing and experiment design
- Strong Python development experience
- Strong SQL skills and experience analyzing large datasets
- Experience building or supporting production LLM or Generative AI applications
- Experience with prompt engineering and systematic prompt evaluation
- Experience using AI evaluation or observability platforms such as Langfuse, LangSmith, MLflow, or Label Studio
- Experience with AWS services such as Bedrock, Lambda, S3, Glue, or SageMaker
- Experience building dashboards using Tableau, Sisense, Power BI, or similar tools
- Knowledge of Responsible AI principles and evaluation methodologies
What We Do
At GovWorx, we help Public Safety tackle one of its toughest challenges: keeping great people in the profession. 911 centers face growing pressure from staffing shortages, burnout, and rising public expectations. Most tech solutions focus on operations, we focus on the people behind the calls. Our AI-powered platform supports the entire telecommunicator career cycle, from hiring and training to QA, coaching, and retention. We help agencies: Hire smarter by identifying high-potential candidates Train faster with insight-driven onboarding Coach better through performance data and suggested prompts Retain longer by reducing burnout and recognizing great work We replace biased, manual QA with fair, scalable, AI-driven reviews, giving supervisors clarity, time, and confidence to lead well. We surface the calls that matter, flag risks early, and provide actionable insights to help teams grow. By putting people first, we’re helping agencies solve today’s staffing crisis and build a stronger, more sustainable future for public safety.
Why Work With Us
At GovWorx, we don’t replace people with AI, we empower them. Our tools help agencies hire smarter, train faster, and retain longer. By putting people first, we’re tackling the staffing crisis and preparing the next generation of public safety pros.
Gallery









