- Develop novel, valid automated evaluations of AI’s impacts on users, including their mental health and decision making.
- Write code to implement and run automated evaluations, such as user simulators or LLM-as-a-judge pipelines.
- Design methods to improve the ecological validity and realism of automated evaluations for specific populations, such as customizing existing user simulation methods to capture the vocabulary used by children.
- Write and revise judge rubrics to evaluate model behaviors related to user wellbeing and decision making, systematizing abstract, socially situated concepts into clear measurement criteria.
- Collaborate with scientists and research engineers to productionize best practices in AI behavioral evaluation.
- Expertise on quantitative generative AI evaluation and measurement. Good intuition about how to systematize and operationalize complex social concepts.
- Relevant experience designing and validating automated AI evaluation methods, such as LLM-as-a-judge systems or multi-turn benchmarks.
- Proficiency in Python to implement analysis and evaluation tooling.
- Meticulous, good experimental design, epistemic self-awareness and transparency.
- Ability to balance between the needs of AI researchers and domain experts, as well as between researchers and senior decision makers.
- Strong communication skills, low ego, openness to giving and receiving feedback.
- Experience running automated evaluations at scale or in a production context.
- Experience conducting controlled human subjects experiments to validate automated evaluation methods.
- Experience in customer-facing, consulting, or forward-deployed roles translating ambiguous stakeholder needs into concrete deliverables.
- Experience or training in human-centered design or HCI research methods, including working with domain experts or impacted communities.
- Experience or demonstrated interest in studying AI’s psychological or social impacts, such as for crisis support, manipulation or sycophancy, political persuasion, or displacing human relationships.
- Experience designing multilingual generative AI evaluations.
- Experience and comfort using AI coding agents at work.
Skills Required
- Expertise in quantitative generative AI evaluation and measurement
- Ability to systematize and operationalize complex social concepts
- Experience designing and validating automated AI evaluation methods, such as LLM-as-a-judge systems or multi-turn benchmarks
- Proficiency in Python for analysis and evaluation tooling
- Strong experimental design skills
- Meticulousness, epistemic self-awareness, and transparency
- Ability to balance the needs of AI researchers, domain experts, and senior decision makers
- Strong communication skills, low ego, and openness to feedback
- Experience running automated evaluations at scale or in production
- Experience conducting controlled human subjects experiments to validate automated evaluation methods
- Experience in customer-facing, consulting, or forward-deployed roles
- Experience or training in human-centered design or HCI research methods
- Experience or demonstrated interest in AI psychological or social impacts
- Experience designing multilingual generative AI evaluations
- Experience using AI coding agents at work
What We Do
Transluce is an independent research lab that builds open, scalable technology for understanding AI systems and steering them in the public interest. Transluce means to shine light through something to reveal its structure. Today’s complex AI systems are difficult to understand—not even experts can reliably predict their behavior once deployed. Given AI's extraordinary consequences on society, we need scalable and open analyses of the capabilities and risks of AI systems. We are building open source, AI-driven tools to understand and analyze AI systems. We will apply these tools to open-weight models, so the world can vet our analyses and improve their reliability. Once our technology has been vetted, we will work with frontier AI labs and governments to ensure that internal assessments reach the same standards as our publicly vetted procedures. Email: [email protected]









