Most current RSI work is highly LLM-centric, treating model weights as the sole unit of improvement. Every step inherits the massive cost of a training run, and compounding progress arrives late. We take a whole-system view: the LLM is just one component of a larger reasoning system that includes code, prompts, search strategies, and tool use. We are building a self-optimizing optimizer—a system where every task it tackles supplies the signal needed to optimize its own orchestration code.
As an AI Scientist, you will propose, explore and hands-on build the core algorithms for our self-improving reasoning engine. You will push the frontier on the data-efficient methods that allow our system to learn how to reason – how to probe LLMs, extract their hidden knowledge, and synthesize fragments into reliable, complex answers.
You are a good fit if you:
Have a deep, first-principles understanding of LLM reasoning, failure modes, and understand their “quirks” through experience.
Are an expert in ML algorithm design, search, or optimization, and know the limitations of common LLM-training methods.
Excel at designing novel, data-efficient methods for discovering optimal, task-specific reasoning strategies.
Are passionate about designing the core self-improvement loops that allow a system to learn from the problems it solves to get better at the next one, autonomously.
Have a strong publication record and thrive bringing crazy ideas to fruition.
Our system-level RSI reaches state-of-the-art (SOTA) performance without modifying a single LLM parameter:
12.3% boost on frontier models for LiveCodeBench Pro, setting a new SOTA at 93.9%, among many other SOTA results we have shared on our blog at poetiq.ai.
Universal improvement: Every model tested improved, proving our harness encodes highly transferable task structure.
Dominated 6 unseen benchmarks spanning competition mathematics, scientific coding, long-horizon planning, agentic tool use, and long-context retrieval—all automatically.
We prioritize explainability. By running our optimization loops at the system level rather than baking them into uninterpretable parameters, every improvement remains human-readable—transparent code, explicit prompts, and clear data. We believe fast, powerful RSI is fully compatible with tighter, more deliberate oversight. This approach values thoughtful, rigorous diagnostic engineering over blindly scaling compute and black-box models.
We are a high-leverage team of 10 engineers and researchers. We thrive in an in-office environment built around high-bandwidth collaboration, rapid whiteboarding, and low-ego problem solving.
Our engineering culture is highly collaborative, mentorship-driven, and deeply inclusive. We value clear communication, rigorous testing, and deliberate architectural design. At Poetiq, you won't just be optimizing weights on the periphery; you will be core to designing the interpretable reasoning architectures of the future.
Skills Required
- Deep, first-principles understanding of LLM reasoning, failure modes, and practical quirks.
- Expertise in ML algorithm design, search, or optimization and knowledge of LLM training limitations.
- Ability to design novel, data-efficient methods for discovering task-specific reasoning strategies.
- Passion and experience designing self-improvement loops enabling autonomous learning from solved problems.
- Strong publication record and demonstrated ability to execute ambitious research ideas.
What We Do
Poetiq’s platform integrates with any major frontier large language model, including ChatGPT, Claude and Gemini, to reduce the data, time and cost required for advanced problem-solving.







