Duration: 3 months, with a possibility of a full-time job afterwards
Start date: immediately
About usWe're building state-of-the-art context compression. Our mission is to become the "Cloudflare for LLMs", a compression layer embedded into most LLM pipelines by default.
We're a team of ex-EPFL MSc/PhDs. We started by publishing papers, then got into YC and started making money helping companies cut their LLM costs.
We run the business like a research lab: form hypotheses, kill the ones that don't work, double down on the ones that do.
Competitive compensation
All the resources you need: GPUs, subscriptions, OpenAI/Anthropic credits
As much responsibility as you can handle. Our goal is to make you an irreplaceable part of the team
A fast-paced environment where you'll learn much faster than usual, surrounded by technical people who push each other
Possibility of a full-time offer based on performance
Hands-on supervision. We're around for brainstorming and high-level guidance, but you own your work and will be the person who knows it best.
A well-defined project. We're early-stage and led by customer and market pull, so we work on several directions at once. You'll navigate this alongside the rest of us.
Training wheels. After a short onboarding, you'll work on hard, customer-facing, time-sensitive problems like everyone else. Not a typical internship.
We're running a tight ship on a rough sea. Not for everyone, but you'll come out the other side a much stronger sailor.
About you1. You love research, read papers and hack on new repos for fun
2. Comfortable training ML models/transformers and doing independent applied research
3. Excellent Claude Code (or similar) user
4. Highly ambitious, ready for high-intensity YC startup culture, self-motivated
5. Strong communicator, fast response time, team player
Preferred1. LLM research experience, shown through publications, open-source contributions, or personal projects
2. BSc or MSc in CS/DS, math, or physics.
3. Startup or research internship experience (industry or academic)
Interview processA 40-minute call: 20 minutes for introductions and motivations, followed by 20 minutes of technical questions (mostly ML/LLM foundational questions)
A paid take-home project designed to take around 6 hours, followed by a 30-minute call to walk us through your work and answer a few questions
A 30-minute culture interview with the whole team
Offer
Skills Required
- Love research, read papers, and experiment with new repositories
- Experience training machine learning models and transformers
- Ability to conduct independent applied research
- Excellent Claude Code or similar tool usage
- High ambition, self-motivation, and readiness for high-intensity startup culture
- Strong communication and teamwork skills
- Fast response time
- LLM research experience demonstrated through publications, open-source contributions, or personal projects
- BSc or MSc in computer science, data science, mathematics, or physics
- Startup or research internship experience in industry or academia
What We Do
Compresr develops an API for compressing context in large language model (LLM) pipelines and AI agents. Its tools reduce context size while preserving information relevant to a request, helping improve model accuracy, speed, and cost efficiency. The platform supports both coarse-grained compression, which selects relevant chunks, and fine-grained, token-level compression, and is designed for agent and retrieval-augmented generation (RAG) workflows.








