Type: Full-time permanent contract or an Internship with a potential follow-up offer
Location: San Francisco or remote with future relocation to San Francisco (sponsored)
Start: ASAP
About usWe're building state-of-the-art context compression. Our mission is to become the "Cloudflare for LLMs" — a compression layer embedded into most LLM pipelines by default.
We're a team of ex-EPFL MSc/PhDs from dlab. We started by publishing papers, then got into YC and started making money helping companies cut their LLM costs.
We run the business like a research lab: form hypotheses, kill the ones that don't work, double down on the ones that do.
A cracked full-stack engineer who enjoys a high-paced startup environment, takes pride in what they build and owns it end to end.
Solid understanding of cloud infrastructure, deployment, and production systems on AWS.
Python/basic ML Ops skills. Experience in scaling AI infra products is a plus.
Proactive, strong communicator with fast response time, team player
Strong backend engineering fundamentals
Experience with concurrency and distributed systems
Experience deploying and scaling production backend services on AWS
Ability to work across systems (Python + light frontend)
Excellent Claude Code (or similar) user
Open-source contributions
Startup experience
OAuth / API auth flows
Backend: Python, FastAPI, PostgreSQL (Supabase), Redis, AWS
Frontend: Next.js, React, TypeScript
Tools: GitHub, Docker, Sentry, GitHub Actions
1. Intro call (20 min)
2. Practical technical interview (60 min)
3. Cultural interview (30 min)
Skills Required
- Strong full-stack engineering ability
- Strong backend engineering fundamentals
- Solid understanding of cloud infrastructure, deployment, and production systems on AWS
- Python and basic MLOps skills
- Experience with concurrency and distributed systems
- Experience deploying and scaling production backend services on AWS
- Ability to work across Python and light frontend development
- Excellent Claude Code or similar AI coding tool usage
- Proactive communication, fast response time, and teamwork
- Experience scaling AI infrastructure products
- Open-source contributions
- Startup experience
- Experience with OAuth and API authentication flows
What We Do
Compresr develops an API for compressing context in large language model (LLM) pipelines and AI agents. Its tools reduce context size while preserving information relevant to a request, helping improve model accuracy, speed, and cost efficiency. The platform supports both coarse-grained compression, which selects relevant chunks, and fine-grained, token-level compression, and is designed for agent and retrieval-augmented generation (RAG) workflows.








