AfterQuery is an applied research lab curating data solutions for foundation model development. We serve every frontier AI lab with the mission of delivering the best data to power the best models. In doing so, we can make expertise that once took a lifetime to build available to anyone who needs it.
Our customers are the ones building the foundation models themselves and our work sits directly in the loop of how those systems improve. This is a rare opportunity to join a company at a defining moment in AI. We are YC's fastest unicorn, valued at $3.2 billion. We're based in San Francisco and backed by leading investors including Altos Ventures, BoxGroup, and Y Combinator and angels from Google DeepMind, OpenAI, Anthropic, Meta Superintelligence Labs, and Microsoft AI.
Why ApplyMassive Opportunity: We are YC's fastest unicorn valued at $3.2 billion and we're not slowing down.
Founding Impact: You will own and architect core infrastructure systems that power our platform from the ground up.
Equity & Growth: Competitive salary and meaningful equity. As we scale, you’ll have the opportunity to shape the engineering organization and lead major technical initiatives.
Strong Team: Our founding team has experience from Citadel Securities, Meta, Google, Silver Lake, and Morgan Stanley — work alongside world-class engineers and researchers.
OverviewAfterQuery builds the data and evaluation systems that power frontier AI models. Every leading AI lab uses our datasets and reinforcement learning environments to encode and scale real-world expertise.
We’re hiring a Machine Learning Engineer, Data Privacy & Anonymization to build the lasting systems that enable us to handle sensitive customer data safely. You'll own anonymization layers, and infrastructure that sits inline with production software, detects identifying information in whatever passes through it, and transforms that data without destroying its usefulness. You’ll support AfterQuery’s mission by curating trainable data from real-world corpuses.
ResponsibilitiesBuild detection models for PII, PHI, and quasi-identifiers across free text, logs, structured payloads, and code
Own the transformation layer: redaction, masking, pseudonymization, tokenization, and format-preserving encryption, chosen per entity and policy
Ship the runtime adaptor: streaming inference, latency budgets, fail-open vs. fail-closed semantics, schema drift
Build evaluation infrastructure with recall-weighted metrics and re-identification attacks against our own output
Make anonymization policy a config surface as customer and jurisdiction requirements diverge
Own high-impact systems from early design through production deployment
3-6 YOE with relevant experiences
Strong software engineering background with experience shipping production systems
Experience building data pipelines at production scale, handling large volumes of data
Applied NLP: NER, sequence labeling, or information extraction on messy text. Fraud or trust/safety detection work where recall on rare events was the objective also counts
Experience with low-latency inference services: streaming pipelines, sidecars, or event systems
Ability to move quickly in a high-ownership, fast-changing environment
Deep care for quality, precision, and customer impact
HIPAA Safe Harbor or Expert Determination, GDPR pseudonymization, differential privacy, k-anonymity, synthetic data, tokenization vaults, or prior health-tech/fintech privacy work.
Health Insurance: Medical, Vision, Dental
401(k) with Employer Match
Daily Meals: Daily UberEats Stipend
Monthly Wellness Stipend
Commute Covered
We are an equal opportunity employer committed to providing a workplace free from discrimination and harassment. Employment decisions are made without regard to legally protected characteristics under applicable federal, state, or local law.
We comply with applicable pay transparency requirements and provide compensation ranges based on the position, qualifications, experience, and other relevant factors. Reasonable accommodations are available to qualified individuals with disabilities and for sincerely held religious beliefs, as required by law. This job description is intended to describe the general nature and level of work performed and is not an exhaustive list of all duties, responsibilities, qualifications, or working conditions associated with the position. We reserve the right to modify this job description as business needs change.
Skills Required
- 3-6 years of experience
- Strong software engineering background with experience shipping production systems
- Experience with applied ML, ranking, recommendations, search quality, marketplace systems, trust/safety, fraud, or data quality systems
- Strong data intuition and ability to work with messy, ambiguous real-world signals
- Comfort working across backend systems, data pipelines, ML models, and internal tools
- Ability to move quickly in a high-ownership, fast-changing environment
- Deep care for quality, precision, and customer impact
AfterQuery Compensation & Benefits Highlights
The following summarizes recurring compensation and benefits themes identified from responses generated by popular LLMs to common candidate questions about AfterQuery and has not been reviewed or approved by AfterQuery.
-
Fair & Transparent Compensation — Pay is considered attractive when work is accepted, with public role postings and materials indicating strong compensation across expert projects and core employee roles. The experts track also highlights transparent pay rates and approval-linked payouts that do land.
-
Healthcare Strength — Employee materials indicate medical, vision, and dental insurance are provided. Job postings reference a comprehensive package consistent with standard coverage.
-
Wellbeing & Lifestyle Benefits — Daily meal stipends, commute support via Uber credits, and a monthly wellness stipend (including gym membership coverage) are prominently advertised. These lifestyle perks suggest attention to day-to-day convenience and wellbeing.
AfterQuery Insights
What We Do
AfterQuery is an applied research lab curating data solutions to accelerate foundation model development.







