Research Engineer, Content Understanding

Posted Yesterday
Be an Early Applicant
San Francisco, CA, USA
In-Office
180K-350K Annually
Junior
Artificial Intelligence • Software
Building embeddings-based search infrastructure
The Role
Develop large-scale machine learning systems for web content understanding and search quality. Responsibilities include parsing web pages, classifying content, extracting metadata, assessing quality and credibility, detecting misinformation and AI-generated content, deduplicating documents, designing supervision for unlabeled problems, and building efficient models that operate across the web in every language and format.
Summary Generated by Built In

Exa is an applied AI lab building a search engine unlike the world has ever seen. We build massive-scale infra to crawl the entire web, train state-of-the-art embedding models to process it, and design super high performant vector databases to retrieve over it. We now power search for Cursor, Cognition, HubSpot, and over 400,000 developers and have raised $350m from Lightspeed, Benchmark, and a16z.

 

Our ultimate goal is to build perfect search over all the world's information, far beyond Google. If you want to build massive-scale ML systems that will define the way the new AI world consumes information, this is the place for you.

 

As a backend engineer, you'd play a critical role in our search architecture. We're pretty flexible on what projects people work on based on their skills and interests.

Search quality is bounded by what we understand about a page. Before anything can be retrieved, something has to work out what the page actually says. That means parsing it into the parts that are content and the parts that are furniture, classifying what kind of page it is and what it is about, telling whether the page is usable at all, extracting when it was published, judging how good it is and whether it can be trusted, and working out whether it says anything that a page we already have does not. All of this has to work on every page on the web, in every language, in every shape the web comes in.

Some of this is classic document understanding. Some of it is much more open. Credibility and misinformation, AI-generated and machine-spun content, and pages written to be found rather than read are all unsolved, and search results are only as trustworthy as our answers to them.

We are looking for a research engineer to work on this. There is a lot of room to do it well.

Desired Experience
  • Graduate-level ML experience (Master’s or PhD with at least 2 years of relevant experience), or an exceptionally strong undergrad

  • You can build a transformer from scratch in PyTorch, and you have trained models that then had to be cheap enough to run everywhere

  • You like building large-scale datasets and living in the data. Most of the wins here are in the supervision rather than the architecture

  • You are comfortable with problems where the ground truth does not exist yet and defining it is part of the job

  • You care about the problem of finding high quality knowledge and recognize how important this is for the world

Example Projects
  • Make parsing work on the pages where it currently does not, and prove the improvement rather than assert it

  • Teach a model to judge page quality, and get everyone to agree on what quality means well enough to supervise it

  • Work on credibility and misinformation as a modelling problem: what a page claims, whether it is a reliable source of it, and whether it was written for a reader or for a crawler

  • Decide whether two documents are semantically the same or genuinely different, so we can deduplicate the web without collapsing pages that a user would want to see separately

  • Build classification and extraction that is accurate at web scale and cheap enough to run on all of it

  • Design the supervision for something nobody has labels for, and find out whether it is learnable at all

  • Trace a bad search result back to the page-level prediction that caused it, and fix it at the source

Exa is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, creed, color, religion, sex, sexual orientation, gender identity or expression, national origin, disability, age, veteran status, marital status, pregnancy or related conditions, criminal histories consistent with applicable law, or any other basis protected by applicable law.

Skills Required

  • Graduate-level machine learning experience through a master's degree or PhD with at least two years of relevant experience
  • Exceptional undergraduate background may be accepted in place of graduate-level experience
  • Ability to build a transformer from scratch in PyTorch
  • Experience training models that are inexpensive enough to run at large scale
  • Experience building large-scale datasets and working extensively with data
  • Ability to define ground truth for problems where established labels do not exist
  • Interest in and understanding of high-quality knowledge discovery
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Burlington, MA
86 Employees
Year Founded: 2021

What We Do

Exa was built with a simple goal — to organize all knowledge. After several years of heads-down research, we developed novel representation learning techniques and crawling infrastructure so that LLMs can intelligently find relevant information.

Similar Jobs

Block Logo Block

Security Engineer

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
153K-270K Annually

Block Logo Block

Senior Data Engineer

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
168K-297K Annually

Block Logo Block

Enterprise Account Executive

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
In-Office or Remote
8 Locations
12000 Employees
224K-395K Annually

BlackLine Logo BlackLine

Customer Success Manager

Cloud • Fintech • Information Technology • Machine Learning • Software • App development • Generative AI
Hybrid
2 Locations
1810 Employees
142K-178K Annually

Similar Companies Hiring

Revel Thumbnail
Aerospace • Hardware • Robotics • Software
Marina Del Rey, California
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account