Web Crawling - Research Engineer

Posted Yesterday
Be an Early Applicant
San Francisco, CA, USA
Hybrid
350K-475K Annually
Senior level
Artificial Intelligence • Information Technology
The Role
Build and own internet-scale web-crawling and ingestion systems for pretraining data. Responsibilities include designing distributed crawlers, extraction and deduplication pipelines, data-quality filtering, specialized crawlers, and petabyte-scale infrastructure. The role partners with pretraining teams to assess model impact, improves system reliability and efficiency, and helps define technical direction. Candidates need extensive experience with web crawlers or large-scale data acquisition, distributed systems, and the practical and legal aspects of web data collection.
Summary Generated by Built In
About Thinking Machines

The mission of Thinking Machines is to build AI that extends human will and judgment. We are training frontier models with Inkling, developing Tinker to let people make models their own, and crafting interfaces that broaden human-AI communication. We believe the future worth building is human, and we're hiring people who want to build it.

About the Role

We're hiring a Software Engineer to build and own our web-crawling systems, from distributed collection at internet scale through filtering, deduplication, and deciding what data we keep.

The ideal candidate has built and scaled a web crawler or large-scale data-acquisition systems. In this role, you'll write and own production systems: the crawler itself, the infrastructure that runs it at scale, and the pipelines that turn raw crawls into usable pretraining data. You'll work closely with our pretraining and data teams to understand what's actually moving model quality, but this is fundamentally an engineering role, not a research one.

What You'll Do
  • Design and scale the web crawler and ingestion infrastructure that sources Inkling's pretraining data

  • Build pipelines for large-scale extraction, deduplication, and data quality filtering

  • Build specialized crawlers for high-value or hard-to-reach data sources

  • Work with the pretraining team to understand how changes in crawled data affect model performance

  • Improve the reliability and efficiency of crawling and ingestion infrastructure at petabyte scale

  • Help set technical direction for this area as it grows, and bring other engineers up to speed on what you've learned

Skills & QualificationsMinimum Qualifications
  • 8+ years designing, building, and scaling web crawlers, scrapers, or large-scale distributed data-acquisition systems

  • A track record of owning crawler or data-acquisition infrastructure at internet scale

  • Strong software engineering skills in a language such as Python, Go, or Rust, with real experience in distributed systems

  • Working knowledge of the practical and legal considerations of large-scale web data collection (robots.txt, rate limiting, licensing)

Preferred Qualifications
  • Experience applying machine learning to crawl selection, extraction, or data quality classification at internet scale

  • Experience setting technical direction for a crawling, data acquisition, or search infrastructure team, whether or not that was your formal title

  • Experience designing systems for petabyte-scale storage and processing

  • Track record of open-source contributions to crawling, scraping, or data infrastructure tools

  • Background at a search engine (crawling, indexing) or a frontier AI lab's data acquisition team

Logistics
  • Location: This role is based in San Francisco, CA.

  • Compensation: Depending on background, skills and experience, the expected annual salary range for this position is $350,000-$475,000 USD.

  • Visa sponsorship: We sponsor visas. While we can't guarantee success for every candidate or role, if you're the right fit, we're committed to working through the visa process together.

  • Benefits: Thinking Machines offers generous health, dental, and vision benefits, unlimited PTO, paid parental leave, and relocation support as needed.

As set forth in Thinking Machines' Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Thinking Machines Lab will consider for employment qualified applicants with criminal histories in a manner consistent with the requirements of the California Fair Chance Act, the San Francisco Fair Chance Ordinance, and any other applicable state or local fair chance ordinance or law.

Skills Required

  • 8+ years designing, building, and scaling web crawlers, scrapers, or large-scale distributed data-acquisition systems
  • Track record of owning crawler or data-acquisition infrastructure at internet scale
  • Strong software engineering skills in Python, Go, Rust, or a similar language
  • Real experience with distributed systems
  • Working knowledge of practical and legal considerations of large-scale web data collection, including robots.txt, rate limiting, and licensing
  • Experience applying machine learning to crawl selection, extraction, or data quality classification at internet scale
  • Experience setting technical direction for a crawling, data acquisition, or search infrastructure team
  • Experience designing systems for petabyte-scale storage and processing
  • Open-source contributions to crawling, scraping, or data infrastructure tools
  • Background at a search engine or frontier AI lab data acquisition team
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Singapore
91 Employees

What We Do

Thinking Machines Lab is an artificial intelligence research and product company. We're building a future where everyone has access to the knowledge and tools to make AI work for their unique needs and goals. While AI capabilities have advanced dramatically, key gaps remain. The scientific community's understanding of frontier AI systems lags behind rapidly advancing capabilities. Knowledge of how these systems are trained is concentrated within the top research labs, limiting both the public discourse on AI and people's abilities to use AI effectively. And, despite their potential, these systems remain difficult for people to customize to their specific needs and values. To bridge the gaps, we're building Thinking Machines Lab to make AI systems more widely understood, customizable and generally capable. We are scientists, engineers, and builders who've created some of the most widely used AI products, including ChatGPT and Character.ai, open-weights models like Mistral, as well as popular open source projects like PyTorch, OpenAI Gym, Fairseq, and Segment Anything.

Similar Jobs

Hybrid
Palo Alto, CA, USA
289097 Employees
Hybrid
Palo Alto, CA, USA
289097 Employees
Hybrid
Foothill Ranch, CA, USA
205000 Employees
27K-41K Hourly
Hybrid
Woodland, CA, USA
205000 Employees
37K-66K Hourly

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees
Vega Thumbnail
Artificial Intelligence • Automotive • Insurance • Transportation
US
43 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account