Web Scraping Specialist

Reposted 22 Days Ago
Be an Early Applicant
Hiring Remotely in USA
Remote
Senior level
Artificial Intelligence • Software
The Role
Lead development and optimization of web scraping pipelines to extract, clean, and store large-scale web data. Handle dynamic content, pagination, distributed scraping, database design with NoSQL, deploy jobs to cloud, and monitor systems for reliability and data quality. Support ML-based data cleaning and categorization.
Summary Generated by Built In

Who We Are:

We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models.

We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.

We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI.

The Role.

We are seeking a Web Scraping Specialist who is proficient and brings significant experience in data extraction and web scraping techniques. You will join a small, specialized team and lead efforts to gather and analyze data, optimize scraping processes, and support our vision for a future where Grass plays a crucial role in transforming internet data accessibility.

Please note: This role requires a work schedule that overlaps sufficiently with EST business hours (min. 3-4 hours) to collaborate effectively with the team.

Who You Are.

  • Demonstrated ability to extract data from complex websites with minimal supervision, with a portfolio or examples of past projects.

  • Proficiency in languages such as Python or JavaScript, with strong skills in libraries and frameworks like BeautifulSoup, Scrapy, or Selenium.

  • Knowledge of asynchronous programming, multithreading, and distributed scraping.

  • In-depth knowledge of HTML, CSS, JavaScript, and the Document Object Model (DOM).

  • Experience with NoSQL databases (MongoDB, Cassandra), capable of designing efficient storage solutions and managing data integrity.

  • Ability to apply machine learning algorithms for data cleaning, categorization, or predictive analysis adds significant value.

  • Experience with cloud services (AWS, Google Cloud, Azure) for deploying and managing scraping jobs at scale.

  • Active participation in open-source projects related to web scraping, data processing, or similar fields.

What You'll Be Doing.

  • Write, test, and refine code that extracts data from various online sources, ensuring reliability and efficiency.

  • Perform data retrieval tasks, handling complexities such as pagination and dynamic content loaded with AJAX.

  • Clean and format extracted data, ensuring it meets quality standards for further analysis or processing.

  • Database management: Store and manage the scraped data in appropriate databases, optimizing for access speed and data integrity.

  • Regularly monitor the scraping processes, identify and resolve any issues to maintain continuous data flow.

Why Work With Us:

  • Opportunity. We are at the forefront of developing a web-scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people.

  • Culture. We're a lean team with a high bar. We come to work not to be comfortable, but to find out what we're capable of and to do work that matters. We're not calling for people who keep things moving. We're calling for people who make everyone around them better.
    We prioritize low ego and high output. This is a fully remote team.

  • Compensation. You’ll receive a competitive salary, benefits and equity package.

Skills Required

  • Demonstrated ability to extract data from complex websites with portfolio or examples
  • Proficiency in Python or JavaScript and frameworks/libraries like BeautifulSoup, Scrapy, Selenium
  • Knowledge of asynchronous programming, multithreading, and distributed scraping
  • In-depth knowledge of HTML, CSS, JavaScript, and the DOM
  • Experience with NoSQL databases such as MongoDB or Cassandra and designing efficient storage
  • Experience deploying and managing scraping jobs at scale on cloud services (AWS, GCP, Azure)
  • Ability to apply machine learning algorithms for data cleaning, categorization, or analysis
  • Active participation in relevant open-source projects (web scraping or data processing)
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
1 Employee

What We Do

Making Public Web Data Accessible for AI.

Similar Jobs

General Motors Logo General Motors

Sales Manager

Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Remote or Hybrid
United States
165000 Employees
81K-109K Annually

General Motors Logo General Motors

Total Rewards Strategy and Portfolio Acceleration Lead

Automotive • Big Data • Information Technology • Robotics • Software • Transportation • Manufacturing
Remote or Hybrid
3 Locations
165000 Employees
186K-259K Annually

Huntress Logo Huntress

Operations Analyst

Information Technology • Cybersecurity
Easy Apply
Remote
United States of America
780 Employees
100K-125K Annually

NBCUniversal Logo NBCUniversal

Sr Staff Threat Intelligence Analyst

AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Remote or Hybrid
New York, NY, USA
140K-180K Annually

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account