Data Crawling - Consultant

Posted 2 Days Ago
Hiring Remotely in Oregon, USA
Remote
Junior
Artificial Intelligence • eCommerce • Information Technology • Retail • Software
The Role
Build and maintain scalable web crawlers to extract, validate, and structure product data from eCommerce sites. Use AI-driven techniques, browser automation, proxies, and APIs to handle dynamic content and anti-bot measures. Monitor performance, clean large datasets, collaborate with Product/Engineering/Data teams, and continuously improve crawling frameworks and automation.
Summary Generated by Built In
Data Crawling ConsultantRole Overview
You will be responsible for building and maintaining Rubick's data acquisition engine by extracting, validating, and structuring product information from eCommerce websites and marketplaces. This role focuses on developing scalable web crawling solutions that power Product Discovery, Search, Catalog Intelligence, and Market Intelligence platforms.

The role involves three areas: Part 1: The Fundamentals | Part 2: AI-Driven Data Crawling Excellence | Part 3: Innovation & Improvements

Part 1: The Fundamentals
  • Develop, maintain, and optimize web crawlers to extract product data from eCommerce websites and marketplaces.
  • Collect structured and unstructured product information, including product details, pricing, images, specifications, and availability.
  • Clean, validate, and organize extracted datasets to ensure accuracy and consistency.
  • Monitor crawler performance, identify failures, and resolve data extraction issues.
  • Collaborate with Product, Engineering, and Data teams to ensure reliable and timely data collection.
  • Maintain documentation for crawling processes, extraction rules, and data quality standards.
Part 2: AI-Driven Data Crawling Excellence
  • Leverage AI-powered extraction techniques to improve data accuracy and extraction efficiency.
  • Build intelligent crawling workflows using Python and modern web automation frameworks.
  • Develop scalable solutions for handling dynamic websites, JavaScript-rendered content, and anti-bot mechanisms.
  • Utilize browser automation, APIs, proxies, and scheduling tools to maximize crawl success rates.
  • Implement automated data validation and monitoring systems to maintain high-quality datasets.
  • Collaborate with AI and Data Engineering teams to support Machine Learning models and Product Intelligence systems.
Part 3: Innovation & Improvements
  • Continuously optimize crawling performance for speed, scalability, and reliability.
  • Identify opportunities to automate repetitive extraction and validation workflows.
  • Improve data collection strategies by adopting new crawling technologies and AI-assisted solutions.
  • Build reusable crawling frameworks and standardized extraction pipelines.
  • Stay updated with the latest web scraping libraries, browser automation tools, and industry best practices.
  • Contribute to knowledge repositories, documentation, and process improvements across the data acquisition function.
To Have
  • 1+ year of experience in Web Crawling, Data Scraping, Product Matching, Data Extraction, or similar roles.
  • Strong proficiency in Python for web scraping and automation.
  • Hands-on experience with BeautifulSoup, Scrapy, Selenium, Playwright, or similar frameworks.
  • Good understanding of HTML, CSS, XPath, JSON, and DOM structures.
  • Familiarity with REST APIs, proxies, browser automation, and dynamic website scraping.
  • Strong analytical and problem-solving skills.
  • Ability to work with large datasets while maintaining high data quality.
  • Experience in eCommerce, Retail, Marketplace Operations, or Catalog Management is preferred.
  • Exposure to AI-assisted data extraction, Product Intelligence, or Machine Learning data pipelines is a strong advantage.
About Rubick : Company Overview
Rubick OS is a leading platform for brands and marketplaces, enabling seamless management of catalog automation, pricing and competitive intelligence, product listing intelligence, CAST systems, and brand and assortment intelligence.
Our solutions are used across five countries by major brands and marketplaces, including Amazon, Myntra, Reliance Group, Nykaa, Flipkart, Rare Rabbit, Decathlon, Celio, The Bay, Kiabi, Myer, Jumbo, and many more.

Work Model
Work from home – Monday to Saturday
Team Size
250+ Employees

Investors
Invested by Innospark US, Betatron Hong Kong, MJV India, 1Crowd India, and other investors.
Rubick.ai is a place to work if you want to build something foundational rather than incremental—it’s aiming to become the AI operating layer for ecommerce, solving complex, real-world problems across cataloging, pricing, and marketplace operations at scale.
You get high ownership, direct impact, and exposure to various parts of the business, making you an entrepreneur in-house. Rubick operates in the AI segment, which means you will be at the forefront of the technological revolution.
Rubick is profitable and growing and is looking to work with 10,000 brands in the next 2–3 years.
Join us in this exciting journey.

Visit Us
www.rubick.ai

Skills Required

  • 1+ year of experience in Web Crawling, Data Scraping, Product Matching, or Data Extraction
  • Strong proficiency in Python for web scraping and automation
  • Hands-on experience with BeautifulSoup, Scrapy, Selenium, Playwright, or similar frameworks
  • Good understanding of HTML, CSS, XPath, JSON, and DOM structures
  • Familiarity with REST APIs, proxies, browser automation, and dynamic website scraping
  • Strong analytical and problem-solving skills
  • Ability to work with large datasets while maintaining high data quality
  • Experience in eCommerce, Retail, Marketplace Operations, or Catalog Management
  • Exposure to AI-assisted data extraction, Product Intelligence, or Machine Learning data pipelines
  • Maintain documentation for crawling processes, extraction rules, and data quality standards
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Bengaluru
78 Employees

What We Do

Rubick.ai is an AI-powered eCommerce enablement platform that specializes in cataloging solutions and digital readiness for marketplaces, brands, and sellers. It provides an AI-powered operating system to automate catalog enrichment, product discovery, and scale growth for e-commerce businesses.

Similar Jobs

Rula Logo Rula

Health System Partnership Manager (Remote)

Healthtech • Social Impact • Software • Telehealth
Remote
United States
620 Employees
149K-166K Annually

CrowdStrike Logo CrowdStrike

Manager, Platform Professional Resident Services (Remote)

Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Remote or Hybrid
2 Locations
11000 Employees
140K-195K Annually

Smartling Logo Smartling

Social Media Manager

Artificial Intelligence • Cloud • Information Technology • Machine Learning • Natural Language Processing • Software
Easy Apply
Remote
US
117 Employees
65K-79K Annually

Immersive Logo Immersive

Enterprise Account Manager

Enterprise Web • HR Tech • Information Technology • Software • Cybersecurity
Remote or Hybrid
United States
330 Employees
130K-160K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account