Data Crawling - Consultant

Reposted Yesterday
2 Locations
In-Office or Remote
Junior
Artificial Intelligence • eCommerce • Information Technology • Retail • Software
The Role
Build and maintain scalable web crawlers to extract, validate, and structure product data from eCommerce sites. Use AI-driven techniques, browser automation, proxies, and APIs to handle dynamic content and anti-bot measures. Monitor performance, clean large datasets, collaborate with Product/Engineering/Data teams, and continuously improve crawling frameworks and automation.
Summary Generated by Built In
Data Crawling ConsultantRole Overview
You will be responsible for building and maintaining Rubick's data acquisition engine by extracting, validating, and structuring product information from eCommerce websites and marketplaces. This role focuses on developing scalable web crawling solutions that power Product Discovery, Search, Catalog Intelligence, and Market Intelligence platforms.

The role involves three areas: Part 1: The Fundamentals | Part 2: AI-Driven Data Crawling Excellence | Part 3: Innovation & Improvements

Part 1: The Fundamentals
  • Develop, maintain, and optimize web crawlers to extract product data from eCommerce websites and marketplaces.
  • Collect structured and unstructured product information, including product details, pricing, images, specifications, and availability.
  • Clean, validate, and organize extracted datasets to ensure accuracy and consistency.
  • Monitor crawler performance, identify failures, and resolve data extraction issues.
  • Collaborate with Product, Engineering, and Data teams to ensure reliable and timely data collection.
  • Maintain documentation for crawling processes, extraction rules, and data quality standards.
Part 2: AI-Driven Data Crawling Excellence
  • Leverage AI-powered extraction techniques to improve data accuracy and extraction efficiency.
  • Build intelligent crawling workflows using Python and modern web automation frameworks.
  • Develop scalable solutions for handling dynamic websites, JavaScript-rendered content, and anti-bot mechanisms.
  • Utilize browser automation, APIs, proxies, and scheduling tools to maximize crawl success rates.
  • Implement automated data validation and monitoring systems to maintain high-quality datasets.
  • Collaborate with AI and Data Engineering teams to support Machine Learning models and Product Intelligence systems.
Part 3: Innovation & Improvements
  • Continuously optimize crawling performance for speed, scalability, and reliability.
  • Identify opportunities to automate repetitive extraction and validation workflows.
  • Improve data collection strategies by adopting new crawling technologies and AI-assisted solutions.
  • Build reusable crawling frameworks and standardized extraction pipelines.
  • Stay updated with the latest web scraping libraries, browser automation tools, and industry best practices.
  • Contribute to knowledge repositories, documentation, and process improvements across the data acquisition function.
To Have
  • 1+ year of experience in Web Crawling, Data Scraping, Product Matching, Data Extraction, or similar roles.
  • Strong proficiency in Python for web scraping and automation.
  • Hands-on experience with BeautifulSoup, Scrapy, Selenium, Playwright, or similar frameworks.
  • Good understanding of HTML, CSS, XPath, JSON, and DOM structures.
  • Familiarity with REST APIs, proxies, browser automation, and dynamic website scraping.
  • Strong analytical and problem-solving skills.
  • Ability to work with large datasets while maintaining high data quality.
  • Experience in eCommerce, Retail, Marketplace Operations, or Catalog Management is preferred.
  • Exposure to AI-assisted data extraction, Product Intelligence, or Machine Learning data pipelines is a strong advantage.
About Rubick : Company Overview
Rubick OS is a leading platform for brands and marketplaces, enabling seamless management of catalog automation, pricing and competitive intelligence, product listing intelligence, CAST systems, and brand and assortment intelligence.
Our solutions are used across five countries by major brands and marketplaces, including Amazon, Myntra, Reliance Group, Nykaa, Flipkart, Rare Rabbit, Decathlon, Celio, The Bay, Kiabi, Myer, Jumbo, and many more.

Work Model
Work from office – Monday to Saturday
Team Size
250+ Employees

Investors
Invested by Innospark US, Betatron Hong Kong, MJV India, 1Crowd India, and other investors.
Rubick.ai is a place to work if you want to build something foundational rather than incremental—it’s aiming to become the AI operating layer for ecommerce, solving complex, real-world problems across cataloging, pricing, and marketplace operations at scale.
You get high ownership, direct impact, and exposure to various parts of the business, making you an entrepreneur in-house. Rubick operates in the AI segment, which means you will be at the forefront of the technological revolution.
Rubick is profitable and growing and is looking to work with 10,000 brands in the next 2–3 years.
Join us in this exciting journey.

Visit Us
www.rubick.ai

Skills Required

  • 1+ year of experience in Web Crawling, Data Scraping, Product Matching, or Data Extraction
  • Strong proficiency in Python for web scraping and automation
  • Hands-on experience with BeautifulSoup, Scrapy, Selenium, Playwright, or similar frameworks
  • Good understanding of HTML, CSS, XPath, JSON, and DOM structures
  • Familiarity with REST APIs, proxies, browser automation, and dynamic website scraping
  • Strong analytical and problem-solving skills
  • Ability to work with large datasets while maintaining high data quality
  • Experience in eCommerce, Retail, Marketplace Operations, or Catalog Management
  • Exposure to AI-assisted data extraction, Product Intelligence, or Machine Learning data pipelines
  • Maintain documentation for crawling processes, extraction rules, and data quality standards
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: Bengaluru
78 Employees

What We Do

Rubick.ai is an AI-powered eCommerce enablement platform that specializes in cataloging solutions and digital readiness for marketplaces, brands, and sellers. It provides an AI-powered operating system to automate catalog enrichment, product discovery, and scale growth for e-commerce businesses.

Similar Jobs

Drata Logo Drata

Senior Director, Global GSI and Audit Firm Alliances

Security • Software • Cybersecurity • Automation
Remote
United States
600 Employees
261K-404K Annually

Liberty Mutual Insurance Logo Liberty Mutual Insurance

Chief Actuary Reserving - North America

Artificial Intelligence • Fintech • Insurance • Marketing Tech • Software • Analytics
Remote or Hybrid
3 Locations
40000 Employees
204K-366K Annually

Pie Insurance Logo Pie Insurance

Manager, Premium Audit

Fintech • Insurance • Machine Learning • Analytics • Financial Services • Automation
Easy Apply
Remote
United States
350 Employees
95K-120K Annually

Ashley Digital Logo Ashley Digital

Copywriter

eCommerce • Retail
Remote
USA
238 Employees
90K-115K Annually

Similar Companies Hiring

Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account