The Role
Build and maintain scalable web crawlers to extract, validate, and structure product data from eCommerce sites. Use AI-driven techniques, browser automation, proxies, and APIs to handle dynamic content and anti-bot measures. Monitor performance, clean large datasets, collaborate with Product/Engineering/Data teams, and continuously improve crawling frameworks and automation.
Summary Generated by Built In
Data Crawling ConsultantRole Overview
Part 1: The Fundamentals
Work Model
Investors
Visit Us
You will be responsible for building and maintaining Rubick's data acquisition engine by extracting, validating, and structuring product information from eCommerce websites and marketplaces. This role focuses on developing scalable web crawling solutions that power Product Discovery, Search, Catalog Intelligence, and Market Intelligence platforms.
The role involves three areas: Part 1: The Fundamentals | Part 2: AI-Driven Data Crawling Excellence | Part 3: Innovation & Improvements
Part 1: The Fundamentals
- Develop, maintain, and optimize web crawlers to extract product data from eCommerce websites and marketplaces.
- Collect structured and unstructured product information, including product details, pricing, images, specifications, and availability.
- Clean, validate, and organize extracted datasets to ensure accuracy and consistency.
- Monitor crawler performance, identify failures, and resolve data extraction issues.
- Collaborate with Product, Engineering, and Data teams to ensure reliable and timely data collection.
- Maintain documentation for crawling processes, extraction rules, and data quality standards.
- Leverage AI-powered extraction techniques to improve data accuracy and extraction efficiency.
- Build intelligent crawling workflows using Python and modern web automation frameworks.
- Develop scalable solutions for handling dynamic websites, JavaScript-rendered content, and anti-bot mechanisms.
- Utilize browser automation, APIs, proxies, and scheduling tools to maximize crawl success rates.
- Implement automated data validation and monitoring systems to maintain high-quality datasets.
- Collaborate with AI and Data Engineering teams to support Machine Learning models and Product Intelligence systems.
- Continuously optimize crawling performance for speed, scalability, and reliability.
- Identify opportunities to automate repetitive extraction and validation workflows.
- Improve data collection strategies by adopting new crawling technologies and AI-assisted solutions.
- Build reusable crawling frameworks and standardized extraction pipelines.
- Stay updated with the latest web scraping libraries, browser automation tools, and industry best practices.
- Contribute to knowledge repositories, documentation, and process improvements across the data acquisition function.
- 1+ year of experience in Web Crawling, Data Scraping, Product Matching, Data Extraction, or similar roles.
- Strong proficiency in Python for web scraping and automation.
- Hands-on experience with BeautifulSoup, Scrapy, Selenium, Playwright, or similar frameworks.
- Good understanding of HTML, CSS, XPath, JSON, and DOM structures.
- Familiarity with REST APIs, proxies, browser automation, and dynamic website scraping.
- Strong analytical and problem-solving skills.
- Ability to work with large datasets while maintaining high data quality.
- Experience in eCommerce, Retail, Marketplace Operations, or Catalog Management is preferred.
- Exposure to AI-assisted data extraction, Product Intelligence, or Machine Learning data pipelines is a strong advantage.
Rubick OS is a leading platform for brands and marketplaces, enabling seamless management of catalog automation, pricing and competitive intelligence, product listing intelligence, CAST systems, and brand and assortment intelligence.
Our solutions are used across five countries by major brands and marketplaces, including Amazon, Myntra, Reliance Group, Nykaa, Flipkart, Rare Rabbit, Decathlon, Celio, The Bay, Kiabi, Myer, Jumbo, and many more.
Work Model
Work from home – Monday to Saturday
Team Size250+ Employees
Invested by Innospark US, Betatron Hong Kong, MJV India, 1Crowd India, and other investors.
Rubick.ai is a place to work if you want to build something foundational rather than incremental—it’s aiming to become the AI operating layer for ecommerce, solving complex, real-world problems across cataloging, pricing, and marketplace operations at scale.
You get high ownership, direct impact, and exposure to various parts of the business, making you an entrepreneur in-house. Rubick operates in the AI segment, which means you will be at the forefront of the technological revolution.
Rubick is profitable and growing and is looking to work with 10,000 brands in the next 2–3 years.
Join us in this exciting journey.
Visit Us
www.rubick.ai
Skills Required
- 1+ year of experience in Web Crawling, Data Scraping, Product Matching, or Data Extraction
- Strong proficiency in Python for web scraping and automation
- Hands-on experience with BeautifulSoup, Scrapy, Selenium, Playwright, or similar frameworks
- Good understanding of HTML, CSS, XPath, JSON, and DOM structures
- Familiarity with REST APIs, proxies, browser automation, and dynamic website scraping
- Strong analytical and problem-solving skills
- Ability to work with large datasets while maintaining high data quality
- Experience in eCommerce, Retail, Marketplace Operations, or Catalog Management
- Exposure to AI-assisted data extraction, Product Intelligence, or Machine Learning data pipelines
- Maintain documentation for crawling processes, extraction rules, and data quality standards
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Rubick.ai is an AI-powered eCommerce enablement platform that specializes in cataloging solutions and digital readiness for marketplaces, brands, and sellers. It provides an AI-powered operating system to automate catalog enrichment, product discovery, and scale growth for e-commerce businesses.
.jpg)








