Senior Data Engineer (Scraping)

Posted Yesterday
Be an Early Applicant
Chennai, Tamil Nādu, IND
In-Office
Senior level
Artificial Intelligence • HR Tech • Software
The Role
Design, develop, and maintain scalable Python web scraping and data harvesting pipelines. Build scrapers for dynamic websites, develop ETL workflows, process large datasets with PySpark, and orchestrate pipelines using Apache Airflow. Manage SQL and NoSQL data stores, implement monitoring and error handling, ensure data quality and regulatory compliance, and collaborate with technical teams and data consumers.
Summary Generated by Built In
Senior Data Engineer (Scraping)

Job Title: Senior Data Engineer (Scraping)

Employment Type: Full-Time

Job Summary

We are looking for an experienced Senior Data Engineer (Scraping) to design, develop, and maintain scalable web scraping and data harvesting pipelines. The ideal candidate should have strong expertise in Python, web scraping frameworks, ETL processes, distributed data processing, and workflow orchestration. This role involves working closely with the Data Harvest team to collect, process, and deliver high-quality data from multiple web sources and APIs.



RequirementsKey Responsibilities
  • Design, develop, and maintain scalable web scraping and data harvesting pipelines.
  • Build and maintain web scrapers using Python frameworks such as Scrapy, BeautifulSoup, Selenium, and Playwright.
  • Handle dynamic JavaScript-rendered websites and overcome anti-bot mechanisms including proxy rotation, IP rotation, user-agent rotation, rate limiting, and CAPTCHA handling.
  • Develop ETL workflows for data extraction, parsing, transformation, and cleaning.
  • Process and transform large-scale datasets using PySpark and distributed computing.
  • Design and schedule workflows using Apache Airflow (Dagster, Prefect, or Luigi experience is an added advantage).
  • Store and manage data in SQL and NoSQL databases.
  • Work with data formats such as CSV, JSON, XML, and Parquet.
  • Implement monitoring, logging, retry mechanisms, and error handling to ensure pipeline reliability.
  • Ensure compliance with website policies, robots.txt, and data privacy regulations such as GDPR.
  • Collaborate with technical teams and data consumers to maintain data quality and timely delivery.


BenefitsRequired Skills
  • Strong programming experience in Python.
  • Knowledge of Node.js or JavaScript is an added advantage.
  • Hands-on experience with:
    • Scrapy
    • BeautifulSoup
    • Selenium
    • Playwright
    • lxml
    • requests/httpx
    • Puppeteer (preferred)
  • Strong understanding of:
    • HTML
    • CSS
    • DOM
    • XPath
    • CSS Selectors
    • HTTP Protocol
  • Experience with REST APIs and GraphQL APIs.
  • Experience parsing JSON, XML, and HTML data.
  • Strong ETL development experience.
  • Experience with PySpark for distributed data processing.
  • Hands-on experience with Apache Airflow.
  • Good knowledge of SQL and NoSQL databases including PostgreSQL, MySQL, and MongoDB.
  • Experience with asynchronous programming, concurrency, and distributed scraping.
  • Knowledge of Git, Docker, and cloud platforms such as AWS, Azure, or GCP.
  • Experience implementing monitoring and alerting for production pipelines.
  • Understanding of legal and compliance requirements related to web scraping.
Preferred Skills
  • Experience with Dagster, Prefect, or Luigi.
  • Knowledge of serverless and cloud-native architectures.
  • Experience with large-scale data engineering projects.
  • Strong debugging, troubleshooting, and analytical skills.
Educational Qualification
  • Bachelor's or Master's degree in Computer Science, Engineering, Information Technology, or a related field.
  • Equivalent practical experience will also be considered.
For more information call @
9952978818
Mail your resume @

Skills Required

  • Strong programming experience in Python
  • Experience building web scrapers with Scrapy, BeautifulSoup, Selenium, Playwright, lxml, and requests or httpx
  • Experience handling JavaScript-rendered websites and anti-bot mechanisms
  • Strong understanding of HTML, CSS, DOM, XPath, CSS selectors, and HTTP protocol
  • Experience with REST APIs and GraphQL APIs
  • Experience parsing JSON, XML, and HTML data
  • Strong ETL development experience
  • Experience with PySpark and distributed data processing
  • Hands-on experience with Apache Airflow
  • Knowledge of SQL and NoSQL databases, including PostgreSQL, MySQL, and MongoDB
  • Experience with asynchronous programming, concurrency, and distributed scraping
  • Knowledge of Git, Docker, and at least one cloud platform such as AWS, Azure, or GCP
  • Experience implementing monitoring and alerting for production pipelines
  • Understanding of legal and compliance requirements related to web scraping
  • Bachelor's or Master's degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience
  • Knowledge of Node.js or JavaScript
  • Experience with Puppeteer
  • Experience with Dagster, Prefect, or Luigi
  • Knowledge of serverless and cloud-native architectures
  • Experience with large-scale data engineering projects
  • Strong debugging, troubleshooting, and analytical skills
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
30 Employees
Year Founded: 2008

What We Do

Verifitech provides AI-powered background verification and screening solutions for enterprises. Its platform supports employment background screening, credit verification, exit interview checks, drug checks, and verification services for housekeeping staff. The company uses modern technology to deliver fast, accurate, and compliant background verification reports, serving organizations across sectors including information technology, banking, logistics, finance, manufacturing, and automobiles.

Similar Jobs

MongoDB Logo MongoDB

Technical Director, Builder Relations

Big Data • Cloud • Software • Database
Easy Apply
Remote or Hybrid
India
5550 Employees

Capco Logo Capco

Product Manager

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

Capco Logo Capco

Visualisation Analyst (Liquidity & Treasury)

Fintech • Professional Services • Consulting • Energy • Financial Services • Cybersecurity • Generative AI
Remote or Hybrid
India
6000 Employees

CSC Logo CSC

Security Engineer

Fintech • Legal Tech • Software • Financial Services • Cybersecurity • Data Privacy
Remote or Hybrid
2 Locations
8500 Employees

Similar Companies Hiring

Hanover Park Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
42 Employees
Kepler  Thumbnail
Fintech • Software
New York, New York
6 Employees
Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account