Job Title: Senior Data Engineer (Scraping)
Employment Type: Full-Time
We are looking for an experienced Senior Data Engineer (Scraping) to design, develop, and maintain scalable web scraping and data harvesting pipelines. The ideal candidate should have strong expertise in Python, web scraping frameworks, ETL processes, distributed data processing, and workflow orchestration. This role involves working closely with the Data Harvest team to collect, process, and deliver high-quality data from multiple web sources and APIs.
RequirementsKey Responsibilities
- Design, develop, and maintain scalable web scraping and data harvesting pipelines.
- Build and maintain web scrapers using Python frameworks such as Scrapy, BeautifulSoup, Selenium, and Playwright.
- Handle dynamic JavaScript-rendered websites and overcome anti-bot mechanisms including proxy rotation, IP rotation, user-agent rotation, rate limiting, and CAPTCHA handling.
- Develop ETL workflows for data extraction, parsing, transformation, and cleaning.
- Process and transform large-scale datasets using PySpark and distributed computing.
- Design and schedule workflows using Apache Airflow (Dagster, Prefect, or Luigi experience is an added advantage).
- Store and manage data in SQL and NoSQL databases.
- Work with data formats such as CSV, JSON, XML, and Parquet.
- Implement monitoring, logging, retry mechanisms, and error handling to ensure pipeline reliability.
- Ensure compliance with website policies, robots.txt, and data privacy regulations such as GDPR.
- Collaborate with technical teams and data consumers to maintain data quality and timely delivery.
BenefitsRequired Skills
- Strong programming experience in Python.
- Knowledge of Node.js or JavaScript is an added advantage.
- Hands-on experience with:
- Scrapy
- BeautifulSoup
- Selenium
- Playwright
- lxml
- requests/httpx
- Puppeteer (preferred)
- Scrapy
- Strong understanding of:
- HTML
- CSS
- DOM
- XPath
- CSS Selectors
- HTTP Protocol
- HTML
- Experience with REST APIs and GraphQL APIs.
- Experience parsing JSON, XML, and HTML data.
- Strong ETL development experience.
- Experience with PySpark for distributed data processing.
- Hands-on experience with Apache Airflow.
- Good knowledge of SQL and NoSQL databases including PostgreSQL, MySQL, and MongoDB.
- Experience with asynchronous programming, concurrency, and distributed scraping.
- Knowledge of Git, Docker, and cloud platforms such as AWS, Azure, or GCP.
- Experience implementing monitoring and alerting for production pipelines.
- Understanding of legal and compliance requirements related to web scraping.
- Experience with Dagster, Prefect, or Luigi.
- Knowledge of serverless and cloud-native architectures.
- Experience with large-scale data engineering projects.
- Strong debugging, troubleshooting, and analytical skills.
- Bachelor's or Master's degree in Computer Science, Engineering, Information Technology, or a related field.
- Equivalent practical experience will also be considered.
Skills Required
- Strong programming experience in Python
- Experience building web scrapers with Scrapy, BeautifulSoup, Selenium, Playwright, lxml, and requests or httpx
- Experience handling JavaScript-rendered websites and anti-bot mechanisms
- Strong understanding of HTML, CSS, DOM, XPath, CSS selectors, and HTTP protocol
- Experience with REST APIs and GraphQL APIs
- Experience parsing JSON, XML, and HTML data
- Strong ETL development experience
- Experience with PySpark and distributed data processing
- Hands-on experience with Apache Airflow
- Knowledge of SQL and NoSQL databases, including PostgreSQL, MySQL, and MongoDB
- Experience with asynchronous programming, concurrency, and distributed scraping
- Knowledge of Git, Docker, and at least one cloud platform such as AWS, Azure, or GCP
- Experience implementing monitoring and alerting for production pipelines
- Understanding of legal and compliance requirements related to web scraping
- Bachelor's or Master's degree in Computer Science, Engineering, Information Technology, or a related field, or equivalent practical experience
- Knowledge of Node.js or JavaScript
- Experience with Puppeteer
- Experience with Dagster, Prefect, or Luigi
- Knowledge of serverless and cloud-native architectures
- Experience with large-scale data engineering projects
- Strong debugging, troubleshooting, and analytical skills
What We Do
Verifitech provides AI-powered background verification and screening solutions for enterprises. Its platform supports employment background screening, credit verification, exit interview checks, drug checks, and verification services for housekeeping staff. The company uses modern technology to deliver fast, accurate, and compliant background verification reports, serving organizations across sectors including information technology, banking, logistics, finance, manufacturing, and automobiles.








