Python Data Extraction Engineer

Posted Yesterday
Be an Early Applicant
Coimbatore, Tamil Nadu, IND
In-Office
Mid level
eCommerce • Information Technology • Marketing Tech • Design
Award winning experts in eCommerce! We have more than ten years of experience with eCommerce solutions.
The Role
Build and maintain Python-based web scraping and data extraction pipelines for government portals, public websites, PDFs, APIs, and open-data sources. Extract, clean, normalize, match, validate, and integrate fragmented records into structured databases. Develop monitoring for schema changes and extraction failures, automate recurring acquisition, and collaborate with GIS, product, and engineering teams to deliver reliable public-data intelligence.
Summary Generated by Built In

Python Data Extraction Engineer – Web Scraping & Government Data
Location: Coimbatore
Experience: 3–7 years
Role Type: Full-time

About the Role
We are building a data intelligence platform that looks to analyze fragmented public information into structured, actionable business data.
We are looking for a strong Python Data Extraction Engineer who can discover, extract, clean, normalize and integrate data from government portals, public websites, PDFs, APIs and other open data sources.
This is not a conventional application-development role. The ideal candidate enjoys solving difficult data-acquisition problems involving poorly structured websites, inconsistent government portals, changing schemas, PDFs, JavaScript-rendered pages and large volumes of semi-structured information.
What You Will Own
You will build and maintain the data-acquisition layer of the platform.
Key responsibilities include:

Identify and evaluate government and public data sources relevant to property, businesses and commercial activity.

Build Python-based crawlers and extraction pipelines for government portals and public websites.

Extract structured information from HTML pages, tables, PDFs, downloadable files and publicly accessible APIs.

Work with JavaScript-rendered websites and multi-step public search interfaces.

Automate recurring extraction from multiple sources while respecting applicable access rules, rate limits, and terms.

Clean, standardize and normalize inconsistent data from different government authorities.

Perform entity matching and record linkage across datasets using fields such as owner name, company name, address, survey number, coordinates and project information.

Develop mechanisms to detect website/schema changes and extraction failures.

Build validation and QA processes to measure completeness and accuracy.

Store extracted information in structured databases and expose clean datasets to downstream applications.

Work closely with GIS, product and engineering teams to combine location-based signals with government/public records.

Research new public and open-data sources that can improve the accuracy and completeness of our intelligence.


Examples of Data Sources
The work may involve sources such as:

Municipal corporation portals

State planning and development authorities

DTCP and similar planning authorities

RERA databases

Building and planning permission records

Land and property records

Tender and procurement portals

Company/business registries

Government open-data portals

Environmental and regulatory approvals

Public notices and downloadable government documents

Maps and geospatial datasets

Other legally accessible public and open-source information

Required Technical Skills
Strong hands-on experience with:

Python

Web scraping and crawling

Requests / HTTP clients

BeautifulSoup / lxml

Selenium and/or Playwright

REST APIs and JSON

HTML/XML parsing

Pandas

SQL

Data cleaning and transformation

Regex and text processing

ETL/data pipelines

Git

What We Are Looking For
We particularly want someone who is a problem solver rather than simply a Python programmer.
The candidate should be able to investigate the available sources, understand how the underlying website works, determine the best extraction approach, build the pipeline and validate the resulting data.
Ideal Background
Candidates may come from backgrounds such as:

Web scraping / data extraction companies

Alternative-data companies

PropTech / real-estate data companies

Market-intelligence companies

OSINT/data intelligence companies

Government-data projects

Data aggregation platforms

Lead/data enrichment companies

GIS/location-intelligence companies

Success in This Role
Within the first few months, the successful candidate should be able to:

Map relevant government/public data sources.

Build reliable extraction pipelines across multiple portals.

Convert fragmented information into standardized records.

Cross-reference records from multiple sources.

Establish automated QA and monitoring.

Continuously discover additional datasets that improve our product's coverage and accuracy.

The objective is not simply to scrape websites. It is to build a scalable public-data acquisition and enrichment engine that becomes a core component of our intelligence platform.




Requirements
Python, pandas, webscraping, web crawling,selenium postman

Skills Required

  • Strong hands-on experience with Python
  • Experience with web scraping and crawling
  • Experience with Requests or HTTP clients
  • Experience with BeautifulSoup or lxml
  • Experience with Selenium and/or Playwright
  • Experience with REST APIs and JSON
  • Experience with HTML/XML parsing
  • Experience with Pandas
  • Experience with SQL
  • Experience with data cleaning and transformation
  • Experience with regex and text processing
  • Experience with ETL/data pipelines
  • Experience with Git
  • Experience with Postman
  • 3-7 years of experience
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
HQ: New York, NY
Year Founded: 2004

What We Do

At Vserve, we are award winning eCommerce specialists! Vserve is one of the leading full-service ecommerce agencies that offers a range of ebusiness solutions, catalog management services, and business intelligence expertise. In short, we specialize in delivering customized ecommerce solutions that make your business work for you so that you can work less!

Gallery

Gallery

Similar Jobs

In-Office
Chennai, Tamil Nadu, IND
175633 Employees
Remote or Hybrid
2 Locations
175633 Employees
Hybrid
Chennai, Tamil Nadu, IND
175633 Employees
In-Office or Remote
2 Locations
175633 Employees

Similar Companies Hiring

NODA AI Thumbnail
Artificial Intelligence • Information Technology • Software • Cybersecurity
Sydney, AU
54 Employees
Golden Pet Brands Thumbnail
Digital Media • eCommerce • Information Technology • Marketing Tech • Pet • Retail • Social Media
El Segundo, California
178 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software
New York, New York
30 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account