The Role
Design, build, and maintain ETL pipelines and data platforms using Python, PySpark, SQL, and Airflow. Support cloud data platforms (AWS/Azure/GCP), maintain relational databases, convert unstructured data into vectors, collaborate with data/ML teams, troubleshoot pipeline issues, and document processes.
Summary Generated by Built In
Key Responsibilities:
- Design, develop, and maintain ETL (Extract, Transform, Load) processes to ensure the seamless integration of raw data from various sources into our data lakes or warehouses.
- Utilize Python, PySpark, SQL and AirFlow etc., to process, analyze, and store large-scale datasets efficiently.
- Write and maintain SQL queries for data retrieval, transformation, and storage in relational databases like Redshift or PostgreSQL.
- Support cloud-based data platforms such as AWS, Azure, or GCP, with a focus on orchestrating AI retraining cycles, versioning, and automated pipeline monitoring.
- Familiarity in converting unstructured data into vectors using frameworks like LangChain or LlamaIndex and storing them.
- Collaborate with cross-functional teams, including data scientists, ML engineers, and domain experts to design and implement scalable solutions.
- Troubleshoot and resolve performance issues, data quality problems, and errors in data pipelines.
- Document processes, code, and best practices for future reference and team training.
Requirements
Additional Information:
- Experience level 3+ years.
- Strong understanding of data governance, security, and compliance principles is preferred.
- Ability to work independently and as part of a team in a fast-paced environment.
- Excellent problem-solving skills with the ability to identify inefficiencies and propose solutions.
- Experience with version control systems (e.g., Git) and scripting languages for automation tasks.
Skills Required
- 3+ years experience in data engineering or related role
- Python
- PySpark
- SQL
- Airflow
- Experience with Redshift or PostgreSQL
- Experience with AWS, Azure, or GCP
- Familiarity with LangChain or LlamaIndex for vectorization
- Experience with version control systems (e.g., Git)
- Scripting for automation tasks
- Strong understanding of data governance, security, and compliance principles
- Ability to work independently and in teams
- Excellent problem-solving skills
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Feathersoft (a ThinkBio.AI company) is a global provider of onshore-offshore software development and IT solutions. It specializes in digital transformation, cloud services, and data analytics, with deep expertise in Machine Learning and AI. The company primarily serves the Fintech, Healthtech, and Agritech industries, delivering scalable technology solutions and consulting services to help enterprises implement impactful IT systems and insight-generating data infrastructure.








