The Role
Design, develop, and maintain scalable ETL pipelines and data infrastructure for large-scale structured and unstructured datasets. Integrate and govern data across sources, optimize pipeline performance, and deploy cloud-based solutions. Collaborate with data scientists, analysts, and engineers to deliver accessible datasets. Implement automation, testing, security, compliance, monitoring, and continuous improvements across data engineering workflows.
Summary Generated by Built In
Responsibilities:
- Design, develop, and maintain scalable ETL pipelines to process and transform large-scale datasets.
- Integrate structured and unstructured data from multiple sources, ensuring quality, security, and consistency.
- Collaborate with data scientists, analysts, and software engineers to deliver well-structured and accessible datasets.
- Build and optimize data infrastructure using big data technologies such as Apache Spark, Hadoop, and Kafka.
- Deploy and manage cloud-based data solutions on AWS, GCP, or Azure.
- Monitor and troubleshoot data pipeline performance, ensuring reliability and efficiency.
- Implement data governance, security, and compliance best practices.
- Drive automation, testing strategies, and continuous improvements in data engineering workflows.
Qualifications:
- 5 years of related experience with a Bachelor’s degree or equivalent work experience.
- Advanced proficiency in SQL and experience with relational and NoSQL databases (PostgreSQL, MySQL, MongoDB, etc.).
- Strong programming skills in Python, Java, or Scala for data processing and automation.
- Deep expertise in ETL processes, data modeling, and data warehousing.
- Hands-on experience with big data frameworks such as Apache Spark, Hadoop, or Kafka.
- Proficiency in cloud platforms (AWS Redshift, Google BigQuery, Azure Synapse) and data infrastructure automation.
- Experience optimizing data pipeline performance and scalability.
- Strong problem-solving skills with the ability to work on complex, large-scale datasets.
- Knowledge of data governance, security, and compliance best practices.
- Excellent leadership, collaboration, and communication skills to work effectively across teams.
Skills Required
- Five years of related experience with a bachelor's degree or equivalent work experience
- Advanced proficiency in SQL
- Experience with relational and NoSQL databases, including PostgreSQL, MySQL, or MongoDB
- Strong programming skills in Python, Java, or Scala
- Deep expertise in ETL processes, data modeling, and data warehousing
- Hands-on experience with Apache Spark, Hadoop, or Kafka
- Proficiency with cloud platforms such as AWS, Google Cloud, or Azure
- Experience with data infrastructure automation
- Experience optimizing data pipeline performance and scalability
- Strong problem-solving skills with complex, large-scale datasets
- Knowledge of data governance, security, and compliance best practices
- Leadership, collaboration, and communication skills
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Numentica is a consulting firm specializing in software product development, business intelligence, data management, data analytics, AI, and cloud services, focusing on digital transformation.








