This is a remote position.
Role OverviewWe are looking for Data Engineers at Senior and Mid-Level to join our team in building a privacy-preserving data platform where data engineering meets production-grade software engineering.
You will work on designing, developing, and maintaining reliable data pipelines that transform operational data into high-quality, secure, and AI-ready datasets.
- Build and maintain production-grade data pipelines on AWS.
- Extract, transform, validate, and curate large-scale Parquet datasets.
- Implement data de-identification, masking, and privacy-preserving transformations.
- Design and maintain data pipeline orchestration, scheduling, retries, and backfill mechanisms.
- Implement comprehensive data quality checks, monitoring, and alerting.
- Work with workflow orchestration tools such as Airflow, Dagster, or AWS Step Functions.
- Contribute to CI/CD pipelines and Infrastructure as Code (IaC) practices.
- Manage schema evolution and schema drift across data sources and pipelines.
- Provide production support, troubleshooting, and root cause analysis for data pipeline issues.
- Maintain data catalogs, metadata, and data lineage.
- Follow software engineering best practices including Git, code reviews, automated testing, and maintainable code.
- Build reliable and idempotent data pipelines capable of handling retries and large-scale backfills.
- Strong proficiency in Python and SQL.
- Hands-on experience with AWS data services and production data pipelines.
- Experience with Apache Spark or equivalent distributed data processing technologies.
- Practical experience with Airflow, Dagster, AWS Step Functions, or similar orchestration tools.
- Strong understanding of data pipeline architecture, ETL/ELT, and data transformation.
- Experience working with Parquet and large-scale datasets.
- Understanding of data quality, schema management, monitoring, and alerting.
- Strong software engineering practices including:
- Git and version control
- Code reviews
- Automated testing
- Idempotency
- Error handling
- Retries and backfills
- Git and version control
- Experience supporting and troubleshooting production data pipelines.
- Ability to work effectively with cross-functional engineering and data teams.
Requirements Required Skills
Python | SQL | AWS | Data Engineering | Data Pipelines | ETL/ELT | Apache Spark | Parquet | Airflow | Dagster | AWS Step Functions | Data Orchestration | Data Quality | Schema Management | Data Transformation | Production Support | Git | CI/CD | Automated Testing | Data De-identification | Data Lineage | Data Catalog
Debezium | AWS DMS | Apache Iceberg | Delta Lake | Apache Hudi | Data Masking | Data Tokenization | Terraform | CloudFormation | ML/AI Training Data | Privacy-Preserving Data
Skills Required
- Strong proficiency in Python and SQL
- Hands-on experience with AWS data services and production data pipelines
- Experience with Apache Spark or equivalent distributed data processing technologies
- Practical experience with Airflow, Dagster, AWS Step Functions, or similar orchestration tools
- Strong understanding of data pipeline architecture, ETL/ELT, and data transformation
- Experience working with Parquet and large-scale datasets
- Understanding of data quality, schema management, monitoring, and alerting
- Experience with Git and version control
- Experience with code reviews and automated testing
- Experience implementing idempotency, error handling, retries, and backfills
- Experience supporting and troubleshooting production data pipelines
- Ability to work effectively with cross-functional engineering and data teams
- Experience with Debezium, AWS DMS, Apache Iceberg, Delta Lake, or Apache Hudi
- Experience with data masking, data tokenization, or privacy-preserving data
- Experience with Terraform or CloudFormation
- Experience preparing ML/AI training data
What We Do
Synthlane Technologies is a deep-tech software company that develops scalable, secure digital solutions for startups, enterprises, and government organizations. Its services include custom software development, consulting, application modernization, AI and machine-learning solutions, cybersecurity and digital forensics, ERP implementation, and managed applications. The company focuses on helping organizations integrate systems, improve workflows, modernize technology, protect digital assets, and achieve measurable business growth.









