Experience
5–8 Years
As per business requirement
We are looking for an experienced Data Engineer with strong expertise in MongoDB, PySpark, and Python to design, develop, and optimize scalable data pipelines. The ideal candidate should have experience working with large datasets, NoSQL databases, distributed data processing, and cloud-based data platforms.
- Design, build, and maintain scalable ETL/ELT data pipelines using PySpark and Python.
- Develop data ingestion frameworks to process structured, semi-structured, and unstructured data.
- Work extensively with MongoDB for data modeling, querying, indexing, aggregation, and performance optimization.
- Optimize Spark jobs for high-performance processing of large datasets.
- Build reusable data transformation and validation frameworks.
- Develop REST API integrations and automate data ingestion using Python.
- Monitor, troubleshoot, and optimize data pipelines for reliability and performance.
- Collaborate with business analysts, data scientists, and application teams to deliver data solutions.
- Implement data quality, governance, and security best practices.
- Participate in code reviews and follow CI/CD and Agile development practices.
- Strong experience in Python programming.
- Hands-on experience with PySpark and Spark SQL.
- Strong knowledge of MongoDB, including:
- CRUD Operations
- Aggregation Framework
- Indexing
- Replication
- Sharding
- Performance Tuning
- CRUD Operations
- Good understanding of data structures and algorithms.
- Experience in developing ETL/ELT pipelines.
- Strong SQL skills.
- Experience with Git version control.
- Knowledge of Linux/Unix commands.
- Experience working with JSON, XML, and Parquet data formats.
- Experience with Databricks.
- Experience with cloud platforms such as Azure, AWS, or GCP.
- Knowledge of Apache Kafka or other streaming technologies.
- Experience with orchestration tools such as Apache Airflow.
- Understanding of Delta Lake and Lakehouse architecture.
- Familiarity with CI/CD pipelines.
Skills Required
- Strong experience in Python programming
- Hands-on experience with PySpark and Spark SQL
- Strong knowledge of MongoDB (CRUD, Aggregation Framework, Indexing, Replication, Sharding, Performance Tuning)
- Experience developing ETL/ELT pipelines and data ingestion frameworks
- Strong SQL skills
- Experience with Git version control
- Knowledge of Linux/Unix commands
- Experience working with JSON, XML, and Parquet data formats
- Experience building REST API integrations and automating ingestion with Python
- Good understanding of data structures and algorithms
- Ability to monitor, troubleshoot, and optimize data pipelines for reliability and performance
- Implement data quality, governance, and security best practices
- Participate in code reviews and follow CI/CD and Agile development practices
- Experience with Databricks
- Experience with cloud platforms (Azure, AWS, or GCP)
- Knowledge of Apache Kafka or other streaming technologies
- Experience with orchestration tools such as Apache Airflow
- Understanding of Delta Lake and Lakehouse architecture
What We Do
Technology, society, economy, policy – all moving at breakneck speed in our 21st century world. You’re feeling the pressure to quickly implement new business models, find new value, make split-second informed decisions and keep one step ahead of customers. How? The answer lies in the ability to make quick, accurate and sustainable business decisions. We believe digital offers a way of doing things better – but the journey to transformation doesn’t have to be painful. At Aligned Automation, we work hard to digitally enable your business strategy – connecting processes, technologies and people to unlock value and drive critical business outcomes.








