Data Engineer
As a Principal Data
Engineer, your responsibilities will include:
- Design and build data pipelines
to process terabytes of data
- Orchestrate in Airflow the data
tasks to run on Kubernetes/Hadoop for the ingestion, processing and
cleaning of data.
- Create Docker images for
various applications and deploy them on Kubernetes
- Design and build best in class
processes to clean and standardize data.
- Troubleshoot production issues in
our Elastic Environment
- Tuning and optimizing data
processes
· Advancing the team’s DataOps culture (CI/CD,
Orchestration, Testing, Monitoring) and building out standard development
patterns
- Drive innovation by
testing new technology and approaches to continually advance the
capability of the data engineering function.
- Drive efficiencies in
current engineering processes via standardization and migration of
existing on-premise processes to the cloud
- Ensuring Data
Quality – building best in class data quality monitoring that
ensure that all data products exceed customer expectations.
Required Qualifications:
- Computer Science bachelor’s
degree or similar.
- Good understanding of Data Modelling techniques i.e.
DataVault, Kimble Star
- Excellent understanding of Column-Store RDBMS
(DataBricks, Snowflake, Redshift, Vertica, Clickhouse)
- Good experience handling real-time, near real-time
and batch data ingestions
- Hands on experience on the
following technologies:
- Developing processes in Spark
- Writing complex SQL queries f
- Building ETL/data pipelines
- Exposure to Kubernetes and
Linux containers (i.e. Docker)
- Related/complementary open
source software platforms and languages (e.g. Scala, Python, Java, Linux)
- Proven track
record of designing effective data strategies and leveraging modern data
architectures that resulted in business value
- Experience
building cloud-native data pipelines on either AWS, Azure or GCP,
following best practices in cloud deployments
· Strong DataOps experience (CI/CD,
Orchestration, Testing, Monitoring)
- Strong
experience leading and developing data engineering teams
· Demonstrated effective interpersonal,
influence, collaboration and listening skills
· Strong stakeholder management skills
· Excellent time management, organizational and
prioritization skills with ability to balance multiple priorities.
Preferred Qualifications:
· Experience with data tokenization and different techniques and
tools i.e. DataVant, Protegrity
· Experience with Azure Data Factory, Databricks and Snowflake
- Experience
with Apache Spark and related Big Data stack and technologies, PySpark
Scala
· Experience working with Apache Kafka, building appropriate
producer/consumer apps
· Experience working with Kubernetes and Docker, and knowledgeable
about cloud infrastructure automation and management (e.g., Terraform)
· Experience working in projects with agile/scrum methodologies
· Familiarity with production quality ML and/or AI model development
and deployment.
· Healthcare industry knowledge and experience with exposure to EDI,
HIPAA, HL7 and FHIR integration standards
Skills Required
- Bachelor’s degree in Computer Science or a similar field
- Understanding of data modeling techniques, including DataVault and Kimball Star Schema
- Strong understanding of column-store relational databases, such as Databricks, Snowflake, Redshift, Vertica, or ClickHouse
- Experience handling real-time, near-real-time, and batch data ingestion
- Hands-on experience developing processes with Apache Spark
- Ability to write complex SQL queries
- Experience building ETL and data pipelines
- Exposure to Kubernetes and Linux containers, including Docker
- Experience with complementary open-source platforms and languages such as Scala, Python, Java, and Linux
- Track record of designing effective data strategies and modern data architectures that deliver business value
- Experience building cloud-native data pipelines on AWS, Azure, or GCP
- Strong DataOps experience, including CI/CD, orchestration, testing, and monitoring
- Strong experience leading and developing data engineering teams
- Effective interpersonal, influence, collaboration, and listening skills
- Strong stakeholder management skills
- Excellent time management, organization, and prioritization skills
- Experience with data tokenization tools and techniques, such as DataVant or Protegrity
- Experience with Azure Data Factory, Databricks, and Snowflake
- Experience with Apache Spark, PySpark, Scala, and related big-data technologies
- Experience with Apache Kafka and producer/consumer application development
- Experience with Kubernetes, Docker, and cloud infrastructure automation tools such as Terraform
- Experience working in Agile/Scrum environments
- Familiarity with production-quality machine learning or artificial intelligence model development and deployment
- Healthcare industry experience and familiarity with EDI, HIPAA, HL7, and FHIR integration standards
What We Do
Test Triangle is a Dublin-headquartered technology services company founded in 2012. It helps organizations transform operations through IT service management, digital transformation, software testing and quality assurance, cloud solutions, consulting, managed services, and recruitment. The company also provides application testing, DevOps, robotic process automation, software and mobile development, technology training, staffing, and resource augmentation across Ireland, the UK, and international markets.






