The Role
Build and maintain reliable cloud-based ETL and ELT pipelines using SQL, Python, Dagster, dbt, and AWS. Model data into business-ready datasets, implement testing and observability, troubleshoot pipeline failures, and deploy containerized solutions. Integrate GenAI capabilities using LLMs, AWS Bedrock, and LangChain. Collaborate with data scientists and stakeholders to support analytics, machine learning, and GenAI products.
Summary Generated by Built In
● Experience: 3+ years of hands-on experience in Data Engineering.
● GenAI Experience: Familiarity with LLMs, AWS Bedrock, or LangChain
● RAG: Experience or exposure to Retrieval-Augmented Generation (RAG).
● AI Automation: Experience automating workflows using AI and integrating LLM capabilities into engineering or data pipelines.
● Engineering Mindset: You treat data pipelines like software products. You are comfortable with Version Control (Git), code reviews, and testing.
● SQL Mastery: You can write complex, efficient queries and understand data modeling concepts (e.g., Joins, Window Functions, Normalization).
● Python Proficiency: You can write clean Python scripts for data manipulation and automation (beyond just "notebook scripting").
● Cloud Native: Familiarity with cloud data warehouses (Redshift, Snowflake, or BigQuery) and core cloud concepts (S3, IAM, Compute).
● Modern ETL: Experience with modern transformation tools (like dbt)
● Problem Solver: Excellent ability to investigate data issues. You don't just restart the job; you dig into the logs to find why it failed.
● CI/CD: Experience automating deployments using GitHub Actions, Jenkins, or similar.
Desired Experience (The "Nice to Have")
● Infrastructure as Code: Familiarity with Terraform or Terragrunt.
● Containerization: Experience running code in Docker or Kubernetes (EKS/ECS).
● Streaming: Exposure to real-time data frameworks (AWS Kinesis, Kafka, SQS/SNS).
● Data orchestration: Experience with orchestration tools like Dagster, Airflow, AWS Step functions, etc.
Requirements
About the Role
We are looking for a Data Engineer with a software engineering mindset to join our Data & AI team. This is not just a role for writing SQL scripts; it is an opportunity to build robust, scalable, and observable data infrastructure on the cloud.
You will work with a modern tech stack (Dagster, dbt, Clickhouse, AWS) to build the pipelines that power our analytics, machine learning, and GenAI products. If you care about code quality, automation, and "Data as a Product," this role is for you.
Our Tech Stack
● Languages: SQL, Python
● Orchestration: Dagster (migrating from Airflow).
● Data Stores: Redshift, Clickhouse, S3.
● Transformation: dbt, Fivetran.
● Cloud & Infra: AWS (ECS/EKS, Glue, Lambda, Athena)
● IaC: Terraform with Terragrunt.
● AI/GenAI: AWS Bedrock, LangChain, LLMs.
Key Responsibilities
● Innovation: Integrate GenAI capabilities (LLMs, LangChain) into our engineering workflows.
● Build & Orchestrate: Develop and maintain reliable ETL/ELT pipelines using SQL and Python. You will use Dagster to orchestrate dependencies, ensuring data flows correctly from source to destination.
● Data Transformation: Use dbt to model raw data into clean, business-ready datasets (Star Schema) that enable stakeholders to self-serve.
● Quality & Observability: Own the quality of your data. Implement tests (dbt tests, unit tests) and monitoring to ensure "silent failures" don't happen. You will troubleshoot pipelines when they break and fix the root cause.
● Cloud Engineering: Work with AWS services (S3, DMS, Glue) and containerized environments (Docker/Kubernetes) to deploy your code.
● Collaborate: Partner with Data Scientists and
Skills Required
- 3+ years of hands-on data engineering experience
- Familiarity with LLMs, AWS Bedrock, or LangChain
- Experience or exposure to Retrieval-Augmented Generation (RAG)
- Experience automating workflows using AI and integrating LLM capabilities into engineering or data pipelines
- Experience with version control, code reviews, and software testing
- Advanced SQL, including complex queries and data modeling concepts
- Proficiency writing Python scripts for data manipulation and automation
- Familiarity with cloud data warehouses and core cloud concepts
- Experience with modern transformation tools such as dbt
- Ability to investigate data issues and troubleshoot pipeline failures using logs
- Experience with CI/CD tools such as GitHub Actions or Jenkins
- Familiarity with Terraform or Terragrunt
- Experience with Docker or Kubernetes, including EKS or ECS
- Exposure to real-time data frameworks such as Kinesis, Kafka, SQS, or SNS
- Experience with data orchestration tools such as Dagster, Airflow, or AWS Step Functions
Am I A Good Fit?
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.
Success! Refresh the page to see how your skills align with this role.
The Company
What We Do
Centro CDX is a global business process outsourcing (BPO) and diversified technology solutions provider. It specializes in digital transformation and data intelligence, offering services that revolutionize business operations through process orchestration and security enablement. Operating across North America, Europe, the GCC, and Africa, the company focuses on streamlining processes and enhancing customer experience to drive sustainable growth for its clients.








