Senior Data EngineerLocation/Region
Latin America | 100% Remote
About CodeRoadCodeRoad provides end-to-end software development services, helping businesses scale with ideal infrastructure solutions. From staff augmentation to dedicated IT teams and general software engineering, our nearshore technology services empower businesses to thrive in an ever-evolving digital landscape.
About the RoleAs a Senior Data Engineer at CodeRoad, you will serve as the technical backbone responsible for architecting and optimizing an agent-ready data foundation and semantic layer. You will model complex analytical views, optimize low-latency query performance for downstream LLM retrieval, and anchor robust access control and data privacy policies across modern database environments.
This role is critical to bridging complex operational data systems with cutting-edge AI frameworks, ensuring that downstream AI agents and analytical engines access accurate, highly secure, and optimized data in real time.
Key ResponsibilitiesDesign and implement modern semantic models (facts, dimensions, and metrics stores using tools like dbt) to abstract complex operational databases for AI querying and analytical consumption.
Build curated, high-performance analytical views and production ELT/ETL pipelines that streamline data retrieval, enhance freshness, and minimize redundant database queries.
Optimize database query performance, indexing, and schema definitions for low-latency analytical workloads through systematic query log analysis.
Lead the implementation of granular data governance frameworks, including Role-Based Access Control (RBAC), Row-Level Security (RLS), and dynamic data masking for PII protection.
Establish comprehensive data quality, pipeline monitoring, and anonymization standards across all analytical layers.
Collaborate with cross-functional AI engineering teams to align data structures with the operational requirements of RAG architectures and Agentic AI workflows.
4+ years of dedicated data engineering experience with strong, production-grade command of SQL (Azure SQL, PostgreSQL, or modern cloud data warehouses).
Solid foundation in dimensional modeling (Kimball methodology) and semantic layer / metrics store architecture (e.g., dbt).
Proven production experience building and maintaining resilient ETL/ELT data pipelines.
Demonstrated experience implementing database security policies, granular role-based permissions, and dynamic data masking / PII protections.
Hands-on expertise tuning complex queries for low-latency analytical workloads.
Practical experience working with RAG Applications and Agentic AI Frameworks.
Soft Skills: Strong ownership mindset, proactive problem-solving ability, and a highly collaborative approach within nearshore team environments.
Language: Advanced English fluency (written and spoken) is strictly required.
Experience with orchestration platforms such as Apache Airflow, Prefect, or Dagster.
Exposure to vector databases (e.g., pgvector, Pinecone, Qdrant) used in generative AI stacks.
Familiarity with streaming or event-driven technologies (e.g., Apache Kafka, AWS Kinesis).
Hands-on experience with cloud infrastructure as code (Terraform) or cloud-native CI/CD automation.
100% Remote culture allowing you to work from anywhere in Latin America.
Holidays off aligned with national and company calendars.
Paid Time Off (PTO) to ensure work-life balance.
Health insurance assistance coverage.
Competitive USD compensation paid directly to you.
Growth opportunities within high-impact, modern tech projects and multi-client ecosystems.
Skills Required
- 4+ years of dedicated data engineering experience
- Production-grade SQL experience with Azure SQL, PostgreSQL, or modern cloud data warehouses
- Experience with dimensional modeling using Kimball methodology
- Experience with semantic layer or metrics store architecture, such as dbt
- Production experience building and maintaining resilient ETL or ELT pipelines
- Experience implementing database security policies and granular role-based permissions
- Experience with dynamic data masking and PII protections
- Hands-on experience tuning complex queries for low-latency analytical workloads
- Practical experience with RAG applications and agentic AI frameworks
- Advanced written and spoken English fluency
- Experience with Apache Airflow, Prefect, or Dagster
- Exposure to vector databases such as pgvector, Pinecone, or Qdrant
- Familiarity with Apache Kafka or AWS Kinesis
- Experience with Terraform or cloud-native CI/CD automation
What We Do
CodeRoad is a nearshore software development and technology-execution partner that helps organizations modernize applications, migrate to the cloud, build digital products, and implement enterprise AI. Its services span the full software development lifecycle, including dedicated engineering teams, platform engineering, data and AI infrastructure, legacy modernization, cloud architecture, quality automation, and DevOps. The company serves startups, enterprises, and private-equity portfolio businesses.








