Data Architect + AI
Overall Stack: Databricks Lakehouse Platform, Apache Spark, Delta Lake, Unity Catalog, and modern cloud data architecture
Must Have
Lakehouse & Medallion Architecture: Expertise in designing end-to-end data architectures (Bronze, Silver, Gold layers) for reliable, production-ready pipelines
Databricks & Spark Internals: Deep understanding of distributed computing. They must know how to troubleshoot and tune large-scale Spark jobs using caching, partitioning, and broadcast joins
Delta Lake: Must understand ACID transactions, schema enforcement, time travel, and optimization operations like Z-ordering
Unity Catalog & Data Governance: Proven ability to design unified governance models for data and AI assets, including role-based access control (RBAC), row/column-level security, and data lineage
Cloud Infrastructure (AWS, Azure, or GCP): Strong grasp of the native cloud ecosystem they work in (e.g., ADLS/Entra for Azure, S3/IAM for AWS), including Virtual Network (VNet) setups and IAM roles
Coding Proficiency: Advanced SQL skills and fluency in Python or Scala
Cost Optimization & Performance Tuning: Ability to monitor DBUs (Databricks Units), right-size serverless and multi-node clusters, and implement best practices for avoiding cloud bill shock
Nice to Have
Databricks Certifications: Candidates holding valid Databricks Certified Data Architect or Databricks Certified Data Engineer Professional badges generally have a proven, up-to-date baseline of the platform's features
Generative AI & MLflow Integration: Experience building, deploying, and monitoring GenAI applications and ML models using Databricks Model Serving, Vector Search, and the Mosaic AI suite
CI/CD & DevOps Practices: Experience automating Databricks workflows using Git (Databricks Repos) and orchestration tools like dbt, Azure Data Factory, or Apache Airflow
Streaming Data: Familiarity with Databricks Structured Streaming and Auto Loader for real-time data ingestion and processing
Data Warehousing & BI: Understanding of Databricks SQL, Serverless Warehouses, and integration with downstream BI tools like Power BI
Remote
Adavenced english
Skills Required
- Expertise designing end-to-end lakehouse and medallion architectures with Bronze, Silver, and Gold layers
- Deep understanding of Databricks and Apache Spark internals, distributed computing, Spark troubleshooting, caching, partitioning, and broadcast joins
- Understanding of Delta Lake ACID transactions, schema enforcement, time travel, and Z-ordering
- Experience designing Unity Catalog data governance models, including RBAC, row-level security, column-level security, and data lineage
- Strong knowledge of AWS, Azure, or GCP cloud infrastructure and related networking and identity services
- Advanced SQL proficiency and fluency in Python or Scala
- Ability to monitor Databricks Units, right-size clusters, and optimize cloud costs and performance
- Advanced English proficiency
- Databricks Certified Data Architect or Databricks Certified Data Engineer Professional certification
- Experience with Generative AI, MLflow, Model Serving, Vector Search, or Mosaic AI
- Experience with Git, Databricks Repos, dbt, Azure Data Factory, or Apache Airflow
- Familiarity with Structured Streaming and Auto Loader
- Understanding of Databricks SQL, Serverless Warehouses, and Power BI integration
What We Do
VALCE Talent Solutions is a recruitment agency and consulting firm specializing in IT talent acquisition and nearshoring, primarily in Mexico. They design customized solutions in IT talent and process optimization to help businesses scale intelligently and profitably. Their expertise focuses on specialized industries including Information Technology, Operations Management, and Supply Chain, providing end-to-end talent offerings to attract and retain top professional profiles.








