Digital Divide Data (DDD) is an AI-enabled digital operations and data services company that delivers ML data solutions and content services to Fortune 500 companies and the world’s leading academic institutions.
Job Description- Design and build robust data ingestion pipelines, transforming raw CSV/TSV data into structured, queryable datasets.
- Model custom entities and map diverse data sources into standardised target schemas.
- Build and manage data environments using SQLite, Google Cloud SQL and BigQuery.
- Develop and maintain Dockerised services for data loading, embedding generation and data-serving applications.
- Deploy and operate production workloads on Google Cloud Run, with a focus on scalability, reliability and high availability.
- Work across Redis, VPC networking, Artifact Registry and Cloud Storage.
- Implement secure cloud architectures using Cloud IAM, VPC and Cloud Run ingress controls.
- Automate data workflows and deployments through GitHub Actions, including scheduled ingestion, ID mapping, embedding generation and container builds.
- Build and integrate AI-powered capabilities, including natural-language querying, semantic search, embeddings and LLM-powered workflows.
- Contribute to emerging agentic architectures and integrations, including technologies such as Model Context Protocol (MCP).
- 2+ years’ experience in data engineering, cloud engineering or backend development.
- Strong Python development skills, including practical experience with Pandas.
- Solid experience with Docker and containerised applications.
- Strong working knowledge of Google Cloud Platform, particularly cloud compute, managed databases, data warehousing, networking and IAM.
- Strong SQL skills and experience working with both relational and analytical data stores.
- Experience designing, building and consuming REST APIs.
- A strong engineering mindset, with the ability to work independently, troubleshoot complex problems and take ownership from development through to production.
It’s a plus if you have
- Experience with AI/ML engineering, particularly embeddings, Sentence Transformers or LLM applications.
- Exposure to agentic frameworks, MCP or other emerging AI integration patterns.
- Experience with Flask and Jinja.
- Familiarity with WSL and Google Cloud Shell.
Why this opportunity?
Because you’ll be building more than pipelines.
You’ll help create the technical foundation for intelligent, data-driven applications, where secure cloud infrastructure, high-quality data and AI come together.
If you’re looking to deepen your expertise in cloud and data engineering while growing into applied AI, this role offers the opportunity to work on that evolution firsthand.
Applications are reviewed on a rolling basis, so apply early.
Skills Required
- 2+ years of experience in data engineering, cloud engineering, or backend development
- Strong Python development skills, including practical Pandas experience
- Strong experience with Docker and containerized applications
- Strong working knowledge of Google Cloud Platform, including compute, managed databases, data warehousing, networking, and IAM
- Strong SQL skills and experience with relational and analytical data stores
- Experience designing, building, and consuming REST APIs
- Ability to work independently, troubleshoot complex problems, and own work from development through production
- AI/ML engineering experience, particularly embeddings, Sentence Transformers, or LLM applications
- Exposure to agentic frameworks, MCP, or emerging AI integration patterns
- Experience with Flask and Jinja
- Familiarity with WSL and Google Cloud Shell
What We Do
Digital Divide Data (DDD) provides end-to-end AI and autonomy solutions, specializing in human-in-the-loop data annotation, validation, and ML model training. Trusted by Fortune 500 companies and government entities, DDD supports the lifecycle of autonomous systems and generative AI. Founded in 2001, the company operates on a unique social impact model, providing professional opportunities and education to talented youth from low-income backgrounds, while ensuring high-quality, reliable AI performance.








