Every day, somewhere in the world, important decisions are made. Whether it is a private equity company deciding to invest millions into a business or a large corporation implementing a new strategic direction, these decisions impact employees, customers and other stakeholders.
Consulting and private equity firms come to proSapient when they need to discover knowledge to help them make great decisions and succeed in their goals. It is our mission to support them in their discovery of knowledge.
We help our clients find industry experts who can provide their knowledge via interview or survey; we curate this knowledge in a market-leading software platform; and we help clients surface knowledge they already have through expansive knowledge management.
We are looking for a Data Engineer to help build and scale the Knowledge Graph at the heart of our AI roadmap.
You will join the Data Foundation squad, building the pipelines that populate, maintain and serve the graph as a reliable Content / Data product. This is a hands-on role for someone who enjoys meaningful data modelling challenges across entities, relationships, taxonomies and versioning.
The key duties of this role will include:
Graph Construction Pipelines:
· Build ingestion and transformation pipelines that populate the Knowledge Graph from internal systems, enrichment outputs and third-party sources.
· Implement entity resolution and linking logic in production.
· Manage late-arriving data, conflicting sources and graph schema changes over time.
Serving the Graph:
· Build query and serving layers used by search, matching and data product teams.
· Model the graph for analytical and application use cases across systems such as BigQuery, PostgreSQL and Elasticsearch / OpenSearch.
· Support external-facing data products with reliable export and delivery pipelines.
Quality & Operations:
· Implement data quality checks, lineage and monitoring across graph pipelines.
· Own production operations, including alerting, backfills and incremental reprocessing.
· Contribute to data contracts so downstream teams can rely on the graph.
Requirements
The key skills needed for this role are:
· 3+ years building production data pipelines.
· Strong Python and SQL skills, with experience writing clean, production-ready code.
· Experience with PostgreSQL or other relational databases.
· Strong data modelling skills, including schema design for evolving requirements.
· Experience integrating messy, multi-source data, including deduplication and normalisation.
· Hands-on experience with Elasticsearch / OpenSearch or a comparable serving layer.
· Comfortable with testing, code review, CI and operating your own pipelines.
Nice to have:
· Experience with graph databases, graph data modelling, taxonomies or ontologies.
· Experience with entity resolution at scale.
· Experience with BigQuery, Kafka, dbt or similar data platforms and frameworks.
· Familiarity with Docker, Kubernetes and cloud environments such as AWS or GCP.
Benefits
What we can offer you:
· Tenure gifts, including vouchers, extra holiday and sabbaticals for each year of employment.
· Health insurance through Vitality.
· Remote working for up to 20 days each year, giving you flexibility and a change of scenery.
· Employee Assistance Programme with personalised health and wellbeing advice from specialist teams.
· Enhanced maternity and paternity pay.
· 25 days’ annual leave plus bank holidays, including a week’s closure over Christmas.
· MyMindPal app for online mental fitness support.
· Corporate events, from quarterly gatherings to annual winter and summer parties.
We are committed to building an inclusive workplace. Marginalised groups are often less likely to apply unless they meet every requirement listed, so if you are interested in this role but do not tick every box, we encourage you to apply anyway — it could still be a great match.
Skills Required
- 3+ years building production data pipelines
- Strong Python skills and experience writing clean, production-ready code
- Strong SQL skills
- Experience with PostgreSQL or other relational databases
- Strong data modeling skills, including schema design for evolving requirements
- Experience integrating messy, multi-source data, including deduplication and normalization
- Hands-on experience with Elasticsearch, OpenSearch, or a comparable serving layer
- Experience with testing, code review, CI, and operating production pipelines
- Experience with graph databases, graph data modeling, taxonomies, or ontologies
- Experience with entity resolution at scale
- Experience with BigQuery, Kafka, dbt, or similar data platforms and frameworks
- Familiarity with Docker, Kubernetes, and cloud environments such as AWS or GCP
What We Do
proSapient provides a comprehensive primary research and expert networking service to institutional investors, consultants and corporate clients around the world. Our advantage is the speed and accuracy we deliver actionable intelligence and relevant experts to our customers, giving them the edge they need to achieve their commercial ambitions. This is made possible by our technology which is driven by artificial intelligence, separating us from other providers of market information








