Graph RAG - Knowledge Engineer

Posted 10 Hours Ago
Be an Early Applicant
Hyderabad, Telangana, IND
In-Office
Senior level
HR Tech • Information Technology • Software
The Role
Build and operate production Python and FastAPI services for extracting entities and relationships from biomedical documents, resolving entities, maintaining large-scale knowledge graphs, and combining graph and vector retrieval. Responsibilities include ontology design, deduplication, provenance, hybrid GraphRAG ranking, monitoring, and production support. The role requires expertise in knowledge-graph construction, applied NLP, entity resolution, graph databases, scalable retrieval, and cloud-native engineering.
Summary Generated by Built In
About Chryselys

Chryselys is a Great Place to Work Certified Pharma Analytics & Business consulting company that delivers data-driven insights leveraging AI-powered, cloud-native platforms to achieve high-impact transformations. We specialize in digital technologies and advanced data science techniques that provide strategic and operational insights.

Role Summary

Design, build and operate the Python/FastAPI services that extract entities and relationships from unstructured documents, resolve them to canonical identifiers, maintain the knowledge graph, and serve graph-augmented retrieval alongside vector search for multi-hop and relational questions.

Responsibilities
  • Build entity and relation extraction services over unstructured documents — molecules, brands, indications, therapeutic areas, endpoints, claims.
  • Build entity resolution: alias handling, blocking and candidate generation, fuzzy and embedding matching, calibrated thresholds, human review routing.
  • Design and maintain the graph schema and ontology; incremental ingest, node and edge deduplication and merging, provenance on every edge.
  • Fuse graph and vector results into a single ranked, cited context for the retrieval service.
  • Instrument, monitor and support the services in production.
Qualifications
  • 5–9 years software engineering, with demonstrable knowledge-graph construction and applied NLP delivered to production.
  • Has built a knowledge graph from unstructured text — not queried an existing one, and not a CRUD application on a graph database.
  • Graph at production scale. Millions of nodes and edges; incremental updates with stable node identity; supernode and traversal-explosion handling with bounded depth and timeouts.
  • Entity resolution at corpus scale. Blocking and candidate generation that avoid O(n²) comparison, with measured precision on a labelled sample.
  • Graph database in production. Neo4j, Amazon Neptune or equivalent; Cypher / openCypher fluency.
  • Ontology and taxonomy modelling. Schema evolution without breaking downstream consumers; judgement on node vs. edge vs. property.
  • Extraction. NER and relation extraction — LLM-based, model-based (spaCy, scispaCy, transformers) or hybrid, with the judgement to choose.
  • Graph vs. vector judgement. Knows where graph retrieval wins — multi-hop, relational, comparative and aggregate questions — and that hybrid is the production norm.
  • Python and FastAPI in production. Python 3.11+, async, Pydantic, Docker, pytest, Git and CI; AWS as a consumer (S3, ECS/EKS, Bedrock, Neptune or self-hosted Neo4j).
Preferred
  • Biomedical ontologies and registries: UMLS, MeSH, SNOMED, RxNorm, ICD-10, DrugBank, ChEMBL.
  • Life sciences or pharma domain experience; RDF/SPARQL alongside property graphs.
  • GraphRAG approaches: community detection for corpus-level summarisation, local vs. global search.

Skills Required

  • 5-9 years of software engineering experience
  • Demonstrable experience building knowledge graphs from unstructured text and delivering applied NLP systems to production
  • Experience operating production-scale graphs with millions of nodes and edges, incremental updates, stable node identity, bounded traversal depth, and timeout handling
  • Experience with corpus-scale entity resolution, including blocking, candidate generation, and measured precision on labeled samples
  • Production experience with a graph database such as Neo4j or Amazon Neptune and fluency with Cypher or openCypher
  • Ontology and taxonomy modeling experience, including schema evolution and node, edge, and property design
  • Experience with NER and relation extraction using LLM-based, model-based, or hybrid approaches
  • Understanding of when graph retrieval outperforms vector retrieval and how to implement hybrid retrieval
  • Production experience with Python 3.11+, FastAPI, asynchronous programming, Pydantic, Docker, pytest, Git, and CI
  • Experience using AWS services such as S3, ECS/EKS, Bedrock, Neptune, or self-hosted Neo4j
  • Knowledge of biomedical ontologies and registries including UMLS, MeSH, SNOMED, RxNorm, ICD-10, DrugBank, or ChEMBL
  • Life sciences or pharmaceutical domain experience
  • Experience with RDF and SPARQL alongside property graphs
  • Experience with GraphRAG approaches such as community detection and local versus global search
Am I A Good Fit?
beta
Get Personalized Job Insights.
Our AI-powered fit analysis compares your resume with a job listing so you know if your skills & experience align.

The Company
15 Employees
Year Founded: 2024

What We Do

Proof-of-Skill is a hiring and talent-matching platform that helps candidates find roles aligned with their skills, values, experience, and goals. It verifies identity, education, work history, and skills through proctored assignments and expert validation, then connects verified candidates with relevant employers and interviews. For businesses, it provides candidate skill-assessment software intended to reduce resume screening and improve recruiting efficiency for modern teams and hiring decisions.

Similar Jobs

Optum Logo Optum

Software Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Software Engineering Lead

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Senior Software Engineering Lead - Python Fullstack, FastAPI, GenAI, AI ML

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Optum Logo Optum

Principal Software Engineer

Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
In-Office
Hyderabad, Telangana, IND
160000 Employees

Similar Companies Hiring

Onshore Thumbnail
Artificial Intelligence • Fintech • Software • Financial Services
New York, New York
60 Employees
Revel.io Thumbnail
Aerospace • Hardware • Robotics • Software
US
50 Employees
Blee Thumbnail
Artificial Intelligence • Marketing Tech • Software • Productivity
US
15 Employees

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account