Thanks, Erik, for publishing this so openly and in such concrete detail.
Disclosure first: I work at Ultipa, where we build a real-time graph database whose native query language is ISO GQL (ISO/IEC 39075). Please read what follows with that in mind. These aren’t a product pitch. Three parts of the document overlap closely with problems we’ve worked on, so here are some notes.
- AGQL: build on GQL rather than beside it
I’d suggest keeping AGQL a strict superset of ISO GQL: archetype-aware syntactic sugar that desugars to plain GQL.
- AQL CONTAINS maps naturally onto GQL path patterns.
- Archetype and template IDs map onto label expressions or node predicates.
- A stored query is translated once, when it’s saved, which matches your point that stored queries are the ones that count in production.
- Any conformant GQL engine could run them, with no AGQL-specific runtime to maintain.
A shared, openly licensed mapping of the RM onto a property graph, plus an AQL→GQL test suite (AQL in, expected results out) on the same Vital Signs data Borut uses, would let every engine be compared on identical queries. Is anyone already working on that?
- One planner for terminology and records, able to run from either side
We agree with your point that the best plan “can’t be determined without looking at both knowledge graph and EHR simultaneously”. Take << 73211009 |Diabetes mellitus| in problem lists:
- For a small cohort, the cheapest plan walks up the IS-A hierarchy from each patient’s recorded codes.
- For a population query, expanding the descendant set first wins.
- An engine can only choose between these when both sides share one planner with statistics on both. Split across two systems, the choice is frozen in application code, usually as “ship the whole descendant list”.
A related point: evaluating subsumption at query time, rather than from materialized closure tables, makes your remark literally true that non-breaking terminology updates can enrich the graph “without the already recorded data instances changing”. A new SNOMED release changes answers without rewriting patient data. The trade-off is that query-time inference costs something on every query, while materialization costs on every release, so most systems end up mixing the two.
- Scope patients inside the traversal and the vector search
For GraphRAG, the patient/cohort restriction has to be a pre-filter inside both the graph traversal and the vector search, not a post-filter on results. Post-filtering either leaks data or quietly returns too little.
One case worth adding to evaluation criteria: approximate vector indexes (e.g. HNSW) under a very selective filter, such as one patient’s notes among millions. Unless the engine falls back to exact search over the allowed set, they can return fewer than k results without any error.
Happy to go deeper on any of these, here or in the performance thread.
Ricky @ Ultipa