Terminology bindings ... again

Hi Mikael,

What efforts are being made to resolve the boundary problem?

I applied to get involved with the SNOMED information modelling group but wasn’t successful, to try to engage on exactly that point.

I’m not aware of any work going on. I’d be very pleased to get involved if I could. It’s a fiendish problem and we need cooperation and collaboration from both sides of the fence.

Regards

Heather

Hi Heather,

In general, anyone is welcome to participate in the work; you don’t need to be one of the small number of Advisory Group members. That helps with travel costs, but most of the real work is done on teleconferences, not so much at the face to face meetings.

I would be very interested to hear people’s articulations of where they think the boundary should be for this boundary line. I’d also be interested to understand better what people think the problem is with having “extra” / unnecessary pre-coordinated concepts; what advantage is to be gained from removing them, and what is the perceived scale of the problem.

michael

HI Michael,

I have no idea what teleconferences are happening. I don’t know how to engage. It’s not easy from the outside!

Heather

Hi Heather,

SNOMED International collaboration site is nowadays located at the address https://confluence.ihtsdotools.org/ . (It seems unfortunately that the link is more clicks away from the start page than before. :frowning: ) At that address, you can click “Spaces” in the upper left corner of the page and then “Space directory”. You will then find a long list with the collaboration space for different groups and other resources. To my knowledge there aren’t any group called exactly “information modelling group”, but there are several groups that work will modelling, so maybe the group has another name? If you only would like to read the material, I think that you don’t need any account on the site, but if you would like to participate a little closer, you can apply for an account on the Confluence site.

Regards

Mikael

One simple rule solves the boundary problem.

In my words.
In principle models we use in the Semantic Stack are autonomous and orthogonal to each other. Meaning that at precisely defines points the intersect,
SNOMED is an ontology defining terms for concepts.
Concept can be primitive or compound.
Primitive concept examples are: Left and Right and Eye, White, Red and Blue. All terms one expects in a dictionary.
Compound concepts are pre-co-ordinated concepts: Left Eye, Right Eye. Eye White coloured Red, ‘Blue eye’. All terms one expects in a pattern/standard phrase using words from a dictionary in a syntax.

The rule is:
Pre-coordination is done via Archetype/Patterns using Primitive concepts.
Pre-coordination is not done via an ontology/terminology.

Off course clinicians on their screens or doing statistics need Compound concepts and therefor need pre-coorodinated terms.
These pre-coordinated terms must never be used to store, retrieve, interpret raw health data inside Health IT-systems.

Gerard Freriks
+31 620347088
  gfrer@luna.nl

Kattensingel 20
2801 CA Gouda
the Netherlands

I have made some attempts to study the problem in the past, not recently, so I don’t know how much the content has changed in the last 5 years. Two points come to mind:

  1. the problem of a profusion of pre-coordinated and post-coordinatable concepts during a lexically-based choosing process (which is often just on a subset).
    this can be simulated by the lexical search in any of the Snomed search engines, as shown in the screen shots below. Now, the returned list is just a bag of lexical matches, not a hierarchy. But - it is clear from just the size of the list that it would take time to even find the right one - usually there are several matches, e.g. ‘blood pressure (obs entity)’, ‘systemic blood pressure’, ‘systolic blood pressure’, ‘sitting blood pressure’, ‘stable blood pressure’ and many more.

I would contend (and have for years) that things like ‘sitting blood pressure’, ‘stable blood pressure’, and ‘blood pressure unrecordable’ are just wrong as atomic concepts, each with a separate argument as to why. I won’t go into any of them now. Let’s assume instead that the lexical search was done on a subset, and that only observables and findings (why are there two?) show up, and that the user clicks through ‘blood pressure (observable entity)’, ignoring the 30 or more other concepts. Then the result is a part of the hierarchy, see the final screenshot. I would have a hard time building any ontology-based argument for even just this one sub-tree, which breaks basic terminology rules such as mutual exclusivity, collective exhaustiveness and so on. How would the user choose from this? If they are recording systolic systemic arterial BP, lying, do they choose ‘systemic blood pressure’, ‘arterial blood pressure’, ‘systolic blood pressure’, ‘lying blood pressure’, or something else.

Most of these terms are pre-coordinated, and the problem would be solved by treating the various factors such as patient position, timing, mathematical function (instant, mean, etc), measurement datum type (systolic, pulse, MAP etc), subsystem (systemic, central venous etc) and so on as post-coordinatable elements that can be attached in specific ways according to the ontological description of measuring blood pressure on a body. This is what the blood pressure archetype does, and we might argue that since that is the model of capturing BP measurements (not an ontological description of course), it is the starting point, and in fact the user won’t ever have to do the lexical choosing above. Now, to achieve the coding that some people say they want, the archetype authors would have the job of choosing the appropriate codes to bind to the elements of the archetype. In theory it would be possible to construct paths and/or expressions in the archetype and bind one of the concepts from the list below to each one. To do so we would need to add 40-50 bindings to that archetype. But why? To what end? I am unclear just who would ever use any of these terms.

The terms that matter are: systemic, systolic/diastolic, terms for body location, terms for body position, terms for exertion, terms for mathematical function, and so on. These should all be available separately, and be usable in combination, preferably via information models like archetypes that put them together in the appropriate way to express BP measurement. Actually creating post-coordinated terms is not generally useful, beyond something like ‘systemic arterial systolic BP’, or just ‘systolic BP’ for short, because you are always going to treat things like exertion and position separately (which is why these are consider ‘patient state’ in openEHR), and you are usually going to ignore things like cuff size and measurement location (things considered as non-meaning modifying ‘protocol’ in openEHR).

  1. similar problems in the authoring phase, i.e. addition of concepts to the terminology in the first place. If more or less any manner of pre-coordinated terms is allowed, with the precoordinations cross-cutting numerous ontological aspects (i.e. concept model attribute types), what rules can even be established as to whether the next proposed concept goes in or not? It is very easy to examine the BP hierarchy, and suggest dozens of new pre-coordinated terms that would fit perfectly alongside the arbitrary and incomprehensible set already there…

(another 3x)

I’ve picked just the most obvious possible example. We can go and look at ‘substances’ or ‘reason for discharge’ or hundreds of other things, and find similar problems. I don’t mind that all these pre-coordinated concepts exist somewhere, but they should not be in the primary hierarchies, which really, in my view should look much more like an ontology, i.e. a description of reality which provides a model of what it is possible to say. If that were the case, the core would be much smaller, and the concept model much larger than it is today.

  • thomas

Hi tom,

I can agree with you that if SNOMED CT was created when all patients in the world already had all information in their health record recorded using cleverly built and structured information models (like archetypes, templates and similar), but that is not the case. Instead SNOMED CT also tries to help healthcare organizations to do something better also with their already recorded health record information, because that information to a large extent still belongs to living patients.

It would be interesting to have your opinion about why it is a real problem with the “extra” pre-coordinated concepts in SNOMED CT in general and not only for the use case of creating archetypes or what would be nicest in theory.

Regards

Mikael

(attachments)

image001.png
image002.png
image003.png

Hi Mikael,

but that would appear to be an argument that since data in systems is dirty, badly designed, so SNOMED will reflect this by being badly designed! I don’t agree with this at all. I think that the majority of SCT concepts which are arbitrary pre-coordinations (presumably due to the fact that someone saw them in real data, or knows they are in use in their hospital) constitute part of a (pretty messy, but practically useful) interface terminology. I remember myself and others arguing for this in the I&I standing committee in about 2011. That should be moved to a whole separate ‘piece’ of SNOMED - leaving a clean core that follows proper rules like mutual exclusion, collective exhaustiveness, coherence in distinction criteria down the IS-A tree and so on. That core would be small, manageable, and the concept model could then be worked on properly to build it out, providing a far richer basis for pre-coordination. This model would be something that certain archetype structures could be harmonised with. (Note though that many archetype structures don’t have a biochemical or biophysical basis, but something like ‘cultural models of recording’). The interface part could then start to be cleaned up and have some discipline of its own; it could be more effectively connected to tools that are actually used in real interfaces, like choosing widgets, or underlying logic for searching. These terms would all be mapped back to the core via post-coordinated expressions - the ultimate test of the approach. The main problem with the current state of affairs is simply that the vast majority of concepts creates incoherent cognitive noise when people are trying to:

I read Thomas’ reply with great interest, and I generally agree that with a well thought out information model, the very detailed precoordinated expressions are redundant. At the same time I understand Mikael’s point of view too. BUT, what I’m often met with is that because these precoordinated expressions exist (like for example “lying blood pressure” and “sitting blood pressure”), we should use them INSTEAD OF using our clever information models (that we do have) for recording new data.

In my opinion this is wrong because it doesn’t take into account that healthcare is unpredictable, and this makes recording more difficult for the clinician. How many different variations would you have to select from? Take the made up example “sitting systolic blood pressure with a medium cuff on the left upper arm”; this will be a lot of possible permutations, especially if you take into account all the different permutations where one or more variable isn’t relevant.

So while I don’t think the existence of these precoordinated terms in itself is a problem, it’s a potential problem that people get a bit overzealous with them.

(attachments)

image001.png
image002.png
image003.png

Thomas,

I agree with your opinions.

In summary:
Model 1- Ontological models define primitive concepts in Terminologies. Concepts that one can expect to see in a dictionary.
Model 2- Archetypes define compound concepts using primitive concepts from terminologies. Archetypes are models that model a concept and its epistemology/context. Archetypes are agreed re-usable patterns and are like sentences constructed using syntax and words from a dictionary. This is the level where post-coordination takes place.
Model 3- Templates are specific models composed of Archetypes defining the content of an interface (database, screen, clinical reasoner, etc) Templates are temporal, locally defined, chapters and paragraphs defining/documenting patient data obtained in the healthcare process.
Model 4- For specific reasons users might want to see pre-coordinated terms on screens, or statistical reporting demand it. This means that raw data obtained using Model 3 can be converted to pre-coordinated terms from a Terminology. It is a User Interface issue.

All models are orthogonal to each other and intersect at well defined places. Models can not be mixed and overlap.
When they do, the problem is intractable.

Next to these 4 models we need additional models:

  • Healthcare Process (ContSys)
  • Documentation process (Observation Process, Evaluation Process, Planning Process, Ordering process, Execution Process, Administrative processes)

Gerard Freriks
+31 620347088
gfrer@luna.nl

Kattensingel 20
2801 CA Gouda
the Netherlands

Maybe a match-table which matches pre coordinated expressions to all possible post coordinated expressions which have the same meaning will be necessary.

How can you else find data-entries of a specific meaning by filtering on SNOMED?

Bert

(attachments)

image003.png
image001.png
image002.png

IMO having both representations (pre and postcordinated) is not bad per se (in fact, knowing that they are equivalent is pretty good). The main problem is that technical people (including myself) shouldn’t just dump the entire snomed ct into a data field and call it a day. To design better and useful systems you need a first “curation” phase where you define your relevant subsets that fit your system. The boundary problem is less of a problem if even if different terms were used when the record was created we can assess that they are in fact the same thing.

I think people are a little unaware of this step and causes problems as the ones you and Thomas mentioned

(attachments)

image003.png
image001.png
image002.png

Dear Silje,

I think we agree.

In my view it is not wise to use pre-coorinated codes that include contextual information. The reason is that the complete why, when, who and how result in too many permutations in order to be tractable.
One must make the distinction between how data is expressed in a generic system interface connected to other services such as: database, clinical reasoner, import/export, etc. and between a specific interface that users connect with for data inspection, data entry, statistical data analysis.
The former must NOT use pre-coordinated terms; the latter depending on user requirements will use pre-coordinated terms that can be constructed using the raw data in the generic system interface.

Gerard Freriks
+31 620347088
gfrer@luna.nl

Kattensingel 20
2801 CA Gouda
the Netherlands

Diego, this is a wise thought!!!
It seems logical, but that is often in good ideas, they seem like: why did no one ever think about this.

No clinician handles the complete medical science/SNOMED repository in his profession. Some even use a very small sub-part, think about a dentist, a physiotherapist, a midwife
For some is it also the case that not only the subjects are different, but also how deep the details must go. For some professions it is not interesting to know if blood-pressure was measured lying or sitting.

It looks like a good idea if the SNOMED repository will be split up for professions, maybe that needs to be done on national level, because the clinical profession-structure can differ in countries.
The rest of the database should be there for second searches, for interpreting codes which come from other professions.

I hope someone will pick up this idea, because it is hardly to be done for individuals. It needs to be done by national SNOMED maintenance teams.

Bert

(attachments)

image001.png
image002.png
image003.png

Totally agreed, Silje. I think preordination for anatomical location is invaluable, but it’s the only use case that I have identified as one we absolutely can’t do without.

But I would love the opportunity to really investigate this properly, and with others who understand SNOMED better than I. That will help with the boundary issue/semantically grey area.

I’d prefer that we could use and reuse simpler, really high quality value sets from in multiple archetypes for different contexts eg a list of diagnoses in the Problem/diagnosis archetype as well as the Family History archetype. The archetype context is invaluable here. And the terminology community focussing on high value terms that would provide great impact.

Regards

Heather

(attachments)

image001.png
image002.png
image003.png

Yes.

The simple rules that solve the Boundary Problem need some exceptions.

  • anatomical structure
  • possibly some aspects of devices

To investigate it properly is a possibility.
The problem I see is the How question.
For me there is only one solution and that is by analysing the problem and thinking about it.

Ingredients that we could take in consideration are notions like:

  • Models in the Interoperability Stack must be autonomous and orthogonal

  • The Models we need in the Stack

  • Closed world assumption of Archetypes versus Open world assumption of the ontology behind SNOMED

  • Requirements for data exposed to users via statistics, screens, forms, and documents

  • Requirements for data stored, retrieved from databases

  • Requirements for data to be interpreted by clinical reasoners

  • Modelling methods

Gerard Freriks
+31 620347088
gfrer@luna.nl

Kattensingel 20
2801 CA Gouda
the Netherlands

Heather,

As you know Brazil has chosen to adopt SNOMED CT on a business case basis, trying to create clinical models that make the better use of structures and meaning.
Therefore my research group has dedicated to study detailed clinical models and to deepen our knowledge of Snomed CT. We agree with GF that we have four ways to do it and that depends on the use cases. As Mikael said there’s no fix boundary between when you use pre or post coordinatinated expressions, although context to me naturally is easier to be represented in the structure ( templates).
TO leverage our knowledge os Snomed CT everyone in our team has taken the Snomed CT Foundation Course. I think that many assumptions made here are explained there.
BRAZIL has just joined Snomed CT International. We intend to propose a creation of a Worlgroup focussed on DCMs, for the openEHR community, as Thomas tried to do years ago.
I will be In London for the next Business meeting, and gladly would have a meeting with others with the same objective.
Jussara Rötzsch

(attachments)

image002.png
image003.png
image001.png

Hi Jussara,

Our approach to building archetypes at present are intended to represent the data as domain experts require it and making no assumptions about the terminology used, including use of pre/post coordination principles etc. The models have to allow end users to choose how they engage with terminology and ensure that it can be represented. As such the international archetypes are maybe overly comprehensive but we have to cater for those using heavily post coordinated SCT terms through to others using only local value sets.

We have deliberately avoided having separate models that assume pre/post coordination as per CIMI equivalents because we view that as being ungovernable, and maintenance of semantic equivalence adds layers of complexity that will likely be impossible to maintain.

However I do see a time when we need to publish associated documentation for models that propose ‘best practice’ use of SNOMED and maybe other terminologies, which will then support interoperability with the same value sets being used in the same way.

It is that ‘best practice’ that I’d love to see start to evolve, which includes requirements gathering that results in parallel archetype AND value set definition/development for use cases and defined contexts.

As clinical program co-lead, to the best of my knowledge, at the moment the formal part of openEHR has no direct influence on the SNOMED CT work/priority and that is frustrating.

Collaborative work is where the magic will start to happen!

Cheers

Heather

(attachments)

image001.png
image002.png
image003.png

Hi Bert,

Most countries (or big organizations) that start to use SNOMED CT creates those kinds of subsets of SNOMED CT to make it more manageable. See for example NHS in England https://isd.digital.nhs.uk/trud3/user/guest/group/0/pack/40 .

Regards

Mikael

(attachments)

image001.png
image002.png
image003.png

Thanks Mikael, very interesting. I will check if they do that in the Netherlands too.

Bert

(attachments)

image003.png
image002.png
image001.png