Suggestions re Term binding in Archetype Editor

Andrew Patterson wrote:

This is a point of view on archetypes that I had never really
considered. I had always assumed that the construction of
the archetype definition and the selection of terminology
binding would be part of the same process, either done by
the same person or the same clinical group. The paper discusses
some of the mapping problems that can occur when this
process is split, but surely that would never be the case?

the process can easily be split and is likely to be so, e.g. in environments like the NHS where there are separate teams of experts for clinical modelling and archetypes. There are other reasons as well:

  • an archetype is instantly usable even without terminology binding being figured out; it may be that prototype testing or a non-terminology enabled rollout is required
  • many things in the ‘ontology’ section of an archetype come after the main modelling exercise: translation, addition of other term bindings; adjustment of term bindings
  • an archetype may be published in a ‘vanilla’ form, with term bindings only being added in local situations (although clearly an internationally agreed binding to internationally use terminologies is preferable).
The intention would be that within some master archetype
repository (be this a local/organisational/national repository),
the archetypes would include a full set of terminology
codes?

the archetypes would only include some cods that are direct maps for the ‘at’ coded terms in the archetype. But for attributes where the value (i.e. the answer) is a term value set from a terminology like snomed, the approach is to us an ‘ac’ code to refer to a query to the terminology service, whose result when run is the subset from which the user is allowed to choose at runtime. This query is stored in a terminology query service, only the id is referenced in the archetype.

(I can understand how one might think this because
none of the sample archetypes in the openehr repository have much
terminology data but that will surely that is a temporary situation
and the intention is that a real official repository i.e. one run
by nehta or nhs etc would have the term codes for their realm?)

correct - but not because of the lack of the archetype repository (one is under construction); the reason is that the community has not yet settled on a way of identifying or expressing subsetting queries to a terminology service; good progress is being made here; at Ocean we have developed a simple language to do it, which turns out to be very nearly the same semantics as the OWL-based expressions described by Alan Rector at a recent Semantic Mining meeting in Paris at which we were both present.

Which brings me onto a related point - at the snomed
workshop in Melbourne late last year there was an
(impressive) demonstration of some of the template building
tools written by Ocean. Part of the demonstration involved
creating a complex binding to snomed based on a
small query language (effectively the query was
"select all 'is_a' children of this snomed code up to a maximum
depth of 5").

right - this is the language I mention above.

 This query binding was placed into the relevant
archetype as a URL reference to a webservice. Doesn't relying
on a URL in the ADL definition
make archetypes quite brittle. i.e. when the archetype definition
is loaded into the clinical system I either have to consult the
URL straight away and store the resulting codes, or else delay
the binding and risk having the terminology codes for my
ADL disappear in the future?

why would that happen? This isn’t semantic brittleness in the archetype - the meaning doesn’t change; it may be that you don’t have a reliable deployment environment. This is why we built the Ocean Terminology Server to provide a query cache for any user site.

It just seems to me that if snomed is indeed the way things
seem to be going, that most terminology references in ADL will
either very simple (this node = 34242343) or need a moderate
level of complexity (all nodes in the 'is_a' 'route of administration'
heirarchy, but not ones with a qualifier 'blah'). Will all these
later style terminology bindings need to be done with URL's?

URLs won’t include the query statement itself (originally we thought they would, but having built a full terminology server and query subsetter has taught us otherwise); the URL will just include an id of the query. Working out an identification system is the next trick…

Isn't it going to be hard to keep these URL's alive for the
lifetime of the archetypes? On the other hand, if the URL's
are bound on archetype entry, how will they keep up with
changes to the terminology?

Should there be a small query language for terminology
built into ADL?

as I say, that’s what I thought we would be doing a year ago; experience has shown that we need to push things apart by one further level, and only allow subset query ids in archetypes (remember - these are what is mapped to the ‘ac’ codes; you can still put snomed or whatever codes in the ‘at’ code binding part if you want).

None of this is simple and it has taken both us (in the industrial context) and the Manchester group over 5 years from the statement of need to get to a solution - which looks quite simple now we seem to have it (or be close), but I have to admit, we were groping in the dark for a long time…

  • thomas beale

Andrew Patterson wrote:

Given Australian government
departments barely keep their names for more than a few years, what are
the chances the URL is still working? If it is a local reference, what are the
chances the machines still have the same IP addresses or names?
  

this points to the need to use safe URLs that will work. Do we think the
URL "http://snomed.org" is safe? Maybe we need "http://terminology.net"
or somesuch. The use of URLs whose meaning will not change over time is
not to be taken lightly...

Can the clinical system still rely on the term codes it cached 10 years
ago?
  

probably not - just consider if they had cached the answer to the query
"types of hepatitis" in 1985 compared to now - the answer now is a
larger set of codes.

I think having URL's in the archetype definition mixes the 'configuration' of
the system with the definition. I guess I would like to see some sort of
query language in this space so that one could say

<"at0004"> = <
   <"snomed"> = <all 'route of medication' where refset('australia')>
  <"icd-10"> = <12312-23>
  

this is where we were about 1-2 years ago in the thinking of this
problem; now we have reached a point of:
* firstly having realised that the query should not itself go into the
archetype, only an id for the query
* secondly we have a basic query language ("subset expression language"
would be closer to the mark); the Manchester group has worked out
something like an OWL-based equivalent from a theoretical perspective (I
don't want people to think we were first here - they have been working
on this a long time as well - as far as I can work out, we have done it
more or less independently over the last 2 years, and the results as at
Paris in early December point toward being able to publish a common
expression syntax/language for this purpose very soon).

I am not sure what we should be doing about ideas like
"refset('australia')" though....
- thomas beale

make archetypes quite brittle. i.e. when the archetype definition
is loaded into the clinical system I either have to consult the
URL straight away and store the resulting codes, or else delay
the binding and risk having the terminology codes for my
ADL disappear in the future?

why would that happen? This isn't semantic brittleness in the archetype -
the meaning doesn't change; it may be that you don't have a reliable
deployment environment. This is why we built the Ocean Terminology Server to
provide a query cache for any user site.

So is the URL just a name/ident for the terminology (ala XML namespaces)
or an actual "I'm a HTTP server URL, send me a GET request for some codes"..
I can understand how the first adds no extra brittleness - but surely you
can see how the second adds more than just meaning into the
archetype. Embedded in the 'meaning' is the configuration for your
deployment environment.. so whether the URL is http://snomed.org/query2323
or http://127.0.0.1:111/query434, the archetype now has embedded within
it a commitment to maintain that particular server, at that particular port
under that particular name running for the life of
the archetype. If I decide to deploy a new local terminology server
5 years down the track, do I need to get my system administrators
to load each archetype in the system and manually change the
port number of the server in all the term binding URLs?

It just smells a bit funny to me - I can't put my finger on exactly why
or propose a better way, but it just gives me the vibe of something
that could come back and bite people on the arse down the track.

as I say, that's what I thought we would be doing a year ago; experience
has shown that we need to push things apart by one further level, and only
allow subset query ids in archetypes (remember - these are what is mapped to
the 'ac' codes; you can still put snomed or whatever codes in the 'at' code
binding part if you want).

None of this is simple and it has taken both us (in the industrial context)
and the Manchester group over 5 years from the statement of need to get to a
solution - which looks quite simple now we seem to have it (or be close),
but I have to admit, we were groping in the dark for a long time....

I admit I am playing catch up here and I accept that your results are
from quite a bit of experience on your part - can you give me a brief
background on the reasons for the change in thinking from
a year ago with embedded queries in the archetype to the current
thinking.

Thanks

Andrew

Thomas Beale wrote
> Andrew Patterson wrote:
> Given Australian government
> departments barely keep their names for more than a few years, what are
> the chances the URL is still working? If it is a local reference, what are the
> chances the machines still have the same IP addresses or names?
>
this points to the need to use safe URLs that will work. Do we think the
URL "http://snomed.org" is safe? Maybe we need "http://terminology.net"
or somesuch. The use of URLs whose meaning will not change over time is
not to be taken lightly...

[...]
A safe URL is needed for each clinical terminology authority. (e.g. standards are managed by ISO, Standards Australia etc.; SNOMED UK and SNOMED AUS may have country specific needs)

> I think having URL's in the archetype definition mixes the 'configuration' of
> the system with the definition. I guess I would like to see some sort of
> query language in this space so that one could say
>
> <"at0004"> = <
> <"snomed"> = <all 'route of medication' where refset('australia')>
> <"icd-10"> = <12312-23>
>

Agreed.

this is where we were about 1-2 years ago in the thinking of this
problem; now we have reached a point of:
* firstly having realised that the query should not itself go into the
archetype, only an id for the query

[...]

I am not sure what we should be doing about ideas like
"refset('australia')" though....
- thomas beale

The query tool needs to manage this, as it should manage the language. I suggest the user (or user environment) should be able to select whether to look at local terminology or that of another country (the default may be where the patient's record was created, and the patient was travelling at the time or subsequently emigrated).

On storing a new event, the term accessed at the time should be stored in the record, not (just) the SNOMED code or URL, because the terminology may be updated on the server, and a future access should show the info known at the time of entry.

Colin Sutton (not an authoritative answer!)

EuroRec, the European Institute for Health Records, has the intention to become the neutral point of reference for several services needed around the EHR.
One of those Services are an Archetype and Template Repository and Inventory.
Eurorec intends to do the same for Coding Systems.

This will create the stable managed environments EHR-systems and EHR’s need.

Gerard

– –
Gerard Freriks, MD
Huigsloterdijk 378
2158 LR Buitenkaag
The Netherlands

T: +31 252544896
M: +31 620347088
E: gfrer@luna.nl

Those who would give up essential Liberty, to purchase a little temporary
Safety, deserve neither Liberty nor Safety. Benjamin Franklin 11 Nov 1755

Andrew Patterson wrote:

make archetypes quite brittle. i.e. when the archetype definition
is loaded into the clinical system I either have to consult the
URL straight away and store the resulting codes, or else delay
the binding and risk having the terminology codes for my
ADL disappear in the future?

why would that happen? This isn't semantic brittleness in the archetype -
the meaning doesn't change; it may be that you don't have a reliable
deployment environment. This is why we built the Ocean Terminology Server to
provide a query cache for any user site.
    
So is the URL just a name/ident for the terminology (ala XML namespaces)
or an actual "I'm a HTTP server URL, send me a GET request for some codes"..
  

The exact structure of the URL is the part that is currently not
determined. Its logical meaning is exemplified by:

terminology = xxxx
subset = a query stored somewhere whose meaning is "any kind of
infection that has-site liver"

for example:

terminology = xxxx
subset = 1234-5678-91bc-def0

Currently we have a way to write such subset expressions (significantly
more complex than this) and evaluate them. If the syntax of such an
expression were standardised and therefore evaluatable by any
terminology service containing an instance of snomed-ct, then we would
probably design the URL to simply indicate the terminology and the query
id. But - since the id of the query will make sense against multiple
terminologies (i.e. "any kind of infection that has-site liver" could
possibly be evaluated against ICD10 or some other disease terminology)
then the namespace of the queries is truly global; if not, then query
namespaces are with respect to terminologies. My guess is that a few
hundred will handle most questions on most forms. There will also be
queries that are subsets of existing subsets. Even with a few hundred,
the ids need to be well-designed (it may just be Guids as I have implied
above, but then there has to be an agreed registry of subset queries,
and an agreement of whether they are globally unique or not).

The problem is to get some international agreement on how to approach
the identification problem. Technically it is easy either way.

I can understand how the first adds no extra brittleness - but surely you
can see how the second adds more than just meaning into the
archetype. Embedded in the 'meaning' is the configuration for your
deployment environment.. so whether the URL is http://snomed.org/query2323
or http://127.0.0.1:111/query434, the archetype now has embedded within
it a commitment to maintain that particular server, at that particular port
under that particular name running for the life of
  

we certainly would not want to do this - there should be no implication
of a server; only of a particular query.

the archetype. If I decide to deploy a new local terminology server
5 years down the track, do I need to get my system administrators
to load each archetype in the system and manually change the
port number of the server in all the term binding URLs?
  

which is the reason for not doing that. URIs properly used should never
identify servers - Berners-Lee got that right years ago, it's just that
almost no-one realised how important he was, and we live in a world of
broken URLs rather than logically sound URIs.

It just smells a bit funny to me - I can't put my finger on exactly why
or propose a better way, but it just gives me the vibe of something
that could come back and bite people on the arse down the track.
  

which is why we need to be careful with this...same as for codes in
terminologies (see e.g. Cimino's desiserata, or rules on good
terminology design).

None of this is simple and it has taken both us (in the industrial context)
and the Manchester group over 5 years from the statement of need to get to a
solution - which looks quite simple now we seem to have it (or be close),
but I have to admit, we were groping in the dark for a long time....
    
I admit I am playing catch up here and I accept that your results are
from quite a bit of experience on your part - can you give me a brief
background on the reasons for the change in thinking from
a year ago with embedded queries in the archetype to the current
thinking.
  

well, one obvious (to us now) reason is that the correct formal query
for "any problem with site liver" might not be gotten right the first
time round. If we embedded the query in every archetype that had an
attribute with that meaning, then you can see the maintenance nightmare
and potential for error. If we instead only have to change it once in an
authoritative service (or the terminology itself, assuming the subset
queries could be added to a chapter of the terminology...) then life is
a lot simpler. Also, newer subsetting syntaxes might come along; again
the logical meaning of the archetype doesn't change, so we don't want it
to have the actual subset query expression in there, just the id of the
subset.

- thomas

Colin Sutton wrote:

The query tool needs to manage this, as it should manage the language. I suggest the user (or user environment) should be able to select whether to look at local terminology or that of another country (the default may be where the patient's record was created, and the patient was travelling at the time or subsequently emigrated).
  

the main thing that needs to be known is if snomed-ct-AUS is a proper
superset of snomed-ct-INT (or however these variants will be
designated); i.e. that there are no incompatibilities between the two.

On storing a new event, the term accessed at the time should be stored in the record, not (just) the SNOMED code or URL, because the terminology may be updated on the server, and a future access should show the info known at the time of entry.

well, snomed is designed never to change the meaning of a code, so
having the code should be safe. But in openEHR we always store the code,
the rubric (in the local language) and the full id of the terminology
(including version if relevant).

- thomas beale

Andrew

I just want to stress that the URLs do not have to be real (as in XML Schema) - but they are IDs for the termset that they define and can be kept in a database somewhere. We could have used anything - but a URL does have the advantage that it might point to something. It doesn’t matter if the website closes down - the URL in the archetype will still get the right query from the Terminology server.

Cheers, Sam

Andrew Patterson wrote:

Not at all - just see the URL as a Query ID - it is completely separate from the query spec and terminology service.
Sam

Andrew Patterson wrote:

Schema) - but they are IDs for the termset that they define and can be kept
in a database somewhere. We could have used anything - but a URL does have
the advantage that it might point to something. It doesn't matter if the
website closes down - the URL in the archetype will still get the right
query from the Terminology server.

Yep, it all makes sense to me now. I had just misunderstood them
to be purely URL's, rather than URI's. Perhaps the documentation
should have some samples showing a non-URL forms as well.

like

urn:www-terminology-org:snomed:1234-5678-91bc-def0

or something.

Andrew