Archetype publication question - implications for implementers

Hi All

Taking this part of the change, I do not see any reason not to add a unit (really symbol change only) and mark the old one as deprecated. The data is unchanged and there is no risk to processing whatsoever.

The location change is a little more complicated and seems to be due to moving a DV_TEXT up a level in the archetype - is that the gist of it? The data appears to be the same. I have read the DIFF but am not sure about the actual motivations here.

Cheers, Sam

Dr Sam Heard
openEHR Board of Governors
Liaison with Management Board

But that will not be v1 anymore...

At this point, anyone who has worked for a time with the archetypes of CKM
knows that the readable archetype ID, including the version number, it is
not a reliable reference to identify the archetypes (this is said somewhere
in the specifications, but should be more clearly stated for newcomers).
The only reliable identifier from a technical point of view is the MD5 hash
of the definition part of the archetype. Any change to the structure will
create a different MD5. Any (correctly implemented) system that uses it
will find that it is a new archetype, call it v1, v1+internal revision, v2
or whatever.

As Diego said, the less complicated solution is to just follow the
versioning rules that already exist.

David

The existing versioning rules allow adding new concepts and opening constraints such as allowing additional units. These change the md5 hash but does require a version /id change.

This is why Sebastian’s suggestion technically works, the existing obsolete concept still exists and the new concepts can be added for those that want to move to the preferred approach.

However, I am concerned about adding new concepts that are in fact the same as an existing just represented differently. This is why I suggested to add new units to the existing tilt element (opening the constraint) rather adding a new concept for tilt with the new units.

I also suggested adding the new representation for body site as a new concept but starting to think this is a bad idea since we are duplicating the concept and have two ways in the same archetype to represent the same concept and worse the concept has two identifiers.

Having said that I suspect the alternative representation is filling a slot with a cluster archetype in a template and hence there is no duplicate concepts in the same archetype, there is a new slot and the alternate representation is in a template instead. Is this any better? Perhaps marginally.

Regards

Heath

Hi, Ian - and everyone else

Yes, I agree on the principle for modellers to evaluate the changes related to potentional downsides. And that’s what happening now. Good. Still, after reading some of the comments, I felt to state that it has to be the clinical need that should have the priority, despite how painful it will be to version up the archetype. If there is a clinical need - for someone - and it’s weighty enough, then it must lead to a new version (or another approach - which others are discussing now). My point were not related to the actual discussion of the BP archetype.

Vebjørn

Hi Heath,

I think the suggested change was from

CLUSTER[at1033] occurrences matches {0..1} matches { – Location
items cardinality matches {1..; unordered} matches {
ELEMENT[at0014] occurrences matches {0..1} matches { – Location of measurement
value matches {
DV_CODED_TEXT matches {
defining_code matches {
[local::
at0025, – Right arm
at0026, – Left arm
at0027, – Right thigh
…
}
}
}
}
ELEMENT[at1034] occurrences matches {0..1} matches { – Specific location
value matches {
DV_TEXT matches {
}
}
}
}
}

to

ELEMENT[at0014] occurrences matches {0..1} matches { – Location of measurement
value matches {
DV_CODED_TEXT matches {
defining_code matches {
[local::
at0025, – Right arm
at0026, – Left arm
at0027, – Right thigh
… }
}
DV_TEXT matches {*} – Specific location

}

which is how we would model it now.

As a side-note, ADL2.0 introduces the capacity to formally deprecate an existing node, which will be very helpful in these sort of situations.

I also favour adding the ‘deg’ unit to the existing Tilt quantity, as I think Sebastian suggested, as an alternative (and not making the change above). That would allow us to add the critical changes in a non-breaking manner.

Thanks to everyone who has contibuted so far - we still need other implementer views!!

Ian

Thanks Ian for explaining the actual proposal.

I can’t see why you can’t add the dv_text to the location of measurement (opening the constraint). Then it just becomes an implementation choice to exclude the specific location inline with the new preferred model.

Regards

Heath

But then what do we do with current implemented adl 1.4 archetypes?

David,

The changes I’m proposing to the v1 archetype are non-breaking, so it will remain a v1 revision and as the changes are not related to content, can be republished.

The changes that are proposed and breaking will be folded into a draft for v2, which will be visible as a potential future version.

Heather

Hi Ian,

These ADL changes don’t represent the non-breaking changes suggested, but do represent one option proposed for consideration in a v2, if we decided it warranted trying to update it to our current modelling pattern for Body site – it is effectively removing a redundant CLUSTER heading that adds no value.

Thanks

Heather

Hi everyone

Thanks for the suggestions and active discussion.

My main issue was with how to proceed with potentially breaking changes to correct technical artefacts and seeking strategies to manage the tension between pure modelling and implementation.

Your suggestions have been helpful. We have developed some strong rules about governance from identification of changes/mods that require either new versions, minor revisions and patches. These rules seem to be quite solid and have withstood this discussion.

However through the discussion we have identified some ways to include these changes in a non-breaking way, and that has been very welcome.

As a result I have just uploaded an updated version that includes all the proposed non-breaking changes. According to our governance rules, it is classified as a minor revision to our existing v1 and it has been republished.

See http://www.openehr.org/ckm/#showArchetype_1013.1.130

We can now consider at leisure whether we want to proceed with the breaking changes, although to a large degree these are not of semantic value and I’m inclined now to note them but not proceed with a proposed v2 at this point. We can do so at any time in the future, if we want.

Kind regards

Heather

Hi

A breaking change should always be new major version.

Then the problem is that changing version number introduces a huge cost. The cost is of course to the implementers – and by this I mean vendors, health care providers, national registries, integrations and so on. The whole ecosystem is influenced by such a change. Which makes it necessary to do a) not eagerly push major changes and b) when needed major changes should not influence earlier entries.

In this concrete change of Archetype there is two major changes:

  1. Introduce UCUM as UNIT

I think we should just add a new unit – the UCUM and make the older deprecated. But keep both of them. This makes it possible to migrate slowly to the new schema for unit. In this case it is important to verify that the magnitude 90 is the same for each of the unit. And it is the only two units used. This makes it somehow safe to compare the magnitude without checking the unit.

  1. Migrate from CLUSTER for Location to a choice between DvCodedText and DvText

This changes is on one side a change in pattern and on the other side a reduction in functionality.

I guess the pattern change is introduced to handle uncertainty in two flavours.

  1. The list of local codes will never be complete – let us introduce the choice to also use free text

  2. The list of local codes is not precise enough – let us introduce the choice to explain the element with free text

This uncertainty will always be present when using coded text to describe a phenomena. It will never be precise enough and cover all use-cases. Given this – what is the criteria to introduce choice between text and coded text? Is this a normal case for this kind of elements? To simplify :

Colours: a) RED, b) BLUE, c) GREEN d)Other. RED, BLUE and GREEEN cover the 80% use-case. Other covers the rest. We have several options to model this:

a) Expand the list of colours and add new colours as new requirements appear

b) Introduce a supporting element to specify colour if other is selected

c) Leave colours as Text field with the possibility to

a. Add list of items in Template (limited or not limited to list)

b. Add list of coded items in Template

c. Bind element to Terminology in Template

Why is pattern a) chosen as the best way to model this kind of features?

The reduction in functionality in the Archetype is because the user now is restricted to 0..1 coded text or text. Before could user choose between 0..1 coded text and an additional text so describe details of the location. I guess this is an wanted reduction in functionality and the intention may be to make entries more precise.

In my, technical, opinion: The changes introduced on this specific archetype should be applied in in such a way that no breaking changes are introduced. Just add UCUM and leave the CLUSTER as is and use the specific location to add specific details. Propose the changes as a possible new major version (v2-ALPHA….) and collect more changes before forcing a new major version.

Hi Heather et al,

Whilst I have followed this thread and agree with many of the observations and conclusions reached so far, I would like to make the following observations, which are restricted to the aspect of the “non-UCUM" unit described in Heather’s original posting. They have more serious and broader implications than this one iconic Blood Pressure Archetype.

Notes on UCUM

According to the UCUM specification at http://unitsofmeasure.org :-

"The Unified Code for Units of Measure provides a single coding system for units that is complete, free of all ambiguities, and that assigns to each defined unit a concise semantics. In communication it is not only important that all communicating parties have the same repertoir of symbols, but also that all attach the same meaning to the symbols they exchange. The common meaning must be computationally verifiable. The Unified Code for Units of Measure assumes a semantics for units based on dimensional analysis.”

  • UCUM introduces ambiguity, despite the above claim.
  • UCUM is complex and comprehensive. It brings together units from various other standards into a single framework. It is designed to support computability and communications interoperability, and hence adopts a highest common denominator 7-bit represention of unit codes as normative for sharing.
  • UCUM does not provide a single code for each unit - it provides 2 normative codes, as well as a non-normative display/print rendition and an ad-hoc name. For each unit, UCUM defines a case-sensitive version, a case-insensitive version, and a version intended for display or printing.
  • Some units have no display/print variant defined.
  • UCUM defines every unit in terms of 7 metric base units and does so in a coherent and consistent fashion. This can support conforming systems to perform conversions from one unit to another.
  • UCUM does not supply normative names of units.
  • Some similar units have quite dissimilar UCUM variants. e.g.

°C Cel for temperature print and case-sensitive variants respectively.

°F [degF] for temperature print and case-sensitive variants respectively.

° deg for plane angle print and case-sensitive variants respectively.

  • Some of UCUM’s names have been US-ised. E.g. litre has been changed to liter, metre to meter, deca to deka.
  • UCUM’s names don’t follow the 7-bit rule. Some names like Ampère and Ångström use 8-bit character representation.
  • UCUM uses [qualifier]s to indicate where a base unit is changed, e.g. mm is a unit for length property whereas mm[Hg] is a unit for pressure property. The syntax is unnecessary and complicates implementations.
  • UCUM provides no simple guide for use, particularly regarding normative components such as c/s, c/i and print.
  • UCUM inconsistently defines print representations of some units as normative and others as non-normative depending on the table.
  • UCUM’s print codes are often 8-bit. UCUM is premised on providing support for the highest common denominator across information systems, by constraining its normative unit strings to 7-bit values. However, unit print and unit name will not work in 7-bit environments.
  • UCUM releases are clearly supported and versioned, although differences between versions are hard to determine.
  • The contents of all UCUM specification tables are published as a single XML file for download.
  • Implementers of UCUM must choose between case insensitive and case-sensitive versions. The two cannot co-exist in the same channel of communication without special additional processing.
  • Even using UCUM, some units are difficult to represent and agree upon e.g. the unit for measuring estimated Glomerular Filtration Rate, quite common in healthcare - see http://unitsofmeasure.org/trac/ticket/98
  • Further useful information: http://motorcycleguy.blogspot.com.au/2009/11/iso-to-ucum-mapping-table.html

Notes on openEHR’s implementation of UCUM within the AOM.

  • openEHR is a standard primarily for building software components for electronic health records. openEHR Archetypes are being created and adopted by national bodies in many countries as canonical models for supporting interoperability of data shared between systems. Whilst these two goals are not mutually exclusive, they do present challenges and compromises for openEHR. EHR systems do need to care about supporting the entering and representation of data to users.
  • openEHR’s implementation of units is a compromise between these to goals that imposes minimal implementation overhead.
  • openEHR is designed to work in an 8-bit unicode/UTF-8 world. All openEHR-based applications are likely to support unicode characters and clearly ‘°’ would be part of that world. There are interesting examples such as the Observation Archetype “Fundoscopic examination of eyes” that constrains Field Angle values to “30º”, “45º", etc.. as Coded Text.
  • The openEHR DataTypes specification defines properties, including units for DV_QUANTITY types. The specification is a little vague on adherance to UCUM for units, stating that unit strings are to be expressed in UCUM unit syntax. This allows for support of units beyond those defined in UCUM.
  • The current openEHR BNF for parsing units appears to have some errors if it were to be considered UCUM-conforming - e.g. presence of ‘μ’ symbol; absence of ‘[‘ and ‘]’ symbols.
  • openEHR is unclear on which variant(s) of UCUM ( case-sensitive, case-insensitive, print ) should be supported or mandated.
    The current openEHR BNF for parsing units cannot support 8-bit UCUM units such as ‘°’ i.e. degree symbol in values conforming to type DV_QUANTITY.

Tooling implementations for openEHR units

( I have only looked briefly at these )

  • ADL Workbench - appears to have support for UCUM beyond the openEHR spec. - e.g. all 7 electronically published fields can be stored within the Workbench. Units appear to be fully parsed against the UCUM spec.
  • Ocean Archetype Editor - constrains creation of archetyped units to the set defined in a units and properties XML file distributed with the AE. This set does not match the UCUM units. Current implementation allows only ‘°’ for degree symbol. Recent versions support UCUM ‘overrides’ for each unit. Where UCUM values have been explictly entered, they have been entered inconsistently - some are case-sensitive, some are case-insensitive. Processing of UCUM overides does not appear to have been implemented yet. See http://lists.openehr.org/pipermail/openehr-clinical_lists.openehr.org/2014-February/003102.html .
  • Clinical Knowledge Manager - Many archetypes on the international CKM use non-valid UCUM units. I have not examined archetype processing, but assume from the archetypes in the international repository that existing archetypes are not validated against UCUM. Even basic archetypes like Body Temperature define units which are not UCUM-conformant. I consider this to be a more serious issue than the tilt angle of a person whose blood pressure is being recorded.

Implications for Archetypes, Archetype repositories and Archetype Governance

  • The CKM Body Temperature archetype constrains values of quantities to have units of “°C” or “°F”. Most archetype tools I suspect, display the value of the units from the ADL verbatim. WYSISYG. This works well in the user interface world. “°C” and“°F” are both valid UCUM print rendition of units (they just happen to be invalid for sharing). Most EHR applications probably behave similarly - units would be displayed directly as stored in data sources.
  • Given openEHR proclaims to support UCUM as the standard for units in DV_QUANTITY, it would seem sensible to constrain to UCUM values in ADL. Most tools don’t do this currently. Most archetypes have invalid values for UCUM. Most tools don’t support both display/print as well as case-sensitive normative UCUM values.
  • Some unit issues, such as the Blood Pressure tilt angle could be “fixed” simply by adding the normative UCUM unit and flagging the deprecated unit as such. However, tools would need to be updated to allow these new, “fixed” units to be entered where they previously could not. Only some of the existing openEHR units could be “fixed” this way.
  • I think units are somewhat akin to terminology systems like SNOMED. There are significant implementation complexities. The main value in standardising units is to ensure safety and quality of data from disparate sources. The main additional value in adopting UCUM more fully is to support unit conversion.
  • Ideally, in order just to support UCUM well, openEHR implementations should support the case-sensitive UCUM value, the print value and the unit definitions, all supplied by UCUM via the published XML file. This does not mean that the DV_QUANTITY type needs to change, but it would mean updating existing archetypes to replace the current archetype units with the correct UCUM case-sensitive one. Let’s call this the UCUM code. openEHR archetype editors would then map these UCUM codes to unit displaynames ( i.e. the UCUM print value ). openEHR applications would also ideally map the UCUM code to unit displaynames. i.e. Applications and archetypes use UCUM codes internally, but those codes aren’t displayed to humans.
  • If there are grounds for changing the Blood Pressure Archetype to “fix” the Tile Angle, then those same grounds dictate that many more Archetypes be changed. This should be undertaken as a major versioning exercise, probably with 6 months warning and with as many ducks lined up before the new archetypes are published. Many deployments will need to change. Testing of those deployments will need to be undertaken. Consideration will need to be given to how to support existing data in live applications.
  • There will always be tension between national and international archetype repositories trying to produce models for an ideal world, rather than for implementations that have to operate in the real world. My observations of how the openEHR world is evolving is that these archetype repositories do generally, and should try to set a gold standard. They should push implementations rather than retard them. That model, in turn, puts a pressure on the repositories to be of high quality, comprehensive, current. That, in turn implies publishing new versions. Implementations don’t have to adopt the new versions. But the new versions need to offer real benefits. The trouble with “fixing” the units of all currently published archetypes is that adopting them in order to really make use of the normative UCUM units would mean pain for implementers. But for what gain?

Notes on current practice regarding unit usage in HL7 laboratory messages

I include the following, simply because it tries to illustrate how units are currently handled in many typical data sharing environments.

  • Legacy, non-openEHR systems using 7-bit databases might use predefined table of units something like UCUM - more likely they would specify their own unit system - perhaps “deg” for values for "°C” in a system in a metric country or “deg” for values in “°F” in a system in the US. In many systems the units are often implicit and not stored with each value.
  • In Australia, diagnostic testing laboratories almost all send test result reports in HL7 v2 messages. Many of these use atomic fields for each observation. HL7 v2 uses an explicit field for transmitting a unit description. The Australian Standards specify ISO+ values to be used to name these units in messages. In practice, messages comply to various levels with these ISO standards. I think the use of these codes is similar in the US, New Zealand and a number of other countries, although I am less familiar with these.
  • Depending on the quantity being sent in a report, the ability to computationally deal with the unit is fraught with implementation issues. Some parameters such as temperatures or weights that have been modelled as such in archetypes can be validated. In many cases there is simply too much variability, either in the units that are allowable for a particular field, or for the variability in quality of the actual units sent.
  • Both of the above shortcomings lead to implementers wanting archetypes with little or no unit constraint on many fields. More often than not this is a result of lack of compliance infrastructure.
  • In the real world of extensive data sharing, highest common denominator ( or often lowest common denominator ) trumps standards and quality, and even safety.

regards,
eric

Eric Browne
eric.browne@montagesystems.com.au
ph 0414 925 845
skype: eric_browne

I’ve skimmed the replies on this thread, and I’m inclined to think everyone could be right. Problem is, they can’t all be right at the same time.

So… considering the issue from a global deployment perspective I had the folllowing idea:

  • in the archetype library, we should stick to proper versioning rules, as Diego, David and Sebastian have said, and correct the error and issue a new v2 archetype
  • => this way there are no surprises in the archetype library for software, or people, by the time we have forgotten about this issue
  • => the paths would stay the same to the various data fields, but the units in the tilt table field would be different, and will break anything that specifically relies on that
  • but we still may need a way of making adding the correction to the v1 version of the BP archetype, for some users, to enable their current queries and software based on the ‘v1’ id to keep working. This could be achieved by:
  • creating a specialisation of the v1 archetype that adds a new node as Sebastian proposed, or via Heath’s proposal
  • => this means that the deployment has to use a new archetype id for data production (i.e. in the CDR), but AQL queries using the old id should continue to work, assuming the query processor correctly finds child archetypes for a given archetype id- OR enable a modification in the template that has the effect of adding the required unit or element
  • => this means that all ids stay the same in data production and querying, but the CDR is being told a white lie, in that what it thinks is the BP archetype is not exactly what it is back in the CKM library

I haven’t tried to analyse the details here that far, but the general idea is to:

  • a) ensure the archetype library follows the versioning and change rules properly
  • b) inject an adjusting fix either as a specialisation, or further along in the processing chain - in a template (or even OPT…)

thoughts?

  • thomas

Eric,

nice summary of issues. If I can take the liberty of pulling out what I think are your key issues to worry about + recommendations. I bolded my own subset of those …

=> conclusion: we should have a PR to correct these issues so that the current specifications are at least clear, even if they still may be ‘wrong’ in some larger analysis. Eric, can you , with the relevant bullet points from this list? I think we could include changes to Release-1.0.3 to make these corrections.

hi Eric

Hi Thomas,

See my email from October 9 regarding how this has been resolved.

There is now an updated v1 Blood Pressure archetype with addition of the correct units for Tilt added to the node and a note that the previous units are no longer to be used. It has been republished as a non-breaking revision.

· It follows the versioning rules that we have firmly established for published archetypes.

· It means that new implementers can use the corrected v1 revision and we don’t have to create a v2 for a relatively trivial problem; existing vendor implementations can remain unchanged or they can choose to update the units if they please. The MD5 changes, but all paths etc are identical. A minimal disruption approach, if you like – thanks Heath.

We can consider other changes that might require a v2 in the future, at our leisure. There is no urgency at this point as the remaining changes that have been proposed are more along the lines of updating the archetype towards ‘more modern’ patterns for anatomical location etc. We don’t need to rush down this path as there will be little to gain, and probably quite a lot of confusion generated.

If we identify other breaking issues in the future, a v2 will be considered again, including the remaining ‘ideal pattern’ proposals. But for now, my advice is to leave the revised v1 as is.

Through this discussion we have identified is another strategy for governance and change management, that I hadn’t considered before. A good outcome from my POV.

Have I missed anything?

Thanks again for all of your contributions.

Regards

Heather

And what happens if a new implementation sends data to an old
implementation? Since the archetype identifier has not changed the receiver
will use its own archetype to validate the received data, and if it
includes the 'deg' unit it will just fail the validation. Breaking
revisions are not only about changing the archetype structure, but also
about generating a different set of possible instances.

surely the obvious approach is that the stored field contains the UCUM case-sensitive code, and that applications / services use UCUM tables to render whatever display form is asked for in a client call? (I realise openEHR archetypes are not doing this; they should be...)

Hence my earlier proposal…

Hi David,

That is clearly a revision change1.0->1.1 but is not a breaking change for data already carried within the system i.e queries for tilt using the degree symbol will still work.

This is is not inherently any different from the situation where we can add codes to an internal codelist, e.g mild/ moderate/severe/ => mild/moderate/severe/fatal

This is considered a non-breaking change since existing data is not invalidated but could cause exactly the same kind of potential mismatch between systems using different minor revisions of the same archetype.

Revision changes can only guarantee that existing data is unaffected but cannot ensure that mis-matches occur between disparate systems using different profiles on the same archetype. This can happen even with existing archetypes e.g the temperature archetype which allows variations of unit. In practice we need to use some form of templating or profiling to resolve these kind of potential variances in real systems and data exchanges.

The good thing in your scenario is that the recipient system would through a validation error, alerting the recipient that an unexpected unit was being sent.

I don’t think there is a problem here. We expect similar variance issues to arise in other circumstances.

Ian