Archetype documentation using XML + XSLT

ummmmmm.......where to even start....oh yes how about...

You left out the fact that XML cannot semantically reproduce an object
model.

Now wrt ADL.....

ADL does semantically describe the AOM.

Cheers,
Tim

The crux might be in the text in red.

Everything can be made to work.
A floating boat made out of match stick. Yes it is possible,
but not the bet engineering choice.
Perhaps XML is not the best choice in all circumstances.

What about ASN.1 isn’t that a good alternative?
How about the tools?
Not as ubiquitous as those for XML.

Gerard

There are many better ways to represent data for large-scale

deployments than XML (even the dADL syntax from ADL does 100% better in space,

and represents all object-oriented constructs unambiguously).

You have got to be kidding me on this one.

Having done XML messaging in very large retail systems (major
supermarket chains in the EU & US), mobile phone systems, home
office/criminal justice system & now the CFH…you simply have got to
be kidding.

ummmmmm…where to even start…oh yes how about…

– –
Gerard Freriks, MD
Huigsloterdijk 378
2158 LR Buitenkaag
The Netherlands

T: +31 252544896
M: +31 620347088
E: gfrer@luna.nl

Those who would give up essential Liberty, to purchase a little temporary
Safety, deserve neither Liberty nor Safety. Benjamin Franklin 11 Nov 1755

Adam Flinton wrote:

  

Other limitations on using XML - it's a no-show for enterprise scale databases
or information processing. All that wasted space starts to count when you have
to buy two £20,000 high availability RAID disk arrays instead of one....and plus
the bandwidth wastage when there are millions of messages rather than just a
few. Yes, binary compression helps, but it just shifts disk and bandwidth loss
to the CPU. There are many better ways to represent data for large-scale
deployments than XML (even the dADL syntax from ADL does 100% better in space,
and represents all object-oriented constructs unambiguously).
  

You have got to be kidding me on this one.

no kidding. You just have to do the maths. 2 years ago I was in a CfH
meeting where it was made clear that the volumetrics on the projected
number x size of HL7v3 prescription (and related) messages was going to
blow out the budget for telecomms expenditure by over 50%. Database
volume estimates made by Oracle for the Spine on the basis of whole of
UK, HL7 XML messages were shown to be simply uneconomic. In our own EHR
product we have had to resort to various kinds of compression, which
impact on performance, and we have to have larger disc arrays than would
be needed if the data were represented in a more efficient way. Our
engineers are currently looking at replacing XML altogether in the
persistence layer.

Of course organisations that have unlimited budgets may not notice this,
but for everyone else it's important.

Having done XML messaging in very large retail systems (major
supermarket chains in the EU & US), mobile phone systems, home
office/criminal justice system & now the CFH....you simply have got to
be kidding.
  

it's not the size of the organisation that matters, it's the amount of
data generated and rate of data generation. The more data (and users),
the more disk space / bandwidth / CPU (if using
compression/decomprssion) you need. Fact of life - if your data gets big
enough, you will hit a wall just using XML.

ummmmmm.......where to even start....oh yes how about...

"XML is a very commonly used standard with thousands of tools in
existence from routing engines, to processing engines to parsers to
database layers ...."

or maybe

"XML is THE standard in enterprise level messaging systems with
standards such as SOAP, EbXML, OAGI BODS etc.etc."

or maybe :

"XML integrates easily with existing web infrastructures by use of such
mechanisms as AJAX, Rest JSON etc".

Now wrt ADL.....

Adam

Although the above are just marketing statements, the standards do
indeed exist - how else would industry even start to cope with XML? My
point is that it is so horribly ineffcient that that there is a general
need for solutions (and these are emerging) to the concomitant costs of
very large amounts of data. That's why all the binary XML work and work
on alternative representations - once you get over a certain size, you
have to use something else, or a serious compression approach.

- thomas

BTW you only need one good open source parser for any language on each
platform. I'm not quite sure of the value of having 20 competing ones,
all with different sets of bugs, maintainers and conformance statements....

This brings up a good point. ASN.1 (much like XML) is a communications protocol.

These work well for flat message transactions where there is no need to be concerned with temporal or spatial context.

This is not the case with healthcare information as a whole.
There are exceptions such as a single order sent in a message format.
This is (to my understanding) how HL7 v2.x grew out of EDI standards.

On ASN.1 (as well as XSD) from an excellent tutorial on the subject at:
http://asn1.elibel.tm.fr/introduction/index.htm

"ASN.1 definition can be contrasted to the concept in ABNF of "valid syntax", or in XSD of a "valid document",
where the focus is entirely on what are valid encodings of data, without concern with any meaning that might be
attached to such encodings. That is, without any of the necessary semantic linkages."

In the 1980's I created some and worked with many EDI message applications. These approaches work great for
asynchronous, context-free DATA exchange. Thomas made a good point in file size as well. It wasn't only a concern 20 years ago.
It remains so today and is a primary reason that EDI (esp. X12 in the US) remains a major message exchange format.
There are many references available that describe EDI msgs. being 10% of the equivalent XML msgs. This adds up very
quickly in time and space when you are talking about MILLIONS of transactions or BILLIONS of records.

But, IMHO, the bottom line still remains the ability of ADL to represent information exchange as archetypes in a compact,
semantically accurate message. If you study the temporal and spatial aspects of clinical information in the context of the
idea that this information should live many years beyond the life of the patient in order to inform medicine for generations,
then you can see where a little inconvenience now can create systems that are very valuable far into the future.

Thanks for reading my rant. :slight_smile:

Cheers,
Tim

Tim Cook wrote:

ummmmmm.......where to even start....oh yes how about...
    
You left out the fact that XML cannot semantically reproduce an object
model.

XML is a form of text markup. It is more analogous to the English
language and can be used to express anything from a Shakespearian sonnet
to a death sentence to this mail.

Any of them could be marked up in XHTML which is a dialect of XML.

Now wrt ADL.....
    
ADL does semantically describe the AOM.
  
No reason why XML could not.

It can suffice for anything from a webform (e.g. XForms) to a vector
graphic (e.g. SVG) to an object model to formatted text (e.g. XHTML) to
an office file (e.g. ODF or OOXML) to a process such as XSLT.

Adam

Gerard Freriks wrote:

The crux might be in the text in red.

Everything can be made to work.
A floating boat made out of match stick. Yes it is possible,

Gee they'll be making boats & bridges out of iron & steel next.

but not the bet engineering choice.
Perhaps XML is not the best choice in all circumstances.

It is the whitworth screw thread of IT.

What about ASN.1 isn't that a good alternative?
How about the tools?
Not as ubiquitous as those for XML.

It is the difference between artisan & industrial.

You can make your own bolts & your own nuts or you can go down a shop &
buy a standard bolts & standard nuts where each may come from different
manufacturers on different continents & yet the nut will fit the bolt.

Adam

Thomas Beale wrote:

Adam Flinton wrote:
  

Other limitations on using XML - it's a no-show for enterprise scale databases
or information processing. All that wasted space starts to count when you have
to buy two £20,000 high availability RAID disk arrays instead of one....and plus
the bandwidth wastage when there are millions of messages rather than just a
few. Yes, binary compression helps, but it just shifts disk and bandwidth loss
to the CPU. There are many better ways to represent data for large-scale
deployments than XML (even the dADL syntax from ADL does 100% better in space,
and represents all object-oriented constructs unambiguously).
  

You have got to be kidding me on this one.

no kidding. You just have to do the maths. 2 years ago I was in a CfH
meeting where it was made clear that the volumetrics on the projected
number x size of HL7v3 prescription (and related) messages was going to
blow out the budget for telecomms expenditure by over 50%. Database
volume estimates made by Oracle for the Spine on the basis of whole of
UK, HL7 XML messages were shown to be simply uneconomic. In our own EHR
product we have had to resort to various kinds of compression, which
impact on performance, and we have to have larger disc arrays than would
be needed if the data were represented in a more efficient way. Our
engineers are currently looking at replacing XML altogether in the
persistence layer.

You are mistaking a format for a design. I agree entirely that HL7 is
vastly inflated. However that is in part mostly down to the huge levels
of duplicate information. e.g. simply to send a "hello world" message in
HL7 you need an astonishing array of elements/objects etc.

For a laugh create a blank CDA document carrying the text "hello world"
& see how large it is.

But that is nothing to do with XML per se as I could create one as small as

<a b="hello world"/>

What is more you are being a little disingenuous on this as when I
proposed using a value attribute to hold element values rather than an
element's text child you were not very keen despite the fact that doing
so shrunk the file sizes by 1/3rd & lead to greater consistency in
accessing such values.

There are a variety of design patterns on XML such as the use of
attributes, the careful layout wrt order of elements etc which can
reduce file sizes and increase speed of data access. i.e. if you want
simply some routing information from a message so as to direct it to the
right place (& that's all you want), that information goes as one of the
first structures (e.g. an attribute on the root element) & you use a
streaming parser such as Sax.

Wrt file sizes per se though you're doomed as it flies in the face of
progress e.g. medical imaging used to be a 2D Xray etc. Now it's a 3d
multi-slice cat scan in vast detail.

Wrt ETP in particular.....I held views on that which stemmed from my
time in "big retail" wrt the only "medical" part of the entire thing was
the initial choice of the clinician wrt what drug. After that it was a
std stock control/management issue & frankly I used to try & hide when
clinicians/clinical IT types were there telling the likes of the large
supermarkets how to do stock control.

Of course organisations that have unlimited budgets may not notice this,
but for everyone else it's important.
  

Having done XML messaging in very large retail systems (major
supermarket chains in the EU & US), mobile phone systems, home
office/criminal justice system & now the CFH....you simply have got to
be kidding.
  

it's not the size of the organisation that matters, it's the amount of
data generated and rate of data generation. The more data (and users),
the more disk space / bandwidth / CPU (if using
compression/decomprssion) you need. Fact of life - if your data gets big
enough, you will hit a wall just using XML.

Fact of life - if your data gets big enough, you will hit a wall period.

2 quick examples:

A) Medical imaging - how will the current network cope with CAT scans
etc being transmitted let alone stored? But hey these are binary files
not XML so from what you say they'll be no problem

B) Codesets. I created a system called TRUD (
http://www.google.co.uk/search?hl=en&q=CFH+TRUD&btnG=Google+Search&meta=
) whereby we distribute our codesets.

The management is done via XML (i.e. you register & have to have a soap
server capable of receiving update notifications. You then have various
stages which you as a user pass through

> New update available & it's here URL
< OK Downloading now
<OK Downloaded
< OK Processing
< OK Processed
< OK ready to go

etc. until all subscribers are ready to go at which point:

> Go live on datetime
< OK will do.

We are talking about large files & we had to set up a completely
different network system of ftp servers etc to cope with the download
demands

So what would you suggest? Shrink Snomed? Distribute each update to
every doctor's surgery? or.....produce some central servers who (using
verbose etc but std) .XML serve up the valuesets etc on demand?

So to recap...
A) HL7 is verbose. Don't confuse that with XML per se.
B) Large date volumes are a fact of life wrt medical IT. Get used to it
& plan how to mitigate it (e.g. wrt Trud/snomed/terminologies per se).

ummmmmm.......where to even start....oh yes how about...

"XML is a very commonly used standard with thousands of tools in
existence from routing engines, to processing engines to parsers to
database layers ...."

or maybe

"XML is THE standard in enterprise level messaging systems with
standards such as SOAP, EbXML, OAGI BODS etc.etc."

or maybe :

"XML integrates easily with existing web infrastructures by use of such
mechanisms as AJAX, Rest JSON etc".

Now wrt ADL.....

Adam

Although the above are just marketing statements, the standards do
indeed exist - how else would industry even start to cope with XML? My
point is that it is so horribly ineffcient that that there is a general
need for solutions (and these are emerging) to the concomitant costs of
very large amounts of data. That's why all the binary XML work and work
on alternative representations - once you get over a certain size, you
have to use something else, or a serious compression approach.

Zip works extremely well due to to the repetitive nature of XML (esp if
you tune the window size). e.g. a repeated string of
<AVeryLongElementName can be compressed to a few bytes in the zip index.

Either way the file size is actually down to a conjunction of the amount
of data to be sent and the design of the document type which is required
to hold that data.

e.g. at one stage we looked at various means of slimming down the
messages & the top one wrt HL7 was to have a templated instance which
contained that information/structures which would not change from
message to message & then an "on the wire" message which only contained
that information/structure which would/could change. Upon arrival the
sink system could either use that message as is or if it required a
"complete" HL7 message it could stitch the 2 together & then process the
"whole" message.

This resulted in massive message size decreases but.......no one ever
said that implementer's concerns were any concern of HL7.....so we have
instead pushed the new XML ITS (XML ITS R2) which includes folding etc
which includes many of these ideas.

However the fact remains that if you have 10 mb of data to send then
neither ADL nor XML will reduce that.

Wrt persistent storage Ronaold Bourret & I wrote the first usable XML
<JavaObjects>SQL engine called XML-DBMS back in Y2K

http://www.rpbourret.com/xmldbms/

It is bi-directional (you can create SQL structures to store existing
XML structures or create XML from existing SQL structures) & formed the
basis for the DB2 & SQL server XML storage engines & influenced the
Oracle one.

http://sourceforge.net/projects/xmldbms/

You will notice I am still a project admin though I have not done much
with it for a while now.

Ron went on to write the IBM RedBook on the issue:

http://www.redbooks.ibm.com/abstracts/sg246994.html

It is worth a read.

- thomas

BTW you only need one good open source parser for any language on each
platform. I'm not quite sure of the value of having 20 competing ones,
all with different sets of bugs, maintainers and conformance statements....

Which is why we use multiple schema checking engines in our schema checker.

Adam

Adam Flinton wrote:

Thomas Beale wrote:

Adam Flinton wrote:


You are mistaking a format for a design. I agree entirely that HL7 is
vastly inflated. However that is in part mostly down to the huge levels
of duplicate information. e.g. simply to send a "hello world" message in
HL7 you need an astonishing array of elements/objects etc.

It is true that a large part of the problem is to do with HL7v3 messages, but I was not trying to pick on HL7. We have enough trouble in openEHR-land keeping the size down - small often used object structures create terrible bloat in XML - in particular, the kind of bloat where the byte count of the tag names is many times greater than the data carried (e.g. Interval in openEHR is a very pedestrian OO type; in XML, its booleans and numeric data values are dwarfed by the huge number of tags).

For a laugh create a blank CDA document carrying the text "hello world"
& see how large it is.

But that is nothing to do with XML per se as I could create one as small as

I would differ on this. It is easy to find thousands of examples where the orthodox XML (let’s stick to .xsd based XML) is many times larger than it should be. My claim is that with XML you have to go around fixing it all the time - this is not an academic debate - it is right at the heart of making more efficient CDAs, HL7 messages, openEHR data, CEN extracts and other artefacts that get or may be in the future created in truly vast numbers. Kinds of fixes regularly employed include:

  • substitute an efficient syntax for a bloated, repetitive bit of XML. E.g. the new proposal for the openEHR .xsds replaces the standard XML representation of Interval with the XMI one (which is like in UML, e.g. “0..1”); this saves many many bytes each time a cardinality / existence / occurrences constraint is mentioned. Another example is representing coded terms either as XML-ised objects or in an efficient mini-syntax. This approach can save significant amounts of data, but requires the use of other parsers / add-ons. These are generally simple, but it is still an add-on, and something else to maintain.

  • use XML attributes instead of elements. This is generally unsystematic but it does reduce space. Changing this also stops Xpath queries from being portable, because aaa/bbb = ‘whatever’ changes to aaa[@bbb = ‘whatever’]

  • do other more complex substitutions on tag names to shorten them, according to various Huffman coding variants

  • use pre- and post-compression

  • go to binary…
    If XML was directly suitable for enterprise / large-scale persistence solutions, none of this would be needed.

<a b="hello world"/>

What is more you are being a little disingenuous on this as when I
proposed using a value attribute to hold element values rather than an
element's text child you were not very keen despite the fact that doing
so shrunk the file sizes by 1/3rd & lead to greater consistency in
accessing such values.

Originally I was completely against this; if I was to stick to any theory of ‘proper’ use, I still would be, because if you have an ad hoc assignment of OO types and attirbutes to XML elements and attributes, you can’t write simple software to do the conversion. Well ok, with more modern tools some of this is taken care of by code generators. What this means is that there is no longer any solid semantics behind the use of element v attribute - why not put everyhting in attributes? But as I note above - your Xpaths are then all dead. So you hve to decide. In ADL, the paths (which are Xpath compatible) remain valid regardless.

There are a variety of design patterns on XML such as the use of
attributes, the careful layout wrt order of elements etc which can
reduce file sizes and increase speed of data access. i.e. if you want
simply some routing information from a message so as to direct it to the
right place (& that's all you want), that information goes as one of the
first structures (e.g. an attribute on the root element) & you use a
streaming parser such as Sax.

sure - what you are saying here is - well designed XML is better than badly designed XML. Clearly this is true. But the reality is that even well-designed XML will eventually hit a wall in terms of size compared to what is available with other approaches - it might get you another 30%, but then it starts costing too much. Using it for enterprise persistence solutions really does not make any sense at all. We may well all end up using EXI XML in the future, but in a way, why bother? It doesn’t fix the semantic limitations of XML. Much better to treat XML as an output syntax on the import/export boundaries of systems. Now, if you are talking about messages, you may still have to use it, but I suspect it will become the option of last resort between any 2 given communicating parties with high-volumes and transactions rates; if two systems can agree on using any one of numerous other more efficient (and therefore cheaper) methods, why wouldn’t they?

Wrt file sizes per se though you're doomed as it flies in the face of
progress e.g. medical imaging used to be a 2D Xray etc. Now it's a 3d
multi-slice cat scan in vast detail.

not really. Just do the maths. Two points:

  • the amount of data put into EHRs in image format is low compared to the total amount of data, represented in XML because the high-definition images nearly always stay on PACS, and reports (and maybe thumbnails) are put in EHRs
  • the number of images created over a large population of patients, over a long time is not necessarily that large - some patients never have any imaging; for many that have routine Xrays, the Xray image is used to write a report, and would server almost no purpose being in the EHR
    So for many - probably most - health care consumers, reducing the XML size will indeed be signficant per EHR. But this is not what matters - it matters to the solution vendors and the institutional customer who is going to pay for more disc, CPU, and networking.
Fact of life - if your data gets big enough, you will hit a wall period.

that is definitely true. But XML makes you hit the wall in persistence much sooner than you need to - and slows you down in the process!

2 quick examples:

A) Medical imaging - how will the current network cope with CAT scans
etc being transmitted let alone stored? But hey these are binary files
not XML so from what you say they'll be no problem

they do pose a problem But large binary images are mostly used within high BW hospital LANs for radiologists to write reports. There are some special networking systems that are put in place to allow them to travel to other locations where consultants do remote report writing. Images are not the main proportion of the data on a large repository of EHRs. They also don’t tend to get repeatedly transmitted around the place and to numerous users. Every nurse looks at a part of the EHR, repeatedly, but they never look at images.

B) Codesets. I created a system called TRUD  (
[http://www.google.co.uk/search?hl=en&q=CFH+TRUD&btnG=Google+Search&meta=](http://www.google.co.uk/search?hl=en&q=CFH+TRUD&btnG=Google+Search&meta=)
) whereby we distribute our codesets.

...

So what would you suggest? Shrink Snomed? Distribute each update to
every doctor's surgery? or.....produce some central servers who (using
verbose etc but std) .XML serve up the valuesets etc on demand?


It all depends on whether doing this in XML is currently posing a problem. If it isn’t then, no - the ‘wall’ has not been hit. But Snomed is reference information and not distributed that frequently. And although Snomed is large, it is nowhere near the size of an EHR repository containing 5,000,000 patient records, or 20,000 prescription messages per week. As I say, it is all in the figures.

  • thomas beale

Thomas Beale wrote:

... In our own EHR
product we have had to resort to various kinds of compression, which
impact on performance, and we have to have larger disc arrays than would
be needed if the data were represented in a more efficient way. Our
engineers are currently looking at replacing XML altogether in the
persistence layer.

And Adam Flinton replied:

You are mistaking a format for a design. I agree entirely that HL7 is
vastly inflated.

Adam, I think you've misunderstood Thomas's point about "our own EHR
product". It is not HL7. It's an XML serialisation of openEHR data. Thomas
was illustrating his contention about XML per se.

- Peter

Peter Gummer wrote:

Thomas Beale wrote:
  

... In our own EHR
product we have had to resort to various kinds of compression, which
impact on performance, and we have to have larger disc arrays than would
be needed if the data were represented in a more efficient way. Our
engineers are currently looking at replacing XML altogether in the
persistence layer.
    
And Adam Flinton replied:

You are mistaking a format for a design. I agree entirely that HL7 is
vastly inflated.
    
Adam, I think you've misunderstood Thomas's point about "our own EHR
product". It is not HL7. It's an XML serialisation of openEHR data. Thomas
was illustrating his contention about XML per se.

Indeed however there are ways of persisting a model & they require at
the end of the day a recognizable document design/format.

I have already noted how using text children of an element to use a
value vs a std "value" attribute in the archetype xml inflates the file
sizes.

<A>
Some value
</A>

&
<A value=""Some value"/>

Are both persisting/serializing the same data.

Adam

Adam Flinton wrote:

I have already noted how using text children of an element to use a value
vs a std "value" attribute in the archetype xml inflates the file sizes.

So you aren't convinced by Thomas's objection that putting the values in XML
attributes would make a mess of the XPath?

- Peter

Peter Gummer wrote:

Adam Flinton wrote:

I have already noted how using text children of an element to use a value
vs a std "value" attribute in the archetype xml inflates the file sizes.
    
So you aren't convinced by Thomas's objection that putting the values in XML
attributes would make a mess of the XPath?
  
No as the existing paths aren't xpaths at all e.g..

<Rule path="/items[at0001]/items[at0004]" max="0" />

let alone something like

<excludedValues>local::at0013</excludedValues>

It's simply a matter of parsing the file into an AOM.

Adam

Adam Flinton wrote:

Peter Gummer wrote:
  

Adam Flinton wrote:

I have already noted how using text children of an element to use a value
vs a std "value" attribute in the archetype xml inflates the file sizes.
    

So you aren't convinced by Thomas's objection that putting the values in XML
attributes would make a mess of the XPath?
  
No as the existing paths aren't xpaths at all e.g..

<Rule path="/items[at0001]/items[at0004]" max="0" />

let alone something like

<excludedValues>local::at0013</excludedValues>

It's simply a matter of parsing the file into an AOM.

Adam,

I was not talking about openEHR paths, I was talking about any Xpaths.
If you define an Xpath based on a given model entity being an XML
attribute, it will be different from the path to the same thing based on
the model entity being an XML element. If that path was shared, e.g. in
a tool, or published for some reason, and you change your XML
representation for efficiency reasons, your published Xpath doesn't work
any more.

Separately, in openEHR, we use Xpath-compatible paths. A path like

    /items[at0001]/items[at0004]

is machine converted to the Xpath

    /items[@archetype_node_id = "at0001"]/items[@archetype_node_id =
"at0004"]

Not hard to guess why we do that;-)

Portable paths are the basis of portable queries; portable queries are
the basis of decision support having any hope of being able to define
queries once and for all into a standardised health record rather than
havig to have queries written not just for each target system type, but
for each installation, if database schemas differ (and they do in the
NHS environment).

- thomas beale

Hi Adam

It will be simple to transform the template statement to a readable form in HTML, but at present the HTML is of the operational template - that is with all the archetype and template constraints in one statement. It is this that will have to be the basis of the transform. It will take a have this so everyone can do it consistently but it should be possible soon.

Cheers, Sam

Adam Flinton wrote:

(attachments)

OceanCsmall.png

Hi All

One senses the different backgrounds of the players here - the more operational the focus the more useful XML is but in the design and modelling world it is of no use at all. I think almost everyone is delighted with what can be done with XML these days - the more it is used the more that is possible on the web. But ASCII was the same when it came out and really OWL as RDF is not very easy to deal with even with all the wonderful tools (because you have to have the model right).

I have been working with Tom long enough to understand the issue when you are formally describing the set of possible constraints that might be necessary in any given model. The UML and XML serialisation come later in this process.

So we all need to accept that value of both approaches - if we are serialising something formal that does the job and can reload and use it then that is fine. I am pushing for the move to documentation via XML archetypes and XSLT because I believe that the power of that approach and its cross platform utility leaves everything else for dead. But, I recognise that it is not suitable for use at design time although it can serialise the UML (sometimes)

You said you have an XSLT for the archetypes - would you like to share it with the community. I think we might see things move quite quickly if we get this out there.

Cheers Sam

Adam Flinton wrote:

(attachments)

OceanCsmall.png

Adam,

Indeed however there are ways of persisting a model & they require at
the end of the day a recognizable document design/format.

I have already noted how using text children of an element to use a
value vs a std "value" attribute in the archetype xml inflates the file
sizes.

<A>
Some value
</A>

&
<A value=""Some value"/>

Are both persisting/serializing the same data.

Adam

Your example here is contrary to your statement, example 1, 17 characters,
example 2, 23 characters. Not sure how you get a 1/3 size reduction.

Heath

Adam Flinton wrote:

I have already noted how using text children of an element to use a
value vs a std "value" attribute in the archetype xml inflates the file
sizes.

<A>
Some value
</A>

&
<A value=""Some value"/>

And Heath Frankel replied:

Your example here is contrary to your statement, example 1, 17 characters,
example 2, 23 characters. Not sure how you get a 1/3 size reduction.

<WellHeathIfYourTagsAreReallyReallyLongAndYourDataValuesAreReallyReallyShortThenYouWouldGetAtLeastAOneThirdSizeReduction>
ok?
</WellHeathIfYourTagsAreReallyReallyLongAndYourDataValuesAreReallyReallyShortThenYouWouldGetAtLeastAOneThirdSizeReduction>

<WellHeathIfYourTagsAreReallyReallyLongAndYourDataValuesAreReallyReallyShortThenYouWouldGetAtLeastAOneThirdSizeReduction
value="ok?"/>

- Peter

With a compression algorithm, the difference may be negligible.
Heath's first example is smaller than the second when zipped due to the
repeated strings.

Of course, the result may be quite different for a larger sample where
"value=" is repeated.

regards,
Colin

My reply was really just an attempt at humour, Colin. Even with really long
tag names, you don't get a 1:3 ratio; not even 1:2, I think.

And of course OPENEHR_ELEMENT_NAMES_ARE_NOT_REALLY_LONG_ANYWAY :wink:

- Peter

Sam Heard wrote:

Hi All

One senses the different backgrounds of the players here - the more
operational the focus the more useful XML is but in the design and
modelling world it is of no use at all. I think almost everyone is
delighted with what can be done with XML these days - the more it is
used the more that is possible on the web. But ASCII was the same when
it came out and really OWL as RDF is not very easy to deal with even
with all the wonderful tools (because you have to have the model right).

I have been working with Tom long enough to understand the issue when
you are formally describing the set of possible constraints that might
be necessary in any given model. The UML and XML serialisation come
later in this process.

So we all need to accept that value of both approaches - if we are
serialising something formal that does the job and can reload and use
it then that is fine. I am pushing for the move to documentation via
XML archetypes and XSLT because I believe that the power of that
approach and its cross platform utility leaves everything else for
dead. But, I recognise that it is not suitable for use at design time
although it can serialise the UML (sometimes)

You said you have an XSLT for the archetypes - would you like to share
it with the community. I think we might see things move quite quickly
if we get this out there.

Cheers Sam

I would love to.

I will have to get clearance from Ravi, Richard & John but then
yup...the sooner the better AFAICS.

Adam