# Archetype IDs starting with a numeric?

**URL:** https://discourse.openehr.org/t/archetype-ids-starting-with-a-numeric/1318
**Category:** Specifications
**Created:** [25 February 2021 08:30 UTC](https://discourse.openehr.org/t/archetype-ids-starting-with-a-numeric/1318 "2021-02-25T08:30:24Z")
**Posts on this page:** 12
**Page:** 1

<div class="post-metadata">

### Author: ![siljelb](https://discourse.openehr.org/user_avatar/discourse.openehr.org/siljelb/32/12_2.png) [@siljelb](https://discourse.openehr.org/u/siljelb)
#### Post date: [25 February 2021 08:30 UTC](https://discourse.openehr.org/t/archetype-ids-starting-with-a-numeric/1318/1 "2021-02-25T08:30:24Z")

</div>

Hi!

We’ve noticed a behaviour both in CKM and AD where we’re not allowed to assign an identifier (the **domain\_concept** part of the ARCHETYPE\_ID) with 0-9 as the first character. We can’t find where this is stated in the specs, and we don’t understand why this limitation exists.

Help please? 😊

---

<div class="post-metadata">

### Author: ![sebastian.garde](https://discourse.openehr.org/user_avatar/discourse.openehr.org/sebastian.garde/32/3070_2.png) [@sebastian.garde](https://discourse.openehr.org/u/sebastian.garde)
#### Post date: [25 February 2021 08:53 UTC](https://discourse.openehr.org/t/archetype-ids-starting-with-a-numeric/1318/2 "2021-02-25T08:53:10Z")

</div>

It is defined here: [Archetype Definition Language 1.4 (ADL1.4)](https://specifications.openehr.org/releases/AM/Release-2.2.0/ADL1.4.html#_symbols_4)  
See the regex for V\_ARCHETYPE\_ID which requires to start each part with a letter before a number can be used.

> ----------/\* V\_ARCHETYPE\_ID _/ ---------------------------------------------  
> [a-zA-Z][a-zA-Z0-9\_]+(-[a-zA-Z][a-zA-Z0-9\_]+){2}.[a-zA-Z][a-zA-Z0-9\_]+(-[azA-Z][a-zA-Z0-9\_]+)_.v[1-9][0-9]\*

---

<div class="post-metadata">

### Author: ![heather.leslie](https://discourse.openehr.org/user_avatar/discourse.openehr.org/heather.leslie/32/1980_2.png) [@heather.leslie](https://discourse.openehr.org/u/heather.leslie)
#### Post date: [25 February 2021 09:09 UTC](https://discourse.openehr.org/t/archetype-ids-starting-with-a-numeric/1318/3 "2021-02-25T09:09:27Z")

</div>

So is there a reason that the regex requires a letter first, or is it just arbitrary?

We have archetypes where the concept name starts with a number, usually well-known scores or scales. It would make sense that the id aligns with the concept name.

---

<div class="post-metadata">

### Author: ![sebastian.iancu](https://discourse.openehr.org/user_avatar/discourse.openehr.org/sebastian.iancu/32/31_2.png) [@sebastian.iancu](https://discourse.openehr.org/u/sebastian.iancu)
#### Post date: [25 February 2021 12:58 UTC](https://discourse.openehr.org/t/archetype-ids-starting-with-a-numeric/1318/4 "2021-02-25T12:58:54Z")

</div>

I don’t know personally the original thoughts and reasons, I guess this happend 15-20 years ago.  
What I can assume is that is related to the fact that **domain\_concept** is an _identifier_, and as programming good practice as well as (perhaps) the parser/lexer perspective, such identifiers does not start with numbers.

I guess the relevant section is in [Archetype Identification](https://specifications.openehr.org/releases/AM/latest/Identification.html#_concept_identifier), but it does not explicitly state my assumptions above. Perhaps @thomas.beale can give more hints / opinions?

---

<div class="post-metadata">

### Author: ![birger.haarbrandt](https://discourse.openehr.org/user_avatar/discourse.openehr.org/birger.haarbrandt/32/24_2.png) [@birger.haarbrandt](https://discourse.openehr.org/u/birger.haarbrandt)
#### Post date: [25 February 2021 13:03 UTC](https://discourse.openehr.org/t/archetype-ids-starting-with-a-numeric/1318/5 "2021-02-25T13:03:49Z")

</div>

The challenge is that within programming languages, often class names are not allowed to start with a number. For example, we use Templates to automatically generate classes. Theoretically we can convert these numbers to their word equivelant (9 becomes “Nine” etc.) but this might become awkward. No execuse to constrain clinicians with these technical details but might be the explanation for this.

---

<div class="post-metadata">

### Author: ![heather.leslie](https://discourse.openehr.org/user_avatar/discourse.openehr.org/heather.leslie/32/1980_2.png) [@heather.leslie](https://discourse.openehr.org/u/heather.leslie)
#### Post date: [26 February 2021 00:30 UTC](https://discourse.openehr.org/t/archetype-ids-starting-with-a-numeric/1318/6 "2021-02-26T00:30:08Z")

</div>

> [@sebastian.iancu](#):
>
> I guess the relevant section is in [Archetype Identification](https://specifications.openehr.org/releases/AM/latest/Identification.html#_concept_identifier), but it does not explicitly state my assumptions above.

Yes, we found this. Hence our question.

> [@birger.haarbrandt](#):
>
> No excuse to constrain clinicians with these technical details but might be the explanation for this.

We have two archetypes at present that are triggering the question - both scores/scales.

- The ‘4 ‘A’s Test (4AT)’. The formal name is the “4 'A’s test” but is commonly known as the ‘4AT’ the ID is currently OBSERVATION.four\_at.v0, which seems a little awkward and misaligned with the concept name - [Observation Archetype: 4AT [openEHR Clinical Knowledge Manager]](https://ckm.openehr.org/ckm/archetypes/1013.1.4304); and
- The '6 Item Cognitive Impairment Test (6CIT) ’ - ID is currently OBSERVATION.six\_cit.v0, commonly known at ‘6CIT’ with the same mismatch - [Observation Archetype: 6 Item Cognitive Impairment Test (6CIT) [openEHR Clinical Knowledge Manager]](https://ckm.openehr.org/ckm/archetypes/1013.1.3341)

It’s not deal-breaking, more just an ugly compromise 🤨, I suppose. We just wanted to be sure it was for a reason, not just someone never thought about it.

---

<div class="post-metadata">

### Author: ![thomas.beale](https://discourse.openehr.org/user_avatar/discourse.openehr.org/thomas.beale/32/35_2.png) [@thomas.beale](https://discourse.openehr.org/u/thomas.beale)
#### Post date: [26 February 2021 11:58 UTC](https://discourse.openehr.org/t/archetype-ids-starting-with-a-numeric/1318/7 "2021-02-26T11:58:30Z")

</div>

I don’t think there is a strong reason for it - we (in IT-land) just have this 50yo habit of not allowing identifiers to start with numbers, which I have replicated in the spec when I created the regexes for Archetype id. I can’t think off-hand of any specific reason not to allow numbers in this case. Where code generation occurs (as Birger mentioned), appropriate conversions can always be made to generate legal class or module names.

However, I don’t know how fast this could be changed in tools, which are likely to have copied the regexes from the specs. Maybe @pieterbos , @yampeku, @borut.fabjan, @sebastian.garde could have a think about that. We can relax the specs easily enough.

---

<div class="post-metadata">

### Author: ![Juha-Pekka.Tolvanen](https://discourse.openehr.org/user_avatar/discourse.openehr.org/juha-pekka.tolvanen/32/908_2.png) [@Juha-Pekka.Tolvanen](https://discourse.openehr.org/u/Juha-Pekka.Tolvanen)
#### Post date: [1 March 2021 12:18 UTC](https://discourse.openehr.org/t/archetype-ids-starting-with-a-numeric/1318/8 "2021-03-01T12:18:25Z")

</div>

This does not happen only openEHR, as other communities defining metamodels tend to follow the same rule, like AUTOSAR in automotive, or EAST-ADL for EE systems. The rule there has been to my best knowledge due to XML serialization as an XML element whose name starts with a number is illegal XML. The same is in modeling support done in some UML profile tools where classes implement it and class names should not start with number. So unfortunately, technology choices dictate the reality - which it obviously should not.

---

<div class="post-metadata">

### Author: ![ian.mcnicoll](https://discourse.openehr.org/user_avatar/discourse.openehr.org/ian.mcnicoll/32/4430_2.png) [@ian.mcnicoll](https://discourse.openehr.org/u/ian.mcnicoll)
#### Post date: [1 March 2021 12:35 UTC](https://discourse.openehr.org/t/archetype-ids-starting-with-a-numeric/1318/9 "2021-03-01T12:35:15Z")

</div>

Thanks - @heather.leslie - I think we maybe found our reason for not messing about with the current rules!! It is a little ugly but I think we can live with that if there are potential tech gotchas out there.

---

<div class="post-metadata">

### Author: ![thomas.beale](https://discourse.openehr.org/user_avatar/discourse.openehr.org/thomas.beale/32/35_2.png) [@thomas.beale](https://discourse.openehr.org/u/thomas.beale)
#### Post date: [1 March 2021 12:55 UTC](https://discourse.openehr.org/t/archetype-ids-starting-with-a-numeric/1318/10 "2021-03-01T12:55:15Z")

</div>

> [@Juha-Pekka.Tolvanen](#):
>
> The rule there has been to my best knowledge due to XML serialization as an XML element whose name starts with a number is illegal XML. The same is in modeling support done in some UML profile tools where classes implement it and class names should not start with number

This is true, but since serial formats are not our primary representation of anything, converters that generate out e.g. XML can always have rules added to synthesise legal names etc. Having said that, it puts the lexical ids of serial format entities out of sync with the original artefact (in ADL, for example), which will always create its own problems.

For sanity’s sake it is arguably better to stick with the less technically painful approach, and put up with a bit of cognitive annoyance in the ids.

---

<div class="post-metadata">

### Author: ![sebastian.garde](https://discourse.openehr.org/user_avatar/discourse.openehr.org/sebastian.garde/32/3070_2.png) [@sebastian.garde](https://discourse.openehr.org/u/sebastian.garde)
#### Post date: [1 March 2021 13:28 UTC](https://discourse.openehr.org/t/archetype-ids-starting-with-a-numeric/1318/11 "2021-03-01T13:28:04Z")

</div>

I agree with the core technical challenges Juha-Pekka and Birger refer to.  
Yes, you wsill be able to work around this in extra steps somehow but that is then also an extra source for errors.

In addition, _changing_ this at this stage has the potential to cause problems in every existing tool defining or managing archetypes including underlying parsers etc, or directly or indirectly consuming/using them including systems, code generators, transforms - various subtle places where it may fail downstream, in sometimes subtle ways.

---

<div class="post-metadata">

### Author: ![heather.leslie](https://discourse.openehr.org/user_avatar/discourse.openehr.org/heather.leslie/32/1980_2.png) [@heather.leslie](https://discourse.openehr.org/u/heather.leslie)
#### Post date: [3 March 2021 23:43 UTC](https://discourse.openehr.org/t/archetype-ids-starting-with-a-numeric/1318/12 "2021-03-03T23:43:18Z")

</div>

🤨 OK

Thanks for the responses. We’ll stick with ugly.
