# archetypes for genomic data

**URL:** https://discourse.openehr.org/t/archetypes-for-genomic-data/14788
**Category:** Clinical (archive)
**Created:** [15 July 2008 08:35 UTC](https://discourse.openehr.org/t/archetypes-for-genomic-data/14788 "2008-07-15T08:35:18Z")
**Posts on this page:** 11
**Page:** 1

<div class="post-metadata">

### Author: ![Koray\_Atalag](https://discourse.openehr.org/user_avatar/discourse.openehr.org/koray_atalag/32/70_2.png) [@Koray\_Atalag](https://discourse.openehr.org/u/Koray_Atalag)
#### Post date: [15 July 2008 08:35 UTC](https://discourse.openehr.org/t/archetypes-for-genomic-data/14788/1 "2008-07-15T08:35:18Z")

</div>

Hi Nicola,

It is great that someone from genomic society is involved with openEHR  
🙂 As an M.D. who has done (thesis not submitted) Ph.D. work in  
Molecular Biology & Genetics and worked on Genomic Databases (I am still  
the curator of Turkish Human Mutation Database - a former HUGO MDI  
initiative and now HGVS) my primary problem was to associate phenotypic  
(i.e. clinical) data with genomic data. It was 1997-1998 and openEHR did  
not exists! However I found out that the fundamental scientific problem  
was not sequencing a whole lot of genes but to handle current vast  
amount of data and extract useful information and discover knowledge by  
Informatics. I have just finished my Ph.D. on Information Systems and my  
thesis topic was related with openEHR archetypes...

The reason that I have given my background is that I am aware of your  
requirements and possible contribution...I have previously tried to  
raise some interest on this in both the openEHR foundation and the  
discussion lists. I was not successful unfortunately in part because  
that was not my primary concern at the time and perhaps in bigger part  
due to the non-existance of a "critical-mass" within this society in  
this area. But I know that some people exist who might also  
contribute...Perhaps your message might be a right trigger for some of  
us to establish an ad-hoc special interest group in genomics?

Your current proposal to represent genomic data with DV\_MULTIMEDIA is  
possibly the only way currently....BUT this is not acceptable in the  
long run as many of the inherent attributes and constraints of genomics  
should be integrated to Reference Model and for example a future  
"DV\_GENOMIC" type may further specialise into "DV\_NUCLEICACIDSEQUENCE"  
or "DV\_RNASEQUENCE" or "DV\_AMINACIDSEQUENCE".....Or the inherent openEHR  
terminology should contain "viral DNA", "bacterial DNA", "mitochondrial  
DNA" and so on.... Most importantly the "Codons" might be integrated in  
the Reference Model and "AUG" might be defined as the universal STOP  
codon for all genomic sequences??

Those are just immediate thoughts....I am sure much more can be set  
forth if we can get together and work.

Cheers,

Koray Atalag, MD, Ph.D.  
skype: atalagk

nicola onano wrote:

---

<div class="post-metadata">

### Author: ![thomas.beale](https://discourse.openehr.org/user_avatar/discourse.openehr.org/thomas.beale/32/35_2.png) [@thomas.beale](https://discourse.openehr.org/u/thomas.beale)
#### Post date: [15 July 2008 09:26 UTC](https://discourse.openehr.org/t/archetypes-for-genomic-data/14788/2 "2008-07-15T09:26:21Z")

</div>

Koray Atalag wrote:

> Your current proposal to represent genomic data with DV\_MULTIMEDIA is  
> possibly the only way currently....BUT this is not acceptable in the  
> long run as many of the inherent attributes and constraints of genomics  
> should be integrated to Reference Model and for example a future  
> "DV\_GENOMIC" type may further specialise into "DV\_NUCLEICACIDSEQUENCE"  
> or "DV\_RNASEQUENCE" or "DV\_AMINACIDSEQUENCE".....Or the inherent openEHR  
> terminology should contain "viral DNA", "bacterial DNA", "mitochondrial  
> DNA" and so on.... Most importantly the "Codons" might be integrated in  
> the Reference Model and "AUG" might be defined as the universal STOP  
> codon for all genomic sequences??

\*I would think that these kind of data types will be needed in the  
future, to allow transparent querying of protein & DNA structured data.  
The DNA / protein information might reside in a special erver, but why  
not make it an openEHR one with the required data types and so on?

- thomas beale

---

<div class="post-metadata">

### Author: ![ian.mcnicoll](https://discourse.openehr.org/user_avatar/discourse.openehr.org/ian.mcnicoll/32/4430_2.png) [@ian.mcnicoll](https://discourse.openehr.org/u/ian.mcnicoll)
#### Post date: [15 July 2008 09:31 UTC](https://discourse.openehr.org/t/archetypes-for-genomic-data/14788/3 "2008-07-15T09:31:11Z")

</div>

Hi Nicola,Koray,

Very interesting. It is perhaps worth thinking about this issue in 2 parts:

1. Developing archetypes for the application of genomic assessment in normal clinical practice e.g for myself as a GP with perhaps a DV\_URI to the source genomic data. For this category of user the genomics is essentially ‘signals’ data akin to raw ECG leads data and I am much more interested in the decison support input and ‘results’ that fall out of the genomic analysis.

Is it possible to create a generic ‘Genomic Assessment’ archetype, which can be further specialised for particular conditions? Some of the work I was doing on the Family History archetype started to intrude on this area.

1. A further look at how the genomic data itself can be represented in openEHR. Does this really need additions to the Reference model as Koray suggests or can much of this be archetyped using existing RM datatypes?

Ian

Dr Ian McNicoll  
office / fax +44(0)141 560 4657  
mobile +44 (0)775 209 7859  
skype ianmcnicoll

Consultant - Ocean Informatics [ian.mcnicoll@oceaninformatics.com](mailto:ian.mcnicoll@oceaninformatics.com)  
Consultant - IRIS GP Accounts

Member of BCS Primary Health Care Specialist Group – [www.phcsg.org](http://www.phcsg.org)

2008/7/15 Koray Atalag \<[atalagk@yahoo.com](mailto:atalagk@yahoo.com)\>:

---

<div class="post-metadata">

### Author: ![system](https://discourse.openehr.org/uploads/default/original/2X/f/f0a1dedb20c42747bddcafd6c7df9db5f34f003c.svg) [@system](https://discourse.openehr.org/u/system)
#### Post date: [16 July 2008 12:17 UTC](https://discourse.openehr.org/t/archetypes-for-genomic-data/14788/4 "2008-07-16T12:17:23Z")

</div>

Hi All

Good to see another genomic.

I am inclined to agree with Ian that the right place to start is with an EVALUATION and see what you come up with. Probably best to even start with a mind map. We need to choose a mindmap product that we all use so that we can all look at a collective set!

Cheers, Sam

Ian McNicoll wrote:

> **(attachments)**
>
> ![OceanInformaticsl.JPG](https://discourse.openehr.org/images/transparent.png)

---

<div class="post-metadata">

### Author: ![system](https://discourse.openehr.org/uploads/default/original/2X/f/f0a1dedb20c42747bddcafd6c7df9db5f34f003c.svg) [@system](https://discourse.openehr.org/u/system)
#### Post date: [16 July 2008 22:21 UTC](https://discourse.openehr.org/t/archetypes-for-genomic-data/14788/5 "2008-07-16T22:21:24Z")

</div>

Hi all,

I got a similar question when I was introducing openEHR and archetypes in Porto, Portugal. My answer (or guess =) was that the generic data types and data structure models could probably be reused to model genomic data but it could be useful to introduce some high level common “container” classes dedicated to genomic data.

Cheers,  
Rong

> **(attachments)**
>
> ![OceanInformaticsl.JPG](https://discourse.openehr.org/images/transparent.png)

---

<div class="post-metadata">

### Author: ![Thilo\_Schuler1](https://discourse.openehr.org/letter_avatar_proxy/v4/letter/t/a4c791/32.png) [@Thilo\_Schuler1](https://discourse.openehr.org/u/Thilo_Schuler1)
#### Post date: [17 July 2008 10:19 UTC](https://discourse.openehr.org/t/archetypes-for-genomic-data/14788/6 "2008-07-17T10:19:55Z")

</div>

Hi,

I agree with Ian and Sam that we definitely need a "genomic  
assessment" EVALUATION archetype and for some genetic disorders  
specialisations of it. It should probably include links (DV\_URI) to  
reference sources on the web and/or encapsulated data (DV\_MULTIMEDIA  
or DV\_PARSEABLE) to optionally store existing genomic information  
formats like MAGE ([http://www.mged.org/Workgroups/MAGE/mage.html](http://www.mged.org/Workgroups/MAGE/mage.html))  
directly in the openEHR instances for research purposes (e.g. in a  
university hospital).

The introduction of new datatypes into the RM is a sensitve issue. The  
RM was explicitely designed to be small and stable by using \*generic\*  
classes/types.  
Although I recognize the increasing importance of genomics/proteomics  
(including phenotype-genotype relations) for medicine at the beginning  
of the age of information-based or personalised medicine, I am  
currently of the opinion that we should not try to model everything in  
openEHR. This is for two main reasons:

1) openEHR excels at "putting the clinician in the driver seat"  
through archetypes. The vast majority of clincians (and IMO this will  
stay like this) is not interested in the the genomic data itself but  
their consequences/conculsions. As Ian said a genomic EVALUATION  
archetype would be needed. The archetype modeller/author should not be  
overwhelmed by a growing number of RM classes it is hard enough to  
properly use the existing ones...

2) By referencing to (DV\_URI) or inclusion of (DV\_MULTIMEDIA or  
DV\_PARSEABLE) we can make reuse the work, experience and tools (! -  
e.g. for analysis) of other groups and can focus on what openEHR is  
really good: let the clinicians decide what they want to record (and  
not many clinicians are interested in DNA-squences etc but they need  
to know e.g. whether a leukaemia is is Philadelphia Chromosome  
(abl-bcr fusion gen) positive because certain drugs work only in that  
instance).

Cheers, Thilo

---

<div class="post-metadata">

### Author: ![thomas.beale](https://discourse.openehr.org/user_avatar/discourse.openehr.org/thomas.beale/32/35_2.png) [@thomas.beale](https://discourse.openehr.org/u/thomas.beale)
#### Post date: [17 July 2008 10:39 UTC](https://discourse.openehr.org/t/archetypes-for-genomic-data/14788/7 "2008-07-17T10:39:23Z")

</div>

In principle I agree Thilo - use openEHR (as it is today) to record the  
'report' based on the 'data', but not necessarily the raw data (image,  
scan, genetic map etc). However - if we see openEHR as an engineering  
approach, there is no reason not to create new pieces of the reference  
model suited to specific purposes such as genomic/proteomic data  
structures, or similarly to say workflow representation (as we have  
considered inthe past). Then the archetypes for this part of the RM  
would define models of content in those spaces. Not that I am suggesting  
doing this right now, but the posibility is certainly technically  
avialable. After all, all such data has to be represented in some kind  
of model and technical framework - the openEHR approach may well offer  
advantages.

- thomas beale

Thilo Schuler wrote:

---

<div class="post-metadata">

### Author: ![ian.mcnicoll](https://discourse.openehr.org/user_avatar/discourse.openehr.org/ian.mcnicoll/32/4430_2.png) [@ian.mcnicoll](https://discourse.openehr.org/u/ian.mcnicoll)
#### Post date: [17 July 2008 10:53 UTC](https://discourse.openehr.org/t/archetypes-for-genomic-data/14788/8 "2008-07-17T10:53:10Z")

</div>

Hi Thilo,

I am in complete agreement with you here.

A couple of things:

1. Can you or anyone else point to/or provide some real-life clinical forms or other documents pertaining to a clinical assessment where genomcs plays a significant part? I tried to google for this without much success.

2. I agree re not further complicating the openEHR reference model but I can see an attraction in using the openEHR 2/3 layer modelling paradigm in other specialist domains such as genomics with, perhaps a different or extended RM which encapsulates some of the basic data structures of that domain which are likely to be rigid and invariant over time such as the suggestions already made.

As I write this, Thomas has just suggested something similar.

Ian

Dr Ian McNicoll  
office / fax +44(0)141 560 4657  
mobile +44 (0)775 209 7859  
skype ianmcnicoll

Consultant - Ocean Informatics [ian.mcnicoll@oceaninformatics.com](mailto:ian.mcnicoll@oceaninformatics.com)  
Consultant - IRIS GP Accounts

Member of BCS Primary Health Care Specialist Group – [www.phcsg.org](http://www.phcsg.org)

2008/7/17 Thilo Schuler \<[thilo.schuler@gmail.com](mailto:thilo.schuler@gmail.com)\>:

---

<div class="post-metadata">

### Author: ![Thilo\_Schuler1](https://discourse.openehr.org/letter_avatar_proxy/v4/letter/t/a4c791/32.png) [@Thilo\_Schuler1](https://discourse.openehr.org/u/Thilo_Schuler1)
#### Post date: [17 July 2008 12:25 UTC](https://discourse.openehr.org/t/archetypes-for-genomic-data/14788/9 "2008-07-17T12:25:02Z")

</div>

Hi Ian,

see inline

> Hi Thilo,
> 
> I am in complete agreement with you here.
> 
> A couple of things:
> 
> 1. Can you or anyone else point to/or provide some real-life clinical forms or other documents pertaining to a clinical assessment where genomcs plays a significant part? I tried to google for this without much success.

I googled two disorders on the top of my head…

As mentioned the Philadelphia Chromosome status is very important in certain leukaemias (can be treated with tyrosine kinase inhibitor) most prominently in CML. Here is a link regarding CML workup and treatment including genetic testing (FISH or PCR):

- [http://www.nccn.org/professionals/physician\_gls/PDF/cml.pdf](http://www.nccn.org/professionals/physician_gls/PDF/cml.pdf)

Factor-V-Leiden mutation is also very important to assess the thrombosis risk. Two more links:

- [http://www.sydpath.stvincents.com.au/tests/ThrombosisRisk.htm](http://www.sydpath.stvincents.com.au/tests/ThrombosisRisk.htm)
- [http://www.doh.wa.gov/hsqa/fsl/Documents/LQA\_Docs/Coag.pdf](http://www.doh.wa.gov/hsqa/fsl/Documents/LQA_Docs/Coag.pdf)

Both these examples are more or less monogenetic disorders, which are well understood with direct clinical relevance and therefore genetic testing is used in routine medicine.

> 1. I agree re not further complicating the openEHR reference model but I can see an attraction in using the openEHR 2/3 layer modelling paradigm in other specialist domains such as genomics with, perhaps a different or extended RM which encapsulates some of the basic data structures of that domain which are likely to be rigid and invariant over time such as the suggestions already made.
> 
> As I write this, Thomas has just suggested something similar.

True, and as Tom said the well desinged (in an engineering way) openEHR approach could be extended to genomics/proteomics to have a uniform environment. But my argument is still that openEHR should focus on the standardised but flexible recording oft “traditional” chart/record information (which turned out to be one of the hardest) and refer to other established specialised formats for raw data (ECG lead readings, microassay data…) via the generic datatypes (DV\_URI, DV\_PAREABLE, DV\_MULTIMEDIA).

Thilo

---

<div class="post-metadata">

### Author: ![Knut\_Bernstein](https://discourse.openehr.org/letter_avatar_proxy/v4/letter/k/7ba0ec/32.png) [@Knut\_Bernstein](https://discourse.openehr.org/u/Knut_Bernstein)
#### Post date: [18 July 2008 09:09 UTC](https://discourse.openehr.org/t/archetypes-for-genomic-data/14788/10 "2008-07-18T09:09:59Z")

</div>

Hi

My interest in the genomic archetypes is from the clinical perspective. I acknowledge that we need a way to represent the genomic information, sequence data etc. But we know that in the clinical setting this data is of little use unless they are accompanied by a clinical assessment. I.e. we need to be able to represent the consequences for the patient (and the family) of any genomic information.

This means that the reference to the genomic information may be rather simple, i.e. the name of a specific mutation, classification of a hereditary disease/variant, reference/link to a specific locus etc.

However what we need is a standardized way of linking genomic information to fenotypic information, i.e. diagnoses, symptoms and findings.

Furthermore we need to be able to express family relations in an explicit way.

There are some experiences in this area from the Danish registry for hereditary non-polyposis colorectal cancer ([www.hnpcc.dk](http://www.hnpcc.dk)) which has been working with data definitions and communication standards for this – among other in the context of the EU project Infobiomed. ([www.infobiomed.org/](http://www.infobiomed.org/))

Regarding the representation of family relations the GEDCOM is a well known format for representation and communication of genealogical data, but newer models and format is also available. There is an overview on [http://xml.coverpages.org/genealogy.html#gdmuml](http://xml.coverpages.org/genealogy.html#gdmuml)

Regards

Knut

---

<div class="post-metadata">

### Author: ![system](https://discourse.openehr.org/uploads/default/original/2X/f/f0a1dedb20c42747bddcafd6c7df9db5f34f003c.svg) [@system](https://discourse.openehr.org/u/system)
#### Post date: [20 July 2008 02:52 UTC](https://discourse.openehr.org/t/archetypes-for-genomic-data/14788/11 "2008-07-20T02:52:43Z")

</div>

Hi Knut

At present we have the ability in openEHR to describe people other than the subject of care, both in terms of their relationship (father/mother etc) and also by an ID, name or other identifiers. I believe that this is all we should do within the EHR. The aggregation of this data from a set of EHRs is what will lead to the information which is of value. It may be that keeping parents Genomic data in the EHR is of value as this may help untangle what is going on genetically when people have obvious problems - although the privacy implications are great as non-biological parents will be detected as will spontaneous mutations.

We have had an archetype as a specialisation of problem for genetic condition. I attach it here…  
Let me know what you think of it clinically.

Cheers, Sam

Knut Bernstein wrote:

> **(attachments)**
>
> ![OceanInformaticsl.JPG](https://discourse.openehr.org/images/transparent.png)  
> [openEHR-EHR-EVALUATION.problem-genetic.v1.adl](https://discourse.openehr.org/404) (8.72 KB)
