# SynPuf: syntetic data (into openEHR)

**URL:** https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802
**Category:** Implementation
**Created:** [4 July 2022 14:21 UTC](https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802 "2022-07-04T14:21:39Z")
**Posts on this page:** 14
**Page:** 1

<div class="post-metadata">

### Author: ![yampeku](https://discourse.openehr.org/user_avatar/discourse.openehr.org/yampeku/32/25_2.png) [@yampeku](https://discourse.openehr.org/u/yampeku)
#### Post date: [4 July 2022 14:21 UTC](https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802/1 "2022-07-04T14:21:39Z")

</div>

I’m thinking in generating a way of load existing SynPuf data into openEHR data repository (probably via openEHR REST API). I have several questions to the community in order to make it more useful/more easily available. First a little summary:  
DE-SynPuf is a realistic looking syntetic data set with information from persons, claims, and prescriptions. SynPuf is organized in a way that you can get a percentage of patients from the total you want loaded into the system (each batch contains a set of patients with all their inpatient, outpatient, and carrier claims, and drug prescription information). SynPuf is used as the go-to dataset in most standards to be able to have a quick way of benchmarking/querying a given system. More info here

> **[CMS 2008-2010 Data Entrepreneurs’ Synthetic Public Use File (DE-SynPUF) | CMS](https://www.cms.gov/Research-Statistics-Data-and-Systems/Downloadable-Public-Use-Files/SynPUFs/DE_Syn_PUF)**
>
> CMS 2008-2010 Data Entrepreneurs’ Synthetic Public Use File (DE-SynPUF) The DE-SynPUF was created with the goal of providing a realistic set of claims data in the public domain while providing the very highest degree of protection to the Medicare...

Now the questions:

- Do we have a way of bulk-loading openEHR data into any of the currently available openEHR CDR?
- Can we create patients by using simplied json format?
- Demographic data storage is probably a big discussion, so for the moment a poll:  
**How should be demographic information be stored in this kind of use case?**

_Poll ([view on site](https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802/1))_

While 3rd option is probably the most correct, I’m not sure of current support for DEMOGRAPHIC model in data repositories. Second one should be easy but not really that “normalized”. First one will probably add quite a bit of data into the ehr\_status. Last one can probably be reused from other available SynPuf → FHIR generation.

---

<div class="post-metadata">

### Author: ![Seref](https://discourse.openehr.org/user_avatar/discourse.openehr.org/seref/32/13_2.png) [@Seref](https://discourse.openehr.org/u/Seref)
#### Post date: [6 July 2022 07:25 UTC](https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802/2 "2022-07-06T07:25:00Z")

</div>

That’s my vote there 🙂 I’d say in this instance, just go with mimicking demographic using compositions. It’s the most pragmatic middle ground for something like this.  
What you have in mind is a great idea, having a significant data set in openEHR form, especially if it’s in EhrBase, would be fantastic. Other approaches to demographic would a) complicate implementation, b) would make querying demographic data problematic/less-convenient and most interesting queries at the population level end up touching demographic one way or another.  
There’s no reason to not to work on a V2 for demographic mapping based on one of the better options above once the most pragmatic one is in place.

I didn’t know about SynPuf, may I be lazy and ask if the claims data is … interesting? I can see there’s prescription/medication data but is there any diagnosis/observation type of data in that set? There’s also a Brazilian hospital data set in openEHR form I think. ORBA? O… something? It was a few years back when it was made available.

---

<div class="post-metadata">

### Author: ![borut.jures](https://discourse.openehr.org/user_avatar/discourse.openehr.org/borut.jures/32/1391_2.png) [@borut.jures](https://discourse.openehr.org/u/borut.jures)
#### Post date: [6 July 2022 07:51 UTC](https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802/3 "2022-07-06T07:51:11Z")

</div>

> [@Seref](#):
>
> Brazilian hospital data set in openEHR form I think. ORBA? O… something?

I believe you have ORBDA in mind: [Geração de uma Base Pública para Avaliação de Mecanismos de Persistência de Sistemas de Registros Eletrônicos de Saúde Baseados nas Especificações da Fundação openEHR – L@MPADA / UERJ](http://www.lampada.uerj.br/en/orbda/)

I also have a link to [Synthea](https://synthetichealth.github.io/synthea/) in my notes: “Synthea data contains a complete medical history, including medications, allergies, medical encounters, and social determinants of health.”  
It is available as FHIR, C-CDA and CSV.

Edit: fixed the name for ORBDA

---

<div class="post-metadata">

### Author: ![Seref](https://discourse.openehr.org/user_avatar/discourse.openehr.org/seref/32/13_2.png) [@Seref](https://discourse.openehr.org/u/Seref)
#### Post date: [6 July 2022 07:52 UTC](https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802/4 "2022-07-06T07:52:36Z")

</div>

thanks, it was orbda indeed 🙂

---

<div class="post-metadata">

### Author: ![yampeku](https://discourse.openehr.org/user_avatar/discourse.openehr.org/yampeku/32/25_2.png) [@yampeku](https://discourse.openehr.org/u/yampeku)
#### Post date: [6 July 2022 10:29 UTC](https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802/5 "2022-07-06T10:29:17Z")

</div>

> [@borut.jures](#):
>
> I also have a link to [Synthea](https://synthetichealth.github.io/synthea/) in my notes

Yes, synthea was also an option, I think @bna commented about it in the past

---

<div class="post-metadata">

### Author: ![thomas.beale](https://discourse.openehr.org/user_avatar/discourse.openehr.org/thomas.beale/32/35_2.png) [@thomas.beale](https://discourse.openehr.org/u/thomas.beale)
#### Post date: [6 July 2022 10:30 UTC](https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802/6 "2022-07-06T10:30:18Z")

</div>

> [@Seref](#):
>
> That’s my vote there 🙂 I’d say in this instance, just go with mimicking demographic using compositions. It’s the most pragmatic middle ground for something like this.

Mine as well. A trick that you can use if you have no dedicated openEHR demographic service is to do the mimicking using Compositions + AdminEntry + Cluster archetypes (the ‘fake’ real ones 😉 and then store the public demographic entities (= professionals + places) in a special EHR in the service - e.g. call it the 0-EHR, with a known EHR id (I don’t know off-hand if a GUID made of all 0s will work, but it would be perfect if it does).

Then you end up with an EHR service with all the usual patient EHRs, containing refs pointing out to demographic entities, but those demographic entities are cunningly hidden in another EHR, whose job is to act as a local demographic registry, or you could think of it as a cache.

Detailed patient demographics could be stored the same way - in another special EHR, since in Synpuf data all patients are fake anyway.

BTW Ricardo Correia’s group at Porto has someone working on an EHR data synthesiser - I don’t remember the details.

---

<div class="post-metadata">

### Author: ![yampeku](https://discourse.openehr.org/user_avatar/discourse.openehr.org/yampeku/32/25_2.png) [@yampeku](https://discourse.openehr.org/u/yampeku)
#### Post date: [6 July 2022 10:32 UTC](https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802/7 "2022-07-06T10:32:20Z")

</div>

> [@Seref](#):
>
> ask if the claims data is … interesting?

They have diagnoses in ICD (ICD9 if I recall correctly) and procedures in HCPCS  
They are provided in a typical database way as “diagnosis1, diagnosis2…procedure1, procedure2…”

---

<div class="post-metadata">

### Author: ![yampeku](https://discourse.openehr.org/user_avatar/discourse.openehr.org/yampeku/32/25_2.png) [@yampeku](https://discourse.openehr.org/u/yampeku)
#### Post date: [6 July 2022 10:39 UTC](https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802/8 "2022-07-06T10:39:12Z")

</div>

> [@thomas.beale](#):
>
> Compositions + AdminEntry + Cluster

This could also be a good question. Is AdminEntry fully supported in all available EHR repositories?

---

<div class="post-metadata">

### Author: ![Seref](https://discourse.openehr.org/user_avatar/discourse.openehr.org/seref/32/13_2.png) [@Seref](https://discourse.openehr.org/u/Seref)
#### Post date: [6 July 2022 11:24 UTC](https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802/9 "2022-07-06T11:24:02Z")

</div>

> [@yampeku](#):
>
> AdminEntry fully supported in all available EHR repositories

I’d be surprised if it was not supported. It’s an `ENTRY` subtype, hence, archetypetypeable. Already some archetypes in CKM, and in (archetypes that are in ) production in Ocean’s archetypes/templates, in case it helps.

---

<div class="post-metadata">

### Author: ![ian.mcnicoll](https://discourse.openehr.org/user_avatar/discourse.openehr.org/ian.mcnicoll/32/4430_2.png) [@ian.mcnicoll](https://discourse.openehr.org/u/ian.mcnicoll)
#### Post date: [6 July 2022 11:42 UTC](https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802/10 "2022-07-06T11:42:11Z")

</div>

ADMIN\_ENRTRY AFAIK yes.

It is also possible to put the the demographics clusters in other\_context, though I think using ADMIN\_ENTRY is preferable.

SynPuf is obviously worth doing but I think Synthea is much richer (for operational data)

---

<div class="post-metadata">

### Author: ![yampeku](https://discourse.openehr.org/user_avatar/discourse.openehr.org/yampeku/32/25_2.png) [@yampeku](https://discourse.openehr.org/u/yampeku)
#### Post date: [6 July 2022 12:22 UTC](https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802/11 "2022-07-06T12:22:52Z")

</div>

Yes, I think both are useful. I was thinking that this could be used as a validation of an openEHR to OMOP transformation , as we already know how synpuf data looks in OMOP CDM

---

<div class="post-metadata">

### Author: ![Seref](https://discourse.openehr.org/user_avatar/discourse.openehr.org/seref/32/13_2.png) [@Seref](https://discourse.openehr.org/u/Seref)
#### Post date: [6 July 2022 12:58 UTC](https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802/12 "2022-07-06T12:58:19Z")

</div>

> [@yampeku](#):
>
> as we already know how synpuf data looks in OMOP CDM

do we? how? where? 🙂

---

<div class="post-metadata">

### Author: ![yampeku](https://discourse.openehr.org/user_avatar/discourse.openehr.org/yampeku/32/25_2.png) [@yampeku](https://discourse.openehr.org/u/yampeku)
#### Post date: [7 July 2022 07:33 UTC](https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802/13 "2022-07-07T07:33:51Z")

</div>

;D  
Here is the ETL

> **[GitHub - OHDSI/ETL-CMS: Workproducts to ETL CMS datasets into OMOP Common...](https://github.com/OHDSI/ETL-CMS)**
>
> Workproducts to ETL CMS datasets into OMOP Common Data Model - GitHub - OHDSI/ETL-CMS: Workproducts to ETL CMS datasets into OMOP Common Data Model

Here is the direct download to the dump

> **[Downloads - LTS Computing LLC](http://www.ltscomputingllc.com/downloads/)**
>
> Go to top Here is the SynPUF 1000 person dataset in OMOP CDM v5.2.2 format: synpuf1k\_omop\_cdm\_5.2.2.zip This zip file contains a 1000 person sample of the CMS 2008-2010 Data Entrepreneurs’ Continue Reading →

---

<div class="post-metadata">

### Author: ![Seref](https://discourse.openehr.org/user_avatar/discourse.openehr.org/seref/32/13_2.png) [@Seref](https://discourse.openehr.org/u/Seref)
#### Post date: [7 July 2022 07:46 UTC](https://discourse.openehr.org/t/synpuf-syntetic-data-into-openehr/2802/14 "2022-07-07T07:46:38Z")

</div>

I’m beginning to suspect you guys have a Jira card open in Veratech titled “Distract Seref” 😃  
Thanks a lot Diego.
