OMOP Data Description

Work in progress. These tutorial materials are still under active development and will continue to change until the tutorial takes place on October 20, 2026. Content, links, and exercises may be incomplete or shift without notice.

The tutorial uses two datasets: a synthetic OMOP CDM population (generated from code in this repository) and a public monthly county PM2.5 dataset for the United States (an external source; see Geo Data Description). Session 1 catalogs the public dataset. In Session 2 you ingest it with Gaia and link it, together with the simulated county SES index (registered through the same catalog), to the synthetic persons, which produces the EXTERNAL_EXPOSURE table. Sessions 3 and 4 query and model that result.


Privacy statement

All persons, clinical events, and SDOH values are synthetic, and no real individuals or addresses appear anywhere. Each person’s location is a random point drawn inside the boundary of a real US county; it does not correspond to a real residence. County identities (FIPS code, name, boundaries, population) and the PM2.5 exposure values are real public data, so the synthetic persons join to real county-level datasets. The county SES index and the SDOH observations derived from it are simulated, not real county statistics.


How the dataset is produced

Source code and build instructions: syntheticDataGIS/. The build runs the Gaia toolchain itself:

  1. Gaia ingests the Census TIGER 2023 counties and the CDC monthly county PM2.5 dataset through the Gaia Catalog.
  2. Stage 1 creates the persons, their residences, and LOCATION / LOCATION_HISTORY, and draws the simulated county SES index. The SES index is then ingested through the Gaia Catalog as a third source (synthetic_county_ses, published as county_ses.csv), the same way as PM2.5, and attached to the TIGER county polygons.
  3. Gaia loads those locations and derives the exposure rows with its own spatial join (working.spatial_join_from_catalog), once for PM2.5 and once for SES.
  4. Stage 2 draws conditions, drugs, procedures, measurements, and SDOH from those exposures (PM2.5 and SES, both as Gaia derived them), and writes the answer key.

The exposures that drive disease risk (PM2.5 and SES) are therefore exactly what you reproduce in Exercise 2. EXTERNAL_EXPOSURE is empty in the published dataset; you populate it from Gaia. A prebuilt copy (external_exposure_fallback.csv.gz) is provided as a fallback in case the pipeline does not run on your machine. Both the dataset (synthetic_omop_gis.sql.gz) and the fallback are in the syntheticDataGIS/data folder. The build is deterministic given the same Gaia image and catalog commit, which are recorded in demo.build_info.


OMOP synthetic dataset

Overview

The dataset is an OMOP CDM v5.4 schema (omopgis) plus the Gaia extension tables (LOCATION_HISTORY, EXTERNAL_EXPOSURE), a demo-only COUNTY_REFERENCE table, and a demo schema holding the generator parameters, the answer key, and the exercise fixtures. See OMOP Data Visualization for charts of the cohort, each paired with its SQL.

  • 10,000 adults, aged 18-90 on 2014-01-01, observed 2014-01-01 to 2019-12-31.
  • Real counties. Persons live in about 2,700 of 3,099 eligible real US counties and county equivalents. Counties are sampled in proportion to the square root of their 2019 population, so populous counties hold more persons while every state is represented.
  • Residence history. Each person has one LOCATION_HISTORY row spanning 2014-2019, except about 10% who move once to a different county between March 2015 and October 2018 and have two consecutive, non-overlapping rows. PERSON.location_id is the current (last) address.
  • Exposure. Once derived in Exercise 2, EXTERNAL_EXPOSURE holds one PM2.5 row per person per month (about 720,000 rows; a move in the middle of a month gives two partial-month rows for that month) and one SES row per residence interval (about 11,000 rows).
  • 14 conditions (5 respiratory, 9 cardiometabolic), about 18,000 occurrences, plus the drugs, procedures, and measurements tied to them.
  • 12 county-level SDOH observations per person, derived from the simulated county SES index.

The county table carries a density category from real population density (Urban Core, Suburban, Small Town, Rural) and a simulated SES index (0-100, higher is more affluent). The SES index is a source dataset in the Gaia catalog like the others: persons receive their SES exposure through the same spatial-temporal join as PM2.5. County SES is correlated with county PM2.5 at about -0.3, so higher-PM2.5 counties tend to have lower SES. That makes SES a genuine confounder of the PM2.5 effect.

How the conditions are generated

Every condition follows the same logistic model, with the true coefficients stored in demo.generator_truth:

logit P(condition) = logit(prevalence_ref)
                   + b_pm25  * (mean PM2.5 - 8 µg/m³)
                   + b_ses   * (SES - 50) / 15
                   + b_age   * (age - 54) / 10
                   + b_female * [female]
                   + county random effect (SD 0.15)

PM2.5 and SES are day-weighted over a person’s residences, so a mover is exposed to the average of both counties. Onset dates are uniform over 2014-2019 and carry no information about exposure timing. The unmeasured county random effect induces spatial clustering.

Condition Category PM2.5 log-odds per µg/m³ (OR) SES per SD Age per decade Female
COPD Respiratory 0.15 (1.16) -0.20 0.45 0.00
Pneumonia Respiratory 0.10 (1.11) -0.30 0.30 -0.10
Bronchitis Respiratory 0.10 (1.11) -0.25 0.20 0.10
Asthma Respiratory 0.08 (1.08) -0.15 0.00 0.25
MI Cardiometabolic 0.08 (1.08) -0.40 0.50 -0.50
CHF Cardiometabolic 0.07 (1.07) -0.45 0.65 -0.20
CAD Cardiometabolic 0.06 (1.06) -0.35 0.60 -0.45
Stroke Cardiometabolic 0.05 (1.05) -0.40 0.55 -0.10
Rhinitis, hypertension, T2DM, obesity, hyperlipidemia, CKD mixed 0 (null) -0.05 to -0.50 varies varies

The six null outcomes have no simulated PM2.5 effect; because of the PM2.5-SES correlation, a crude analysis still shows an apparent association with them, which is useful as negative controls and for illustrating confounding.

The PM2.5 coefficients are deliberately exaggerated. Real county PM2.5 varies little across 10,000 people (SD about 1.4 µg/m³), so effects at epidemiologic sizes would not be recoverable. The simulated effects (for example COPD, OR about 1.16 per µg/m³) are much larger than published estimates. Treat them as a teaching device, not as real-world effect sizes. The null outcomes are simulated as null; real-world literature may report associations.

Fixtures

Seven persons are built to exercise specific checks. They are listed in demo.fixture_person, with a matching answer key in demo.expected_result. Person IDs may change between dataset versions, so always look them up in the tables rather than hardcoding them.

Fixture What it contains Used in
PREGNANCY_STATIC Female, one pregnancy episode (2016-03-18 to 2016-12-09, 267 days) at a single residence; the first and last months overlap the episode only partially Exercise 3
PREGNANCY_MOVER A pregnancy episode (2016-05-20 to 2017-02-10) during which the person moves counties on 2016-08-15, so the episode spans two residences Exercise 3
DUPLICATE_SOURCE_ROW A staging row that duplicates a loaded 2017-03 row Exercise 2
UNIT_MISMATCH A staging row for 2017-04 in mg/m³ instead of µg/m³ Exercise 2
NON_OVERLAPPING_INTERVAL A staging row for January 2020, outside the residence and observation period Exercise 2
MISSING_EXPOSURE_VALUE A staging row for 2017-05 with a NULL value Exercise 2
MISSING_SDOH_VALUE One SDOH observation (neighborhood disadvantage) is missing for this person Exercise 3

The four bad exposure rows are in demo.rejected_exposure_row. They never appear in the clean tables, because the clean exposure comes only from the Gaia pipeline; they are for the QA drill in Exercise 2. The pregnancy episodes are in the OMOP EPISODE table. demo.expected_result holds, for each pregnancy, the correct day-weighted mean PM2.5 and the naive mean of the overlapping rows, plus the expected row counts for the exposure checks.

Grain and uniqueness

A person has 72 monthly PM2.5 rows (73 if they move in the middle of a month, because that month is split into two rows). The expected grain of EXTERNAL_EXPOSURE is one row per person, location, exposure concept, and interval, so no two rows share:

person_id + location_id + exposure_concept_id + exposure_start_date + exposure_end_date

Tables

Standard OMOP (populated): person, location, observation_period, condition_occurrence, drug_exposure, procedure_occurrence, measurement, observation, episode, cdm_source

GIS extension: location_history (populated); external_exposure (empty, populated in Exercise 2)

Demo-only - county_reference: one row per real county with FIPS, name, state, centroid, land area, 2019 population, density category, mean 2014-2019 PM2.5, and the simulated SES index. location.county_ref_id is a foreign key into it and is the generator’s ground truth for each person’s county. - demo.generator_params, demo.generator_truth: the parameters and true coefficients of the generator. - demo.fixture_person, demo.expected_result, demo.rejected_exposure_row: the fixtures and answer key described above. - demo.build_info: Gaia image digest, Gaia Catalog commit, source datasets, and build date.

Mini vocabulary (populated): concept, concept_ancestor, concept_relationship, concept_synonym, vocabulary, domain, concept_class, relationship. It holds the 73 standard concepts the data use (SNOMED, RxNorm, LOINC, UCUM and a few CDM and type concepts) plus the OMOP GIS, SDoH, and Exposome concepts, so Exercises 3 and 4 and the OHDSI tools that resolve concept sets work without a separate vocabulary download. It is not a substitute for the full vocabulary; see the vocabulary notes.

Provenance and known limitations

  • cdm_source records the CDM version (5.4) and the generator version. demo.build_info records the exact Gaia image and catalog commit, so a build can be reproduced.
  • On the exposure rows it derives, Gaia leaves unit_concept_id and dose_unit_source_value empty. The exposure concept comes from the catalog (2052499839, “Annual Mean Of Particulate Matter (PM2.5) Concentration”), which does not match the monthly grain of the data; the tutorial keeps it on purpose as a vocabulary discussion point in Exercise 3. exposure_type_concept_id is the data-source type (2052499878, Air Quality Database), passed explicitly to the join in Exercise 2. exposure_source_value is the Gaia variable_source_id of the variable, exposure_source_concept_id repeats the exposure concept, and exposure_relationship_source_value is the verbatim operator (ST_Within). value_as_concept_id and unit_concept_id follow the OMOP MEASUREMENT meaning (a categorical result and the unit of the numeric value).
  • Every concept id used by the data is checked against the mini vocabulary at build time (it must exist, be a standard concept, and be in the right domain). Residences use the OMOP GIS concept Patient Residence as the LOCATION_HISTORY relationship, and pregnancies are EPISODE rows of the Disease Episode concept whose object is the Pregnancy condition.
  • Persons are placed uniformly at random inside a county polygon, so there is no within-county clustering; everyone in a county shares the same exposure value.