OHDSI GIS
WGThe tutorial uses two datasets: a synthetic OMOP CDM population
(generated from code in this repository) and a public monthly county
PM2.5 dataset for the United States (an external source; see Geo Data Description).
Session 1 catalogs the public dataset. In Session 2 you ingest it with
Gaia and link it, together with the simulated county SES index
(registered through the same catalog), to the synthetic persons, which
produces the EXTERNAL_EXPOSURE table. Sessions 3 and 4
query and model that result.
All persons, clinical events, and SDOH values are synthetic, and no real individuals or addresses appear anywhere. Each person’s location is a random point drawn inside the boundary of a real US county; it does not correspond to a real residence. County identities (FIPS code, name, boundaries, population) and the PM2.5 exposure values are real public data, so the synthetic persons join to real county-level datasets. The county SES index and the SDOH observations derived from it are simulated, not real county statistics.
Source code and build instructions: syntheticDataGIS/.
The build runs the Gaia toolchain itself:
LOCATION / LOCATION_HISTORY, and draws the
simulated county SES index. The SES index is then ingested through the
Gaia Catalog as a third source (synthetic_county_ses,
published as county_ses.csv), the same way as PM2.5, and
attached to the TIGER county polygons.working.spatial_join_from_catalog), once
for PM2.5 and once for SES.The exposures that drive disease risk (PM2.5 and SES) are therefore
exactly what you reproduce in Exercise 2.
EXTERNAL_EXPOSURE is empty in the published
dataset; you populate it from Gaia. A prebuilt copy
(external_exposure_fallback.csv.gz) is provided as a
fallback in case the pipeline does not run on your machine. Both the
dataset (synthetic_omop_gis.sql.gz) and the fallback are in
the syntheticDataGIS/data
folder. The build is deterministic given the same Gaia image and catalog
commit, which are recorded in demo.build_info.
The dataset is an OMOP CDM v5.4 schema (omopgis) plus
the Gaia extension tables (LOCATION_HISTORY,
EXTERNAL_EXPOSURE), a demo-only
COUNTY_REFERENCE table, and a demo schema
holding the generator parameters, the answer key, and the exercise
fixtures. See OMOP Data
Visualization for charts of the cohort, each paired with its
SQL.
LOCATION_HISTORY row spanning 2014-2019, except about 10%
who move once to a different county between March 2015 and October 2018
and have two consecutive, non-overlapping rows.
PERSON.location_id is the current (last) address.EXTERNAL_EXPOSURE holds one PM2.5 row per person per month
(about 720,000 rows; a move in the middle of a month gives two
partial-month rows for that month) and one SES row per residence
interval (about 11,000 rows).The county table carries a density category from real population density (Urban Core, Suburban, Small Town, Rural) and a simulated SES index (0-100, higher is more affluent). The SES index is a source dataset in the Gaia catalog like the others: persons receive their SES exposure through the same spatial-temporal join as PM2.5. County SES is correlated with county PM2.5 at about -0.3, so higher-PM2.5 counties tend to have lower SES. That makes SES a genuine confounder of the PM2.5 effect.
Every condition follows the same logistic model, with the true
coefficients stored in demo.generator_truth:
logit P(condition) = logit(prevalence_ref)
+ b_pm25 * (mean PM2.5 - 8 µg/m³)
+ b_ses * (SES - 50) / 15
+ b_age * (age - 54) / 10
+ b_female * [female]
+ county random effect (SD 0.15)
PM2.5 and SES are day-weighted over a person’s residences, so a mover is exposed to the average of both counties. Onset dates are uniform over 2014-2019 and carry no information about exposure timing. The unmeasured county random effect induces spatial clustering.
| Condition | Category | PM2.5 log-odds per µg/m³ (OR) | SES per SD | Age per decade | Female |
|---|---|---|---|---|---|
| COPD | Respiratory | 0.15 (1.16) | -0.20 | 0.45 | 0.00 |
| Pneumonia | Respiratory | 0.10 (1.11) | -0.30 | 0.30 | -0.10 |
| Bronchitis | Respiratory | 0.10 (1.11) | -0.25 | 0.20 | 0.10 |
| Asthma | Respiratory | 0.08 (1.08) | -0.15 | 0.00 | 0.25 |
| MI | Cardiometabolic | 0.08 (1.08) | -0.40 | 0.50 | -0.50 |
| CHF | Cardiometabolic | 0.07 (1.07) | -0.45 | 0.65 | -0.20 |
| CAD | Cardiometabolic | 0.06 (1.06) | -0.35 | 0.60 | -0.45 |
| Stroke | Cardiometabolic | 0.05 (1.05) | -0.40 | 0.55 | -0.10 |
| Rhinitis, hypertension, T2DM, obesity, hyperlipidemia, CKD | mixed | 0 (null) | -0.05 to -0.50 | varies | varies |
The six null outcomes have no simulated PM2.5 effect; because of the PM2.5-SES correlation, a crude analysis still shows an apparent association with them, which is useful as negative controls and for illustrating confounding.
Seven persons are built to exercise specific checks. They are listed
in demo.fixture_person, with a matching answer key in
demo.expected_result. Person IDs may change between dataset
versions, so always look them up in the tables rather than hardcoding
them.
| Fixture | What it contains | Used in |
|---|---|---|
PREGNANCY_STATIC |
Female, one pregnancy episode (2016-03-18 to 2016-12-09, 267 days) at a single residence; the first and last months overlap the episode only partially | Exercise 3 |
PREGNANCY_MOVER |
A pregnancy episode (2016-05-20 to 2017-02-10) during which the person moves counties on 2016-08-15, so the episode spans two residences | Exercise 3 |
DUPLICATE_SOURCE_ROW |
A staging row that duplicates a loaded 2017-03 row | Exercise 2 |
UNIT_MISMATCH |
A staging row for 2017-04 in mg/m³ instead of µg/m³ | Exercise 2 |
NON_OVERLAPPING_INTERVAL |
A staging row for January 2020, outside the residence and observation period | Exercise 2 |
MISSING_EXPOSURE_VALUE |
A staging row for 2017-05 with a NULL value | Exercise 2 |
MISSING_SDOH_VALUE |
One SDOH observation (neighborhood disadvantage) is missing for this person | Exercise 3 |
The four bad exposure rows are in
demo.rejected_exposure_row. They never appear in the clean
tables, because the clean exposure comes only from the Gaia pipeline;
they are for the QA drill in Exercise 2. The pregnancy episodes are in
the OMOP EPISODE table. demo.expected_result
holds, for each pregnancy, the correct day-weighted mean PM2.5 and the
naive mean of the overlapping rows, plus the expected row counts for the
exposure checks.
A person has 72 monthly PM2.5 rows (73 if they move in the middle of
a month, because that month is split into two rows). The expected grain
of EXTERNAL_EXPOSURE is one row per person, location,
exposure concept, and interval, so no two rows share:
person_id + location_id + exposure_concept_id + exposure_start_date + exposure_end_date
Standard OMOP (populated): person,
location, observation_period,
condition_occurrence, drug_exposure,
procedure_occurrence, measurement,
observation, episode,
cdm_source
GIS extension: location_history
(populated); external_exposure (empty,
populated in Exercise 2)
Demo-only - county_reference: one row
per real county with FIPS, name, state, centroid, land area, 2019
population, density category, mean 2014-2019 PM2.5, and the simulated
SES index. location.county_ref_id is a foreign key into it
and is the generator’s ground truth for each person’s county. -
demo.generator_params, demo.generator_truth:
the parameters and true coefficients of the generator. -
demo.fixture_person, demo.expected_result,
demo.rejected_exposure_row: the fixtures and answer key
described above. - demo.build_info: Gaia image digest, Gaia
Catalog commit, source datasets, and build date.
Mini vocabulary (populated): concept,
concept_ancestor, concept_relationship,
concept_synonym, vocabulary,
domain, concept_class,
relationship. It holds the 73 standard concepts the data
use (SNOMED, RxNorm, LOINC, UCUM and a few CDM and type concepts) plus
the OMOP GIS, SDoH, and Exposome concepts, so Exercises 3 and 4 and the
OHDSI tools that resolve concept sets work without a separate vocabulary
download. It is not a substitute for the full vocabulary; see the vocabulary
notes.
cdm_source records the CDM version (5.4) and the
generator version. demo.build_info records the exact Gaia
image and catalog commit, so a build can be reproduced.unit_concept_id and dose_unit_source_value
empty. The exposure concept comes from the catalog
(2052499839, “Annual Mean Of Particulate Matter (PM2.5)
Concentration”), which does not match the monthly grain of the data; the
tutorial keeps it on purpose as a vocabulary discussion point in
Exercise 3. exposure_type_concept_id is the data-source
type (2052499878, Air Quality Database), passed explicitly
to the join in Exercise 2. exposure_source_value is the
Gaia variable_source_id of the variable,
exposure_source_concept_id repeats the exposure concept,
and exposure_relationship_source_value is the verbatim
operator (ST_Within). value_as_concept_id and
unit_concept_id follow the OMOP MEASUREMENT meaning (a
categorical result and the unit of the numeric value).Patient Residence as the LOCATION_HISTORY
relationship, and pregnancies are EPISODE rows of the
Disease Episode concept whose object is the
Pregnancy condition.