OHDSI GIS
WGWelcome to the OHDSI GIS tutorial materials. This section walks through hands-on exercises for working with the Gaia toolchain and GIS-related OMOP CDM extensions.
This tutorial takes a single research question — the relationship between an environmental exposure and a clinical outcome — and carries it through four stages: finding and evaluating a source dataset, ingesting and spatially/temporally linking it with Gaia, integrating the result into the OMOP CDM, and turning the validated exposure fact into an analysis-ready covariate.
Each session builds directly on the output of the one before it. The metadata record you produce in Session 1 is the input to the ingestion you run in Session 2; the exposure row you trace in Session 2 is the row you query and validate in Session 3; the validated exposure fact from Session 3 becomes the feature you specify in Session 4.
| Session | Focus | Learning outcome |
|---|---|---|
| 1. Cataloging and data discovery | Turning a research question into a computable data requirement | Produce one valid, machine-actionable metadata record and a dataset-fitness decision |
| 2. The Gaia pipeline | Ingestion and spatial-temporal linkage | Run the ingestion, check it against the answer key, and explain a derived exposure row and its provenance |
| 3. OMOP integration | The person-place-time model | Query the model, validate vocabulary roles, calculate temporally aligned exposure metrics, and build exposure-stratified cohorts |
| 4. Analytical applications | From exposure fact to covariate | Design an analysis-ready exposure feature, fit a crude vs adjusted model, and identify threats to validity |
You will need basic command line computing skills (powershell on windows, terminal on mac, or bash on linux). If you are not familiar with these tools, we suggest you complete the Software Carpentry lesson on the unix shell (this applies to all command line computing). The lesson can be used in a self-learning tutorial environment.
We will use docker and the docker compose container orchestration system for the tutorial (and for the Gaia toolchain). While no knowledge is needed, familiarity will be helpful. If you are not familiar and have a little time, the docker maintains basic documentation and several get started tutorials, and the build and share tutorial is a good place to start.
No prior GIS or spatial-analysis experience is assumed; familiarity with Structured Query Language (SQL) and the OMOP CDM (PERSON, OBSERVATION_PERIOD, CONCEPT) is helpful for Sessions 3-4.
Docker Note: docker is proprietary software, but there is a community edition, Docker CE. There are several ways to install Docker, which is best depends on your compute environment and your experience.
Windows note: to run docker you need a linux subsystem installed forst. We recommend the WSL2 linux subsystem for windows.
After installing docker, please complete the following steps before coming to the workshop
docker run -d --name gaia-core --platform linux/amd64 --network gaiadocker_default -p 8787:8787 \
-e USER=ohdsi -e PASSWORD=<choose a password> ohdsi/gaia-core:main
Then open http://localhost:8787 and sign in with
ohdsi and that password. In R, connect with
server = "gaia-db/gaiacore" and the Postgres password from
gaiaDocker/secrets/gaia/POSTGRES_PASSWORD;
DATABASECONNECTOR_JAR_FOLDER is already set, so
pathToDriver is optional. (gaiadocker_default
is the network docker compose creates for a checkout in a folder called
gaiaDocker; list yours with
docker network ls.)
PGAdmin and QGIS are standard components in an analytic toolbox for working with GIS data. They can be installed on windows, macOS, or linux. We encourage all participants to have these two software tools installed before the workshop.
There are many other tools available as well, but we find these two work well with the Gaia toolchain.
VERY IMPORTANT: If you have trouble with any installation or setup step above, contact the workshop organizers immediately ( houghtaling@ohdsi.org or tnorris@miami.edu ).
All exercises run against a synthetic cohort designed for this tutorial: 10,000 adults with residence histories (about 10% move once), 14 respiratory and cardiometabolic conditions, and county-level social determinants. No real individuals, addresses, or coordinates are used. Persons are placed at random points inside real US counties, and the exposure is real public data: monthly county PM2.5 from the CDC, 2014-2019, which you ingest and link to the persons yourself in Session 2 (the exposure table ships empty). A handful of fixtures (a pregnancy that partially overlaps exposure months, a pregnancy that spans a residential move, and four bad exposure rows, plus a missing SDOH value) give each session’s checks something real to catch. The simulated PM2.5 effects are deliberately exaggerated so they can be recovered with 10,000 persons; they are not real-world effect sizes. See the OMOP Data Description page for details.
Choose your track. Exercises 2 to 4 are written in
SQL (the reference path for the workshop). An R version that uses the gaiaCore package is available as a beta: start
with the R setup, then R Exercise 2, 3 and 4. Both tracks use the same data,
the same answer key and the same ohdsi/gaia-core:main
image.
If you get stuck, see Getting Help or reach out on the GIS Working Group Teams channel.