Tutorial

Work in progress. These tutorial materials are still under active development and will continue to change until the tutorial takes place on October 20, 2026. Content, links, and exercises may be incomplete or shift without notice.

Welcome to the OHDSI GIS tutorial materials. This section walks through hands-on exercises for working with the Gaia toolchain and GIS-related OMOP CDM extensions.


Overview

This tutorial takes a single research question — the relationship between an environmental exposure and a clinical outcome — and carries it through four stages: finding and evaluating a source dataset, ingesting and spatially/temporally linking it with Gaia, integrating the result into the OMOP CDM, and turning the validated exposure fact into an analysis-ready covariate.

Each session builds directly on the output of the one before it. The metadata record you produce in Session 1 is the input to the ingestion you run in Session 2; the exposure row you trace in Session 2 is the row you query and validate in Session 3; the validated exposure fact from Session 3 becomes the feature you specify in Session 4.

Session Focus Learning outcome
1. Cataloging and data discovery Turning a research question into a computable data requirement Produce one valid, machine-actionable metadata record and a dataset-fitness decision
2. The Gaia pipeline Ingestion and spatial-temporal linkage Run the ingestion, check it against the answer key, and explain a derived exposure row and its provenance
3. OMOP integration The person-place-time model Query the model, validate vocabulary roles, calculate temporally aligned exposure metrics, and build exposure-stratified cohorts
4. Analytical applications From exposure fact to covariate Design an analysis-ready exposure feature, fit a crude vs adjusted model, and identify threats to validity


Prerequisites

You will need basic command line computing skills (powershell on windows, terminal on mac, or bash on linux). If you are not familiar with these tools, we suggest you complete the Software Carpentry lesson on the unix shell (this applies to all command line computing). The lesson can be used in a self-learning tutorial environment.

We will use docker and the docker compose container orchestration system for the tutorial (and for the Gaia toolchain). While no knowledge is needed, familiarity will be helpful. If you are not familiar and have a little time, the docker maintains basic documentation and several get started tutorials, and the build and share tutorial is a good place to start.

No prior GIS or spatial-analysis experience is assumed; familiarity with Structured Query Language (SQL) and the OMOP CDM (PERSON, OBSERVATION_PERIOD, CONCEPT) is helpful for Sessions 3-4.

Required software

  • Docker Note: docker is proprietary software, but there is a community edition, Docker CE. There are several ways to install Docker, which is best depends on your compute environment and your experience.

    Windows note: to run docker you need a linux subsystem installed forst. We recommend the WSL2 linux subsystem for windows.

    • Docker Desktop is perhaps the easiest install and comes with a relatively easy to use user interface. This runs well on Windows with WSL2 installed. It also can run on a mac.
    • Colima + Docker is good on a mac as long as you are comfortable on the command line. This is a very lightweight install with no bells and whistles. This does not work on windows.

After installing docker, please complete the following steps before coming to the workshop

  • docker run -d --name gaia-core --platform linux/amd64 --network gaiadocker_default -p 8787:8787 \
      -e USER=ohdsi -e PASSWORD=<choose a password> ohdsi/gaia-core:main

    Then open http://localhost:8787 and sign in with ohdsi and that password. In R, connect with server = "gaia-db/gaiacore" and the Postgres password from gaiaDocker/secrets/gaia/POSTGRES_PASSWORD; DATABASECONNECTOR_JAR_FOLDER is already set, so pathToDriver is optional. (gaiadocker_default is the network docker compose creates for a checkout in a folder called gaiaDocker; list yours with docker network ls.)

Additional useful software (not required)

PGAdmin and QGIS are standard components in an analytic toolbox for working with GIS data. They can be installed on windows, macOS, or linux. We encourage all participants to have these two software tools installed before the workshop.

  • PGAdmin - an open source postgres database management tool
  • QGIS - an open source geographic infomration system (GIS) and visualization tool

There are many other tools available as well, but we find these two work well with the Gaia toolchain.

VERY IMPORTANT: If you have trouble with any installation or setup step above, contact the workshop organizers immediately ( or ).

Additional resources (not required)


Demo dataset

All exercises run against a synthetic cohort designed for this tutorial: 10,000 adults with residence histories (about 10% move once), 14 respiratory and cardiometabolic conditions, and county-level social determinants. No real individuals, addresses, or coordinates are used. Persons are placed at random points inside real US counties, and the exposure is real public data: monthly county PM2.5 from the CDC, 2014-2019, which you ingest and link to the persons yourself in Session 2 (the exposure table ships empty). A handful of fixtures (a pregnancy that partially overlaps exposure months, a pregnancy that spans a residential move, and four bad exposure rows, plus a missing SDOH value) give each session’s checks something real to catch. The simulated PM2.5 effects are deliberately exaggerated so they can be recovered with 10,000 persons; they are not real-world effect sizes. See the OMOP Data Description page for details.


Exercises

  1. Exercise 1: Cataloging and data discovery - Evaluate a source dataset and produce a dataset-level and variable-level metadata record with a fitness decision.
  2. Exercise 2: The Gaia pipeline - Run the ingestion and spatial-temporal linkage, check it against the answer key, and trace one exposure row back to its source dataset and geometry.
  3. Exercise 3: OMOP integration - Query the person-place-time model, validate vocabulary roles, calculate temporally aligned exposure metrics, and build exposure-stratified cohorts.
  4. Exercise 4: Analytical applications - Draft an analysis-ready feature specification and its validity/QA checklist, then fit and critique a crude vs adjusted exposure model.

Choose your track. Exercises 2 to 4 are written in SQL (the reference path for the workshop). An R version that uses the gaiaCore package is available as a beta: start with the R setup, then R Exercise 2, 3 and 4. Both tracks use the same data, the same answer key and the same ohdsi/gaia-core:main image.


Getting Help

If you get stuck, see Getting Help or reach out on the GIS Working Group Teams channel.