R/conceptSearch.R
rankedSearchConcepts.RdA two-phase search that first narrows candidates with ILIKE, then ranks them using the best similarity function available for the connected DBMS:
PostgreSQL: similarity() from pg_trgm (requires the extension).
Snowflake: JAROWINKLER_SIMILARITY().
Spark / Databricks: normalized levenshtein().
All other dialects: positional boost scoring (exact > prefix > contains).
Synonym matching and mapping-count enrichment are scoped to the candidate set, avoiding a full-table synonym scan.
rankedSearchConcepts(
keyword,
connection,
vocabularyDatabaseSchema,
domain = NULL,
standardOnly = TRUE,
limit = 50L,
offset = 0L
)Character string to search for.
A DatabaseConnector connection to an OMOP CDM.
Schema containing the OMOP vocabulary tables. Required.
Optional character vector restricting by domain_id (e.g. "Condition").
If TRUE (default), return only standard concepts.
Maximum rows to return. Default 50L.
Pagination offset. Default 0L.
A tibble with columns: concept_id, concept_name, concept_code,
vocabulary_id, domain_id, concept_class_id, standard_concept,
relevance (0–1.5 float), mapping_count.
For PostgreSQL the GIN trigram index dramatically improves performance but is
not required: CREATE INDEX ON concept USING GIN (concept_name gin_trgm_ops);.
if (FALSE) {
connection <- DatabaseConnector::connect(Eunomia::getEunomiaConnectionDetails())
rankedSearchConcepts("atrial fibrillation", connection, vocabularyDatabaseSchema = "main")
rankedSearchConcepts("metformin", connection, vocabularyDatabaseSchema = "main",
domain = "Drug", limit = 20L)
}