A two-phase search that first narrows candidates with ILIKE, then ranks them using the best similarity function available for the connected DBMS:

  • PostgreSQL: similarity() from pg_trgm (requires the extension).

  • Snowflake: JAROWINKLER_SIMILARITY().

  • Spark / Databricks: normalized levenshtein().

  • All other dialects: positional boost scoring (exact > prefix > contains).

Synonym matching and mapping-count enrichment are scoped to the candidate set, avoiding a full-table synonym scan.

rankedSearchConcepts(
  keyword,
  connection,
  vocabularyDatabaseSchema,
  domain = NULL,
  standardOnly = TRUE,
  limit = 50L,
  offset = 0L
)

Arguments

keyword

Character string to search for.

connection

A DatabaseConnector connection to an OMOP CDM.

vocabularyDatabaseSchema

Schema containing the OMOP vocabulary tables. Required.

domain

Optional character vector restricting by domain_id (e.g. "Condition").

standardOnly

If TRUE (default), return only standard concepts.

limit

Maximum rows to return. Default 50L.

offset

Pagination offset. Default 0L.

Value

A tibble with columns: concept_id, concept_name, concept_code,

vocabulary_id, domain_id, concept_class_id, standard_concept,

relevance (0–1.5 float), mapping_count.

Details

For PostgreSQL the GIN trigram index dramatically improves performance but is not required: CREATE INDEX ON concept USING GIN (concept_name gin_trgm_ops);.

See also

Examples

if (FALSE) {
connection <- DatabaseConnector::connect(Eunomia::getEunomiaConnectionDetails())
rankedSearchConcepts("atrial fibrillation", connection, vocabularyDatabaseSchema = "main")
rankedSearchConcepts("metformin", connection, vocabularyDatabaseSchema = "main",
                   domain = "Drug", limit = 20L)
}