Defining Cohort Phenotypes logo

Defining Cohort Phenotypes

CommunityPopular
maziyarpanahi
defining-cohort-phenotypes

Authors computable phenotype and cohort definitions in the OHDSI ATLAS / CIRCE style over the OMOP CDM, combining standard concept sets with NLP-derived features that OpenMed extracts. Use when the user wants to define a patient cohort, write a computable phenotype, reuse PheKB or OHDSI Phenotype Library logic, build concept sets, or augment code-based criteria with text features. Trigger keywords: phenotype, cohort definition, OHDSI, ATLAS, CIRCE, OMOP CDM, concept set, PheKB, Phenotype Library, eMERGE, computable phenotype. Pairs adjacent to OpenMed: NLP features from openmed.analyze_text augment code-based phenotypes for entities that are poorly captured by structured codes. OMOP CDM and OHDSI tools are open source; restricted vocabularies (SNOMED, CPT) are user-supplied.

Overview

Publishermaziyarpanahi
Repositoryopenmed
Skill namedefining-cohort-phenotypes
Stars
5.3K
Forks
677
Bundled files
Instructions only
LicenseApache-2.0
Links
  • Markdown instructions

    A SKILL.md file the model loads on demand, so it only costs tokens when a request actually matches.

  • Works with any LLM

    AI skills are plain Markdown, not provider-specific code, so this works with GPT, Claude, Gemini, Grok, or a local model.

  • Self-contained

    Everything the model needs lives in the instructions — no extra files to sync.

  • Open source

    Published by maziyarpanahi on GitHub. Read the source before you install it.

Installation

Install the Defining Cohort Phenotypes AI skill in TypingMind to use it with any LLM, or drop it into another agent that reads SKILL.md.

1

Install in TypingMind

TypingMind installs a skill straight from its GitHub folder — it reads SKILL.md, bundles the resource files, and stores the result locally.

  1. Open the app and go to Plugins → Skills.
  2. Choose "Install from GitHub".
  3. Paste the skill folder URL below and confirm.
  4. Enable the skill in any chat where you want it available.
Plugins → Skills → Add skill → From GitHub URL, then paste the folder URL and press Continue.
2

Install in another agent

Any agent that reads the Agent Skills format can use this skill — copy the folder into that agent's skills directory.

Claude Code — .claude/skills
git clone --depth 1 https://github.com/maziyarpanahi/openmed.git /tmp/openmed
mkdir -p .claude/skills
cp -r /tmp/openmed/skills/defining-cohort-phenotypes .claude/skills/defining-cohort-phenotypes
Restart Claude Code after copying so it picks up the new skill.

Use it in TypingMind

Enable Defining Cohort Phenotypes in any TypingMind chat and the model takes it from there. Its name and description sit in the system prompt, and the moment a request matches, the model loads the full instructions itself — you never invoke it by hand, and it costs no tokens until it is actually used.

The model loads Defining Cohort Phenotypes on its own as soon as a request matches it.

Works with any AI model

AI skills are plain Markdown instructions rather than provider-specific code, so Defining Cohort Phenotypes is not tied to the model it was written for. Install it once in TypingMind and use it with GPT-5, Claude, Gemini, Grok, DeepSeek, Mistral, Llama, or a local model you run yourself — all on your own API keys.

  • Loaded only when it is needed

    The system prompt carries just the name and description. The instructions are fetched on the first matching request, so an idle skill costs nothing.

  • Switch models mid-chat

    Because the skill is instructions rather than code, changing model does not break it — the next model reads the same SKILL.md.

Skill instructions

This is the SKILL.md content the model loads. Read it before installing — a skill is instructions your model will follow.

Defining cohort phenotypes (OHDSI / OMOP CDM)

A computable phenotype is a portable, executable definition of "which patients have condition X" — concept sets plus inclusion logic that runs against any OMOP CDM-compliant database. In the OHDSI stack, ATLAS authors these visually, CIRCE serializes them to a standardized JSON representation, and that JSON compiles to database-specific SQL. This skill helps you author such definitions and augment them with NLP features that OpenMed extracts from clinical text — exactly the signals that structured codes miss.

OMOP CDM, ATLAS, CIRCE, and the OHDSI Phenotype Library are open source. The vocabulary content you reference (SNOMED CT, CPT4, ICD) is user-supplied — do not bundle restricted terminologies; load them into your own OMOP vocabulary tables with your own licenses.

When to use

  • You need a reproducible cohort definition for analytics or research.
  • You want to reuse an existing PheKB or OHDSI Phenotype Library definition and adapt it.
  • A phenotype depends on facts that live only in free text (e.g. smoking status, symptom severity, social context) and code-based logic alone is weak.

For terminology grounding of individual entities, see coding-icd10, normalizing-rxnorm, mapping-loinc; this skill is about composing them into a cohort.

Anatomy of a CIRCE cohort definition

A CIRCE cohort definition JSON has two parts: ConceptSets (the code lists) and an expression (entry event + inclusion rules). Shape (abridged):

jsonc
{
  "ConceptSets": [{
    "id": 0, "name": "Type 2 diabetes",
    "expression": { "items": [{
      "concept": { "CONCEPT_ID": 201826,           // OMOP standard concept
                   "CONCEPT_CODE": "44054006",      // SNOMED (user vocab)
                   "VOCABULARY_ID": "SNOMED" },
      "includeDescendants": true                    // pull the hierarchy
    }] }
  }],
  "PrimaryCriteria": {                              // entry event
    "CriteriaList": [{ "ConditionOccurrence": { "CodesetId": 0 } }],
    "ObservationWindow": { "PriorDays": 0, "PostDays": 0 },
    "PrimaryCriteriaLimit": { "Type": "First" }
  },
  "InclusionRules": [{
    "name": "Adult at index",
    "expression": { "Type": "ALL", "CriteriaList": [{
      "Criteria": { "ConditionEra": { "AgeAtStart": { "Value": 18, "Op": "gte" } } }
    }] }
  }]
}

You author this in ATLAS (recommended) or by hand. The OHDSI Phenotype Library ships hundreds of vetted definitions as exactly this JSON; reuse before you write.

Augmenting with OpenMed NLP features

Code-based phenotypes are blind to facts that only appear in notes. The pattern is materialize an NLP feature as OMOP rows, then reference it like any concept set.

python
import openmed

# 1) Extract the text feature OpenMed is good at (e.g. tobacco use, symptom)
note = "Patient is a current smoker, ~1 pack/day, with worsening dyspnea."
res = openmed.analyze_text(note, model_name="disease_detection_superclinical",
                           output_format="dict")

# 2) Write a derived OBSERVATION (or a custom cohort attribute) per patient,
#    mapping each extracted entity to a standard concept (grounded out-of-process).
#    e.g. Observation: "Current smoker" -> a SNOMED concept in your vocab.

# 3) Reference that concept in a CIRCE ConceptSet, so the phenotype combines
#    structured codes AND the NLP-derived flag in one inclusion rule.

This mirrors how eMERGE and PheKB phenotypes mix structured codes with NLP: the NLP step contributes high-recall flags for concepts that ICD/CPT capture poorly, and CIRCE composes them with the rest of the logic.

Workflow

  1. Start from a library definition if one exists (OHDSI Phenotype Library / PheKB) and adapt; otherwise design entry event + inclusion rules.
  2. Build concept sets from standard OMOP concepts; set includeDescendants to capture hierarchies. Vocabulary content comes from your own licensed tables.
  3. Identify text-only criteria the codes miss; extract them with openmed.analyze_text and materialize as OMOP rows / cohort attributes.
  4. Assemble the CIRCE JSON (concept sets + expression) — in ATLAS or directly.
  5. Validate against OMOP CDM: generate SQL, run on a (synthetic/de-identified) database, review cohort counts; iterate with PheValuator-style checks.
  6. Document human-readable logic alongside the JSON for portability.

Hand-off to / from OpenMed

  • OpenMed → phenotype features. openmed.analyze_text over notes yields Disease, Pharmaceutical, Genomics, Oncology, and social/behavioral spans. Ground each to a standard concept (coding-icd10, normalizing-rxnorm, mapping-loinc, or your SNOMED map) and write it into OMOP so CIRCE can reference it.
  • Phenotype → OpenMed scope. A cohort definition tells you which notes to process: run OpenMed only on the cohort's documents to extract the features the phenotype needs, keeping compute and PHI exposure minimal.
  • Run locally on de-identified or synthetic OMOP data. De-identify notes with openmed.deidentify before they enter any shared analytics environment.

Edge cases & gotchas

  • Standard vs source concepts. OMOP maps source codes (ICD-10-CM) to standard concepts (usually SNOMED). Build concept sets on standard concepts and let the source-to-standard map do the translation, or you will miss rows.
  • Descendants matter. Forgetting includeDescendants silently drops the hierarchy (e.g. all diabetes subtypes). Forgetting nothing can over-capture — review the resolved concept list.
  • NLP feature provenance. Tag NLP-derived OMOP rows distinctly (e.g. a type_concept indicating "derived from NLP") so analysts know the signal is probabilistic, not adjudicated.
  • Vocabulary licensing. SNOMED CT, CPT4, and similar require their own licenses and are not redistributed here — load them into your OMOP vocab.
  • Portability ≠ equivalence. The same JSON runs everywhere, but data capture differs by site; validate cohort counts per source before trusting them.
  • Not clinical advice. Phenotype membership supports research/analytics; it is not a diagnosis.

Standards & references

Frequently asked questions

What does the Defining Cohort Phenotypes AI skill do?

Authors computable phenotype and cohort definitions in the OHDSI ATLAS / CIRCE style over the OMOP CDM, combining standard concept sets with NLP-derived features that OpenMed extracts. Use when the user wants to define a patient cohort, write a computable phenotype, reuse PheKB or OHDSI Phenotype Library logic, build concept sets, or augment code-based criteria with text features. Trigger keywords: phenotype, cohort definition, OHDSI, ATLAS, CIRCE, OMOP CDM, concept set, PheKB, Phenotype Library, eMERGE, computable phenotype. Pairs adjacent to OpenMed: NLP features from openmed.analyze_text...

Why use Defining Cohort Phenotypes on TypingMind?

Because you install it once and use it with any model. Defining Cohort Phenotypes is plain Markdown rather than provider-specific code, so the same skill runs on GPT-5, Claude, Gemini, Grok, or a local model — and you can switch model mid-chat without it breaking. TypingMind runs on your own API keys, so you pay providers directly instead of a per-seat subscription, and your skills and chats stay in your own storage.

How do I install Defining Cohort Phenotypes in TypingMind?

Open Plugins → Skills → Install from GitHub in TypingMind and paste https://github.com/maziyarpanahi/openmed/tree/master/skills/defining-cohort-phenotypes. TypingMind reads its SKILL.md and installs it as a skill you can enable per chat.

Which AI models can use Defining Cohort Phenotypes?

Any model you connect in TypingMind. AI skills are plain Markdown instructions rather than provider-specific code, so GPT, Claude, Gemini, Grok, and local models can all load this skill when a request matches it.

How many AI models can I use with Defining Cohort Phenotypes?

As many as you like. As long as a model supports skills, you can use Defining Cohort Phenotypes with it — GPT, Claude, Gemini, Grok, DeepSeek, Mistral, Llama and more — all on TypingMind with your own API keys.

Is the Defining Cohort Phenotypes AI skill free?

Yes. It is published on GitHub by maziyarpanahi under the Apache-2.0 license. You only pay your own AI provider for the tokens you use.

What are AI skills?

An AI skill is a reusable instruction bundle that teaches an AI model how to do one specific task. It follows the open Agent Skills format: a SKILL.md file with a name and description, plus any scripts, templates or reference files the model may need. The model reads the instructions only when your request matches the skill, so an installed skill costs nothing until it is used.

How are AI skills different from plugins or MCP servers?

A plugin or MCP server gives a model new tools to call — code that runs somewhere and returns a result. An AI skill gives the model knowledge and process instead: how to approach a task, which steps to follow, what good output looks like. Skills are plain Markdown, so they need no server, no API key and no runtime, and they work with any model.

View all

Set up your own AI workspace now

Get notified about new features and future giveaways by subscribing to our newsletter 👇