Adoption

Agent Skills are supported by leading AI development tools.

VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory

mims-harvard/tooluniverse-dataset-discovery

Name: tooluniverse-dataset-discovery
Author: mims-harvard

skills/tooluniverse-dataset-discovery/SKILL.md

npx skillsauth add mims-harvard/tooluniverse tooluniverse-dataset-discovery

Clean

TrivyContainer and dependency vulnerability scanner

Clean

SemgrepStatic code analysis for vulnerabilities

Clean

mcp-scan (Snyk)Model Context Protocol security validation

Skipped

Snyk (dep)Open source security scanning

Skipped

Socket.devSupply chain security analysis

Skipped

VirusTotalMulti-engine malware detection

Skipped

CrowdStrikeAdvanced threat intelligence

Skipped

OSV-ScannerOpen Source Vulnerability database check

Skipped

OWASP Dep-Check

Dataset Discovery

When to Use

User asks "find me data about X" or "where can I get data on Y"
User wants to analyze a relationship between variables
User needs specific study designs (longitudinal, cross-sectional, experimental)
User asks about specific surveys or cohorts

Step 1: Understand What the Research Question Requires

Before searching, determine the minimum data requirements:

Study design needed:

"Does X predict CHANGES in Y over time?" → longitudinal (same people measured repeatedly). Cross-sectional data CANNOT answer this — don't settle for it.
"Is X associated with Y?" → cross-sectional is sufficient (one-time measurement)
"Does intervention X cause outcome Y?" → experimental (clinical trial with controls)
"What genes/proteins are involved in X?" → omics (sequencing, expression, proteomics)

Variables needed:

List the specific exposure, outcome, and confounder variables
For each variable, note the measurement type (continuous, categorical, biomarker vs self-report)
Identify minimum confounders needed (age, sex are almost always required; domain-specific confounders depend on the question)

Population needed:

Age range, geography, clinical status, sample size requirements
Power analysis: to detect a small effect (r=0.1), you need ~800 subjects at 80% power

Step 2: Search Strategy

Search from broadest to most specific. Use find_tools to discover available dataset search tools — don't rely on memorized tool names.

Layer 1 — Cross-repository search (cast wide net): Search tools that index datasets across thousands of repositories. These find datasets you didn't know existed.

Search by: research topic keywords, variable names, population descriptors
Look for: DOI-registered datasets, repository listings, government data portals

Layer 2 — Domain-specific repositories: Search repositories specialized for your data type.

Health surveys: CDC, NHANES (search by variable name, not topic keywords)
Genomics: SRA, ENA, ArrayExpress, GEO
Proteomics: PRIDE, MassIVE
Metabolomics: MetaboLights, Metabolomics Workbench
Clinical: ClinicalTrials.gov (for trial data with results)

Layer 3 — Literature-based discovery: Many datasets aren't in any repository — they're described in paper methods sections.

Search PubMed/EuropePMC for papers that analyzed the relationship you're interested in
Read their methods: "We used data from [DATASET NAME]" tells you exactly what exists
Check supplementary materials for deposited data (GEO/SRA accession numbers)
This is often the MOST effective strategy for finding niche datasets

Step 3: Evaluate Dataset Fitness

For each candidate dataset, assess these dimensions:

Variables:

Does it contain your SPECIFIC exposure and outcome variables?
Are they measured the way you need? (biomarker vs self-report, continuous vs categorical)
Are key confounders available? (missing confounders = biased analysis)

Design match:

If you need longitudinal: does it follow the SAME individuals over time? How many waves? What's the follow-up interval?
Beware: "repeated cross-sections" (different people each wave) are NOT longitudinal
If you need experimental: is there a proper control group? Randomization?

Sample:

Is the sample large enough for your analysis? (logistic regression needs ~10 events per predictor)
Does the population match? (age range, geography, clinical characteristics)
Are there subgroups you need? (stratified by sex, race, disease status)

Access:

Publicly downloadable (best) vs registration required (days) vs collaboration agreement (months) vs restricted (may be impossible)
Data format: CSV/TSV (easy), XPT/SAS (need conversion), proprietary database (may need special software)

Quality:

Is it from a well-known study with published methods? (NHANES, HRS, UK Biobank = high quality)
Has it been used in peer-reviewed publications? (indicates data is usable)
What's the response rate / missingness pattern?

Step 4: Download and Analyze

Don't stop at finding datasets — download and analyze them. Write and run Python code via Bash. Never describe what you "would do" — execute it.

Data Loading Cookbook

Choose the loader that matches your data source. When unsure of the format, download a small sample first and inspect.

import requests, io, pandas as pd

# --- Tabular files (most common) ---
df = pd.read_csv("data.csv")                                # CSV / TSV (use sep="\t" for TSV)
df = pd.read_excel("data.xlsx")                              # Excel
df = pd.read_stata("data.dta")                               # Stata
df = pd.read_sas("data.xpt", format="xport")                # SAS transport (XPT)
df = pd.read_sas("data.sas7bdat", format="sas7bdat")        # SAS native
df = pd.read_parquet("data.parquet")                         # Parquet
df = pd.read_json("data.json")                               # JSON (records or columnar)
df = pd.read_fwf("data.dat")                                 # Fixed-width (some legacy surveys)

# --- Download from URL first, then parse ---
resp = requests.get(url, timeout=120)
content = resp.content
# Detect format from URL or content header
if url.endswith(".XPT") or url.endswith(".xpt"):
    df = pd.read_sas(io.BytesIO(content), format="xport")
elif url.endswith(".csv") or url.endswith(".csv.gz"):
    df = pd.read_csv(io.BytesIO(content))
elif url.endswith(".tsv") or url.endswith(".tsv.gz"):
    df = pd.read_csv(io.BytesIO(content), sep="\t")
elif url.endswith(".json"):
    df = pd.read_json(io.BytesIO(content))
else:
    # Try CSV first, then inspect
    df = pd.read_csv(io.BytesIO(content))

# --- REST API pagination (common for GDC, ClinicalTrials.gov, etc.) ---
import json
all_records = []
offset = 0
while True:
    resp = requests.get(f"{api_url}?offset={offset}&limit=100", timeout=30)
    batch = resp.json().get("data", [])
    if not batch:
        break
    all_records.extend(batch)
    offset += len(batch)
df = pd.DataFrame(all_records)

Merge, Clean, Analyze

# Merge multiple files on participant/sample ID
merged = df1.merge(df2, on="id_col", how="inner")

# Filter population
subset = merged[(merged["age"] >= 60) & (merged["age"] <= 80)].copy()

# Handle missing values
missing_pct = subset.isnull().mean() * 100
print("Missing % per variable:\n", missing_pct[missing_pct > 0].sort_values(ascending=False))
subset = subset.dropna(subset=["exposure_var", "outcome_var"])

# Quick regression
import statsmodels.formula.api as smf
model = smf.ols("outcome ~ exposure + age + sex", data=subset).fit()
print(model.summary())

# Visualization
import matplotlib.pyplot as plt
plt.scatter(subset["exposure"], subset["outcome"], alpha=0.3)
plt.xlabel("Exposure"); plt.ylabel("Outcome")
plt.savefig("/tmp/scatter.png", dpi=150, bbox_inches="tight")

Always run the code and report actual numbers (β, p-value, CI, N).

Step 5: Report Honestly

Structure the report as:

Best available dataset — name, what it contains, access method, key limitation
Analysis results — actual statistics (β, p-value, CI, N) from running the code
Alternative datasets — ranked by fitness, with tradeoffs
What CANNOT be answered — if no dataset matches the study design needed, say so clearly
Recommended next steps — apply for access to longitudinal data, replicate in other cohorts

Critical honesty rules:

Never claim a dataset answers a temporal question if it's cross-sectional
Distinguish "data exists but needs registration" from "data doesn't exist"
Report actual computed statistics, not hypothetical analyses
State the strongest analysis possible with available data, even if it's weaker than what was asked

LOOK UP, DON'T GUESS

Never assume a dataset exists — search for it. Never assume access is public — check. Never assume variables are measured the way you need — verify the codebook.

mims-harvard/tooluniverse-dataset-discovery

skills/tooluniverse-dataset-discovery/SKILL.md

Find and evaluate research datasets for any scientific question. Maps research questions to required study designs (longitudinal vs cross-sectional, observational vs experimental, single-cohort vs multi-cohort). Use when the user asks 'find data about X', 'where can I get data on Y', or needs a specific cohort/survey/repository. Covers GEO, ArrayExpress, dbGaP, NHANES, UK Biobank, ClinicalTrials.gov, GWAS Catalog, and 30+ scientific repositories.

1,384 stars

tools

Updated May 28, 2026

$ install --global

skillsauth

npx skillsauth add mims-harvard/tooluniverse tooluniverse-dataset-discovery

Install this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.

Security Scan Results

3 of 9 scanners reported clean

Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.

Scanners Passed

Scanners in report

Clean

TrivyContainer and dependency vulnerability scanner

95%

Clean

SemgrepStatic code analysis for vulnerabilities

95%

Clean

mcp-scan (Snyk)Model Context Protocol security validation

95%

Skipped

Snyk (dep)Open source security scanning

50%

Skipped

Socket.devSupply chain security analysis

50%

Skipped

VirusTotalMulti-engine malware detection

50%

Skipped

CrowdStrikeAdvanced threat intelligence

50%

Skipped

OSV-ScannerOpen Source Vulnerability database check

50%

Skipped

OWASP Dep-Check

50%

Last scanned: May 22, 2026, 6:16 AM142.6s1 file scanned

SKILL.md

name:: tooluniverse-dataset-discovery
description:: Find and evaluate research datasets for any scientific question. Maps research questions to required study designs (longitudinal vs cross-sectional, observational vs experimental, single-cohort vs multi-cohort). Use when the user asks 'find data about X', 'where can I get data on Y', or needs a specific cohort/survey/repository. Covers GEO, ArrayExpress, dbGaP, NHANES, UK Biobank, ClinicalTrials.gov, GWAS Catalog, and 30+ scientific repositories.
disable-model-invocation:: true

Dataset Discovery

When to Use

User asks "find me data about X" or "where can I get data on Y"
User wants to analyze a relationship between variables
User needs specific study designs (longitudinal, cross-sectional, experimental)
User asks about specific surveys or cohorts

Step 1: Understand What the Research Question Requires

Before searching, determine the minimum data requirements:

Study design needed:

"Does X predict CHANGES in Y over time?" → longitudinal (same people measured repeatedly). Cross-sectional data CANNOT answer this — don't settle for it.
"Is X associated with Y?" → cross-sectional is sufficient (one-time measurement)
"Does intervention X cause outcome Y?" → experimental (clinical trial with controls)
"What genes/proteins are involved in X?" → omics (sequencing, expression, proteomics)

Variables needed:

List the specific exposure, outcome, and confounder variables
For each variable, note the measurement type (continuous, categorical, biomarker vs self-report)
Identify minimum confounders needed (age, sex are almost always required; domain-specific confounders depend on the question)

Population needed:

Age range, geography, clinical status, sample size requirements
Power analysis: to detect a small effect (r=0.1), you need ~800 subjects at 80% power

Step 2: Search Strategy

Search from broadest to most specific. Use find_tools to discover available dataset search tools — don't rely on memorized tool names.

Layer 1 — Cross-repository search (cast wide net): Search tools that index datasets across thousands of repositories. These find datasets you didn't know existed.

Search by: research topic keywords, variable names, population descriptors
Look for: DOI-registered datasets, repository listings, government data portals

Layer 2 — Domain-specific repositories: Search repositories specialized for your data type.

Health surveys: CDC, NHANES (search by variable name, not topic keywords)
Genomics: SRA, ENA, ArrayExpress, GEO
Proteomics: PRIDE, MassIVE
Metabolomics: MetaboLights, Metabolomics Workbench
Clinical: ClinicalTrials.gov (for trial data with results)

Layer 3 — Literature-based discovery: Many datasets aren't in any repository — they're described in paper methods sections.

Search PubMed/EuropePMC for papers that analyzed the relationship you're interested in
Read their methods: "We used data from [DATASET NAME]" tells you exactly what exists
Check supplementary materials for deposited data (GEO/SRA accession numbers)
This is often the MOST effective strategy for finding niche datasets

Step 3: Evaluate Dataset Fitness

For each candidate dataset, assess these dimensions:

Variables:

Does it contain your SPECIFIC exposure and outcome variables?
Are they measured the way you need? (biomarker vs self-report, continuous vs categorical)
Are key confounders available? (missing confounders = biased analysis)

Design match:

If you need longitudinal: does it follow the SAME individuals over time? How many waves? What's the follow-up interval?
Beware: "repeated cross-sections" (different people each wave) are NOT longitudinal
If you need experimental: is there a proper control group? Randomization?

Sample:

Is the sample large enough for your analysis? (logistic regression needs ~10 events per predictor)
Does the population match? (age range, geography, clinical characteristics)
Are there subgroups you need? (stratified by sex, race, disease status)

Access:

Publicly downloadable (best) vs registration required (days) vs collaboration agreement (months) vs restricted (may be impossible)
Data format: CSV/TSV (easy), XPT/SAS (need conversion), proprietary database (may need special software)

Quality:

Is it from a well-known study with published methods? (NHANES, HRS, UK Biobank = high quality)
Has it been used in peer-reviewed publications? (indicates data is usable)
What's the response rate / missingness pattern?

Step 4: Download and Analyze

Don't stop at finding datasets — download and analyze them. Write and run Python code via Bash. Never describe what you "would do" — execute it.

Data Loading Cookbook

Choose the loader that matches your data source. When unsure of the format, download a small sample first and inspect.

import requests, io, pandas as pd

# --- Tabular files (most common) ---
df = pd.read_csv("data.csv")                                # CSV / TSV (use sep="\t" for TSV)
df = pd.read_excel("data.xlsx")                              # Excel
df = pd.read_stata("data.dta")                               # Stata
df = pd.read_sas("data.xpt", format="xport")                # SAS transport (XPT)
df = pd.read_sas("data.sas7bdat", format="sas7bdat")        # SAS native
df = pd.read_parquet("data.parquet")                         # Parquet
df = pd.read_json("data.json")                               # JSON (records or columnar)
df = pd.read_fwf("data.dat")                                 # Fixed-width (some legacy surveys)

# --- Download from URL first, then parse ---
resp = requests.get(url, timeout=120)
content = resp.content
# Detect format from URL or content header
if url.endswith(".XPT") or url.endswith(".xpt"):
    df = pd.read_sas(io.BytesIO(content), format="xport")
elif url.endswith(".csv") or url.endswith(".csv.gz"):
    df = pd.read_csv(io.BytesIO(content))
elif url.endswith(".tsv") or url.endswith(".tsv.gz"):
    df = pd.read_csv(io.BytesIO(content), sep="\t")
elif url.endswith(".json"):
    df = pd.read_json(io.BytesIO(content))
else:
    # Try CSV first, then inspect
    df = pd.read_csv(io.BytesIO(content))

# --- REST API pagination (common for GDC, ClinicalTrials.gov, etc.) ---
import json
all_records = []
offset = 0
while True:
    resp = requests.get(f"{api_url}?offset={offset}&limit=100", timeout=30)
    batch = resp.json().get("data", [])
    if not batch:
        break
    all_records.extend(batch)
    offset += len(batch)
df = pd.DataFrame(all_records)

Merge, Clean, Analyze

# Merge multiple files on participant/sample ID
merged = df1.merge(df2, on="id_col", how="inner")

# Filter population
subset = merged[(merged["age"] >= 60) & (merged["age"] <= 80)].copy()

# Handle missing values
missing_pct = subset.isnull().mean() * 100
print("Missing % per variable:\n", missing_pct[missing_pct > 0].sort_values(ascending=False))
subset = subset.dropna(subset=["exposure_var", "outcome_var"])

# Quick regression
import statsmodels.formula.api as smf
model = smf.ols("outcome ~ exposure + age + sex", data=subset).fit()
print(model.summary())

# Visualization
import matplotlib.pyplot as plt
plt.scatter(subset["exposure"], subset["outcome"], alpha=0.3)
plt.xlabel("Exposure"); plt.ylabel("Outcome")
plt.savefig("/tmp/scatter.png", dpi=150, bbox_inches="tight")

Always run the code and report actual numbers (β, p-value, CI, N).

Step 5: Report Honestly

Structure the report as:

Best available dataset — name, what it contains, access method, key limitation
Analysis results — actual statistics (β, p-value, CI, N) from running the code
Alternative datasets — ranked by fitness, with tradeoffs
What CANNOT be answered — if no dataset matches the study design needed, say so clearly
Recommended next steps — apply for access to longitudinal data, replicate in other cohorts

Critical honesty rules:

Never claim a dataset answers a temporal question if it's cross-sectional
Distinguish "data exists but needs registration" from "data doesn't exist"
Report actual computed statistics, not hypothetical analyses
State the strongest analysis possible with available data, even if it's weaker than what was asked

LOOK UP, DON'T GUESS

Never assume a dataset exists — search for it. Never assume access is public — check. Never assume variables are measured the way you need — verify the codebook.

Related Skills

mims-harvard/tooluniverse-self-review

tools

VerifiedTrustedCommunity

Generate the success criteria for a task or question, then review work against them. Given a task, goal, or open-ended question, decompose it into scenarios, evaluation perspectives, and fine-grained weighted YES/NO criteria using the Recursive Expansion Tree (RET) method; if work is supplied, score it criterion-by-criterion and surface what is missing or could be better. Use when asked to self-review or check your own work, judge whether a task is done well or completely, build a definition-of-done or completeness checklist, create an evaluation rubric or grading criteria, score or grade answers to a question, set up an LLM-as-judge rubric, or when the user mentions self-review, completeness check, success criteria, evaluation criteria, scoring rubric, Qworld, or the RET algorithm.

1,583SKILL.mdUpdated Jul 22, 2026

mims-harvard/tooluniverse-self-review

mims-harvard/tooluniverse-peptide-target-deorphanization

tools

VerifiedTrustedCommunity

Find the real protein target(s) of a peptide from its sequence — peptide target deorphanization / off-target identification, for ANY target class (GPCR, ion channel, protease, cytokine/growth-factor receptor, enzyme, integrin), not only GPCRs. Use when a peptide has a phenotype but does not bind its hypothesized target, when a peptide binds a target in one species or assay but not another, or to screen candidate targets for an orphan peptide. A target-class router steers a multi-route keyless pipeline (PROSITE/ELM motif, BLAST homology, HGNC/InterPro/GPCRdb/GtoPdb target-family enumeration, OpenTargets phenotype anchor, EnsemblCompara/Alliance cross-species reconciliation) plus optional NVIDIA-NIM co-folding (Boltz2, AlphaFold2-Multimer, OpenFold3) for structural confirmation.

1,583SKILL.mdUpdated Jul 22, 2026

mims-harvard/tooluniverse-peptide-target-deorphanization

mims-harvard/tooluniverse-cs-setup

tools

VerifiedTrustedCommunity

Install or update ToolUniverse in Claude Science — create the conda env, install the tooluniverse pip package, and (re)build the tooluniverse-research skill by fetching the current workflow library from GitHub. Use for first-time setup, upgrading the ToolUniverse version, refreshing the bundled workflows after an upstream release, or reinstalling on a new machine.

1,583SKILL.mdUpdated Jul 22, 2026

mims-harvard/tooluniverse-cs-setup

mims-harvard/tooluniverse-codex-plugin

tools

VerifiedTrustedCommunity

Install, set up, verify, update, pin, uninstall, or troubleshoot the ToolUniverse plugin on OpenAI Codex. ALWAYS consult this skill for any of those — don't answer from memory, because the exact marketplace name (mims-harvard/ToolUniverse), the "codex plugin marketplace add" then "codex plugin add -m tooluniverse" flow, Codex's startup auto-upgrade behavior, the uvx tooluniverse MCP server, and the API-key env vars are easy to get wrong. Use it whenever someone wants to get ToolUniverse (or "the 1000+ scientific tools" / "the harvard tools") working on Codex, says the Codex plugin or its tools/skills won't load, hits a uvx or MCP-server startup error, asks how Codex updates it, wants to pin or remove it, or finds it running an old tool version — even if they never say the word "plugin". Not for the Claude Code plugin (use tooluniverse-claude-code-plugin), for running research with the tools, or for authoring new tools or skills.

1,583SKILL.mdUpdated Jul 22, 2026

mims-harvard/tooluniverse-codex-plugin

Download

For Claude Desktop. Download once, then upload the file in the app — no terminal needed.

Need help? View full Cowork setup guide →

Install manually

Choose your platform

# Clone the repo
git clone https://github.com/mims-harvard/tooluniverse.git

# Copy into Claude Code skills folder (global)
cp -r tooluniverse/skills/tooluniverse-dataset-discovery ~/.claude/skills/

Claude Code Skills — official skills path docs.

Repository

mims-harvard/tooluniverse

1,384 stars

Compatible with

Claude Code

OpenAI Codex CLI

ChatGPT