plugin/skills/tooluniverse-protein-structural-annotation-pdb/SKILL.md
Given a PDB structure, produce a per-residue annotation table: which residues sit at a binding interface (vs a partner chain), which line a ligand pocket, which are buried (core) vs solvent-exposed (surface), and optionally secondary structure. This is the structural track drawn under a DMS heatmap and the structural prior SAE feature drops are read against. Use when you need to anchor a variant-interpretation or DMS analysis to the protein's actual physical context.
npx skillsauth add mims-harvard/tooluniverse tooluniverse-protein-structural-annotation-pdbInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
For each residue of a target protein chain, classify whether it sits at a binding interface, in a ligand pocket, is buried vs solvent-exposed, and (optionally) which secondary-structure element it belongs to. This is the annotation track that anchors any DMS heatmap or per-residue interpretation to the protein's actual physical context.
Not for:
tooluniverse-computational-biophysics| Input | Format | Example |
|---|---|---|
| PDB ID | 4 characters | 6VJJ (KRAS-RAF1-GTP analogue) |
| Target chain | single character | A |
| Partner chain(s) | list of chain IDs | ["B"] |
| Ligand resnames | 3-letter PDB names | ["GNP", "MG"] |
Optional:
distance_cutoff (default 5.0 Å)core_rsa_cutoff (default 0.25)include_secondary_structure (default false; uses PDBe REST if true)pdb_content instead of pdb_id for local / predicted structuresIf you only have a UniProt accession or gene symbol, pick a structure first:
# PDBe's curated UniProt→PDB mapping (recommended; ranks by coverage + resolution)
PDBeSIFTS_get_best_structures(uniprot_accession="P01116")
# Returns a ranked list of PDB IDs for KRAS with chain mapping
# Or full list (unranked)
PDBeSIFTS_get_all_structures(uniprot_accession="P01116")
# RCSB advanced search (free-text, when you don't have a UniProt yet)
RCSBAdvSearch_search_structures(query="KRAS GTP complex")
Pick the structure that contains the right complex: include the binding partner chain you care about, the relevant ligand, and a resolution adequate for distance-based classification (≤ 3 Å is a safe default).
Structure_annotate_per_residue(
pdb_id="6VJJ",
target_chain="A",
partner_chains=["B"],
ligand_resnames=["GNP", "MG"],
distance_cutoff=5.0,
core_rsa_cutoff=0.25,
include_secondary_structure=False,
)
Returns annotations: List[{position, aa, dist_partner, dist_ligand, rsa, region, is_core, ss_element?}] for every residue of the target chain. For
KRAS in 6VJJ, this yields 168 rows.
PDB residue numbers carry silent offsets — crystal constructs add N-terminal cloning residues, and published figures sometimes shift the track relative to the panel sequence. Always verify with a landmark:
# Get the canonical reference sequence
UniProt_get_sequence_by_accession(accession="P01116")
# Then spot-check: KRAS canonical position 12 should be glycine
assert annotations[11]["aa"] == "G" # 1-indexed position 12, 0-indexed index 11
If the landmark mismatches, record the offset explicitly (e.g. pdb_pos = uniprot_pos + offset) before any downstream join. Do not silently rebase
positions.
If you set include_secondary_structure=True, the tool fetches per-residue
helix/strand/coil from PDBe REST. Alternatively, use the dedicated PDBe
secondary-structure tool separately:
pdbe_get_entry_secondary_structure(pdb_id="6VJJ")
# Returns per-chain helix + strand ranges
The returned table is keyed by 1-based canonical residue number. Typical downstream uses:
| Use case | Field to read |
|---|---|
| Is variant X in a pocket? | by_pos = {a["position"]: a for a in annotations}; by_pos[X]["region"] in ("ligand", "both") — index by position field, NOT list index (PDB residue numbers may not start at 1 or be contiguous) |
| Build a DMS heatmap annotation track | [(r["position"], r["region"], r["is_core"], r.get("ss_element"))] |
| Filter SAE hotspot features to ligand-binding residues | filter clusters by region == "ligand" |
| Compare buried vs surface signal | group statistics by is_core |
| Region label | Biological meaning | Common functional role |
|---|---|---|
| interface | Within distance_cutoff of a partner chain | Protein-protein binding residue; variants often disrupt complex formation |
| ligand | Within distance_cutoff of a ligand heavy atom | Pocket residue; variants often disrupt substrate / cofactor / drug binding |
| both | Both | Allosteric or shared-surface residue |
| other | Neither | Surface (if not is_core) or core (if is_core) — variants impact through stability or distal effects |
| is_core=true | RSA < core_rsa_cutoff (0.25 by default) | Buried residue; variants often destabilize the fold |
partner_chains=[] is permitted but then all dist_partner values
are null — interface analysis is skipped entirely.| Tool | Role | Use it for |
|---|---|---|
| Structure_annotate_per_residue | This skill's atomic tool | The annotation itself |
| PDBeSIFTS_get_best_structures | UniProt → ranked PDB list | Step 1 |
| PDBeSIFTS_get_all_structures | UniProt → full PDB list | Step 1 |
| RCSBAdvSearch_search_structures | Free-text RCSB search | Step 1 |
| UniProt_get_sequence_by_accession | Canonical sequence | Step 3 (numbering verification) |
| pdbe_get_entry_secondary_structure | SS alone | Step 4 alternative |
| tooluniverse-residue-functional-mechanism-interpretation | Downstream consumer | Use this annotation as the structural evidence layer when interpreting DMS hotspots; the skill also plots an annotated DMS heatmap in its Step 7 |
tools
Generate the success criteria for a task or question, then review work against them. Given a task, goal, or open-ended question, decompose it into scenarios, evaluation perspectives, and fine-grained weighted YES/NO criteria using the Recursive Expansion Tree (RET) method; if work is supplied, score it criterion-by-criterion and surface what is missing or could be better. Use when asked to self-review or check your own work, judge whether a task is done well or completely, build a definition-of-done or completeness checklist, create an evaluation rubric or grading criteria, score or grade answers to a question, set up an LLM-as-judge rubric, or when the user mentions self-review, completeness check, success criteria, evaluation criteria, scoring rubric, Qworld, or the RET algorithm.
tools
Find the real protein target(s) of a peptide from its sequence — peptide target deorphanization / off-target identification, for ANY target class (GPCR, ion channel, protease, cytokine/growth-factor receptor, enzyme, integrin), not only GPCRs. Use when a peptide has a phenotype but does not bind its hypothesized target, when a peptide binds a target in one species or assay but not another, or to screen candidate targets for an orphan peptide. A target-class router steers a multi-route keyless pipeline (PROSITE/ELM motif, BLAST homology, HGNC/InterPro/GPCRdb/GtoPdb target-family enumeration, OpenTargets phenotype anchor, EnsemblCompara/Alliance cross-species reconciliation) plus optional NVIDIA-NIM co-folding (Boltz2, AlphaFold2-Multimer, OpenFold3) for structural confirmation.
tools
Install or update ToolUniverse in Claude Science — create the conda env, install the tooluniverse pip package, and (re)build the tooluniverse-research skill by fetching the current workflow library from GitHub. Use for first-time setup, upgrading the ToolUniverse version, refreshing the bundled workflows after an upstream release, or reinstalling on a new machine.
tools
Install, set up, verify, update, pin, uninstall, or troubleshoot the ToolUniverse plugin on OpenAI Codex. ALWAYS consult this skill for any of those — don't answer from memory, because the exact marketplace name (mims-harvard/ToolUniverse), the "codex plugin marketplace add" then "codex plugin add -m tooluniverse" flow, Codex's startup auto-upgrade behavior, the uvx tooluniverse MCP server, and the API-key env vars are easy to get wrong. Use it whenever someone wants to get ToolUniverse (or "the 1000+ scientific tools" / "the harvard tools") working on Codex, says the Codex plugin or its tools/skills won't load, hits a uvx or MCP-server startup error, asks how Codex updates it, wants to pin or remove it, or finds it running an old tool version — even if they never say the word "plugin". Not for the Claude Code plugin (use tooluniverse-claude-code-plugin), for running research with the tools, or for authoring new tools or skills.