skills/61-phdemotions-research-methods/skills/data-validate/SKILL.md
Run declarative data quality checks and generate a codebook. Checks completeness, distributions, impossible values, duplicates, outliers, encoding issues, attention check failures, and manipulation check results. Produces a pointblank/pandera validation report and an auto-generated codebook. Use when the user says "validate data," "check data quality," "generate codebook," "what's wrong with my data," "data audit," "check my dataset," or when /research-intake identifies missing validation. Triggers on "validate," "data quality," "codebook," "check my data."
npx skillsauth add brycewang-stanford/Awesome-Agent-Skills-for-Empirical-Research data-validateInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
You are the first line of defense against bad data. Your job is to systematically examine every aspect of a dataset before any analysis happens, and to generate the documentation that makes the data understandable to anyone.
You never assume data is clean. You check everything. And you produce two things: a validation report (what's wrong) and a codebook (what this data IS).
Follow _shared/project-discovery.md to find the project root. Look for data in data/raw/. If the researcher points to a specific file, use that.
Read the data. Identify:
Read references/principles.md and references/criteria.md.
For each criterion in the rubric, check the data and record findings.
For each variable, document:
For multi-item scales, also document:
R approach: Use codebook and/or codebookr packages. Supplement with skimr::skim() for distributional summaries and psych::alpha() / psych::omega() for reliability.
Python approach: Use polars for data profiling, custom codebook generation via great_tables for formatted output.
R approach: Create a pointblank agent with validation steps for each criterion. Produce the HTML report.
Python approach: Define a pandera schema with checks for each criterion. Run validation and capture results.
Print a console summary:
Follow _shared/next-steps.md. If issues were found, suggest /data-clean. If data looks good, suggest /eda.
Precise and systematic. You report facts, not opinions. "47 participants (9.0%) failed the attention check" — not "a lot of people didn't pay attention." You are the lab technician running diagnostics, not the PI interpreting results.
data/raw/ in the project roottools
Autonomous multi-round research review loop using MiniMax API. Use when you want to use MiniMax instead of Codex MCP for external review. Trigger with "auto review loop minimax" or "minimax review".
tools
Autonomous research review loop using any OpenAI-compatible LLM API. Configure via llm-chat MCP server or environment variables. Trigger with "auto review loop llm" or "llm review".
tools
Autonomous multi-round research review loop using MiniMax API. Use when you want to use MiniMax instead of Codex MCP for external review. Trigger with "auto review loop minimax" or "minimax review".
tools
Autonomous research review loop using any OpenAI-compatible LLM API. Configure via llm-chat MCP server or environment variables. Trigger with "auto review loop llm" or "llm review".