skills/de-ai-revise/SKILL.md
Revise prose to read less AI-generated using corpus-validated scorers. Use when the user asks to 'de-AI this', 'make it sound less like AI', 'remove AI-isms', 'de-tic this draft', 'humanize the prose', 'fix AI writing tells', or 'less AI-sounding'. Also the standard AI-prose pass inside /writing-review and /writing-revise.
npx skillsauth add edwinhu/workflows de-ai-reviseInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
A writing-improvement tool. It audits a draft with three corpus-validated scorers, then rewrites only the flagged spans so the prose reads less like an LLM wrote it — plainer diction, burstier rhythm, fewer machine tics — while leaving already-human passages untouched.
This is the GENERATION side of the AI-writing apparatus, not detection. Detecting polished AI was proven near-impossible (60%+ false-positive rates on real human writing); this skill never renders a verdict on authorship. It improves readability for a human reader. The scorers GUIDE which spans to revise; they are not a target to maximize.
This is a BACKSTOP, not the main event. The primary lever for human-reading prose is the GENERATION contract upstream — writing-draft now drafts topic-sentence-led and proportional (varied paragraph/sentence length), which is what produces human burstiness in the first place. A draft generated well needs little here. If de-ai-revise is finding a lot, the fix usually belongs upstream (the outline's POINTs aren't real topic sentences, or the draft padded uniformly), not in a heavy span-by-span rewrite here. Use this to catch residue, not to manufacture rhythm a flat draft never had.
<law> ## The Iron Law of GoodhartTHE SCORERS GUIDE; THEY DO NOT GRADE. NO EDIT THAT IMPROVES A NUMBER BUT NOT THE READING. This is not negotiable.
A human reads the output. Mechanically maxing burstiness (chop every sentence), nuking every em-dash, or swapping every flagged word degrades prose to win a composite — that is the failure this skill exists to prevent. Revise a span only when the rewrite reads better to a person. Leave a flagged span alone when the author's choice is the right one (see Preserve-Human below). </law>
scripts/de_ai_audit.py folds them into one line-anchored span list. Every signal
was gated against a 14.3M-sentence law+finance corpus, so flags are AI defaults
real scholars don't write — not generic "fancy word" lint.
| Scorer | Catches | Remedy |
|--------|---------|--------|
| Scored AI-tics (ai-anti-patterns/references/scored-tics-patterns.py) | phrase/structure tics that passed the ~0-human-rate gate (sev1-5) | rewrite the construction; these have no honest use |
| Tiered diction (references/diction.yaml) | fancy→plain words, tiered by corpus rate | always_flag → swap on sight; cluster → fix when 2+/para; density → vary at saturation; dropped → never touch (legal-normal) |
| Stylometrics (ai-anti-patterns/scripts/style_metrics.py) | rhythm/structure: composite_human_likeness 0-100, em-dash, metronomic runs, opener transitions, nominalization, burstiness/passive advisories | vary sentence length toward bursty; em-dash → semicolon/period; plainer Latinate→Anglo-Saxon |
| Mode | Trigger | Behavior |
|------|---------|----------|
| rewrite (default) | "de-AI this", "make it less AI" | audit → rewrite flagged spans → one corrective 2nd pass → return an edits-made + verification report (NOT the whole file) |
| detect-only | "just flag", "scan", "what AI tells are in this", "audit only" | audit only; report flagged spans + composite/tic-density; no edits |
| edit-in-place | "fix draft.md directly", "clean the file in place" | minimal targeted Edits to the file; preserve already-human paragraphs; re-audit after |
Default to rewrite when unspecified.
START
│
├─ Step 1: AUDIT — run de_ai_audit.py --json on the target
│ uv run --with pyyaml python3 ${CLAUDE_SKILL_DIR}/scripts/de_ai_audit.py --json <file>
│ Read: composite_human_likeness, tic_density, spans[], advisories[]
│
├─ detect-only? → report spans + signals, STOP.
│
├─ Step 2: REWRITE the flagged spans (NOT the whole draft)
│ - tic spans → rewrite the construction (no honest use)
│ - diction:always_flag → swap for the listed plain replacement
│ - diction:cluster → fix enough of the cluster to drop below 2/para
│ - style:em_dash → recast as semicolon / period / comma — but NOT all (see Preserve)
│ - advisories (burstiness) → vary sentence length where it reads flat; do NOT chop for chop's sake
│ PRESERVE already-human passages (no spans) untouched.
│ PRESERVE quoted material, block quotes, code, footnote citations.
│
├─ Step 3: ONE corrective 2nd pass
│ Re-run de_ai_audit.py. Fix spans the first pass introduced or missed.
│ STOP at 2 passes — a 3rd rarely finds more and costs a full regeneration.
│
└─ Step 4: REPORT (edits-made + verification), NOT the whole file
- what changed and why (span → before → after, grouped by scorer)
- before/after composite + tic-density (must improve or hold; if it dropped, you over-edited)
- spans deliberately LEFT (author's voice / quoted / domain term) and why
If text and flowchart disagree, the flowchart wins.
The composite penalizes em-dashes hard, and real legal scholarship — including this user's own published prose — uses them deliberately. Do NOT zero them out.
dropped-tier diction (significant, robust, leverage, comprehensive, …): NEVER
flag or swap — these are legal/finance-normal; the audit already excludes them.^[...] and markdown [^id]:
footnotes before scoring, so findings never land inside them (citation/legal-normal text). You
will not see footnote spans to triage; if you ever do, do not edit them. (--keep-footnotes
disables masking for debugging the raw signal only.)diction.yaml dropped tier exists because "significant/robust/leverage" fire on
every real law-review article; a linter that flags them is worse than none. The audit
omits them — if you hand-flag one anyway, you reintroduced the false positive.de_ai_audit.py on every draft as a standard audit step; its
always_flag + sev≥4 tic spans become AI-ism findings in REVIEW.md (advisory minors
unless they cluster into a major).development
Build the meeting-level proxy-voting × ownership panel on the WRDS SGE grid — ISS N-PX fund votes reduced to (item × block) direction cells, joined to institutional and mutual-fund ownership. Use when working with risk.voteanalysis_npx, N-PX fund-level votes, ISS→CRSP fund linking, index/passive/active voting blocks, or a proxy-voting panel that needs ownership attached.
development
Use when "CRSP CIZ", "CRSP v2", "CRSP flat file format 2.0", "crsp.dsf_v2 / msf_v2", "StkDlySecurityData", "StkMthSecurityData", "StkSecurityInfoHist", "stocknames_v2", "DlyRet / MthRet / DlyPrc / MthPrc", "SHRCD or EXCHCD equivalent in new CRSP", "SIZ to CIZ migration", "CRSP data after 2024", "CRSP delisting returns", "CRSP cumulative adjustment factors", "CRSP index INDNO / INDFAM", or any CRSP stock/index query where the legacy SIZ column names no longer exist.
development
Use when linking or deduping datasets by entity name rather than a shared key — 'fuzzy match', 'fuzzy name matching', 'entity resolution', 'record linkage', 'match company/person names', 'dedupe entity names', 'name-based join', 'bridge identifiers' (CIK ↔ permno ↔ gvkey ↔ wficn ↔ EIN ↔ personid), or any use of char n-gram TF-IDF, cosine similarity on names, `sparse_dot_topn`, or RapidFuzz at scale.
development
Use when building a publication-quality table in Python — 'regression table', 'results table', 'summary statistics table', 'etable', 'coefplot', 'great_tables', 'GT', 'gt table', 'format a table for the paper', 'export table to LaTeX/HTML', significance stars, spanners, or column formatting for a table headed into a paper, slide deck, or notebook.