bundled/skills/preprocessing-data-with-automated-pipelines/SKILL.md
Design and implement repeatable preprocessing pipelines for cleaning, encoding, transforming, and validating ML input data.
npx skillsauth add foryourhealth111-pixel/vco-skills-codex preprocessing-data-with-automated-pipelinesInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Use this skill as the direct owner for ML input-preparation pipelines.
It covers preprocessing-heavy tasks where the requested deliverable is a repeatable pipeline for cleaning, encoding, transforming, and validating input data.
Use this skill when:
scikit-learn or ml-pipeline-workflowml-data-leakage-guardscientific-data-preprocessingml-data-leakage-guard before trusting fitted preprocessing stepssplitting-datasets when the next narrow problem is partition strategydevelopment
Chunked N-D arrays for cloud storage. Compressed arrays, parallel I/O, S3/GCS integration, NumPy/Dask/Xarray compatible, for large-scale scientific computing pipelines.
tools
Use only when the user explicitly asks to stage, commit, push, and open a GitHub pull request in one flow using the GitHub CLI (`gh`).
tools
Spreadsheet toolkit (.xlsx/.csv). Create/edit with formulas/formatting, analyze data, visualization, recalculate formulas, for spreadsheet processing and analysis.
tools
High-performance CSV processing with xan CLI for large tabular datasets, streaming transformations, and low-memory pipelines.