bundled/skills/scrapling/SKILL.md
CLI-first web scraping & content extraction with optional MCP server. Use when you have target URLs and need clean, selector-based outputs (html/md/txt).
npx skillsauth add foryourhealth111-pixel/vco-skills-codex scraplingInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Scrapling is a Python-based web scraping / extraction toolkit that exposes:
scrapling ...) for fetching + extracting content into filesscrapling mcp) so an agent can call structured scraping toolsThis skill is CLI-first. Prefer it when you already have URLs and need reliable, repeatable extraction (CSS selector → file).
Use scrapling when you need:
.txt / .md / .htmlplaywrightscrapling: best for “get URL → extract selector → write file” workflows; simpler, faster iterationplaywright: best for interactive UI flows (login, multi-step navigation, downloads, complex JS actions, stateful sessions)If you must navigate or click through a UI, use playwright.
If you can directly fetch the target page and just need extraction, use scrapling.
scrapling is for acquisition + extraction once you already know the URL(s).A common pipeline:
python --version
scrapling --help
Scrapling’s CLI and MCP features are enabled via extras.
Recommended (CLI + MCP + fetchers):
python -m pip install "scrapling[ai]"
If you only want CLI fetch/extract without MCP:
python -m pip install "scrapling[fetchers]"
If you use browser-based fetchers, you may need browser binaries:
# Option A: via Scrapling helper (after install)
scrapling install
# Option B: directly via Playwright
python -m playwright install
This skill ships a thin PowerShell wrapper:
C:/Users/羽裳/.codex/skills/scrapling/scripts/scrapling.ps1It checks whether scrapling exists and prints install hints if missing.
scrapling extract get "https://example.com" out.md
scrapling extract get "https://example.com" out.txt --css-selector "main article"
scrapling extract get "https://example.com" out.html --css-selector "#content"
scrapling extract fetch "https://example.com" out.md --css-selector "main"
Tip: keep outputs in files and only feed the smallest relevant snippet to the LLM.
Scrapling can run as an MCP server. This is useful when:
Start MCP server (stdio transport by default):
scrapling mcp
Optional: run MCP server with HTTP transport:
scrapling mcp --http --host 127.0.0.1 --port 8765
{
"servers": {
"scrapling": {
"mode": "stdio",
"command": "scrapling",
"args": ["mcp"],
"required": false,
"note": "Requires: python -m pip install \"scrapling[ai]\""
}
}
}
playwright.development
Chunked N-D arrays for cloud storage. Compressed arrays, parallel I/O, S3/GCS integration, NumPy/Dask/Xarray compatible, for large-scale scientific computing pipelines.
tools
Use only when the user explicitly asks to stage, commit, push, and open a GitHub pull request in one flow using the GitHub CLI (`gh`).
tools
Spreadsheet toolkit (.xlsx/.csv). Create/edit with formulas/formatting, analyze data, visualization, recalculate formulas, for spreadsheet processing and analysis.
tools
High-performance CSV processing with xan CLI for large tabular datasets, streaming transformations, and low-memory pipelines.