openclaw-skills/markdown-tools/SKILL.md
Converts documents to markdown with multi-tool orchestration for best quality. Supports Quick Mode (fast, single tool) and Heavy Mode (best quality, multi-tool merge). Use when converting PDF/DOCX/PPTX files to markdown, extracting images from documents, validating conversion quality, or needing LLM-optimized document output.
npx skillsauth add seaworld008/commonly-used-high-value-skills markdown-toolsInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Convert documents to high-quality markdown with intelligent multi-tool orchestration.
Use this skill when the user wants to:
Recommended flow:
choose quick or heavy mode
-> convert with best-fit tool(s)
-> validate output quality
-> extract images if needed
-> merge or refine outputs
| Mode | Speed | Quality | Use Case | |------|-------|---------|----------| | Quick (default) | Fast | Good | Drafts, simple documents | | Heavy | Slower | Best | Final documents, complex layouts |
# Required: PDF/DOCX/PPTX support
uv tool install "markitdown[pdf]"
pip install pymupdf4llm
brew install pandoc
# Quick Mode (default) - fast, single best tool
uv run --with pymupdf4llm --with markitdown scripts/convert.py document.pdf -o output.md
# Heavy Mode - multi-tool parallel execution with merge
uv run --with pymupdf4llm --with markitdown scripts/convert.py document.pdf -o output.md --heavy
# Check available tools
uv run scripts/convert.py --list-tools
| Format | Quick Mode Tool | Heavy Mode Tools | |--------|----------------|------------------| | PDF | pymupdf4llm | pymupdf4llm + markitdown | | DOCX | pandoc | pandoc + markitdown | | PPTX | markitdown | markitdown + pandoc | | XLSX | markitdown | markitdown |
Heavy Mode runs multiple tools in parallel and selects the best segments:
| Segment Type | Selection Criteria | |--------------|-------------------| | Tables | More rows/columns, proper header separator | | Images | Alt text present, local paths preferred | | Headings | Proper hierarchy, appropriate length | | Lists | More items, nested structure preserved | | Paragraphs | Content completeness |
# Extract images with metadata
uv run --with pymupdf scripts/extract_pdf_images.py document.pdf -o ./assets
# Generate markdown references file
uv run --with pymupdf scripts/extract_pdf_images.py document.pdf --markdown refs.md
Output:
assets/img_page1_1.png, assets/img_page2_1.jpgassets/images_metadata.json (page, position, dimensions)# Validate conversion quality
uv run --with pymupdf scripts/validate_output.py document.pdf output.md
# Generate HTML report
uv run --with pymupdf scripts/validate_output.py document.pdf output.md --report report.html
| Metric | Pass | Warn | Fail | |--------|------|------|------| | Text Retention | >95% | 85-95% | <85% | | Table Retention | 100% | 90-99% | <90% | | Image Retention | 100% | 80-99% | <80% |
# Merge multiple markdown files
python scripts/merge_outputs.py output1.md output2.md -o merged.md
# Show segment attribution
python scripts/merge_outputs.py output1.md output2.md -o merged.md --verbose
# Windows → WSL conversion
python scripts/convert_path.py "C:\Users\name\Documents\file.pdf"
# Output: /mnt/c/Users/name/Documents/file.pdf
"No conversion tools available"
# Install all tools
pip install pymupdf4llm
uv tool install "markitdown[pdf]"
brew install pandoc
FontBBox warnings during PDF conversion
Images missing from output
scripts/extract_pdf_images.pyTables broken in output
scripts/validate_output.py| Script | Purpose |
|--------|---------|
| convert.py | Main orchestrator with Quick/Heavy mode |
| merge_outputs.py | Merge multiple markdown outputs |
| validate_output.py | Quality validation with HTML report |
| extract_pdf_images.py | PDF image extraction with metadata |
| convert_path.py | Windows to WSL path converter |
references/heavy-mode-guide.md - Detailed Heavy Mode documentationreferences/tool-comparison.md - Tool capabilities comparisonreferences/conversion-examples.md - Batch operation examplestools
飞书审批:查询和处理审批待办/已办/实例,搜索可发起审批定义、查看定义详情并发起原生审批实例。当用户要处理审批任务、查看审批实例、搜索或发起审批时使用。审批待办不是飞书任务;非审批类待办走 lark-task。不负责创建审批定义;三方审批定义不走原生提单。
development
Use when a user needs reproducible repository sizing, language composition, file counts, or code-versus-comment ratios with pygount; record exclusions and verify measurement scope before interpreting results.
development
Route a development task to the official Hermes Agent skill, Graphify Codex artifact set, Open GSD Core bundle, or optional GSD Pi bundle without duplicating their installers or state machines.
development
飞书 / Lark 通讯录:按姓名 / 邮箱解析成 open_id,或按 open_id 反查姓名 / 部门 / 邮箱 / 联系方式 / 个人状态 / 签名,以及按关键词搜索当前用户可见的机器人 / 智能体(agent)。当用户提到一个名字要下一步发消息 / 排日程,或拿到 open_id 想查具体信息时使用。不负责部门树遍历、按部门列员工、组织架构图,这类需求走原生 OpenAPI。