offensive-tools/recon/katana/SKILL.md
Auth/lab ref: ProjectDiscovery web crawler for endpoint and JS-endpoint discovery.
npx skillsauth add aeondave/malskill katanaInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
ProjectDiscovery web crawler — endpoints, JS paths, XHR calls, and API routes from modern web apps.
# Single URL crawl
katana -u https://target.com
# Crawl with JS parsing (extract endpoints from JS files)
katana -u https://target.com -jc
# Headless mode (renders JS — required for SPA/React/Angular apps)
katana -u https://target.com -hl
# Output to file
katana -u https://target.com -jc -o endpoints.txt
# Multiple targets from file
katana -list urls.txt -jc -o all_endpoints.txt
# Limit to same domain (default behavior — no external crawl)
katana -u https://target.com -jc
# Include subdomains
katana -u https://target.com -jc -cs target.com
# Depth control
katana -u https://target.com -jc -d 5 # max depth 5 (default: 3)
# Concurrency
katana -u https://target.com -jc -c 20 -p 20 # 20 concurrent crawlers, 20 parallelism
# Rate limit (requests per second)
katana -u https://target.com -jc -rl 50
# Agent-safe JSONL baseline with bounded depth and static asset filtering
katana -u https://target.com -d 3 -jc -kf robotstxt -c 10 -p 10 -rl 50 -timeout 10 -retry 1 -ef png,jpg,jpeg,gif,svg,css,woff,woff2,ttf,eot,map -silent -j -o katana.jsonl
Katana's JS crawling (-jc) uses JSLuice and regex patterns to extract endpoints from JavaScript files. This is the primary value over generic crawlers.
# JS crawl + XHR/fetch call tracing
katana -u https://target.com -jc -xhr
# Deeper JS parsing (memory intensive)
katana -u https://target.com -d 5 -jc -jsl -kf all -c 10 -p 10 -rl 50 -o katana_urls.txt
# Extract all JS file URLs only
katana -u https://target.com -jc | grep "\.js$" > js_files.txt
# Show discovered endpoints from JS (exclude static assets)
katana -u https://target.com -jc | grep -v "\.(png|jpg|gif|svg|ico|woff|css)$"
# Filter for API paths
katana -u https://target.com -jc | grep -E "(/api/|/v[0-9]+/|/graphql|/rest/)"
Standard crawler misses dynamically rendered content. Use headless for apps that require JavaScript execution.
# Headless Chrome/Chromium required
katana -u https://target.com -hl -jc -d 3
# Headless with wait (allow JS to execute before capture)
katana -u https://target.com -hl -jc -nos
# Headless with system Chrome and XHR extraction
katana -u https://target.com -hl -sc -nos -xhr -j -o katana_headless.jsonl
# Authenticated crawl — provide session cookie
katana -u https://target.com -hl -H "Cookie: session=<token>" -jc
# Bearer token
katana -u https://api.target.com -H "Authorization: Bearer <token>" -jc
# Multiple headers
katana -u https://target.com -H "X-Api-Key: abc123" -H "Accept: application/json" -jc
# POST requests (for apps requiring login state)
katana -u https://target.com -X POST -H "Content-Type: application/json" \
-body '{"email":"[email protected]","password":"test"}' -jc
Katana integrates with the ProjectDiscovery ecosystem:
# httpx → katana: crawl all live web hosts
httpx -l live_hosts.txt -silent | katana -jc -o all_endpoints.txt
# subfinder → httpx → katana: full passive-to-crawl pipeline
subfinder -d target.com -silent | httpx -silent | katana -jc -o endpoints.txt
# Feed into nuclei for vuln scanning
katana -u https://target.com -jc -o endpoints.txt
nuclei -l endpoints.txt -t exposures/ -t vulnerabilities/
# JSON output (structured for parsing)
katana -u https://target.com -jc -jsonl -o katana.jsonl
# Known files (robots.txt and sitemap.xml)
katana -u https://target.com -kf all -d 3 -silent
# Filter by extension (exclude static assets)
katana -u https://target.com -jc -ef png,jpg,jpeg,gif,svg,css,woff,woff2,ttf,eot,map
# Filter by response code
katana -u https://target.com -jc -fsc 200,301,302
# Show only unique paths (deduplicate)
katana -u https://target.com -jc | sort -u > unique_endpoints.txt
# Extract parameters from found URLs
katana -u https://target.com -jc | grep "?" | cut -d"?" -f2 | tr "&" "\n" | cut -d"=" -f1 | sort -u
-rl rate limiting on sensitive or production scopes.robots.txt boundaries unless explicitly authorized to ignore them: -cr to crawl despite robots.| Scenario | Tool |
|----------|------|
| Modern SPA/React/Angular | katana -headless |
| API endpoint extraction from JS | katana -jc |
| Fast directory brute-force | feroxbuster |
| Historical URL collection | gau |
| Fast link extraction from static HTML | hakrawler |
development
Auth/lab ref: Unicorn Engine CPU-only emulation for shellcode, decryptors, custom VM handlers, instruction tracing, memory hooks, and register-level experiments.
development
Auth/lab ref: Renode board and SoC simulation for MCU/RTOS firmware, UART/GPIO/peripheral modeling, GDB remote debugging, REPL platforms, and RESC scripts.
development
Auth/lab ref: Qiling OS-layer binary emulation for PE/ELF/Mach-O/UEFI/shellcode with rootfs, syscall/API hooks, filesystem mapping, and runtime patching.
databases
Auth/lab ref: QEMU user-mode and full-system emulation for cross-arch binaries, firmware, kernels, disks, serial consoles, networking, and GDB stubs.