openclaw-skills/web-scraper/SKILL.md
Use when users need webpage scraping, structured data extraction, crawling strategy, anti-bot handling, selector design, or repeatable web data collection workflows.
npx skillsauth add seaworld008/commonly-used-high-value-skills web-scraperInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
waitForSelector, waitForResponse, waitForTimeout 以确保数据加载完成。from playwright.sync_api import sync_playwright
def run(playwright):
browser = playwright.chromium.launch(headless=True)
context = browser.new_context(user_agent="Mozilla/5.0 ...")
page = context.new_page()
page.goto("https://example.com")
# 模拟滚动到底部触发加载
page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
page.wait_for_selector(".item-list")
items = page.query_selector_all(".item-list .title")
results = [item.inner_text() for item in items]
print(results)
browser.close()
with sync_playwright() as playwright:
run(playwright)
from bs4 import BeautifulSoup
import requests
response = requests.get(url, headers=my_headers)
soup = BeautifulSoup(response.text, 'lxml')
# 使用 CSS Selector
price = soup.select_one('.product-price').get_text(strip=True)
# 使用 XPath (需配合 lxml.etree)
# tree.xpath('//div[@id="title"]/h1/text()')
robots.txt。禁止抓取非公开个人信息。遵循 GDPR 和反不正当竞争法。注:本技能适用于合规、合理的网页公开数据采集场景。
tools
飞书审批:查询和处理审批待办/已办/实例,搜索可发起审批定义、查看定义详情并发起原生审批实例。当用户要处理审批任务、查看审批实例、搜索或发起审批时使用。审批待办不是飞书任务;非审批类待办走 lark-task。不负责创建审批定义;三方审批定义不走原生提单。
development
Use when a user needs reproducible repository sizing, language composition, file counts, or code-versus-comment ratios with pygount; record exclusions and verify measurement scope before interpreting results.
development
Route a development task to the official Hermes Agent skill, Graphify Codex artifact set, Open GSD Core bundle, or optional GSD Pi bundle without duplicating their installers or state machines.
development
飞书 / Lark 通讯录:按姓名 / 邮箱解析成 open_id,或按 open_id 反查姓名 / 部门 / 邮箱 / 联系方式 / 个人状态 / 签名,以及按关键词搜索当前用户可见的机器人 / 智能体(agent)。当用户提到一个名字要下一步发消息 / 排日程,或拿到 open_id 想查具体信息时使用。不负责部门树遍历、按部门列员工、组织架构图,这类需求走原生 OpenAPI。