agent-web-fetch: web scraping for AI agents
Web scraping for AI agents: cheapest request first, then report captchas, logins, and paywalls. An MCP server for Claude Code and Cursor.
TL;DR: agent-web-fetch (awf) is an MIT-licensed web-reading toolkit for AI agents. It tries the cheapest request first, escalates only when blocked, and says why a page could not be read. Nothing to install and nothing to configure if you have uv: uvx --from git+https://github.com/mrlong0129/agent-web-fetch awf fetch <url>.
I built this after one WeChat article. A normal fetcher returned WeChat's "环境异常" verification wall. The same short link, sent as the iPhone WeChat in-app browser, returned the article, and the full text of that short post was in the og:title share card. The focused tool for that job is wechat-article-fetcher. This one is the general lesson.
Most fetch tools for a model do one thing. They send a request, turn the HTML into text, and stop. When the site refuses, the agent is handed an error, or the text of a "please verify" page, and it will summarize that page as if it were the article. The same URL can be a full article, a bot check, or an empty JavaScript shell, depending on who is asking.
Web scraping for AI agents
The rule is one sentence: when a path fails, take the next one, cheapest first, and check at every step that you actually got the article.
Layer 0 is a public API, when the site already has one that needs no login. A GitHub repo is read from api.github.com as Markdown, then from raw.githubusercontent.com if the API allowance is spent. A single X post uses the public syndication endpoint, then official oEmbed. Profiles, timelines, and search need a login, so the tool reports login_required and does not scrape them.
Layer 1 is one plain HTTP request with a coherent browser header set. User-Agent, Accept, Accept-Language, and the sec-ch-ua fields agree with each other. A bare python-requests User-Agent is a common reason a site says no. Most public pages succeed here, in tens or hundreds of milliseconds.
Layer 2 is another client, if that visitor was refused. The order is desktop Chrome, mobile Safari, then Android Chrome. A WeChat article starts as wechat_ios instead. Later tries send a same-site Referer. A 429 or 503 reads Retry-After and backs off. Tries inside one fetch share cookies.
Layer 3 is a check after every response, with a machine-readable blocked_reason: cloudflare_challenge, captcha, verification_wall, login_required, paywall, rate_limited, forbidden, not_found, soft_404, content_removed, js_required, robots_disallowed. Keyword checks count only when the page is thin, so an article about captchas is not flagged. A Cloudflare challenge-platform script on its own is not a block. Captcha, login, paywall, not found, and removed stop the ladder immediately.
Layer 4 is several extractors, all local, and the richest wins. Trafilatura, readability, JSON-LD articleBody, embedded state (__NEXT_DATA__, __NUXT__, window.__INITIAL_STATE__, __APOLLO_STATE__, ytInitialPlayerResponse), og: and twitter: meta, and the site adapter's own rules. Score is length times a trust weight. The winner is extraction_method. Every candidate and its length is extraction_candidates. Metadata is merged in order: adapter, then JSON-LD, then og tags, then trafilatura. A thin page also tries the AMP URL and a matching RSS or Atom entry.
Layer 5 is a real browser, last. Optional Playwright renders the page only when a cheaper step failed for a reason worth retrying. If Playwright is not installed, that step is recorded as skipped and the tool does not error. The browser does not solve a captcha and does not click through an interactive challenge. It is the slow step, usually 10–20 seconds.
uvx --from git+https://github.com/mrlong0129/agent-web-fetch awf fetch <url>
uvx --from git+https://github.com/mrlong0129/agent-web-fetch awf fetch <url> --format jsonExit code 0 means the page was read. Exit code 2 means it was blocked or empty, and stderr carries blocked_reason, for example login_required, plus what each step tried. Exit code 1 is any other error. awf extract page.html --url <original> runs detection and extraction on HTML you already have. awf reasons lists the block reasons. Python 3.10+.
An MCP server for Claude Code and Cursor
Claude Code, one line. If you already installed the package, claude mcp add agent-web-fetch -- awf-mcp is enough. The tools are fetch_page and extract_from_html.
claude mcp add agent-web-fetch -- uvx --from git+https://github.com/mrlong0129/agent-web-fetch awf-mcpCursor: ~/.cursor/mcp.json, or the project's .cursor/mcp.json. The same JSON works as .mcp.json for Claude Code.
{
"mcpServers": {
"agent-web-fetch": {
"command": "uvx",
"args": ["--from", "git+https://github.com/mrlong0129/agent-web-fetch", "awf-mcp"]
}
}
}The skill works even when the package is not installed. The agent follows the playbook with curl or its own browser, and prefers awf when it is there. The handbook is SKILL.md. Instructions for a coding agent are in AGENTS.md.
mkdir -p ~/.claude/skills/agent-web-fetch && curl -fsSL https://raw.githubusercontent.com/mrlong0129/agent-web-fetch/main/skill/SKILL.md -o ~/.claude/skills/agent-web-fetch/SKILL.md
mkdir -p ~/.cursor/skills/agent-web-fetch && curl -fsSL https://raw.githubusercontent.com/mrlong0129/agent-web-fetch/main/skill/SKILL.md -o ~/.cursor/skills/agent-web-fetch/SKILL.mdCaptchas are a hard stop
A captcha stops the ladder. That includes reCAPTCHA, hCaptcha, Turnstile, GeeTest, DataDome, and Tencent's captcha. The tool does not solve it, does not rotate clients to slip past it, and does not click through it in the browser. A login wall stops. A paywall stops, including JSON-LD isAccessibleForFree set to false. A 404, a soft 404, and a removed post stop. The agent is told the reason, so it can ask the person instead of inventing a summary.
A verification wall can be worth one more client. WeChat's "环境异常" page is that kind of wall. The WeChat adapter therefore starts as the iPhone in-app browser, with Accept-Language: zh-CN. A long article is read from #js_content, images from data-src. A short text post (item_show_type 10) has an empty #js_content; the body is in content_noencode, and usually in og:title as well. Account, author, and time come from nick_name, author, and ori_create_time (or create_time), written as ISO 8601 with +08:00.
When to use agent-web-fetch
- Your agent is handed arbitrary links and needs clean Markdown, plus a machine-readable
blocked_reasonwhen the page is a captcha, a login wall, a paywall, or a soft 404, instead of summarizing a verification page. - You want a record of what was tried:
fetch_attemptslists every step. - For a WeChat-only workflow with image download and front matter, use wechat-article-fetcher. For crawling a whole site, use a crawl framework: agent-web-fetch is built to read one link at a time.
Limits and responsible use
Public content only. Login, paid, and fans-only pages are out of reach on purpose. Detection is heuristic, so each step also stores blocked_detail. Sites change, and the WeChat and X adapters will need maintenance when they do. Some sites refuse a datacenter IP no matter which client you send. Reuters' DataDome, Zhihu, and Bilibili's risk check are examples. The tool reports captcha or forbidden and stops.
GitHub's unauthenticated API allows 60 requests an hour per IP. GITHUB_TOKEN raises that. It is not required. The headless browser is usually 10–20 seconds and is a separate install: agent-web-fetch[browser], then playwright install chromium. An installed Google Chrome is used when it is already there.
WeChat's robots.txt disallows /s. agent-web-fetch checks robots.txt on every fetch and records it on robots and warnings. The default mode is warn, for one link a person handed you, the same kind of read as opening it in a browser. For a crawl, a batch job, or a schedule, pass --robots strict. The fetch is then refused with blocked_reason robots_disallowed.
There is per-host pacing and backoff. Don't poll a site with it. The words you read belong to their author. Quote and link. Don't republish the piece, and don't use it against a site's terms. Switching User-Agent picks the client a page was built for, such as WeChat's in-app browser. It does not impersonate Googlebot or any other crawler, and it does not pretend to be a specific person.
fetch_attempts records every step: strategy, result, status, blocked reason, detail, time, bytes, and how much text was extracted. That is the difference I wanted from a fetch tool that returns either text or an error and nothing else. It is not a crawl framework. It does not keep a proxy pool. The job is the one an agent actually gets: here is a link, read it, and if you cannot, say why.
FAQ
How do I give Claude Code a web-reading tool?
Run claude mcp add agent-web-fetch -- uvx --from git+https://github.com/mrlong0129/agent-web-fetch awf-mcp. Nothing else is installed or configured. Claude Code then has fetch_page and extract_from_html. If the package is already installed, claude mcp add agent-web-fetch -- awf-mcp is enough.
How do I add agent-web-fetch to Cursor?
Put an agent-web-fetch server in ~/.cursor/mcp.json or the project's .cursor/mcp.json: command uvx, args --from, git+https://github.com/mrlong0129/agent-web-fetch, awf-mcp. The same JSON works as .mcp.json for Claude Code. You can also drop skill/SKILL.md into ~/.cursor/skills/agent-web-fetch/.
How do I read a WeChat article as Markdown?
For a WeChat Official Account link, awf fetch <url> uses the WeChat adapter: iPhone WeChat User-Agent first, then #js_content or content_noencode / og:title. The focused CLI for that one job is wechat-article-fetcher: uvx --from git+https://github.com/mrlong0129/wechat-article-fetcher wxfetch <url>. Use the short /s/<id> link. Long ?__biz= links are challenged outside WeChat.
What happens on a captcha, login wall, or paywall?
The tool stops and reports blocked_reason: captcha, login_required, or paywall. It does not solve captchas, pass login walls, or bypass paywalls, including JSON-LD isAccessibleForFree set to false. The CLI exits 2 and writes the reason to stderr.
Does it respect robots.txt?
It checks robots.txt on every fetch and records the result on robots and warnings. The default mode, warn, is for one link a person asked to read. For a crawl, a batch job, or a schedule, use --robots strict, which stops with robots_disallowed. WeChat's robots.txt disallows /s, so a strict fetch of an article link is refused and the warning is the honest outcome.
When does it open a real browser?
Only after cheaper steps failed for a reason another method might fix, and only if the optional Playwright extra is installed. Otherwise that step is skipped. The browser does not solve captchas. It is the slow step, usually 10–20 seconds.