wechat-article-fetcher: read WeChat articles as Markdown
Read public WeChat articles as Markdown or JSON. One uvx command, or an MCP server for Claude Code and Cursor.
TL;DR: wechat-article-fetcher (wxfetch) is a small MIT-licensed Python CLI and library for a person, or an AI agent, who needs a public WeChat Official Account article as Markdown. Nothing to install and nothing to configure if you have uv: uvx --from git+https://github.com/mrlong0129/wechat-article-fetcher wxfetch <url>.
I asked my AI assistant to summarize a WeChat article. The ordinary fetcher came back with WeChat's "环境异常" page, the one that tells you to finish a verification before you can continue. I sent the same short link again, this time as the iPhone WeChat in-app browser, and the real page came back. The post was a short text post. Its full text was not in the article body. WeChat had stored the whole thing in the og:title tag it uses for the share card.
That became wechat-article-fetcher. The wider lesson is agent-web-fetch: try the cheapest request first, rotate the client identity, report a captcha or a login wall instead of bypassing it, run several extractors and keep the richest text, and open a real browser only at the end.
Read WeChat articles as Markdown
wxfetch reads a public article on mp.weixin.qq.com and prints Markdown, or JSON with -f json. The JSON includes the title, account name, author, publish time as ISO 8601 in +08:00, the Markdown body, the plain text, the image list, which extraction method won, and which User-Agent fetched the page.
uvx --from git+https://github.com/mrlong0129/wechat-article-fetcher wxfetch <url>
uvx --from git+https://github.com/mrlong0129/wechat-article-fetcher wxfetch <url> -f jsonIt accepts a short /s/<id> link, a long ?__biz= link, a link with no https://, and a link copied out of HTML with &. You can pass several URLs, or a file, or stdin. --from-html saved_page.html parses a page you saved and does not touch the network. --download-images stores the images beside the note and rewrites the links. --front-matter adds a YAML header for Obsidian or a static blog.
From Python, after a local install: fetch_article(url) on the network, extract_article(html) offline. The library needs Python 3.8+. Its dependencies are requests, beautifulsoup4, and lxml. Exit code 0 means every article succeeded, 1 means at least one failed and the reason is on stderr, 2 means the command was used wrong.
An MCP server for Claude Code and Cursor
Claude Code, one line. The server needs Python 3.10+. The tools are fetch_wechat_article and parse_wechat_html.
claude mcp add wechat-article-fetcher -- uvx --from git+https://github.com/mrlong0129/wechat-article-fetcher wxfetch-mcpCursor: the same server in ~/.cursor/mcp.json, or in the project's .cursor/mcp.json.
{
"mcpServers": {
"wechat-article-fetcher": {
"command": "uvx",
"args": ["--from", "git+https://github.com/mrlong0129/wechat-article-fetcher", "wxfetch-mcp"]
}
}
}You can also drop in the skill and skip the package. The agent follows the handbook with its own tools. Full instructions for a coding agent are in AGENTS.md.
mkdir -p ~/.claude/skills/wechat-article-fetcher && curl -fsSL https://raw.githubusercontent.com/mrlong0129/wechat-article-fetcher/main/skill/SKILL.md -o ~/.claude/skills/wechat-article-fetcher/SKILL.md
mkdir -p ~/.cursor/skills/wechat-article-fetcher && curl -fsSL https://raw.githubusercontent.com/mrlong0129/wechat-article-fetcher/main/skill/SKILL.md -o ~/.cursor/skills/wechat-article-fetcher/SKILL.mdHow a WeChat page gets read
The request goes out as WeChat's in-app browser: iPhone first, then Android, then desktop Chrome, with Referer https://mp.weixin.qq.com/ and Accept-Language: zh-CN. That is the page a person already sees by opening a public link in WeChat. It does not log in, pay, or solve a captcha.
The response is classified before anything is treated as the article. The "环境异常 / 去验证" page (wappoc_appmsgcaptcha, secitptpage) moves on to the next User-Agent. A deleted article, a violation page, or a bad link stops with a reason. Network errors, 429, and 5xx are retried with exponential backoff and jitter. Between articles the default pause is about 2 seconds.
Four sources are tried, and the richest wins. When the lengths are close, the richer structure wins, and the winner is recorded as extraction_method:
- Long articles:
#js_content, parsed as HTML. Images are lazy. The real URL isdata-src, not the placeholder insrc. - Short text posts (
item_show_type10): there is no#js_content. The body is thecontent_noencodeJavaScript string insidewindow.cgiDataNew, with escapes such as\x0aturned back into characters. - Image posts (
item_show_type8):picture_page_info_list. - Share-card metadata:
og:titleandog:description. On a short text post, WeChat puts the entire body inog:title, with newlines written as the two characters\n.
Title, account name, author, and publish time are taken the same way, from the first non-empty source among the page's JavaScript variables, cgiDataNew, the DOM, and the og tags. The publish timestamp is Unix time, written as Beijing time +08:00.
The Markdown keeps headings, paragraphs, lists, quotes, bold, links, images, code, and simple tables. Hidden nodes, account cards, mini-program cards, and empty paragraphs are dropped. Punctuation is moved outside **, because CommonMark will not render bold when ** sits against a Chinese quotation mark. A line break inside a paragraph is written as a Markdown hard break, since a single newline would otherwise become a space.
Images live on mmbiz.qpic.cn. With no Referer, or with Referer https://mp.weixin.qq.com/, the server returns the real image. Another site's Referer gets a small placeholder that says the image may not be quoted. Downloads send the WeChat Referer.
When to use wxfetch, and when to use agent-web-fetch
- Use wxfetch when the link is a public WeChat Official Account article and you want Markdown or JSON with the account, author, and publish time, offline parsing of a saved page, or the images downloaded beside the note.
- Use agent-web-fetch when links come from many sites (GitHub, X, news, blogs, and WeChat too) and you want one tool that says why a page could not be read, checks robots.txt, and can fall back to a headless browser.
- Use neither for login, paid, or fans-only content. Both stop and report the reason.
What it will not do
Public articles only. Login, paid, fans-only, deleted, and violation pages are reported. This tool does not bypass them.
Outside WeChat, long ?__biz= links are almost always challenged. In October 2026 I checked the same article both ways. The short /s/<id> link was readable with all three User-Agents. The long link was sent to the verification page every time, whichever User-Agent I sent. Copy the short link from WeChat.
WeChat can tighten detection at any time. Frequent requests and a datacenter IP are more likely to be asked to verify. The tool does not run JavaScript, so a video-account card, a poll, or some interactive blocks will not appear. Video is kept as a placeholder link. A headless-browser fallback is not built. If HTTP is blocked, open the article, save the HTML, and run wxfetch --from-html.
WeChat's robots.txt disallows /s. wxfetch does not ship a robots mode. It reads a public link you already have. The general reader, agent-web-fetch, checks robots.txt on every fetch, warns by default, and can refuse the fetch with --robots strict.
The layout of a WeChat article varies a lot. The Markdown is meant to be readable. It is not a pixel copy of the original.
Responsible use
The article belongs to its author and the account. I use this for personal reading, notes, and research. Don't bulk-copy accounts, and don't republish a piece without permission. Keep the default pace of about 2 seconds between articles. A verification page means stop: slow down, try later, or read it in WeChat. Don't solve the captcha.
FAQ
How do I read a WeChat article as Markdown?
If you have uv, install nothing and run uvx --from git+https://github.com/mrlong0129/wechat-article-fetcher wxfetch <url>. That prints Markdown. Add -f json for JSON. Use the short https://mp.weixin.qq.com/s/<id> link. After a local install the same command is wxfetch <url> (Python 3.8+).
How do I add wechat-article-fetcher to Claude Code?
One line: claude mcp add wechat-article-fetcher -- uvx --from git+https://github.com/mrlong0129/wechat-article-fetcher wxfetch-mcp. The MCP server needs Python 3.10+. Its tools are fetch_wechat_article and parse_wechat_html.
How do I use it in Cursor?
Add a wechat-article-fetcher entry to ~/.cursor/mcp.json or the project's .cursor/mcp.json, with command uvx and args --from, git+https://github.com/mrlong0129/wechat-article-fetcher, wxfetch-mcp. Or copy skill/SKILL.md into ~/.cursor/skills/wechat-article-fetcher/SKILL.md and let the agent follow it.
Why do long WeChat links hit a verification page?
Outside WeChat, long ?__biz= links are almost always challenged. In October 2026 the same article's short /s/<id> link was readable with the iPhone, Android, and desktop User-Agents, and the long link was sent to the verification page every time. Copy the short link from WeChat.
Does it bypass WeChat's captcha or a login wall?
No. The "环境异常 / 去验证" page is detected, and the next User-Agent is tried. The captcha is never solved. Login, paid, fans-only, deleted, and violation pages are reported and not bypassed.
Where is the text of a short WeChat post?
A short text post (item_show_type 10) has no #js_content body. The full text is the content_noencode JavaScript string, and WeChat also puts that body in the og:title share-card tag. The tool tries all four sources and keeps the richest.