wechat-article-fetcher: read WeChat articles as Markdown

WeChatMarkdownMCPClaude CodeCursorPython

Read public WeChat articles as Markdown or JSON. One uvx command, or an MCP server for Claude Code and Cursor.

View on GitHub →

TL;DR: wechat-article-fetcher (wxfetch) is a small MIT-licensed Python CLI and library for a person, or an AI agent, who needs a public WeChat Official Account article as Markdown. Nothing to install and nothing to configure if you have uv: uvx --from git+https://github.com/mrlong0129/wechat-article-fetcher wxfetch <url>.

I asked my AI assistant to summarize a WeChat article. The ordinary fetcher came back with WeChat's "环境异常" page, the one that tells you to finish a verification before you can continue. I sent the same short link again, this time as the iPhone WeChat in-app browser, and the real page came back. The post was a short text post. Its full text was not in the article body. WeChat had stored the whole thing in the og:title tag it uses for the share card.

That became wechat-article-fetcher. The wider lesson is agent-web-fetch: try the cheapest request first, rotate the client identity, report a captcha or a login wall instead of bypassing it, run several extractors and keep the richest text, and open a real browser only at the end.

Read WeChat articles as Markdown

wxfetch reads a public article on mp.weixin.qq.com and prints Markdown, or JSON with -f json. The JSON includes the title, account name, author, publish time as ISO 8601 in +08:00, the Markdown body, the plain text, the image list, which extraction method won, and which User-Agent fetched the page.

uvx --from git+https://github.com/mrlong0129/wechat-article-fetcher wxfetch <url>
uvx --from git+https://github.com/mrlong0129/wechat-article-fetcher wxfetch <url> -f json

It accepts a short /s/<id> link, a long ?__biz= link, a link with no https://, and a link copied out of HTML with &amp;. You can pass several URLs, or a file, or stdin. --from-html saved_page.html parses a page you saved and does not touch the network. --download-images stores the images beside the note and rewrites the links. --front-matter adds a YAML header for Obsidian or a static blog.

From Python, after a local install: fetch_article(url) on the network, extract_article(html) offline. The library needs Python 3.8+. Its dependencies are requests, beautifulsoup4, and lxml. Exit code 0 means every article succeeded, 1 means at least one failed and the reason is on stderr, 2 means the command was used wrong.

An MCP server for Claude Code and Cursor

Claude Code, one line. The server needs Python 3.10+. The tools are fetch_wechat_article and parse_wechat_html.

claude mcp add wechat-article-fetcher -- uvx --from git+https://github.com/mrlong0129/wechat-article-fetcher wxfetch-mcp

Cursor: the same server in ~/.cursor/mcp.json, or in the project's .cursor/mcp.json.

{
  "mcpServers": {
    "wechat-article-fetcher": {
      "command": "uvx",
      "args": ["--from", "git+https://github.com/mrlong0129/wechat-article-fetcher", "wxfetch-mcp"]
    }
  }
}

You can also drop in the skill and skip the package. The agent follows the handbook with its own tools. Full instructions for a coding agent are in AGENTS.md.

mkdir -p ~/.claude/skills/wechat-article-fetcher && curl -fsSL https://raw.githubusercontent.com/mrlong0129/wechat-article-fetcher/main/skill/SKILL.md -o ~/.claude/skills/wechat-article-fetcher/SKILL.md
mkdir -p ~/.cursor/skills/wechat-article-fetcher && curl -fsSL https://raw.githubusercontent.com/mrlong0129/wechat-article-fetcher/main/skill/SKILL.md -o ~/.cursor/skills/wechat-article-fetcher/SKILL.md

How a WeChat page gets read

The request goes out as WeChat's in-app browser: iPhone first, then Android, then desktop Chrome, with Referer https://mp.weixin.qq.com/ and Accept-Language: zh-CN. That is the page a person already sees by opening a public link in WeChat. It does not log in, pay, or solve a captcha.

The response is classified before anything is treated as the article. The "环境异常 / 去验证" page (wappoc_appmsgcaptcha, secitptpage) moves on to the next User-Agent. A deleted article, a violation page, or a bad link stops with a reason. Network errors, 429, and 5xx are retried with exponential backoff and jitter. Between articles the default pause is about 2 seconds.

Four sources are tried, and the richest wins. When the lengths are close, the richer structure wins, and the winner is recorded as extraction_method:

  • Long articles: #js_content, parsed as HTML. Images are lazy. The real URL is data-src, not the placeholder in src.
  • Short text posts (item_show_type 10): there is no #js_content. The body is the content_noencode JavaScript string inside window.cgiDataNew, with escapes such as \x0a turned back into characters.
  • Image posts (item_show_type 8): picture_page_info_list.
  • Share-card metadata: og:title and og:description. On a short text post, WeChat puts the entire body in og:title, with newlines written as the two characters \n.

Title, account name, author, and publish time are taken the same way, from the first non-empty source among the page's JavaScript variables, cgiDataNew, the DOM, and the og tags. The publish timestamp is Unix time, written as Beijing time +08:00.

The Markdown keeps headings, paragraphs, lists, quotes, bold, links, images, code, and simple tables. Hidden nodes, account cards, mini-program cards, and empty paragraphs are dropped. Punctuation is moved outside **, because CommonMark will not render bold when ** sits against a Chinese quotation mark. A line break inside a paragraph is written as a Markdown hard break, since a single newline would otherwise become a space.

Images live on mmbiz.qpic.cn. With no Referer, or with Referer https://mp.weixin.qq.com/, the server returns the real image. Another site's Referer gets a small placeholder that says the image may not be quoted. Downloads send the WeChat Referer.

When to use wxfetch, and when to use agent-web-fetch

  • Use wxfetch when the link is a public WeChat Official Account article and you want Markdown or JSON with the account, author, and publish time, offline parsing of a saved page, or the images downloaded beside the note.
  • Use agent-web-fetch when links come from many sites (GitHub, X, news, blogs, and WeChat too) and you want one tool that says why a page could not be read, checks robots.txt, and can fall back to a headless browser.
  • Use neither for login, paid, or fans-only content. Both stop and report the reason.

What it will not do

Public articles only. Login, paid, fans-only, deleted, and violation pages are reported. This tool does not bypass them.

Outside WeChat, long ?__biz= links are almost always challenged. In October 2026 I checked the same article both ways. The short /s/<id> link was readable with all three User-Agents. The long link was sent to the verification page every time, whichever User-Agent I sent. Copy the short link from WeChat.

WeChat can tighten detection at any time. Frequent requests and a datacenter IP are more likely to be asked to verify. The tool does not run JavaScript, so a video-account card, a poll, or some interactive blocks will not appear. Video is kept as a placeholder link. A headless-browser fallback is not built. If HTTP is blocked, open the article, save the HTML, and run wxfetch --from-html.

WeChat's robots.txt disallows /s. wxfetch does not ship a robots mode. It reads a public link you already have. The general reader, agent-web-fetch, checks robots.txt on every fetch, warns by default, and can refuse the fetch with --robots strict.

The layout of a WeChat article varies a lot. The Markdown is meant to be readable. It is not a pixel copy of the original.

Responsible use

The article belongs to its author and the account. I use this for personal reading, notes, and research. Don't bulk-copy accounts, and don't republish a piece without permission. Keep the default pace of about 2 seconds between articles. A verification page means stop: slow down, try later, or read it in WeChat. Don't solve the captcha.

FAQ

How do I read a WeChat article as Markdown?

If you have uv, install nothing and run uvx --from git+https://github.com/mrlong0129/wechat-article-fetcher wxfetch <url>. That prints Markdown. Add -f json for JSON. Use the short https://mp.weixin.qq.com/s/<id> link. After a local install the same command is wxfetch <url> (Python 3.8+).

How do I add wechat-article-fetcher to Claude Code?

One line: claude mcp add wechat-article-fetcher -- uvx --from git+https://github.com/mrlong0129/wechat-article-fetcher wxfetch-mcp. The MCP server needs Python 3.10+. Its tools are fetch_wechat_article and parse_wechat_html.

How do I use it in Cursor?

Add a wechat-article-fetcher entry to ~/.cursor/mcp.json or the project's .cursor/mcp.json, with command uvx and args --from, git+https://github.com/mrlong0129/wechat-article-fetcher, wxfetch-mcp. Or copy skill/SKILL.md into ~/.cursor/skills/wechat-article-fetcher/SKILL.md and let the agent follow it.

Why do long WeChat links hit a verification page?

Outside WeChat, long ?__biz= links are almost always challenged. In October 2026 the same article's short /s/<id> link was readable with the iPhone, Android, and desktop User-Agents, and the long link was sent to the verification page every time. Copy the short link from WeChat.

Does it bypass WeChat's captcha or a login wall?

No. The "环境异常 / 去验证" page is detected, and the next User-Agent is tried. The captcha is never solved. Login, paid, fans-only, deleted, and violation pages are reported and not bypassed.

Where is the text of a short WeChat post?

A short text post (item_show_type 10) has no #js_content body. The full text is the content_noencode JavaScript string, and WeChat also puts that body in the og:title share-card tag. The tool tries all four sources and keeps the richest.