
1. 公众号历史文章采集为什么总卡在鉴权这一步公众号历史文章采集这件事真正难的不是翻页也不是正文抽取而是鉴权链路。你打开mp.weixin.qq.com的图文素材编辑器点超链接、搜公众号、F12 抓包会看到cgi-bin/searchbiz和cgi-bin/appmsg两个接口它们都要求带上token和Cookie。这两个值绑定当前登录用户换台机器、换个浏览器、隔几小时就失效脚本跑一半报ret: 200013或直接返回空列表是采集里最常见的翻车点。我试过把 token 硬编码在脚本里结果每次手动改代码采集任务一多就乱。后来把鉴权部分抽出来交给 TaoToken 统一管理 Key 和 API 通道脚本只负责 XPath 解析和翻页逻辑配置和代码解耦维护成本一下降下来。这篇就按这个思路走先讲清楚 Cookie 和 XPath 各自负责什么再给一份可复制的config.toml骨架接着是 Cookie 注入示例最后用一次真实采集请求验证整条链路跑通。适合谁看写过一点 Python、用过 requests 或 lxml、想把手动抓包改成配置化采集的人。不需要你懂逆向也不需要你研究加密参数核心就是把 token、Cookie、XPath 三件事拆开管好。先说结论公众号采集的稳定性取决于鉴权信息能不能被安全、可替换地注入而不是取决于你写多复杂的爬虫框架。TaoToken 在这里的角色是统一 Key/API 通道让 token 的获取和轮换有地方落脚本侧只读配置。2. TaoToken 在采集链路里承担什么角色TaoToken 是一个统一的大模型与 API 通道管理平台官网在 https://taotoken.net/?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content 。放到公众号采集这个场景里它不直接帮你抓公众号而是解决「鉴权配置散落各处」的问题把 Key、通道地址、模型或接口配置集中管理脚本通过读取配置拿到需要的凭证而不是把敏感值写死在代码里。你可以这样理解分工。公众号侧负责返回token和Cookie这是微信自己的登录态TaoToken 侧负责把这些凭证和你的采集任务配置统一收口提供 API 通道和 Key 管理。采集脚本启动时先从配置里读通道信息再注入 Cookie最后发请求。这样换账号、换任务、换环境只改配置不改代码。具体到操作你需要先拿到 TaoToken 的 API Key。进入控制台创建https://taotoken.net/console?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content 然后在 API Keys 页面生成密钥https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content 。API 基础地址是 https://taotoken.net/api 注意这个地址不带 UTM 参数配置里直接写它。如果你后面想把采集结果做语义清洗、标题分类、正文摘要可以走模型对话通道https://taotoken.net/models?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content 。如果是要长期跑编码类采集任务、写 Agent 自动翻页可以看 Coding Planhttps://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content 。接入细节查文档https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content 。注意TaoToken 管的是你的 API 通道和 Key公众号的token/Cookie仍然来自微信登录态两者不要混为一谈。配置里要分开存放。3. 可复制的 config.toml 骨架与 Cookie 注入先给配置文件。用 TOML 是因为它可读性好Python 用tomllib3.11或tomli都能读。下面这份骨架把通道配置、公众号鉴权、XPath 规则、翻页参数分开改哪块一目了然。# config.toml [taotoken] api_base https://taotoken.net/api api_key sk-你的TaoToken密钥 timeout 30 [wechat] # 以下两项来自 F12 抓包绑定登录用户会过期 token 你的token cookie 你的cookie user_agent Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 [wechat.search] action search_biz begin 0 count 5 query 目标公众号名称 type 9 lang zh_CN f json ajax 1 [wechat.list] action list_ex begin 0 count 5 sleep_every 3 # 每3页休眠一次 sleep_seconds 60 # 休眠60秒降低触发风控概率 [xpath] title //h2[idactivity-name]/text() content //*[idjs_content]读取配置的代码import tomllib def load_config(pathconfig.toml): with open(path, rb) as f: return tomllib.load(f) cfg load_config() print(cfg[taotoken][api_base]) print(cfg[wechat][token])Cookie 注入是重点。公众号接口要求Cookie和token同时出现在请求里缺一个就返回鉴权失败。下面这个函数把配置里的 Cookie 字符串转成请求头并拼好公共参数import requests def build_session(cfg): headers { User-Agent: cfg[wechat][user_agent], Cookie: cfg[wechat][cookie], Referer: https://mp.weixin.qq.com/, } session requests.Session() session.headers.update(headers) return session def base_params(cfg, action): w cfg[wechat] return { action: action, begin: w[search][begin], count: w[search][count], query: w[search][query], fakeid: None, type: w[search][type], token: w[token], lang: w[search][lang], f: w[search][f], ajax: w[search][ajax], }搜索公众号拿到fakeid这是后续翻页的关键参数。fakeid是公众号的唯一标识做 base64 编码后传给接口。如果固定采集几个号建议把名称到fakeid的映射存成表减少重复搜索请求import json def search_biz(session, cfg): url https://mp.weixin.qq.com/cgi-bin/searchbiz params base_params(cfg, search_biz) resp session.get(url, paramsparams, timeoutcfg[taotoken][timeout]) data resp.json() if not data.get(list): raise RuntimeError(f搜索失败返回{data}) fakeid data[list][0][fakeid] return fakeid拿到fakeid后翻页接口用list_ex每次begin递增count。下面这段把翻页和 XPath 解析串起来import time from lxml import etree def fetch_article_links(session, cfg, fakeid): url https://mp.weixin.qq.com/cgi-bin/appmsg params base_params(cfg, list_ex) params[fakeid] fakeid links [] begin 0 page 0 while True: params[begin] begin resp session.get(url, paramsparams, timeoutcfg[taotoken][timeout]) data resp.json() msg_list data.get(app_msg_list, []) if not msg_list: break for item in msg_list: links.append(item.get(link)) page 1 if page % cfg[wechat][list][sleep_every] 0: time.sleep(cfg[wechat][list][sleep_seconds]) begin cfg[wechat][list][count] return links正文抽取用 XPath配置里已经写好title和content两条规则。解析函数如下def parse_article(html, cfg): tree etree.HTML(html) title tree.xpath(cfg[xpath][title]) content_nodes tree.xpath(cfg[xpath][content]) html_parts [] for node in content_nodes: html_parts.append( etree.tostring(node, encodingutf-8).decode(utf-8) ) return { title: title[0].strip() if title else , content_html: .join(html_parts), }提示etree.tostring输出的是修正后的 HTML比直接取text_content()保留更多结构适合后续做正文清洗。4. 验证一次采集请求是否跑通配置和代码都齐了接下来做一次最小验证。目标搜索一个公众号拿到fakeid翻一页解析出标题和正文。整个过程不需要全量跑一页就够确认链路通。def verify_once(): cfg load_config() session build_session(cfg) fakeid search_biz(session, cfg) print(fakeid , fakeid) url https://mp.weixin.qq.com/cgi-bin/appmsg params base_params(cfg, list_ex) params[fakeid] fakeid params[begin] 0 resp session.get(url, paramsparams, timeoutcfg[taotoken][timeout]) data resp.json() print(app_msg_cnt , data.get(app_msg_cnt)) print(本页条数 , len(data.get(app_msg_list, []))) first data[app_msg_list][0] print(首篇标题 , first.get(title)) print(首篇链接 , first.get(link)) detail session.get(first[link], timeoutcfg[taotoken][timeout]) parsed parse_article(detail.text, cfg) print(解析标题 , parsed[title]) print(正文长度 , len(parsed[content_html])) if __name__ __main__: verify_once()成功时你会看到类似输出fakeid MzA5MTc0NjUxNw app_msg_cnt 128 本页条数 5 首篇标题 某篇历史文章标题 首篇链接 https://mp.weixin.qq.com/s/xxxxxx 解析标题 某篇历史文章标题 正文长度 4821几个判断点。fakeid有值说明搜索接口鉴权通过app_msg_cnt大于 0说明翻页接口正常本页条数等于配置里的count说明分页参数生效正文长度明显大于 0说明 XPath 规则命中了正文容器。四项都过链路就算跑通。如果app_msg_cnt是 0 但fakeid有值通常是该公众号没有历史文章或接口返回被限制如果fakeid为空回到搜索接口看返回体里的base_resp.ret常见是200013频率限制或-6token 失效。5. 本篇常见错误排查采集过程中报错集中在鉴权、XPath、频率三类。下面按现象、原因、处理列出来方便对照。现象可能原因处理方式返回ret: 200013请求频率过高被限制增大sleep_seconds降低count分时段跑返回ret: -6或空列表token 或 Cookie 过期重新登录抓包更新 config.toml 里的两项fakeid为 None搜索接口未命中或参数错检查query是否精确、type是否为 9正文长度始终为 0XPath 规则不匹配用浏览器 F12 复制实际节点路径更新[xpath]KeyError: app_msg_list接口返回结构变化或鉴权失败先打印resp.text看原始返回再判断Cookie 里有换行导致请求头异常复制时带了换行把 Cookie 拼成单行或读取时strip()关于 token 和 Cookie 的更新建议单独写一个小脚本抓包后直接写入 config.toml避免手改出错def update_wechat_auth(token, cookie, pathconfig.toml): import re with open(path, r, encodingutf-8) as f: text f.read() text re.sub(rtoken .*?, ftoken {token}, text) text re.sub(rcookie .*?, fcookie {cookie}, text) with open(path, w, encodingutf-8) as f: f.write(text)注意Cookie 属于登录态凭证不要提交到公开仓库。config.toml 加进 .gitignore或者用环境变量覆盖。还有一个容易忽略的点begin和count的配合。微信翻页接口的begin是偏移量count是每页条数两者必须和实际返回条数对齐。如果你把count设成 5但接口实际返回 5 条begin每次加 5 就对了如果接口返回条数不稳定建议用返回列表长度动态累加而不是固定加count。6. 把鉴权配置收口到 TaoToken 通道整条链路跑通后回头看最开始的问题采集脚本的稳定性取决于鉴权信息能不能被安全、可替换地注入。公众号侧的token和Cookie会过期这是微信登录态决定的改不了但你可以把「配置从哪来、怎么更新、脚本怎么读」这三件事管好。TaoToken 在这里提供的是统一 Key/API 通道让通道配置和密钥管理有固定入口。采集脚本启动时读 config.toml通道信息走 TaoToken公众号鉴权走微信登录态XPath 规则独立成段。换任务只改配置不动代码。如果你要长期跑采集建议把 API Key 和通道配置放在 TaoToken 控制台统一管理https://taotoken.net/console?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content 密钥在 API Keys 页面生成https://taotoken.net/api-keys?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content 。采集结果要做语义处理时走模型对话通道https://taotoken.net/models?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content 。长期编码和 Agent 自动化看 Coding Planhttps://taotoken.net/coding-plan?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content 。接入细节以文档为准https://taotoken.net/doc?utm_sourcetaotoken_aicg_blog_endutm_mediumcsdnutm_campaignrewriteutm_content 。最后留一个实用习惯每次采集任务开始前先跑一遍第 4 节的verify_once()确认fakeid、app_msg_cnt、正文长度三项正常再启动全量。这一步花不到十秒能省掉大量中途报错重跑的时间。