fix(weibo): 修复搜索翻页到底时崩溃退出 - #977
Open
ybwbqg9379 wants to merge 1 commit into
Open
ybwbqg9379 wants to merge 1 commit into
ybwbqg9379 wants to merge 1 commit into
Conversation
微博对「已无更多结果」的响应是 {'ok': 0, 'msg': '这里还没有内容'},
与真实接口错误共用 ok=0。request() 无差别抛 DataFetchError,叠加
未排除该情况的 @Retry,导致正常翻到最后一页会重试 5 次后以
tenacity.RetryError 打穿整个爬虫进程。
在 request() 中区分这一信号并抛 NoMoreResultsError,将其排除出重试
(与 xhs client 既有做法一致);搜索与创作者主页两个分页循环捕获后
跳出,改为继续下一个关键词或返回已采集结果。
Refs NanmiCoder#907
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
问题
微博在「已无更多结果」时返回
{'ok': 0, 'msg': '这里还没有内容'},与真实接口错误共用ok=0。WeiboClient.request()对ok=0无差别抛DataFetchError,而它的@retry(stop_after_attempt(5), wait_fixed(3))又没有排除这种情况。结果是搜索正常翻到最后一页时,会白白重试 5 次(约 12 秒),最后以tenacity.RetryError打穿整个爬虫进程:WeiboCrawler.search()的分页循环唯一出口是CRAWLER_MAX_NOTES_COUNT,异常会一路冒泡到顶层,所以后面的关键词也不会再被采集。get_all_notes_by_creator_id()的循环同样会被这个异常打穿。改动
exception.py:新增NoMoreResultsErrorclient.py:request()区分「无更多结果」信号并抛NoMoreResultsError,用retry_if_not_exception_type将其排除出重试。这与media_platform/xhs/client.py中已有的做法一致client.py:get_all_notes_by_creator_id()捕获后跳出,返回已采集结果core.py:search()捕获后跳出当前关键词的分页,继续下一个关键词只匹配
这里还没有内容这一条已被日志证实的消息,其余ok=0的行为完全不变,仍然按错误处理并重试。测试
新增
tests/test_weibo_search_pagination_end.py,纯离线 mock,不访问真实平台:NoMoreResultsError且只请求 1 次(不重试)ok=0错误仍然重试满 5 次并抛RetryError,内层异常仍是DataFetchError关于 #907
这个 PR 只解决 #907 中暴露出来的崩溃问题,没有解决该 issue 标题所问的「抓取全部搜索结果(含被折叠的相似结果)」——那需要调整微博搜索接口的去重参数,且必须对真实平台验证,不适合放在同一个改动里。所以这里用
Refs而非Fixes,issue 请保持开启。