Sitelet https://web.archive.org/web/20201105012010/https://github.com/scrapy/scrapy/issues/4680
Skip to content
New issue

Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.

By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.

Already on GitHub? Sign in to your account

twisted.internet.error.TimeoutError causing slowness in the crawler #4680

Closed
svnshikhil opened this issue Jul 16, 2020 · 7 comments
Closed

twisted.internet.error.TimeoutError causing slowness in the crawler #4680

svnshikhil opened this issue Jul 16, 2020 · 7 comments

Comments

@svnshikhil
Copy link

@svnshikhil svnshikhil commented Jul 16, 2020 •

We have a crawler running on our 2GB, 2 CPU ubuntu machine, for the past few months we have noticed a large number of twisted.internet.error.TimeoutError in the crawler stats, which is causing slowness in our entire crawl process. Also, we noticed that this is happening only when we crawling sites with large contents In it and crawler is using only 50% (1 core ) of the CPU in the system.

I tried changing the DOWNLOAD_DELAY settings and I couldn't make any progress in that.  Please let me know if you have a better way to solve this problem.

Attaching a sample stats from the crawler

{
  "memusage/startup": 337682432,
  "downloader/exception_type_count/twisted.internet.error.TimeoutError": 58457,
  "splash/execute/response_count/500": 1507,
  "downloader/response_status_count/500": 1507,
  "downloader/request_count": 1081283,
  "memusage/warning_notified": 1,
  "finish_time": "2020-07-10 06:38:44.584240",
  "item_scraped_count": 11468349,
  "downloader/response_bytes": 31299223551,
  "error_item_count": 2,
  "retry/reason_count/500 Internal Server Error": 1454,
  "memusage/warning_reached": 1,
  "splash/execute/response_count/504": 17,
  "splash/execute/response_count/502": 379,
  "scheduler/enqueued": 2102353,
  "splash/execute/response_count/400": 29,
  "log_count/WARNING": 1794,
  "memusage/max": 1925443584,
  "splash/execute/response_count/200": 1020130,
  "siteclarity_item_count": 1021043,
  "downloader/response_status_count/502": 379,
  "downloader/exception_type_count/twisted.web._newclient.ResponseNeverReceived": 60,
  "retry/reason_count/twisted.web._newclient.ResponseNeverReceived": 59,
  "start_time": "2020-06-19 19:56:17.613957",
  "downloader/response_status_count/504": 17,
  "ignore/unhandled_exception": 1,
  "retry/max_reached": 472,
  "memusage/limit_notified": 1,
  "downloader/exception_count": 58517,
  "type/canonical": 1136,
  "downloader/response_status_count/302": 49,
  "type/doc": 1019907,
  "downloader/response_status_count/503": 265,
  "dupefilter/filtered": 13580051,
  "retry/reason_count/504 Gateway Time-out": 17,
  "log_count/INFO": 1063157,
  "splash/execute/response_count/301": 327,
  "robots_txt/recording/blocked/doc": 55911322,
  "retry/reason_count/twisted.internet.error.TimeoutError": 58092,
  "scheduler/enqueued/disk": 2102353,
  "downloader/response_status_count/404": 61,
  "splash/execute/response_count/302": 49,
  "splash/execute/response_count/404": 61,
  "splash/execute/request_count": 1021070,
  "retry/reason_count/503 Service Unavailable": 256,
  "response_received_count": 1020619,
  "downloader/response_status_count/307": 2,
  "downloader/request_method_count/POST": 1081283,
  "memusage/limit_reached": 1,
  "downloader/response_count": 1022766,
  "scheduler/dequeued": 2102353,
  "downloader/request_bytes": 1231343961,
  "scheduler/dequeued/disk": 2102353,
  "downloader/response_status_count/301": 327,
  "finish_reason": "finished",
  "retry/count": 60213,
  "request_depth_max": 4,
  "splash/execute/response_count/503": 265,
  "downloader/response_status_count/400": 29,
  "splash/execute/response_count/307": 2,
  "log_count/ERROR": 1990,
  "retry/reason_count/502 Bad Gateway": 335,
  "downloader/response_status_count/200": 1020130
}

(edited for code formatting)

@Gallaecio
Copy link
Member

@Gallaecio Gallaecio commented Jul 16, 2020

Do the URLs that cause the timeout errors work if run separately?

@svnshikhil
Copy link
Author

@svnshikhil svnshikhil commented Aug 3, 2020

Yes, they are not an actual timeout. I tried them separately and they are working fine

@Gallaecio
Copy link
Member

@Gallaecio Gallaecio commented Aug 5, 2020

Oh, I see you are using Splash. Is the issue reproducible if you do not use Splash?

If so, it’s possible that your concurrency is too high for a single Splash instance, or that you need to configure your Splash instance to support higher concurrency. In any case, the issue would not be related to Scrapy itself.

@svnshikhil
Copy link
Author

@svnshikhil svnshikhil commented Aug 19, 2020

Oh, I see you are using Splash. Is the issue reproducible if you do not use Splash?

Yes

@svnshikhil
Copy link
Author

@svnshikhil svnshikhil commented Aug 19, 2020

I think this is not something related to splash
Because the splash is a configurable mechanism in my crawler and I tried some other sites without splash and they are also showing this error
This issue is only happening when the site contains large html

@Gallaecio
Copy link
Member

@Gallaecio Gallaecio commented Aug 19, 2020

Without a minimal, reproducible example it will be hard to pinpoint the root cause 🙁

@wRAR wRAR added the needs more info label Sep 2, 2020
@Gallaecio
Copy link
Member

@Gallaecio Gallaecio commented Oct 2, 2020

Closing due to a lack of feedback.

@Gallaecio Gallaecio closed this Oct 2, 2020
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Projects
None yet
Linked pull requests

Successfully merging a pull request may close this issue.

None yet
3 participants
You can’t perform that action at this time.