Join GitHub today
GitHub is home to over 50 million developers working together to host and review code, manage projects, and build software together.
Sign uptwisted.internet.error.TimeoutError causing slowness in the crawler #4680
Comments
|
Do the URLs that cause the timeout errors work if run separately? |
|
Yes, they are not an actual timeout. I tried them separately and they are working fine |
|
Oh, I see you are using Splash. Is the issue reproducible if you do not use Splash? If so, it’s possible that your concurrency is too high for a single Splash instance, or that you need to configure your Splash instance to support higher concurrency. In any case, the issue would not be related to Scrapy itself. |
|
Yes |
|
I think this is not something related to splash |
|
Without a minimal, reproducible example it will be hard to pinpoint the root cause |
|
Closing due to a lack of feedback. |
We have a crawler running on our 2GB, 2 CPU ubuntu machine, for the past few months we have noticed a large number of twisted.internet.error.TimeoutError in the crawler stats, which is causing slowness in our entire crawl process. Also, we noticed that this is happening only when we crawling sites with large contents In it and crawler is using only 50% (1 core ) of the CPU in the system.
I tried changing the DOWNLOAD_DELAY settings and I couldn't make any progress in that. Please let me know if you have a better way to solve this problem.
Attaching a sample stats from the crawler
(edited for code formatting)