Repository navigation
let v8 use libuv thread pool #11855
Description
Activity
If that were possible, doing so would further reduce the performance of
fs,dns.lookup()(default method used internally by node), etc.- addedlibuvIssues and PRs related to the libuv dependency or the uv binding.Issues and PRs related to the libuv dependency or the uv binding.questionIssues asking questions about Node.js.Issues asking questions about Node.js.v8 engineIssues and PRs related to the V8 dependency.Issues and PRs related to the V8 dependency.
on Mar 15, 2017 @mscdex We can increase libuv thread number.
Also, doing so we can make thread pool load more balance.Sure, you can increase it, but I would guess most people use the default (currently 4).
Well, and we also can increase default libuv thread number...
I'm not 100% sure but I think 4 may have been chosen as that is/was the typical number of cores/cpus on most machines?
Reacted by James M SnellWhile increasing the default will hopefully be something we can do at some point, 4 is still the most common average if I'm remembering correctly.
In any case, having V8 use threads outside of the thread pool seems wrong. If 4 is the right default, you would in fact use more than 4 if V8 spins up its own threads. If you expect 4 + additional V8 threads to be a good number, then 4 is chosen too conservatively.
@hashseed I'm not sure what the performance difference is (waiting for a libuv thread vs. OS scheduler), but if V8 were to use the libuv thread pool, some node requests would/could get blocked (even more so than they may be currently), whereas they may not have before.
FWIW V8 hands over thread management to Chrome when embedded in Chrome.
@hashseed Could you please explain the reason what Chrome did ?
Node is quite a different thing than Chrome though ;-)
@jeisinger probably knows more.
Iiuc V8 simply prefers to let the embedder take control. In case of Chrome, page start-up can be very busy wrt threading, and Chrome's scheduler likely has a better grasp than the OS scheduler.
10 remaining items
That enum is not used by v8. We only added it because chrome used to use Windows worker pool reflecting WT_EXECUTELONGFUNCTION
The problem is that a single threadpool is used for IO and CPU threads.
If there were two threadpools, one could be used for IO (and have many threads) and the other could be used for CPU (and have threads === cores).
This would make tuning possible. Currently, there's no way to tune the threadpool for both use cases.
If libuv knows that all its threads are IO only, then the above can be rolled out as follows:
-
All AsyncWorker, Node and other userland threads continue to use the current threadpool. This becomes known as the CPU threadpool and keeps the existing 4 thread default.
-
Add a new "IO" threadpool with a saner, conservative yet tuneable default (perhaps 16).
-
Libuv moves only its own known IO ops into this new threadpool and exposes a way for userland to do the same if it wants to.
-
@jorangreef We've been discussing such schemes in libuv since the thread pool was first added. You can find at least two attempts in my fork, maybe more.
Both attempts stranded on not performing significantly better most of the time and significantly worse some of the time. :-/
Thanks @bnoordhuis
Were you benchmarking latency or throughput? "Significantly worse" latency (say 1.1x) may not be that bad if it means significantly more throughput.
Was the workload mostly IO bound or mostly CPU bound?
The idea with two separate threadpools is to delegate these choices to the user. Everyone's benchmark requirements are different.
And if there is no advantage gained by separating CPU and IO intensive operations into their own theadpools, then the arguments against increasing the default 4 thread limit should no longer hold.
Are you 100% confident that the current hardcoded 4 thread default is the optimal solution?
Are you 100% confident that the current hardcoded 4 thread default is the optimal solution?
Hah, I don't think I ever claimed it was. Good enough most of the time, but optimal? No, sir!
Were you benchmarking latency or throughput? [...] Was the workload mostly IO bound or mostly CPU bound?
A bit of both. Node and libuv's benchmarks test both ends of the spectrum.
A bit of both. Node and libuv's benchmarks test both ends of the spectrum.
Just to double-check, when you benchmarked using two separate threadpools (one for IO tasks, one for CPU tasks), did you let both threadpools contend for the same set of cores or did you pin them to separate sets of cores?
In a sufficiently warmed-up webapp (assuming the most common node use case), do the v8 threads incur any significant CPU work? I thread-profiled a client server app with ~20% CPU consumption by the server process, and could not find v8 threads as contributing. Any suggested (v8) tunables to get more conclusive info?
@gireeshpunathil the v8 threads mostly handle i/o, I've only seen an uptick in thread cpu with heavy garbage collection activity. Nodejs is single-threaded, so the main thread will normally account for almost all the cpu usage. One can introduce some concurrency with the
clustermodule.@andrasq - thanks. But sorry, your explanation seems orthogonal to my understanding:
v8 threads handle mostly i/o
which I/O? Are you talking about the primordial (main) thread? I doubt that is classified under v8 thread.
Nodejs is single-threaded, ...
in this discussion and in #14001 we are definitely talking about a number of background threads from v8 module, that is different from the priomordial thread and libuv worker threads. And this discussion focusses on the tradeoff of tenanting the v8 threads with libuv worker threads.
So, my question stands as:
(i) What are those v8 threads which will potentially contend for time slice with more critical work from libuv? Execution tracers? CPU profilers? GC helpers? JIT helpers?
(ii) How do we bring those threads to the forefront to be seen as eating up cycles? (any tunables in the commandline and tunables in the code). This will help us make observations in terms of time-slice distribution variations between the threads and make inferences on throughput, latency as a function of some of these tunables.
(iii) How significant these threads would be in terms of CPU consumption in a fully warmed up production server (no tracing and no new scripts)?What are those v8 threads which will potentially contend for time slice with more critical work from libuv?
Compiler and GC threads.
How do we bring those threads to the forefront to be seen as eating up cycles?
perf(1)?
How significant these threads would be in terms of CPU consumption in a fully warmed up production server
Depends on the application.
thanks @bnoordhuis - that explains.
In terms of measurements, perf(1) did not help to get thread-wise split up (there is a-tflag, but behavior is weird) so I was using AIXtprofand the result, as I mentioned earlier, suggests that v8 threads contribute very less (not displayed in the top consumers)I agree this depends on the workload characteristics of the application. Given:
(i) every new page brings in new script (chrome) vs. relatively boot-time-only script (node)
(ii) highly transactional based web workload means most objects collected in the scavenge phasev8 threads may not consume much slice for either GC or JIT, and is in-line with my tprof observation.
I wish if I could modify the test to manifest the contrasting characteristics and was /am looking for pointers on that line - so that we know the extent to which these threads are insignificant, and what is the tipping point beyond which they show up.
This discussion seems to have run its course. I'm going to close, but if you think that's wrong and there's something concrete, feel free to re-open (if GitHub allows) or comment (requesting it be re-opened if you wish) or open another issue as appropriate.
And FWIW In #22631 I am working on uniting the V8 and libuv threadpools in Node-land using my pluggable threadpool PR in libuv.
Reacted by Joran Dirk Greef
Is it possible to let v8 use libuv thread pool ? this could reduce extra thread numbers.