Repository navigation
ENOMEM with exec/spawn - child process tries to reserve as much mem as parent #25382
Description
Activity
This could be the underlying cause of this very popular SO issue for some users: https://stackoverflow.com/questions/26193654/node-js-catch-enomem-error-thrown-after-spawn
I think something like this would be handled/solved at the libuv level, not node. The libuv issue tracker is here.
Duplicate of
#14917?Not really a dupe. Both are caused by using
fork(), but different symptoms (ENOMEMvs. blocked loop), and this issue isn't addressed byMADV_DONTFORKas suggested in #14917. I just tried my repro with a buffer created frommadvise(..., MADV_DONTFORK)'ed memory, and thefork()syscall still failed withENOMEM. Even if the kernel doesn't make the memory available, it apparently still requests the allocation (or goes through some of the same assertions).Edit May 23, 2019 tried
madvise(MADV_DONTFORK)with a 5.1+ kernel and it still causes ENOMEM.Edit: in Aug 2019 it does seem to work: #25382 (comment)
Reacted by Joran Dirk GreefWhat does
sysctl vm.overcommit_memoryprint on your system?As you've probably gathered from the issues linked in libuv/libuv#2133, this is not something that is solvable, easily or at all.
vfork()andclone()have their own share of issues.- addedchild_processIssues and PRs related to the child_process subsystem.Issues and PRs related to the child_process subsystem.memoryIssues and PRs related to Node.js memory management or memory footprint.Issues and PRs related to Node.js memory management or memory footprint.
on Jan 8, 2019 @bnoordhuis Can you point me in the right direction for the issues you've encountered with using
vfork? I've tried to search the issue tracker and commit history, but seemed hard to discover back that far. I've seen a couple references to cases where it was seen to have issues (such as libc developers arguing about it https://sourceware.org/bugzilla/show_bug.cgi?id=10354), butforkseems to share most of the same limitations (posix also specifies that "Consequently, [after fork] to avoid errors, the child process may only execute async-signal-safe operations until such time as one of the exec functions is called."), but additionally (on linux) has been observed to be slower and to fail in low-memory situations. Additionally on linux, libuv could useclonedirectly to avoid some of the issues (with the vfork flag, but providing a separate stack). I did find this interesting graph though, which suggests that the problems withforkmight be linux-specific: https://github.com/rtomayko/posix-spawnIt's been a couple years, but when I looked into the linux kernel code for
fork, my recollection is that it does not respect the overcommit_memory flag(s) and just enforces that process memory < physical memory (plus probably some extra constant fudge factors). I couldn't really say what memory it counts against the limit however (e.g. if it includes MADV_DONTFORK).Reacted by Joran Dirk GreefThe biggest issue for libuv is that
vfork()doesn't runpthread_atfork()handlers. That might be surmountable for Node.js but add-ons might be a problem.It's been a couple years, but when I looked into the linux kernel code for fork, my recollection is that it does not respect the overcommit_memory flag(s)
Linux's overcommit logic has been overhauled several times in the last years. I wouldn't know where exactly it stands now without checking the source (and also of past releases.) :-)
They aren't run though because they shouldn't often be needed. Although I don't know what Node.js does (or promises to do) in this case. From the rationale of pthread_atfork, running them is necessary to work around design issues with
forkarising due to the COW semantics in multi-threaded programs that try to also use "unsafe" functions. But we can audit uv_spawn and know that it only uses the async-safe functions thenexecv(perforkdocumentation). Did someone actually need theirpthread_atforkhandler to run, or is that just a hypothetical concern that someone might try to do something more in theirpthread_atforkhandler?OTOH, libuv now does register a
pthread_atforkhandler which does some extra operations that adds a couple of additional unnecessary syscalls, and thus may increase the cost of launching a process withoutvfork. I have no idea whether the impact of that code is measurable though.But if the overcommit-related logic has changed, perhaps it might be faster now too? I see the OP is on an older Ubuntu 16.04. But if the kernel can do
fork(almost) as fast and now asvforkand without hitting ENOMEM limits (as it seems may be true on mach, so perhaps possible at the hardware level), that would great! (and would save me from actually needing to ever upstream thevforksupport to libuv, haha)Reacted by Joran Dirk Greef@bnoordhuis the default
overcommit_memorywas 0. The OP repro run with the three options:0 - ENOMEM
1 - okay
2 - ENOMEM (with additional difficulty allocating the buffer in the first place)Aside from that workaround (possibly viable) or adding swap (not viable), for my specific situation I'm looking into a PR to make https://github.com/googleapis/google-auth-library-nodejs either not exec at all, or exec earlier before the process grows.
Reacted by Dmitry Poroh and Jesper EkAlthough I don't know what Node.js does (or promises to do) in this case.
@vtjnash There is (or was) at least one shared memory add-on module that uses
pthread_atfork()so I don't think it's out of the question that other add-ons exist that, directly or indirectly, rely on the current behavior.I say "indirectly" because add-on X might depend on shared library Y, which in turn depends on a functioning
pthread_atfork(). Since that's impossible to audit for, I'd be uncomfortable changing the default, but, as discussed in libuv/libuv#141, an opt-in should be acceptable.@zbjornson Thanks for checking! I figured that DWIM mode (
vm.overcommit_memory=1) would end up working around this.It's curious that
madvise(MADV_DONTFORK)didn't work and I wonder if that's a kernel bug. It's not the behavior I'd expect at any rate.(Random data point: also seeing this on a 2GB server trying to execute
imageminvia CLI)Hrm, twice previously I tried using
MADV_DONTFORKand was still getting ENOMEMs, but looking at the kernel source to figure out why, I saw that it should work. I just tried again and, lo and behold, the ENOMEM goes away. (Maybe I didn't have a page-aligned allocation or maybe I had two advice flags OR'ed together, and wasn't checking the return value? /shrug, and sorry for the wasted cycles thinking about alternate solutions.)Anyway, getting https://chromium-review.googlesource.com/c/v8/v8/+/1101679 landed would be great. Looks like there's just a typo pending.
(Also, on my system an
madvisecall only takes ~6 μs.)can we push https://chromium-review.googlesource.com/c/v8/v8/+/1101679 forward?
19 remaining items
- added a commit that references this issue
on Jul 31, 2023 - added 2 commits that reference this issue
on Oct 9, 2023 - added 2 commits that reference this issue
on Nov 26, 2023 - added a commit that references this issue
on Feb 18, 2024 - added a commit that references this issue
on Apr 12, 2026
A tiny
execcall when Node.js is using at least half of the otherwise-available memory causesspawn ENOMEM:Apparently this is a common problem when using
fork()/clone(),and using.posix_spawn(3)avoids itI don't see any upstream issues in libuv for this.