Sitelet https://github.com/nodejs/node/issues/48578
Skip to content

cluster.worker.on('message', (msg) => ...) fails to register a callback if ESM file extensions is used #48578

Description

@jerome-benoit

Version

v20.3.1

Platform

All supported platform

Subsystem

No response

What steps will reproduce the bug?

After a migration of code to ESM by using .mjs file extensions, the poolifier project has encountered issues at running internal benchmarks code.

Test case:

  • main.mjs:
import cluster from 'cluster'

cluster.setupPrimary({ exec: './worker.mjs' })

const worker = cluster.fork()

worker.on('message', message => {
  console.info('message received from worker:', message)
})

worker.on('online', () => {
  console.info('worker is online')
})

worker.on('error', (error) => {
  console.info('worker error', error)
})

worker.on('disconnect', () => {
  console.info('worker disconnected')
})

worker.on('exit', () => {
  console.info('worker exited')
})

worker.send('hello')
  • worker.mjs:
import cluster from 'cluster'

cluster.worker.on('message', message => {
  console.info('echo message received from main:', message)
  cluster.worker.send(message)
})

node main.mjs is frozen.

How often does it reproduce? Is there a required condition?

100% reproducible

What is the expected behavior? Why is that the expected behavior?

No response

What do you see instead?

Callback is never called if a message is sent from the primary.

Additional information

No response

Activity

  1. jerome-benoit commented on Jun 27, 2023

    @jerome-benoit
    ContributorAuthor

    Additional information: if example code for IPC with cluster is put the same ESM file, the callback for 'message' event is properly registered and called. Poolifier code is using the exec option to specify the worker file path. And in that case, cluster.worker.on('',() => {}) seems to have issues with ESM.

  2. aduh95 commented on Jun 27, 2023

    @aduh95
    Contributor

    Can you share a repro?

  3. jerome-benoit commented on Jun 27, 2023

    @jerome-benoit
    ContributorAuthor

    Please read the detailed bug report on poolifier repo, it contains all the bits to reproduce it reliably with poolifier. I do not have the time currently to extract the code from it to reproduce with two files using .mjs extension.

  4. aduh95 commented on Jun 27, 2023

    @aduh95
    Contributor

    I don't have the time either, I don't plan on working on it without a repro that does not involve external code.

  5. bnoordhuis commented on Jun 28, 2023

    @bnoordhuis
    Member

    I'll close this until there is a standalone test case. Let me know when I should reopen.

  6. jerome-benoit commented on Jun 28, 2023

    @jerome-benoit
    ContributorAuthor

    The way that bug 100% reproducible has been handled by the node.js project is below the common sense standards expected if the project goal is stability.
    When I receive a confirmed bug report 100% reproducible on one of the FOSS project I maintain, I would not even dare to put the burden on the bug reporter to solve it: it's unrespectful of the work done to identify the issue and pinpoint the root cause.

  7. bnoordhuis commented on Jun 28, 2023

    @bnoordhuis
    Member

    The reason we don't usually accept bug reports that manifest in third-party code is that 9 out of 10 times the bug is in said third-party code, not Node.js.

    Eventually comes the point where you decide you've wasted enough unpaid hours of your life on other people's crappy code.

    Long story short: happy to take a look if you have a minimal test case. But if you are not willing to put in the time, then neither are we.

  8. jerome-benoit commented on Jun 28, 2023

    @jerome-benoit
    ContributorAuthor

    Please read carefully the bug report I've made before applying blindly rules that do not apply to it:

    • CommonJS benchmarking code: everything works, callback registered with cluster.worker.on(), called in poolifier
    • benchmarking code migrated to ESM: identical code, except the import part, using the very same (not even the ESM bundle, but the ESM bundle has the same issue) poolifier code previously working: callback not registered with cluster.worker.on() or called

    => Bug confirmed, 100% reproducible. Extracting a standalone test case is then becoming a corollary, not a prerequisite to its resolution. And then the burden is expected to be shared between project maintainers and the bug reporter.

  9. bnoordhuis commented on Jun 28, 2023

    @bnoordhuis
    Member

    Unless they're on your payroll, then no, you don't get to decide how and where other people invest their time.

    Either put together a test case or don't, but stop arguing.

  10. jerome-benoit commented on Jun 28, 2023

    @jerome-benoit
    ContributorAuthor

    So the node.js project is not interested in releasing a stable ESM support in the cluster module by trying to welcome a proven 100% reproducible confirmed bug? And prefers to apply blindly rules instead of collaborating with the bug reporter to help planning fixing it and makes node.js better?

    I will do the standalone test case when time permits. But given the way my first confirmed 100% reproducible bug report on the node.js project has been welcomed, I do not think I will redo it unless it impacts deeply a FOSS project I comaintain.

  11. aduh95 commented on Jun 28, 2023

    @aduh95
    Contributor

    What about you? Are you not interested in releasing a stable ESM support in the cluster module? I think there's a misunderstanding here, the Node.js project is run by volunteers, you can't expect that someone else would be interested in investing their time in fixing your bug, especially when you yourself are not. Or you need to pay them (send me an email, I can make myself available).
    The Node.js project would gladly accept a PR to fix it, but the Node.js project cannot open PRs by itself, only contributors can; and you are unlikely to find someone to spend time on it if they are not affected by the bug themselves. If you provide a repro – which is as you probably know is often the hardest part of fixing a bug – you're way more likely to find a volunteer.

    But given the way my first confirmed 100% reproducible bug report on the node.js project has been welcomed, I do not think I will redo it unless it impacts deeply a FOSS project I comaintain.

    You'll be deeply missed. If I may suggest to tune down the entitlement and show a little more respect of other people's personal time, you could be amazed what difference it makes.

  12. jerome-benoit commented on Jun 28, 2023

    @jerome-benoit
    ContributorAuthor

    I do not expect any fix to be done, I do not expect any answer, I expect nothing from a bug report on a FOSS project, except one thing: that if taken into account, it's done fairly in depth by a volunteer that matters at fixing it and has the time to handle it. A sign of respect of the time spent by another volunteer to pointpoint the bug. If that's not the view of the node.js project volunteers, that's fine by me. But easily understandable that the volunteer in question will be reluctant at doing another detailed bug reports on the same project.
    If the node.js project thinks that the present bug report has been handled in a respectful way and fairly, that discussion is going nowhere and there's no point at continuing it.

    Back to the subject: I'll do the test case reproducing the issue when time permits and attach the files to the bug report. And when times permit probably do a PR to try to fix it, as I usually do as a long time FOSS hacker on the open bug reports I do.

  13. jerome-benoit commented on Jul 29, 2023

    @jerome-benoit
    ContributorAuthor

    The test case:

    • main.mjs:
    import cluster from 'cluster'
    
    cluster.setupPrimary({ exec: './worker.mjs' })
    
    const worker = cluster.fork()
    
    worker.on('message', message => {
      console.info('message received from worker:', message)
    })
    
    worker.on('online', () => {
      console.info('worker is online')
    })
    
    worker.on('error', (error) => {
      console.info('worker error', error)
    })
    
    worker.on('disconnect', () => {
      console.info('worker disconnected')
    })
    
    worker.on('exit', () => {
      console.info('worker exited')
    })
    
    worker.send('hello')
    • worker.mjs:
    import cluster from 'cluster'
    
    cluster.worker.on('message', message => {
      console.info('echo message received from main:', message)
      cluster.worker.send(message)
    })

    node main.mjs is frozen.

    Please reopen the issue.

  14. jerome-benoit commented on Jul 29, 2023

    @jerome-benoit
    ContributorAuthor
    • main.js:
    const cluster = require('cluster')
    
    cluster.setupPrimary({ exec: './worker.js' })
    
    const worker = cluster.fork()
    
    worker.on('message', message => {
      console.info('message received from worker:', message)
    })
    
    worker.on('online', () => {
      console.info('worker is online')
    })
    
    worker.on('error', (error) => {
      console.info('worker error', error)
    })
    
    worker.on('disconnect', () => {
      console.info('worker disconnected')
    })
    
    worker.on('exit', () => {
      console.info('worker exited')
    })
    
    worker.send('hello')
    • worker.js:
    const cluster = require('cluster')
    
    cluster.worker.on('message', message => {
      console.info('echo message received from main:', message)
      cluster.worker.send(message)
    })

    node main.js just works.

  15. bnoordhuis commented on Jul 30, 2023

    @bnoordhuis
    Member

    Thanks, I'll reopen. Can you update your original report with the test cases?

    I can make an educated guess as to why it doesn't work:

    1. cluster.fork() is implemented on top of child_process.fork()
    2. child_process.fork() spawns the child process with NODE_CHANNEL_FD=<num> set in the environment
    3. NODE_CHANNEL_FD makes node's bootstrap code call child_process._forkChild() to set up IPC
    4. the asynchronous loading of worker.mjs makes that malfunction somehow

    Interestingly, the worker sends the 'online' internal message. That's done by cluster._setupWorker() and it's called from lib/internal/process/pre_execution.js so at least some of the steps are working, just not all.

    Anyway, you know where to start looking now. Pull request welcome.

  16. added
    confirmed-bugIssues and PRs for confirmed bugs.
    clusterIssues and PRs related to the cluster subsystem.
    on Jul 30, 2023
  17. jerome-benoit commented on Aug 21, 2023

    @jerome-benoit
    ContributorAuthor

    That will take quite a while before I will be able to work on it, so if someone starts to work on that issue, please say so and share your progress.

  18. jerome-benoit commented on Sep 22, 2023

    @jerome-benoit
    ContributorAuthor

    @fabiancook: thanks for your interest. You do need to use poolifier to fix that issue.

  19. ronag commented on Oct 7, 2023

    @ronag
    Member

    @mcollina i think this is quite a significant issue in terms of esm stability.

    See @bnoordhuis last comment.

  20. jerome-benoit commented on Oct 7, 2023

    @jerome-benoit
    ContributorAuthor

    @mcollina i think this is quite a significant issue in terms of esm stability.

    The issue is only met if the worker file is an ESM one. The main file type has no impact.

  21. st3ffgv4 commented on Jun 13, 2024

    @st3ffgv4

    I'm facing the same issue

  22. electriquo commented on Mar 7, 2025

    @electriquo

    The issue is around for very long time. Would you please share an update?


    Anyhow, I am including a simple example of the issue for reproduction.
    Given the following

    const cluster = require("cluster");
    // import cluster from "cluster";
    
    if (cluster.isPrimary) {
      const worker = cluster.fork();
      worker.on("message", (msg) => {
        console.log(`Worker to master: ${msg}`);
      });
      worker.send("Hey worker")
    } else {
      process.on("message", (msg) => {
        console.log(`Master to worker: ${msg}`);
      });
      process.send("Hey master");
    }

    When the first line is modified from const cluster = require("cluster"); to import cluster from "cluster";, the message "Master to worker" is no longer received by the worker process, where as the expected behavior is that the worker should receive the "Master to worker" message regardless of whether the cluster module is imported using require or import.

    Worth noting that the issue exists in NodeJS v22.14.0 and doesn't exist in Bun 1.2.4 :(

  23. electriquo commented on Jun 24, 2025

    @electriquo

    This issue has been open for nearly two years.
    Does this mean it’s unlikely to be fixed?
    Should we consider migrating away from Node.js in favor of a more actively maintained JavaScript runtime?

  24. mcollina commented on Jun 24, 2025

    @mcollina
    SponsorMember

    Node.js is maintained by volunteers. If something is not fixed in a timely fashion, it's usually because no one volunteered to do it. You can always send a pull request to fix it yourself.

  25. mhayk commented on Jul 20, 2026

    @mhayk
    Contributor

    I attempted to reproduce this issue on the current main branch (v27.0.0-pre) and was unable to do so in any of the reported scenarios:

    • process.on('message') inside a .mjs worker — works ✅
    • cluster.worker.on('message') inside a .mjs worker — works ✅
    • Single .mjs file acting as both primary and worker — works ✅
    • worker.send() called immediately after cluster.fork() (before 'online') — works ✅

    It appears this was silently fixed at some point between v22.14.0 and v27, likely as a side-effect of ESM loader refactoring.

    Could someone verify whether this still affects the v22 or v20 LTS lines? If so, a backport may be needed. If it is confirmed fixed on main, it would be good to close this issue or at least document the fix.

  26. electriquo commented on Jul 21, 2026

    @electriquo

    @mhayk what are your results for #48578 (comment)?

  27. mhayk commented on Jul 21, 2026

    @mhayk
    Contributor

    Testing the specific scenario from your March 2025 comment on main (v27.0.0-pre):

    import cluster from "cluster";
    
    if (cluster.isPrimary) {
      const worker = cluster.fork();
      worker.on("message", (msg) => {
        console.log(`Worker to primary: ${msg}`);
      });
      worker.send("Hey worker");
    } else {
      process.on("message", (msg) => {
        console.log(`Primary to worker: ${msg}`);
        process.send("Hey primary");
      });
    }

    Output on main:

    Primary to worker: Hey worker
    Worker to primary: Hey primary
    

    Both message callbacks are invoked correctly ✅. The process does not exit because there is no explicit process.exit() call, the primary keeps running while the worker is alive. This is identical behaviour to the CJS equivalent using require("cluster").

    The original bug (callbacks never being called) appears to be fixed on main. The process hanging indefinitely is expected and not specific to ESM, it reproduces the same way with CJS.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    clusterIssues and PRs related to the cluster subsystem.confirmed-bugIssues and PRs for confirmed bugs.

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions