Repository navigation
test-child-process-fork-regr-gh-2847 - fails on pLinux BE and AIX #3245
Description
Activity
- addedchild_processIssues and PRs related to the child_process subsystem.Issues and PRs related to the child_process subsystem.testIssues and PRs related to Node.js core tests and test infrastructure.Issues and PRs related to Node.js core tests and test infrastructure.
on Oct 8, 2015 @mhdawson you've probably already checked -- but in case you didn't, have a look and see if there are trailing processes from prior runs. It's the most common cause of ECONNREFUSED.
@gireeshpunathil you'd see both; one being workers not being able to connect and one being master not being able to spawn. Not saying this is the issue here though.
@jbergstroem, I believe trailing processes from prior runs would case EADDRINUSE, not ECONNREFUSED.
The scenario this test case covers is a cluster topology, wherein the worker exits on the first message from the master, and the master, sends a dummy message to the worker after connectivity is established with a server, which is done 100 times back-to-back. The intention is to make sure that the queued up data in the channel causes the process.disconnect() to wait until they are flushed.
I believe there are two issues with the test case:
- It assumes that all 100 messaged are dispatched before the worker recieves its first message. If this does not happen due to timing differences, worker.send() call will fail, as the worker is dead when some messages are still being dispatched:
#Error: channel closed
at ChildProcess.target.send (internal/child_process.js:510:16)
at Worker.send (cluster.js:47:21)
at Socket. (test/parallel/test-child-process-fork-regr-gh-2847.js:27:14)
at Socket.g (events.js:260:16)
at emitNone (events.js:67:13)
at Socket.emit (events.js:166:7)
at TCPConnectWrap.afterConnect as oncomplete2.It assumes that when the first message is received in the worker and subsequently it's close callback is invoked in the master, next message is already dispatched by the master. If this does not happen due to timing differences, the server gets closed, and net.connect() will fail, as the server is already closed:
Error: connect ECONNREFUSED 127.0.0.1:12346
at Object.exports._errnoException (util.js:837:11)
at exports._exceptionWithHostPort (util.js:860:20)
at TCPConnectWrap.afterConnect as oncomplete@jbergstroem, thanks for the quick reply. I confirmed that is not the case here with a clean run, that too with new, free port number.
Also, with a bit of timing change (changing the for loop to a setInterval or delaying the second worker.send call etc.) I am able to reproduce these two errors in Linux IA32 as well.
Gireesh could you add an update of your investigation so far
Further debugging revealed the following:
- AIX loopback socket is slower than that of Linux. Only a handful of requests are made before the worker gets the first message, and subsequently shutdown itself.
- If I use IP address instead of localhost, the test passes consistently. All the 100 requests are dispatched before the worker gets the first message.
- I am not able to reproduce the error in 1000 iterations in PLinux. I will run for some more time to make any inference.
- However, as mentioned in the previous comment, the test case depends on strict timing assumptions based on Linux behavior, pLinux failure can be inferred as occurred under heavy load.
- If I change the tight loop to setInterval, the ECONNREFUSED failure is reproducible in all platforms.
- If I add a slight delay inside the send function for subsequent sends, the "channel closed" error is reproducible in all platforms.
I propose these changes to the test:
- Install an error handler on the connected socket, and ignore error if any.
- Implement a callback for the message send function, and absorb error if any.
- To make sure at least initial few connections go through successfully, use a counter.
Please let me know about this approach. If agree, I will come up with a PR.
@indutny , request you to share your thoughts.It seems to fail on PPC BE consistently in the CI
https://ci.nodejs.org/job/node-test-commit-plinux/17/nodes=ppcbe-fedora20/console
not ok 45 test-child-process-fork-regr-gh-2847.js #events.js:141 # throw er; // Unhandled 'error' event # ^ # #Error: connect ECONNREFUSED 127.0.0.1:12346 # at Object.exports._errnoException (util.js:889:11) # at exports._exceptionWithHostPort (util.js:912:20) # at TCPConnectWrap.afterConnect [as oncomplete] (net.js:1063:14) --- duration_ms: 2.712
The suggested approach seems reasonable. Validating that at least some number of request get through will catch the case when we might otherwise ignore real errors.
Gireesh if you can put together a pull request or even just a branch that includes your changes we can start node-test-commit-plinux pointed at that in order to see if it resolves the problems seen on PPC in the CI runs
@mhdawson , thanks. Here is the test in a branch with the proposed modification for your validation in CI. A pull request is on its way.
https://github.com/gireeshpunathil/node/blob/clustermasterworkersync/test/parallel/test-child-process-fork-regr-gh-2847.js
/cc @indutnyCI run here L https://ci.nodejs.org/job/node-test-commit/883/
Shows that failure is resolved for PPC, there are some issues with formatting reported by the linter though. You should be able to catch these with a complete make test run and fix up when putting together the PR
The other CI failures look unrelated to the test change to me.
Since Gireesh is on vacation this week went ahead and resolved lint issues and created PR #3459
PR closed, closing
https://ci.nodejs.org/job/iojs+pr+ppc/nodes=ppcbe-fedora20/21/console
Failure on pLinux in CI https://ci.nodejs.org/job/iojs+pr+ppc/nodes=ppcbe-fedora20/21/console:
I see the AIX failures in our internal builds but not the one on pLinux