Repository navigation
IPC channel stops delivering messages to cluster workers #9706
Description
Activity
- addedclusterIssues and PRs related to the cluster subsystem.Issues and PRs related to the cluster subsystem.
on Nov 20, 2016 What's ulimit value for
open files? You can see the value by executingulimit -aIt could be you're hitting that value. Does increasing that value help? See: http://superuser.com/a/303058
@santigimeno I'm able to repro with the hard and soft file descriptor limits set to 1M.
@davidvetrano yes, I could reproduce the issue on
OS XandFreeBSD.What I have observed is that at some point the
masterdoesn't receive anNODE_HANDLE_ACKmessage in response to aNODE_HANDLEcommand sent to the workers causing that the ping messages are being stored in the child_process._handleQueueand they're never actually sent. In fact, the problem is that, for some reason, theNODE_HANDLEmessage is actually sent (apparently with success) to the workers, but the workers never receive it. I thought this was not possible in theIPCchannel as it's anAF_UNIXSOCK_STREAMconnection. Any idea why this could be happening?/cc @bnoordhuis
- addedchild_processIssues and PRs related to the child_process subsystem.Issues and PRs related to the child_process subsystem.freebsdIssues and PRs related to the FreeBSD platform.Issues and PRs related to the FreeBSD platform.osIssues and PRs related to the os subsystem.Issues and PRs related to the os subsystem.
on Nov 22, 2016 I lose IPC communication on Linux too, having 41 workers and rather heavy DB access on each. No error is given on console.
Update: After testing again, v6.1 does in fact have has the bug. Please disregard.
I've taken the above snippet and did a manual "git bisect" on all minor versions on macOS (10.12.2) and the behavior is not exhibited on v6.1.0 but is introduced in ^v6.2.0, and is still prevalent in v7.I'm currently running a real git bisect on the commits between v6.1 and v6.2 to hopefully identify the commit where this regression took place.Yeah, I have also reproduced it in
4.7.1.@santigimeno This should remain open?
/cc @bnoordhuis, @cjihrig, @mcollina
@santigimeno #13235 fixed this, didn't it?
That PR is on track for v6.x and it seems reasonable to me to also target v4.x since it's a rather insidious bug.
I think so. I still can reproduce it with current master on
FreeBSD@bnoordhuis sorry I hadn't read your comment before answering...
#13235 fixed this, didn't it?
I had forgotten about this one but it certainly looks like it could have been solved by #13235, but from a quick check it doesn't look it's solved so it may be a different issue.
- added a commit that references this issue
on Aug 15, 2017 20 remaining items
- addedmacosIssues and PRs related to the macOS platform.Issues and PRs related to the macOS platform.
on May 21, 2018 @gireeshpunathil I'm on Windows, and it seems like your comment is mostly about Linux?
@pitaj, thanks. I was following the code and the platform from the original postings.
Are you using the same code on Windows, or something different? if so, please pass it on. Also, what is the observation - similar to mac os, same as mac os, or different?
I too tested in windows, and I got some surprising result (certain tunings to the original test,and we get complete hang!) . We need to separate that issue from this, so let me hear from you.
looked at the windows hang, and understood the reason too.
every time a client connects, 25 messages (
fromEndpoint) go from the node to the master.
every time the master receives a message of that type, it sends 25 messages back (625 messages per worker)So depending on the number of concurrent requests, performance can really vary, and the dependancy between the requests and the latency is exponential (s you already observed earlier):
but it might be exponential time as very small changes in message number (like 990 vs 1000) result in very large changes in the amount of time required.
There is nothing Windows specific issue observed here from Node.js perspective, other than potential difference in the system configuration / resources. So my original proposal on horizontal scaling stands.
Not a Node.js bug, closing. Exponential stress in the tcp layer causes process to slow down, suggested to share work between multiple hosts.
@gireeshpunathil I have some repro code that doesn't use any TCP AFAIK, unless the IPC channel uses TCP itself.
I can throw that up on a gist later today.
Reacted by Gireesh PunathilThanks @pitaj for the repro. Turns out that the windows issue is unrelated (compelte hang) to the originally posted issue (slow response) on macos and family.
I am able to reproduce the hang. Looking at multiple dumps, I see that the main thread of different processes (including the master) are engaged in:
node.exe!uv_pipe_write_impl(uv_loop_s * loop, uv_write_s * req, uv_pipe_s * handle, const uv_buf_t * bufs, unsigned int nbufs, uv_stream_s * send_handle, void(*)(uv_write_s *, int) cb) Line 1347 C node.exe!uv_write(uv_write_s * req, uv_stream_s * handle, const uv_buf_t * bufs, unsigned int nbufs, void(*)(uv_write_s *, int) cb) Line 139 C node.exe!node::LibuvStreamWrap::DoWrite(node::WriteWrap * req_wrap, uv_buf_t * bufs, unsigned __int64 count, uv_stream_s * send_handle) Line 345 C++ node.exe!node::StreamBase::Write(uv_buf_t * bufs, unsigned __int64 count, uv_stream_s * send_handle, v8::Local<v8::Object> req_wrap_obj) Line 222 C++ node.exe!node::StreamBase::WriteString<1>(const v8::FunctionCallbackInfo<v8::Value> & args) Line 300 C++ node.exe!node::StreamBase::JSMethod<node::LibuvStreamWrap,&node::StreamBase::WriteString<1> >(const v8::FunctionCallbackInfo<v8::Value> & args) Line 408 C++ node.exe!v8::internal::FunctionCallbackArguments::Call(v8::internal::CallHandlerInfo * handler) Line 30 C++ node.exe!v8::internal::`anonymous namespace'::HandleApiCallHelper<0>(v8::internal::Isolate * isolate, v8::internal::Handle<v8::internal::HeapObject> new_target, v8::internal::Handle<v8::internal::HeapObject> fun_data, v8::internal::Handle<v8::internal::FunctionTemplateInfo> receiver, v8::internal::Handle<v8::internal::Object> args, v8::internal::BuiltinArguments) Line 110 C++ node.exe!v8::internal::Builtin_Impl_HandleApiCall(v8::internal::BuiltinArguments args, v8::internal::Isolate * isolate) Line 138 C++ node.exe!v8::internal::Builtin_HandleApiCall(int args_length, v8::internal::Object * * args_object, v8::internal::Isolate * isolate) Line 126 C++ [External Code]
this is a known issue with libuv where multiple parties attempt to write to the same pipe, from either sides, under rare situations.
#7657 posted this originally, and libuv/libuv#1843 fixed it recently. It will be sometime before Node.js consumes it.
When many IPC messages are sent between the master process and cluster workers, IPC channels to workers stop delivering messages. I have not been unable to restore working functionality of the workers and so they must be killed to resolve the issue. Since IPC has stopped working, simply using
Worker.destroy()does not work since the method will wait for thedisconnectevent which never arrives (because of this issue).I am able to repro on OS X by running the following script:
and using ApacheBench to place the server under load as follows:
I see the following, for example:
As I alluded to earlier, I have seen an issue on Linux which I believe is related but I have been so far unable to repro using this technique on Linux.