Repository navigation
Node.js v14.15.5 segfault in v8::internal::ConcurrentMarking::Run #37553
Description
Activity
Any update on this?
It can be reproduced in MacOS, with exception
node:events:346 throw er; // Unhandled 'error' event ^ Error: write EINVAL at afterWriteDispatched (node:internal/stream_base_commons:160:15) at writevGeneric (node:internal/stream_base_commons:143:3) at Socket._writeGeneric (node:net:771:11) at Socket._writev (node:net:780:8) at doWrite (node:internal/streams/writable:412:12) at clearBuffer (node:internal/streams/writable:565:5) at onwrite (node:internal/streams/writable:467:7) at WriteWrap.onWriteComplete [as oncomplete] (node:internal/stream_base_commons:106:10) Emitted 'error' event on Socket instance at: at emitErrorNT (node:internal/streams/destroy:188:8) at emitErrorCloseNT (node:internal/streams/destroy:153:3) at processTicksAndRejections (node:internal/process/task_queues:81:21) { errno: -22, code: 'EINVAL', syscall: 'write' }I also observe that the memory usage is huge (~2.0GB real memory usage ~4.0 GB memory usage, according to activity monitor). Could it be related to memory exhaustion or leaking?
Because I'm not very familar with C++ backend, I'm going to label this with c++ to identify problem further.
- addedc++Issues and PRs that require attention from people who are familiar with C++.Issues and PRs that require attention from people who are familiar with C++.
on Mar 22, 2021 I'm sorry it's not more useful, but this bug is candidate for continuous segfaults in our own system since upgrading from node 12 to 14:
- Download a series of large files from AWS S3
- Decompress them using
zlib.gunzip - Call
JSON.parse()on the result
The segfault always happens on
JSON.parse()(which I can verify using console.log either side), and typically when memory use is getting towards 1GB.Unfortunately I'm running within an AWS Lambda environment, so there are no dumps. I'm working on recreating it locally, but it seems like this is the same issue (which is what gave me the hint to isolate the JSON parsing in the first place).
At the moment I know the following:
- I'm not out of memory (the lambda environment fails differently in this case)
- There's nothing wrong with the files; the nature of the retries mean that the same files are later processed fine
- This exact same code work(s/ed) on Node 12
I'm currently working on a local / smaller reproduction, although I regret my C++ is unlikely to be good enough to provide any real insights here; I just wanted to register my 'it is not just you'.
ed. now with somewhat identical back trace:
General information about node instance
(llnode) v8 nodeinfo Information for process id 14854 (process=0x2ae4cf281d81) Platform = linux, Architecture = x64, Node Version = v14.16.0 Component versions (process.versions=0x112a563b3f29): ares = 1.16.1 brotli = 1.0.9 cldr = 37.0 icu = 67.1 llhttp = 2.1.3 modules = 83 napi = 7 nghttp2 = 1.41.0 node = 14.16.0 openssl = 1.1.1j tz = 2020a unicode = 13.0 uv = 1.40.0 v8 = 8.4.371.19-node.18 zlib = 1.2.11 Release Info (process.release=0x112a563b4161): name = node lts = Fermium sourceUrl = https://nodejs.org/download/release/v14.16.0/node-v14.16.0.tar.gz headersUrl = https://nodejs.org/download/release/v14.16.0/node-v14.16.0-headers.tar.gz Executable Path = /home/ubuntu/.nvm/versions/node/v14.16.0/bin/node Command line arguments (process.argv=0x112a563b4039): [0] = '/home/ubuntu/.nvm/versions/node/v14.16.0/bin/node' [1] = '/home/ubuntu/persist/code/backend/packages/segfault/.webpack/.script.js' Node.js Command line arguments (process.execArgv=0x112a563b4101):List of all threads
(llnode) thread list Process 14854 stopped * thread #1: tid = 14857, 0x0000000000cff994 node`v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) + 1364, name = 'node', stop reason = signal SIGSEGV thread #2: tid = 14855, 0x00007f829681e5ce libc.so.6`epoll_wait + 94, stop reason = signal 0 thread #3: tid = 14854, 0x00007f829688a8e9 libc.so.6`___lldb_unnamed_symbol1092$$libc.so.6 + 633, stop reason = signal 0 thread #4: tid = 14870, 0x00007f82968fe376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13, stop reason = signal 0 thread #5: tid = 14860, 0x00007f82969013f4 libpthread.so.0`do_futex_wait at futex-internal.h:320:13, stop reason = signal 0 thread #6: tid = 14856, 0x0000000000cfdb4c node`v8::internal::ConcurrentMarkingVisitor::ShouldVisit(v8::internal::HeapObject) + 60, stop reason = signal 0 thread #7: tid = 14868, 0x00007f82968fe376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13, stop reason = signal 0 thread #8: tid = 14867, 0x00007f82968fe376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13, stop reason = signal 0 thread #9: tid = 14869, 0x00007f82968fe376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13, stop reason = signal 0 thread #10: tid = 14859, 0x0000000000cff942 node`v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) + 1282, stop reason = signal 0 thread #11: tid = 14858, 0x0000000000cffec1 node`v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) + 2689, stop reason = signal 0Threads' backtrace
(llnode) bt all * thread #1, name = 'node', stop reason = signal SIGSEGV * frame #0: 0x0000000000cff994 node`v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) + 1364 frame #1: 0x0000000000c6c9bb node`non-virtual thunk to v8::internal::CancelableTask::Run() + 59 frame #2: 0x0000000000a71405 node`node::(anonymous namespace)::PlatformWorkerThread(void*) + 405 frame #3: 0x00007f82968f7609 libpthread.so.0`start_thread(arg=<unavailable>) at pthread_create.c:477:8 frame #4: 0x00007f829681e293 libc.so.6`__clone + 67 thread #2, stop reason = signal 0 frame #0: 0x00007f829681e5ce libc.so.6`epoll_wait + 94 frame #1: 0x000000000138e994 node`uv__io_poll at linux-core.c:324:14 frame #2: 0x000000000137c438 node`uv_run(loop=0x0000000005fb4828, mode=UV_RUN_DEFAULT) at core.c:385:5 frame #3: 0x0000000000a75f4b node`node::WorkerThreadsTaskRunner::DelayedTaskScheduler::Start()::'lambda'(void*)::_FUN(void*) + 123 frame #4: 0x00007f82968f7609 libpthread.so.0`start_thread(arg=<unavailable>) at pthread_create.c:477:8 frame #5: 0x00007f829681e293 libc.so.6`__clone + 67 thread #3, stop reason = signal 0 frame #0: 0x00007f829688a8e9 libc.so.6`___lldb_unnamed_symbol1092$$libc.so.6 + 633 frame #1: 0x00000000014650c3 node`Builtins_TypedArrayPrototypeSet + 835 frame #2: 0x000037143acc7cd8 frame #3: 0x000000000139a5a2 node`Builtins_InterpreterEntryTrampoline + 194 frame #4: 0x000037143acc3263 frame #5: 0x000037143acc3780 frame #6: 0x00000000013944d9 node`Builtins_ArgumentsAdaptorTrampoline + 185 frame #7: 0x000000000139a5a2 node`Builtins_InterpreterEntryTrampoline + 194 frame #8: 0x000037143acd54e0 frame #9: 0x00000000013982ba node`Builtins_JSEntryTrampoline + 90 frame #10: 0x0000000001398098 node`Builtins_JSEntry + 120 frame #11: 0x0000000000cc2cc1 node`v8::internal::(anonymous namespace)::Invoke(v8::internal::Isolate*, v8::internal::(anonymous namespace)::InvokeParams const&) + 449 frame #12: 0x0000000000cc3b2f node`v8::internal::Execution::Call(v8::internal::Isolate*, v8::internal::Handle<v8::internal::Object>, v8::internal::Handle<v8::internal::Object>, int, v8::internal::Handle<v8::internal::Object>*) + 95 frame #13: 0x0000000000b8ba24 node`v8::Function::Call(v8::Local<v8::Context>, v8::Local<v8::Value>, int, v8::Local<v8::Value>*) + 324 frame #14: 0x000000000096ad61 node`node::InternalCallbackScope::Close() + 1233 frame #15: 0x000000000096b357 node`node::InternalMakeCallback(node::Environment*, v8::Local<v8::Object>, v8::Local<v8::Object>, v8::Local<v8::Function>, int, v8::Local<v8::Value>*, node::async_context) + 647 frame #16: 0x0000000000978f69 node`node::AsyncWrap::MakeCallback(v8::Local<v8::Function>, int, v8::Local<v8::Value>*) + 121 frame #17: 0x0000000000ac0bbf node`non-virtual thunk to node::(anonymous namespace)::CompressionStream<node::(anonymous namespace)::ZlibContext>::AfterThreadPoolWork(int) + 255 frame #18: 0x00000000009d8475 node`node::ThreadPoolWork::ScheduleWork()::'lambda0'(uv_work_s*, int)::_FUN(uv_work_s*, int) + 341 frame #19: 0x000000000137750d node`uv__work_done(handle=0x000000000446c870) at threadpool.c:313:5 frame #20: 0x000000000137bb06 node`uv__async_io.part.1 at async.c:163:5 frame #21: 0x000000000138e5e5 node`uv__io_poll at linux-core.c:462:11 frame #22: 0x000000000137c438 node`uv_run(loop=0x000000000446c7c0, mode=UV_RUN_DEFAULT) at core.c:385:5 frame #23: 0x0000000000a44974 node`node::NodeMainInstance::Run() + 580 frame #24: 0x00000000009d1e15 node`node::Start(int, char**) + 277 frame #25: 0x00007f82967230b3 libc.so.6`__libc_start_main + 243 frame #26: 0x00000000009694cc node`_start + 41 thread #4, stop reason = signal 0 frame #0: 0x00007f82968fe376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13 frame #1: 0x00007f82968fe359 libpthread.so.0`__pthread_cond_wait at pthread_cond_wait.c:508 frame #2: 0x00007f82968fe290 libpthread.so.0`__pthread_cond_wait(cond=0x000000000446c760, mutex=0x000000000446c720) at pthread_cond_wait.c:638 frame #3: 0x000000000138a4a9 node`uv_cond_wait at thread.c:780:7 frame #4: 0x0000000001376ea4 node`worker(arg=0x0000000000000000) at threadpool.c:76:7 frame #5: 0x00007f82968f7609 libpthread.so.0`start_thread(arg=<unavailable>) at pthread_create.c:477:8 frame #6: 0x00007f829681e293 libc.so.6`__clone + 67 thread #5, stop reason = signal 0 frame #0: 0x00007f82969013f4 libpthread.so.0`do_futex_wait at futex-internal.h:320:13 frame #1: 0x00007f82969013ca libpthread.so.0`do_futex_wait(sem=0x0000000004465600, abstime=0x0000000000000000, clockid=0) at sem_waitcommon.c:112 frame #2: 0x00007f82969014e8 libpthread.so.0`__new_sem_wait_slow(sem=0x0000000004465600, abstime=0x0000000000000000, clockid=0) at sem_waitcommon.c:184:10 frame #3: 0x000000000138a2e2 node`uv_sem_wait at thread.c:626:9 frame #4: 0x000000000138a2d0 node`uv_sem_wait(sem=0x0000000004465600) at thread.c:682 frame #5: 0x0000000000afbd45 node`node::inspector::(anonymous namespace)::StartIoThreadMain(void*) + 53 frame #6: 0x00007f82968f7609 libpthread.so.0`start_thread(arg=<unavailable>) at pthread_create.c:477:8 frame #7: 0x00007f829681e293 libc.so.6`__clone + 67 thread #6, stop reason = signal 0 frame #0: 0x0000000000cfdb4c node`v8::internal::ConcurrentMarkingVisitor::ShouldVisit(v8::internal::HeapObject) + 60 frame #1: 0x0000000000cffe82 node`v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) + 2626 frame #2: 0x0000000000c6c9bb node`non-virtual thunk to v8::internal::CancelableTask::Run() + 59 frame #3: 0x0000000000a71405 node`node::(anonymous namespace)::PlatformWorkerThread(void*) + 405 frame #4: 0x00007f82968f7609 libpthread.so.0`start_thread(arg=<unavailable>) at pthread_create.c:477:8 frame #5: 0x00007f829681e293 libc.so.6`__clone + 67 thread #7, stop reason = signal 0 frame #0: 0x00007f82968fe376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13 frame #1: 0x00007f82968fe359 libpthread.so.0`__pthread_cond_wait at pthread_cond_wait.c:508 frame #2: 0x00007f82968fe290 libpthread.so.0`__pthread_cond_wait(cond=0x000000000446c760, mutex=0x000000000446c720) at pthread_cond_wait.c:638 frame #3: 0x000000000138a4a9 node`uv_cond_wait at thread.c:780:7 frame #4: 0x0000000001376ea4 node`worker(arg=0x0000000000000000) at threadpool.c:76:7 frame #5: 0x00007f82968f7609 libpthread.so.0`start_thread(arg=<unavailable>) at pthread_create.c:477:8 frame #6: 0x00007f829681e293 libc.so.6`__clone + 67 thread #8, stop reason = signal 0 frame #0: 0x00007f82968fe376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13 frame #1: 0x00007f82968fe359 libpthread.so.0`__pthread_cond_wait at pthread_cond_wait.c:508 frame #2: 0x00007f82968fe290 libpthread.so.0`__pthread_cond_wait(cond=0x000000000446c760, mutex=0x000000000446c720) at pthread_cond_wait.c:638 frame #3: 0x000000000138a4a9 node`uv_cond_wait at thread.c:780:7 frame #4: 0x0000000001376ea4 node`worker(arg=0x0000000000000000) at threadpool.c:76:7 frame #5: 0x00007f82968f7609 libpthread.so.0`start_thread(arg=<unavailable>) at pthread_create.c:477:8 frame #6: 0x00007f829681e293 libc.so.6`__clone + 67 thread #9, stop reason = signal 0 frame #0: 0x00007f82968fe376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13 frame #1: 0x00007f82968fe359 libpthread.so.0`__pthread_cond_wait at pthread_cond_wait.c:508 frame #2: 0x00007f82968fe290 libpthread.so.0`__pthread_cond_wait(cond=0x000000000446c760, mutex=0x000000000446c720) at pthread_cond_wait.c:638 frame #3: 0x000000000138a4a9 node`uv_cond_wait at thread.c:780:7 frame #4: 0x0000000001376ea4 node`worker(arg=0x0000000000000000) at threadpool.c:76:7 frame #5: 0x00007f82968f7609 libpthread.so.0`start_thread(arg=<unavailable>) at pthread_create.c:477:8 frame #6: 0x00007f829681e293 libc.so.6`__clone + 67 thread #10, stop reason = signal 0 frame #0: 0x0000000000cff942 node`v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) + 1282 frame #1: 0x0000000000c6c9bb node`non-virtual thunk to v8::internal::CancelableTask::Run() + 59 frame #2: 0x0000000000a71405 node`node::(anonymous namespace)::PlatformWorkerThread(void*) + 405 frame #3: 0x00007f82968f7609 libpthread.so.0`start_thread(arg=<unavailable>) at pthread_create.c:477:8 frame #4: 0x00007f829681e293 libc.so.6`__clone + 67 thread #11, stop reason = signal 0 frame #0: 0x0000000000cffec1 node`v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) + 2689 frame #1: 0x0000000000c6c9bb node`non-virtual thunk to v8::internal::CancelableTask::Run() + 59 frame #2: 0x0000000000a71405 node`node::(anonymous namespace)::PlatformWorkerThread(void*) + 405 frame #3: 0x00007f82968f7609 libpthread.so.0`start_thread(arg=<unavailable>) at pthread_create.c:477:8 frame #4: 0x00007f829681e293 libc.so.6`__clone + 67This happens if you use the
--no-concurrent-marking(blocking GC) too:* thread #1: tid = 18979, 0x0000000000d7be8d node`unsigned long v8::internal::MarkCompactCollector::ProcessMarkingWorklist<(v8::internal::MarkCompactCollector::MarkingWorklistProcessingMode)0>(unsigned long) + 173, name = 'node', stop reason = signal SIGSEGV thread #2: tid = 18981, 0x00007f60e59b7376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13, stop reason = signal 0 thread #3: tid = 18980, 0x00007f60e58d75ce libc.so.6`epoll_wait + 94, stop reason = signal 0 thread #4: tid = 18985, 0x00007f60e59ba3f4 libpthread.so.0`do_futex_wait at futex-internal.h:320:13, stop reason = signal 0 thread #5: tid = 18992, 0x00007f60e59b7376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13, stop reason = signal 0 thread #6: tid = 18984, 0x00007f60e59b7376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13, stop reason = signal 0 thread #7: tid = 18990, 0x00007f60e59b7376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13, stop reason = signal 0 thread #8: tid = 18991, 0x00007f60e59b7376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13, stop reason = signal 0 thread #9: tid = 18982, 0x00007f60e59b7376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13, stop reason = signal 0 thread #10: tid = 18989, 0x00007f60e59b7376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13, stop reason = signal 0 thread #11: tid = 18983, 0x00007f60e59b7376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13, stop reason = signal 0I wonder if #37106 (SIGSEGV for address: 0x0 in ConcurrentMarking::Run) is related; similar story, upgrading from 12 to 14 results in sporadic segfaults, seemingly when causing garbage collection and using concurrent socket requests.
@thomasmichaelwallace thank you for your confirmation. From my point of view #37106 may be related or even the same issue. However, from the issue description I was not able to derive a direct connection between the two problems.
As far as I can tell, the segfault from our use-case is caused by an error in the Node.js/v8 internal data-structure/memory management that gets triggered on high load in the main thread while reading big data chunks from sockets.
Since this seams to be a legit use case for Node.js I was wondering why nobody else had this problem, or even why nobody seems to be worried about this issues, as errors in memory-management may lead to serious security problems in some cases.From my research i found out, that in the past @gireeshpunathil solved some issues with similar context (#25814). Maybe we should ask him for advice?
Reacted by thomas michael wallace, Denis Frenademetz, Gireesh Punathil and IamgabrielsoftReacted by Gireesh Punathillet us start by understanding the failing context a little deeper:
- select the failing thread:
thread listfollowed bythread select <thread number>
- get the instruction pointer:
reg read rip
- dump few instruction
behindthe faulty onedi -s <current rip value - 80> -c 20
just want to state that it is going to be an iterative process!
alternatively, if you have a standalone recreate, let me know - I can reproduce and debug myself!Reacted by Ivan HellReacted by thomas michael wallace- select the failing thread:
Hi @gireeshpunathil, thank you for your fast reply.
Right now I have no pc at hand. I will try to provide you with the requestes information later today, as soon as I get home. Do you have any preferences for the debugger (gdb or llnode)?Besides: In the issue description I referenced a repositor that contains a sample application, that eventually should recreate the issue.
Reacted by Gireesh Punathil- gdb is fine
- the code in the referenced repo - will try
never mind, I am able to recreate with your sample program! thanks for the nice setup for the recreate!
#node ./new_server_test.js & [1] 89290 #Server listening on port 5673 #node ./new_client_test.js Got startsegment! 0: 2.147s 1: 1.820s 2: 2.043s 3: 2.315s 4: 2.168s Connection closed [1]+ Segmentation fault: 11 node ./new_server_test.js #lI will debug and let you know!
Reacted by thomas michael wallaceThank you for taking a look at this @gireeshpunathil!
It would seem too much of a coincidence for mine and hellivan's socket+gc segfaults to have different underlying causes, so I'm going to trust that his reproduction repo will be enough.
For what it's worth, here are the results of my following those commands; just in case it becomes immediately obvious to you that mine is a different issue, which I should separately raise:
(llnode) thread list Process 21256 stopped * thread #1: tid = 21261, 0x0000000000cff994 node`v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) + 1364, name = 'node', stop reason = signal SIGSEGV thread #2: tid = 21256, 0x00007f352d5288e9 libc.so.6`___lldb_unnamed_symbol1092$$libc.so.6 + 633, stop reason = signal 0 thread #3: tid = 21272, 0x00007f352d59c376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13, stop reason = signal 0 thread #4: tid = 21262, 0x00007f352d59f3f4 libpthread.so.0`do_futex_wait at futex-internal.h:320:13, stop reason = signal 0 thread #5: tid = 21257, 0x00007f352d4bc5ce libc.so.6`epoll_wait + 94, stop reason = signal 0 thread #6: tid = 21259, 0x0000000000cff957 node`v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) + 1303, stop reason = signal 0 thread #7: tid = 21273, 0x00007f352d59c376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13, stop reason = signal 0 thread #8: tid = 21270, 0x00007f352d59c376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13, stop reason = signal 0 thread #9: tid = 21258, 0x00007f352d43174b libc.so.6`___lldb_unnamed_symbol378$$libc.so.6 + 43, stop reason = signal 0 thread #10: tid = 21260, 0x0000000000cfdb7d node`v8::internal::ConcurrentMarkingVisitor::ShouldVisit(v8::internal::HeapObject) + 109, stop reason = signal 0 thread #11: tid = 21271, 0x00007f352d59c376 libpthread.so.0`__pthread_cond_wait at futex-internal.h:183:13, stop reason = signal 0 (llnode) thread select 1 * thread #1, name = 'node', stop reason = signal SIGSEGV frame #0: 0x0000000000cff994 node`v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) + 1364 node`v8::internal::ConcurrentMarking::Run: -> 0xcff994 <+1364>: addb %al, (%rax) 0xcff996 <+1366>: addb %al, (%rax) 0xcff998 <+1368>: addb %al, (%rax) 0xcff99a <+1370>: addb %al, (%rax) (llnode) reg read rip rip = 0x0000000000cff994 node`v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) + 1364 (llnode) di -s (0x0000000000cff994-80) -c 20 node`v8::internal::ConcurrentMarking::Run: 0xcff944 <+1284>: addb %al, (%rax) 0xcff946 <+1286>: addb %al, (%rax) 0xcff948 <+1288>: addb %al, (%rax) 0xcff94a <+1290>: addb %al, (%rax) 0xcff94c <+1292>: addb %al, (%rax) 0xcff94e <+1294>: addb %al, (%rax) 0xcff950 <+1296>: addb %al, (%rax) 0xcff952 <+1298>: addb %al, (%rax) 0xcff954 <+1300>: addb %al, (%rax) 0xcff956 <+1302>: addb %al, (%rax) 0xcff958 <+1304>: addb %al, (%rax) 0xcff95a <+1306>: addb %al, (%rax) 0xcff95c <+1308>: addb %al, (%rax) 0xcff95e <+1310>: addb %al, (%rax) 0xcff960 <+1312>: addb %al, (%rax) 0xcff962 <+1314>: addb %al, (%rax) 0xcff964 <+1316>: addb %al, (%rax) 0xcff966 <+1318>: addb %al, (%rax) 0xcff968 <+1320>: addb %al, (%rax) 0xcff96a <+1322>: addb %al, (%rax)Hello! I am creator of #37106 and it looks like my issue is exactly the same.
Reacted by thomas michael wallace and Gireesh PunathilThank you very much for your efforts @gireeshpunathil.
If it still helps, this would be the output of my debug session in
llnode:(llnode) thread list Process 27845 stopped * thread #1: tid = 27847, 0x0000000000cff9c4 node`v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) + 1364, name = 'node', stop reason = signal SIGSEGV thread #2: tid = 27848, 0x0000000000cfc324 node`v8::internal::ConcurrentMarkingVisitor::VisitPointersInSnapshot(v8::internal::HeapObject, v8::internal::SlotSnapshot const&) + 68, stop reason = signal 0 thread #3: tid = 27849, 0x0000000000cff987 node`v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) + 1303, stop reason = signal 0 thread #4: tid = 27851, 0x00007f2db79c39ba libpthread.so.0`__futex_abstimed_wait_common64 + 202, stop reason = signal 0 thread #5: tid = 27850, 0x00007f2db786c8e2 libc.so.6`malloc + 770, stop reason = signal 0 thread #6: tid = 27845, 0x0000000000d49001 node`v8::internal::IncrementalMarking::RecordWriteSlow(v8::internal::HeapObject, v8::internal::FullHeapObjectSlot, v8::internal::HeapObject) + 65, stop reason = signal 0 thread #7: tid = 27853, 0x00007f2db79c39ba libpthread.so.0`__futex_abstimed_wait_common64 + 202, stop reason = signal 0 thread #8: tid = 27852, 0x00007f2db79c39ba libpthread.so.0`__futex_abstimed_wait_common64 + 202, stop reason = signal 0 thread #9: tid = 27846, 0x00007f2db78e039e libc.so.6`epoll_wait + 94, stop reason = signal 0 thread #10: tid = 27854, 0x00007f2db79c39ba libpthread.so.0`__futex_abstimed_wait_common64 + 202, stop reason = signal 0 thread #11: tid = 27855, 0x00007f2db79c39ba libpthread.so.0`__futex_abstimed_wait_common64 + 202, stop reason = signal 0 (llnode) thread select 1 * thread #1, name = 'node', stop reason = signal SIGSEGV frame #0: 0x0000000000cff9c4 node`v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) + 1364 node`v8::internal::ConcurrentMarking::Run: -> 0xcff9c4 <+1364>: addb %al, (%rax) 0xcff9c6 <+1366>: addb %al, (%rax) 0xcff9c8 <+1368>: addb %al, (%rax) 0xcff9ca <+1370>: addb %al, (%rax) (llnode) reg read rip rip = 0x0000000000cff9c4 node`v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) + 1364 (llnode) di -s (0x0000000000cff9c4-80) -c 20 node`v8::internal::ConcurrentMarking::Run: 0xcff974 <+1284>: addb %al, (%rax) 0xcff976 <+1286>: addb %al, (%rax) 0xcff978 <+1288>: addb %al, (%rax) 0xcff97a <+1290>: addb %al, (%rax) 0xcff97c <+1292>: addb %al, (%rax) 0xcff97e <+1294>: addb %al, (%rax) 0xcff980 <+1296>: addb %al, (%rax) 0xcff982 <+1298>: addb %al, (%rax) 0xcff984 <+1300>: addb %al, (%rax) 0xcff986 <+1302>: addb %al, (%rax) 0xcff988 <+1304>: addb %al, (%rax) 0xcff98a <+1306>: addb %al, (%rax) 0xcff98c <+1308>: addb %al, (%rax) 0xcff98e <+1310>: addb %al, (%rax) 0xcff990 <+1312>: addb %al, (%rax) 0xcff992 <+1314>: addb %al, (%rax) 0xcff994 <+1316>: addb %al, (%rax) 0xcff996 <+1318>: addb %al, (%rax) 0xcff998 <+1320>: addb %al, (%rax) 0xcff99a <+1322>: addb %al, (%rax)and
gdb(since the results fromllnodedo not make any sense to me) :(gdb) info registers rip 0xcff9c4 0xcff9c4 <v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*)+1364> (gdb) x/20i $rip-80 0xcff974 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1284>: subb $0x1,(%rax) 0xcff977 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1287>: add %al,(%rax) 0xcff979 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1289>: mov 0x80(%rax),%rsi 0xcff980 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1296>: mov -0x11a8(%rbp),%r15 0xcff987 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1303>: lea -0x1(%r15),%rax 0xcff98b <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1307>: cmp %rcx,%rax 0xcff98e <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1310>: setae %cl 0xcff991 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1313>: cmp %rdx,%rax 0xcff994 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1316>: setb %dl 0xcff997 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1319>: test %dl,%cl 0xcff999 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1321>: jne 0xd02d48 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+14552> 0xcff99f <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1327>: cmp %rsi,%rax 0xcff9a2 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1330>: je 0xd02d48 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+14552> 0xcff9a8 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1336>: mov -0x1(%r15),%r13 0xcff9ac <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1340>: cmpb $0x0,-0x11c8(%rbp) 0xcff9b3 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1347>: jne 0xd032e0 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+15984> 0xcff9b9 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1353>: mov -0x11a8(%rbp),%r15 0xcff9c0 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1360>: lea 0xa(%r13),%r14 => 0xcff9c4 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1364>: movzbl (%r14),%eax 0xcff9c8 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1368>: cmp $0x47,%al@bolt-juri-gavshin - thanks. @hellivan -thanks for the detailed data. As I said, I have the repro now! an interim update:
(gdb) set disassembly-flavor intel (gdb) x/i $rip => 0xcff9c4 <_ZN2v88internal17ConcurrentMarking3RunEiPNS1_9TaskStateE+1364>: movzx eax,BYTE PTR [r14] (gdb) i r r14 r14 0x3938363632303039 4123105065255841849 (gdb) x/b $r14 0x3938363632303039: Cannot access memory at address 0x3938363632303039 (gdb)
- the immediate cause of the crash is memory overwrite
- the issue is reproducible in
v14.xlines, not inv15.xlines - the issue vanishes if I run under a debugger, or
valgrind
Reacted by thomas michael wallace and Ivan Hell22 remaining items
Thanks for the prompt:
➜ valgrind --track-origins=yes node direct.js==955794== Thread 6: ==955794== Invalid read of size 1 ==955794== at 0xD1E474: v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) (in /home/ubuntu/.nvm/versions/node/v14.17.0/bin/node) ==955794== by 0xC8B13A: non-virtual thunk to v8::internal::CancelableTask::Run() (in /home/ubuntu/.nvm/versions/node/v14.17.0/bin/node) ==955794== by 0xA8FAF4: node::(anonymous namespace)::PlatformWorkerThread(void*) (in /home/ubuntu/.nvm/versions/node/v14.17.0/bin/node) ==955794== by 0x445D608: start_thread (pthread_create.c:477) ==955794== by 0x4FCE292: clone (clone.S:95) ==955794== Address 0x30312e313a is not stack'd, malloc'd or (recently) free'd ==955794== ==955794== ==955794== Process terminating with default action of signal 11 (SIGSEGV) ==955794== at 0x4469229: raise (raise.c:46) ==955794== by 0x9ECF32: node::TrapWebAssemblyOrContinue(int, siginfo_t*, void*) (in /home/ubuntu/.nvm/versions/node/v14.17.0/bin/node) ==955794== by 0x44693BF: ??? (in /usr/lib/x86_64-linux-gnu/libpthread-2.31.so) ==955794== by 0xD1E473: v8::internal::ConcurrentMarking::Run(int, v8::internal::ConcurrentMarking::TaskState*) (in /home/ubuntu/.nvm/versions/node/v14.17.0/bin/node) ==955794== by 0xC8B13A: non-virtual thunk to v8::internal::CancelableTask::Run() (in /home/ubuntu/.nvm/versions/node/v14.17.0/bin/node) ==955794== by 0xA8FAF4: node::(anonymous namespace)::PlatformWorkerThread(void*) (in /home/ubuntu/.nvm/versions/node/v14.17.0/bin/node) ==955794== by 0x445D608: start_thread (pthread_create.c:477) ==955794== by 0x4FCE292: clone (clone.S:95) ==955794== ==955794== HEAP SUMMARY: ==955794== in use at exit: 619,939,748 bytes in 156,655 blocks ==955794== total heap usage: 1,748,136 allocs, 1,591,481 frees, 4,705,060,432 bytes allocated ==955794== ==955794== LEAK SUMMARY: ==955794== definitely lost: 4,362 bytes in 3 blocks ==955794== indirectly lost: 0 bytes in 0 blocks ==955794== possibly lost: 123,032 bytes in 18 blocks ==955794== still reachable: 619,812,354 bytes in 156,634 blocks ==955794== of which reachable via heuristic: ==955794== stdstring : 41,928 bytes in 859 blocks ==955794== suppressed: 0 bytes in 0 blocks ==955794== Rerun with --leak-check=full to see details of leaked memory ==955794== ==955794== For lists of detected and suppressed errors, rerun with: -s ==955794== ERROR SUMMARY: 1 errors from 1 contexts (suppressed: 0 from 0)(I'll update with the suggested output suggested
-s --leak-check-full) once it completes.As I previously pointed out, I always had a
differencein passing / failing versions with that of @thomasmichaelwallace . Now I am able to recreate a crash with reasonable consistency, but with a different stack. I can open a different issue, but once we progress a bit more and conclude that we are seeing two different things. As of now, I don't know about it:v8::internal::ConcurrentMarking::Runversusv8::internal::ScavengingTask::RunInParallelwithv8::internal::CancelableTask::Runcommon in the stack.(gdb) where #0 0x00005564483eda05 in v8::internal::MemoryChunk::InYoungGeneration ( this=0x0) at ../deps/v8/src/heap/spaces.h:837 837 return (GetFlags() & kIsInYoungGenerationMask) != 0; #1 v8::internal::Heap::InYoungGeneration (heap_object=...) at ../deps/v8/src/heap/heap-inl.h:389 #2 0x000055644849ac41 in v8::internal::Scavenger::ScavengeObject<v8::internal::FullHeapObjectSlot> (this=this@entry=0x55644cbcb890, p=p@entry=..., object=object@entry=...) at ../deps/v8/src/objects/heap-object.h:219 #3 0x000055644849e0ea in v8::internal::ScavengeVisitor::VisitHeapObjectImpl<v8::internal::FullObjectSlot> (this=0x7ffe4d4a8d00, heap_object=..., slot=...) at ../deps/v8/src/base/macros.h:365 #4 v8::internal::ScavengeVisitor::VisitPointersImpl<v8::internal::FullObjectSlot> (end=..., start=..., this=<optimized out>, host=...) at ../deps/v8/src/heap/scavenger-inl.h:474 #5 v8::internal::ScavengeVisitor::VisitPointers (end=..., start=..., host=..., this=<optimized out>) at ../deps/v8/src/heap/scavenger-inl.h:427 #6 v8::internal::BodyDescriptorBase::IteratePointers<v8::internal::ScavengeVisitor> (obj=..., obj@entry=..., end_offset=end_offset@entry=112, v=v@entry=0x7ffe4d4a8d00, start_offset=8) at ../deps/v8/src/objects/objects-body-descriptors-inl.h:127 --Type <RET> for more, q to quit, c to continue without paging-- #7 0x000055644849fcda in v8::internal::FlexibleBodyDescriptor<8>::IterateBody<v8::internal::ScavengeVisitor> (v=0x7ffe4d4a8d00, object_size=112, obj=..., map=...) at ../deps/v8/src/objects/objects-body-descriptors.h:118 #8 v8::internal::HeapVisitor<int, v8::internal::ScavengeVisitor>::VisitStruct (object=..., map=..., this=0x7ffe4d4a8d00) at ../deps/v8/src/heap/objects-visiting-inl.h:154 #9 v8::internal::HeapVisitor<int, v8::internal::ScavengeVisitor>::Visit ( this=this@entry=0x7ffe4d4a8d00, map=..., object=object@entry=...) at ../deps/v8/src/heap/objects-visiting-inl.h:59 #10 0x00005564484a3fd9 in v8::internal::HeapVisitor<int, v8::internal::ScavengeVisitor>::Visit (object=..., this=0x7ffe4d4a8d00) at /usr/include/x86_64-linux-gnu/bits/string_fortified.h:34 #11 v8::internal::Scavenger::Process (this=0x55644cbcb890, barrier=<optimized out>) at ../deps/v8/src/heap/scavenger.cc:547 #12 0x00005564484a49ca in v8::internal::ScavengingTask::ProcessItems ( this=0x55644cbf3b10) at ../deps/v8/src/heap/scavenger.cc:70 #13 v8::internal::ScavengingTask::RunInParallel (this=0x55644cbf3b10, runner=<optimized out>) at ../deps/v8/src/heap/scavenger.cc:49 #14 0x000055644842eebf in v8::internal::ItemParallelJob::Task::RunInternal ( --Type <RET> for more, q to quit, c to continue without paging-- this=<optimized out>) at ../deps/v8/src/heap/item-parallel-job.cc:34 #15 v8::internal::CancelableTask::Run (this=<optimized out>) at ../deps/v8/src/tasks/cancelable-task.h:155 #16 v8::internal::ItemParallelJob::Run (this=this@entry=0x7ffe4d4a9010) at ../deps/v8/src/heap/item-parallel-job.cc:103 #17 0x00005564484a20ed in v8::internal::ScavengerCollector::CollectGarbage ( this=0x55644cb46880) at ../deps/v8/src/heap/scavenger.cc:303 #18 0x00005564483f1083 in v8::internal::Heap::Scavenge ( this=this@entry=0x55644caddb90) at /usr/include/c++/9/bits/unique_ptr.h:360 #19 0x000055644841e454 in v8::internal::Heap::PerformGarbageCollection ( this=this@entry=0x55644caddb90, collector=collector@entry=v8::internal::SCAVENGER, gc_callback_flags=gc_callback_flags@entry=v8::kNoGCCallbackFlags) at ../deps/v8/src/heap/heap.cc:2028 #20 0x000055644841ed62 in v8::internal::Heap::CollectGarbage ( this=this@entry=0x55644caddb90, space=space@entry=v8::internal::NEW_SPACE, gc_reason=gc_reason@entry=v8::internal::GarbageCollectionReason::kAllocationFailure, gc_callback_flags=gc_callback_flags@entry=v8::kNoGCCallbackFlags) at ../deps/v8/src/heap/heap.cc:1587 --Type <RET> for more, q to quit, c to continue without paging-- #21 0x0000556448421daf in v8::internal::Heap::AllocateRawWithLightRetrySlowPath (this=this@entry=0x55644caddb90, size=size@entry=16, allocation=v8::internal::AllocationType::kYoung, origin=origin@entry=v8::internal::AllocationOrigin::kRuntime, alignment=alignment@entry=v8::internal::kDoubleUnaligned) at ../deps/v8/include/v8-internal.h:223 #22 0x0000556448421f45 in v8::internal::Heap::AllocateRawWithRetryOrFailSlowPath (this=this@entry=0x55644caddb90, size=size@entry=16, allocation=<optimized out>, origin=origin@entry=v8::internal::AllocationOrigin::kRuntime, alignment=alignment@entry=v8::internal::kDoubleUnaligned) at ../deps/v8/src/heap/heap.cc:5000 #23 0x00005564483c55a1 in v8::internal::Heap::AllocateRawWith<(v8::internal::Heap::AllocationRetryMode)1> (this=this@entry=0x55644caddb90, size=size@entry=16, allocation=allocation@entry=v8::internal::AllocationType::kYoung, origin=origin@entry=v8::internal::AllocationOrigin::kRuntime, alignment=alignment@entry=v8::internal::kDoubleUnaligned) at ../deps/v8/src/objects/heap-object.h:108 #24 0x00005564483c56e8 in v8::internal::Factory::AllocateRaw ( --Type <RET> for more, q to quit, c to continue without paging-- this=this@entry=0x55644cad4880, size=size@entry=16, allocation=allocation@entry=v8::internal::AllocationType::kYoung, alignment=alignment@entry=v8::internal::kDoubleUnaligned) at ../deps/v8/src/execution/isolate.h:913 #25 0x00005564483ab919 in v8::internal::FactoryBase<v8::internal::Factory>::AllocateRaw (this=this@entry=0x55644cad4880, size=size@entry=16, allocation=allocation@entry=v8::internal::AllocationType::kYoung, alignment=alignment@entry=v8::internal::kDoubleUnaligned) at ../deps/v8/src/heap/factory-base.cc:236 #26 0x00005564483ab930 in v8::internal::FactoryBase<v8::internal::Factory>::AllocateRawWithImmortalMap (this=this@entry=0x55644cad4880, size=size@entry=16, allocation=allocation@entry=v8::internal::AllocationType::kYoung, map=..., alignment=alignment@entry=v8::internal::kDoubleUnaligned) at ../deps/v8/src/heap/factory-base.cc:227 #27 0x00005564483b24d1 in v8::internal::Factory::NewHeapNumber<(v8::internal::AllocationType)0> (this=this@entry=0x55644cad4880) at /usr/include/x86_64-linux-gnu/bits/string_fortified.h:34 #28 0x00005564483b42a2 in v8::internal::Factory::NewHeapNumber<(v8::internal::AllocationType)0> (value=6.9529954760082175e-310, this=0x55644cad4880) --Type <RET> for more, q to quit, c to continue without paging-- at ../deps/v8/src/heap/factory-inl.h:66 #29 v8::internal::Factory::NewNumber<(v8::internal::AllocationType)0> ( this=0x55644cad4880, value=value@entry=0.33400000000000002) at ../deps/v8/src/heap/factory.cc:2029 #30 0x000055644857fe19 in v8::internal::JsonParser<unsigned short>::ParseJsonNumber (this=this@entry=0x7ffe4d4aa580) at ../deps/v8/src/execution/isolate.h:1059 #31 0x0000556448581278 in v8::internal::JsonParser<unsigned short>::ParseJsonValue (this=this@entry=0x7ffe4d4aa580) at /usr/include/c++/9/ext/new_allocator.h:89 #32 0x0000556448581be2 in v8::internal::JsonParser<unsigned short>::ParseJson ( this=this@entry=0x7ffe4d4aa580) at ../deps/v8/src/json/json-parser.cc:309 #33 0x00005564481e2558 in v8::internal::JsonParser<unsigned short>::Parse ( reviver=..., source=..., isolate=0x55644cad4880) at ../deps/v8/src/handles/handles.h:108 #34 v8::internal::Builtin_Impl_JsonParse (args=..., isolate=isolate@entry=0x55644cad4880) at ../deps/v8/src/builtins/builtins-json.cc:24 #35 0x00005564481e3bd0 in v8::internal::Builtin_JsonParse (args_length=6, args_object=0x7ffe4d4aa690, isolate=0x55644cad4880) --Type <RET> for more, q to quit, c to continue without paging-- at ../deps/v8/src/builtins/builtins-json.cc:16 #36 0x0000556448fa1ba0 in Builtins_CEntry_Return1_DontSaveFPRegs_ArgvOnStack_BuiltinExit () at ../../deps/v8/../../deps/v8/src/builtins/promise-misc.tq:91 #37 0x0000556448d9f458 in Builtins_InterpreterEntryTrampoline () at ../../deps/v8/../../deps/v8/src/objects/string.tq:72
(gdb) p this $1 = (const v8::internal::MemoryChunk * const) 0x0
ok - another thread too seems to have entered the scavenge cycle, which is in the process of re-arranging the chunks / objects. Can scavenge run in parallel? If so, how do the threads co-ordinate?
(gdb) t 3 (gdb) where #0 0x0000557bca39e3f0 in v8::base::List<v8::internal::MemoryChunk>::Contains ( this=0x557bce712f70, element=0x49b03cc0000) at ../deps/v8/src/base/list.h:60 #1 v8::base::List<v8::internal::MemoryChunk>::Remove (element=0x49b03cc0000, this=0x557bce712f70) at ../deps/v8/src/base/list.h:43 #2 v8::internal::PagedSpace::RemovePage (this=this@entry=0x557bce712f50, page=page@entry=0x49b03cc0000) at ../deps/v8/src/heap/spaces.cc:1841 #3 0x0000557bca3acba0 in v8::internal::PagedSpace::RefillFreeList ( this=0x557bce7d0590) at ../deps/v8/src/heap/spaces.cc:1700 #4 0x0000557bca3a9f5e in v8::internal::PagedSpace::RawSlowRefillLinearAllocationArea (this=0x557bce7d0590, size_in_bytes=32, origin=v8::internal::AllocationOrigin::kGC) at ../deps/v8/src/heap/spaces.cc:3809 #5 0x0000557bca29857a in v8::internal::PagedSpace::EnsureLinearAllocationArea ( origin=v8::internal::AllocationOrigin::kGC, size_in_bytes=32, this=0x557bce7d0590) at ../deps/v8/src/heap/spaces-inl.h:387 #6 v8::internal::PagedSpace::EnsureLinearAllocationArea ( origin=v8::internal::AllocationOrigin::kGC, size_in_bytes=32, this=0x557bce7d0590) at ../deps/v8/src/heap/spaces-inl.h:382 #7 v8::internal::PagedSpace::AllocateRawUnaligned ( this=this@entry=0x557bce7d0590, size_in_bytes=size_in_bytes@entry=32, origin=origin@entry=v8::internal::AllocationOrigin::kGC) at ../deps/v8/src/heap/spaces-inl.h:419 #8 0x0000557bca2994b4 in v8::internal::PagedSpace::AllocateRaw ( this=0x557bce7d0590, size_in_bytes=32, alignment=<optimized out>, origin=v8::internal::AllocationOrigin::kGC) at ../deps/v8/src/heap/spaces-inl.h:483 #9 0x0000557bca37f06b in v8::internal::LocalAllocator::Allocate ( alignment=v8::internal::kWordAligned, origin=v8::internal::AllocationOrigin::kGC, object_size=32, space=v8::internal::OLD_SPACE, this=0x557bce7d0578) at ../deps/v8/src/heap/spaces.h:3107 #10 v8::internal::Scavenger::PromoteObject<v8::internal::FullHeapObjectSlot> ( object_fields=v8::internal::ObjectFields::kDataOnly, object_size=32, object=..., slot=..., map=..., this=0x557bce7d04e0) at ../deps/v8/src/heap/scavenger-inl.h:174 #11 v8::internal::Scavenger::EvacuateObjectDefault<v8::internal::FullHeapObjectSlot> (this=this@entry=0x557bce7d04e0, map=map@entry=..., slot=slot@entry=..., object=..., object_size=object_size@entry=32, object_fields=v8::internal::ObjectFields::kDataOnly) at ../deps/v8/src/heap/scavenger-inl.h:259 #12 0x0000557bca37fae8 in v8::internal::Scavenger::EvacuateObject<v8::internal::FullHeapObjectSlot> (source=..., map=..., slot=..., this=0x557bce7d04e0) at ../deps/v8/src/objects/map.h:814 #13 v8::internal::Scavenger::ScavengeObject<v8::internal::FullHeapObjectSlot> ( this=0x557bce7d04e0, p=p@entry=..., object=...) at ../deps/v8/src/heap/scavenger-inl.h:396 #14 0x0000557bca38019f in v8::internal::IterateAndScavengePromotedObjectsVisitor::HandleSlot<v8::internal::FullHeapObjectSlot> (this=this@entry=0x7f385a7fbb20, --Type <RET> for more, q to quit, c to continue without paging-- host=host@entry=..., slot=slot@entry=..., target=..., target@entry=...) at ../deps/v8/src/base/atomic-utils.h:149 #15 0x0000557bca380635 in v8::internal::IterateAndScavengePromotedObjectsVisitor::VisitPointersImpl<v8::internal::FullObjectSlot> (end=..., start=..., host=..., this=0x7f385a7fbb20) at ../deps/v8/src/base/macros.h:365 #16 v8::internal::IterateAndScavengePromotedObjectsVisitor::VisitPointers ( end=..., start=..., host=..., this=0x7f385a7fbb20) at ../deps/v8/src/heap/scavenger.cc:94 #17 v8::internal::BodyDescriptorBase::IteratePointers<v8::internal::IterateAndScavengePromotedObjectsVisitor> (obj=obj@entry=..., end_offset=end_offset@entry=112, v=v@entry=0x7f385a7fbb20, start_offset=8) at ../deps/v8/src/objects/objects-body-descriptors-inl.h:127 #18 0x0000557bca380eef in v8::internal::FlexibleBodyDescriptor<8>::IterateBody<v8::internal::IterateAndScavengePromotedObjectsVisitor> (v=0x7f385a7fbb20, object_size=112, obj=..., map=...) at ../deps/v8/src/objects/objects-body-descriptors.h:118 --Type <RET> for more, q to quit, c to continue without paging-- #19 v8::internal::CallIterateBody::apply<v8::internal::FlexibleBodyDescriptor<8>, v8::internal::IterateAndScavengePromotedObjectsVisitor> (v=0x7f385a7fbb20, object_size=112, obj=..., map=...) at ../deps/v8/src/objects/objects-body-descriptors-inl.h:1088 #20 v8::internal::BodyDescriptorApply<v8::internal::CallIterateBody, void, v8::internal::Map, v8::internal::HeapObject, int, v8::internal::IterateAndScavengePromotedObjectsVisitor*> (p4=0x7f385a7fbb20, p3=112, p2=..., p1=..., type=<optimized out>) at ../deps/v8/src/objects/objects-body-descriptors-inl.h:1056 #21 v8::internal::HeapObject::IterateBodyFast<v8::internal::IterateAndScavengePromotedObjectsVisitor> (this=<synthetic pointer>, v=0x7f385a7fbb20, object_size=112, map=...) at ../deps/v8/src/objects/objects-body-descriptors-inl.h:1094 #22 v8::internal::Scavenger::IterateAndScavengePromotedObject ( this=this@entry=0x557bce7d04e0, target=target@entry=..., map=map@entry=..., size=size@entry=112) at ../deps/v8/src/heap/scavenger.cc:471 --Type <RET> for more, q to quit, c to continue without paging-- #23 0x0000557bca389193 in v8::internal::Scavenger::Process ( this=0x557bce7d04e0, barrier=<optimized out>) at ../deps/v8/src/heap/scavenger.cc:559 #24 0x0000557bca389b8a in v8::internal::ScavengingTask::ProcessItems ( this=0x557bd0014c70) at ../deps/v8/src/heap/scavenger.cc:70 #25 v8::internal::ScavengingTask::RunInParallel (this=0x557bd0014c70, runner=<optimized out>) at ../deps/v8/src/heap/scavenger.cc:54 #26 0x0000557bca313a91 in v8::internal::ItemParallelJob::Task::RunInternal ( this=0x557bd0014c70) at ../deps/v8/src/heap/item-parallel-job.cc:34 #27 0x0000557bca16d121 in non-virtual thunk to v8::internal::CancelableTask::Run() () at ../deps/v8/src/heap/heap-write-barrier-inl.h:213 #28 0x0000557bc9dc28d7 in node::(anonymous namespace)::PlatformWorkerThread ( data=0x557bce693670) at ../src/node_platform.cc:43 #29 0x00007f3861162609 in start_thread (arg=<optimized out>) at pthread_create.c:477 #30 0x00007f3861089293 in clone ()
/cc @addaleax @nodejs/v8
Is this issue what https://chromium-review.googlesource.com/c/v8/v8/+/2988414 fixes?
Reacted by Gireesh PunathilReacted by thomas michael wallacehighly probable - as the context seems similar.
I was so excited by this that I immediately did the following:
- confirm the bug still exists on master (yes, my reproduction still works)
- apply exactly the patch @vlovich linked to (lit. changed three files)
- confirm the bug no longer exists on my build (yup, fixed! 😄)
- apply patch to last v14 because that's where I/we actually need it (n.b. files are not identical at v14, but patch still 'fits')
- confirm the bug still doesn't exist with my patched v14 build (yup, also fixes! 😃 )
As @gireeshpunathil says, it fits the narrative. The bug was [probably] introduced when v8 was updated, it seems to be caused by some combination of parsing json, from a buffer, in a threaded way, where the garbage collector sets off; which is exactly what the patch addresses.
So we have the fix!
My problem is getting it anywhere. I'm happy to (and probably will, unless stopped :P) submit this patch as a PR against master. But for it to be useful to myself (and anyone else enjoying this problem in AWS lambda) it's got to make its way back to 14.What's the best way of making this happen? Is there someone better to get this done than me?Reacted by Gireesh Punathilthanks for confirming this @thomasmichaelwallace !!
the patch will be consumed here naturally, but takes its own sweet time. Pinging @targos to know the standard procedure v8 changes to Node: do we pro-actively PR in master, or cherry-pick a bunch of v8 patches occasionally, or consume v8 only on version boundaries.
- added 2 commits that reference this issue
on Jul 15, 2021 - added a commit that references this issue
on May 22, 2026
14.15.5Linux WorkMachine 5.11.1-arch1-1 #1 SMP PREEMPT Tue, 23 Feb 2021 14:05:30 +0000 x86_64 GNU/LinuxWhat steps will reproduce the bug?
As far as we found out, the segfault happens if
Node.jssends/receives lots of data via sockets and processes it in an expensive synchronous method (e.g.JSON.parse).The original problem involved some basic
JSONdata processing where the data was received from aRabbitMQusing the amqplib npm package. Meanwhile we were able to recreate the problem by only usingNode.jsinternal mechanisms (netpackage) in this sample repository:https://github.com/hellivan/nodejs-14.15.5-ConcurrentMarking-segfault
How often does it reproduce? Is there a required condition?
The error only reproduces under uncertain conditions that are difficult to replicate. Under normal circumstances, it may possible that the application runs for hours and then crashes without a reason. However it may also happen that it crashes right after the start.
What is the expected behavior?
Node.jsruntime should executeJSapplication without interruptions.What do you see instead?
Node.jscrashes with aSIGSEGV.Additional information
During the analysis of the original application crashes, we were able to extract some coredumps which are listed below. Due to privacy reasons we replaced some paths in the results. Due to the complexity of the original application, we created a reduced sample application, which we hope reproduces the same segmentation fault as the original one. During our tests, we found out that other
Node.jsversions may be affected by this bug, too. We were able to sporadically reproduce the issue forNode.jsversions14.16.0and15.10.0.If you need any help or information regarding the coredumps please let me know.
1. Coredump
General information about node instance
List of all threads
Threads' backtrace
2. Coredump
List of all threads
Threads' backtrace
3. Coredump
List of all threads
Threads' backtrace