Skip to content

Concurrent PGlite WebAssembly initialization intermittently SIGSEGVs Node.js #64500

Description

@gadicc

Version

v26.5.0 (also reproduced on Node 22, 24, 25)

Platform

The primary Docker reproduction ran on this host:


Linux dragon-pro18 6.18.35-1-cachyos-lts #1 SMP PREEMPT_DYNAMIC Sat, 13 Jun 2026 12:39:40 +0000 x86_64 GNU/Linux
Intel Core Ultra 9 285HX, 24 cores online


Docker replaces the userspace and Node installation but shares the host kernel and CPU. The same failure occurs with the host Node binary.

Subsystem

V8 / WebAssembly (observed through child_process)

What steps will reproduce the bug?

The minimal repository is:

https://github.com/gadicc/node-pglite-wasm-sigsegv-repro

It has one runtime dependency, @electric-sql/pglite@0.5.4. PGlite ships its PostgreSQL engine as WebAssembly; there is no native addon.

On a machine with at least 20 GiB available:

git clone https://github.com/gadicc/node-pglite-wasm-sigsegv-repro.git
cd node-pglite-wasm-sigsegv-repro
docker build -t pglite-node-sigsegv-repro .
docker run --rm pglite-node-sigsegv-repro

Equivalently, with a local Node installation:

npm ci
npm run repro

Each of 16 child processes independently creates an in-memory PGlite client, runs SELECT 1, awaits client.close() in finally, and exits. The parent repeats that wave up to 50 times and stops on the first failure. There is no Vitest, Vite, test runner, application code, native addon, IPC dependency, or database persistence.

The memory warning matters: the observed cgroup peak was approximately 19 GiB. The recorded failing run had oom=0 and oom_kill=0.

How often does it reproduce? Is there a required condition?

It is intermittent and requires concurrent processes in the measured setup.

Two clean Node 26 official-image runs failed at waves 13 and 8. Other results
without V8 CLI overrides were:

Node Result
25 1/20 waves failed
24 4/45 waves failed
22 1/10 waves failed
20 0/40 waves failed

The Node 20 result is a non-reproduction in an intermittent test, not evidence that it is unaffected. Official Node 20, 22, 24, and 26 images expose the same relevant defaults: Liftoff, lazy compilation, dynamic tiering, and WASM tier-up are enabled, with up to 128 compilation tasks. Node 20 did reproduce with --no-liftoff, which forces the optimizing compiler.

The Node 27 V8 canary failed at wave 17 with child_process.fork() and failed 2/50 waves with ordinary child_process.spawn(). A sequential one-child control passed 50/50 waves.

What is the expected behavior? Why is that the expected behavior?

Every child should complete the query, close the PGlite client, and exit with code 0.

What do you see instead?

An intermittent child is terminated by SIGSEGV:

node=v26.5.0 v8=14.6.202.34-node.24 platform=linux arch=x64 children=16 waves=50
wave=1 passed=16/16
...
wave=8 passed=15/16
child=10 code=null signal=SIGSEGV elapsedMs=4315
failedWaves=1 completedWaves=8 requestedWaves=50

Additional information

Native diagnostics from a sanitized-environment reproduction:

  1. strace recorded an initial SIGSEGV with SEGV_MAPERR at a large guarded WASM address.
  2. Node's node::TrapWebAssemblyOrContinue called raise(SIGSEGV) after V8 declined to handle the address as a valid WASM trap.
  3. The retained core records the resulting SI_TKILL signal.
  4. GDB placed another crash in the lazy WASM compilation path through LiftoffAssembler::PrepareCall, ExecuteLiftoffCompilation, CompileLazy, and Runtime_WasmCompileLazy, with background Turboshaft compilation active. A separate retained crash was in v8::internal::MarkingBarrier::MarkValueLocal.

Relevant flag controls on Node 25:

  • default: reproduced;
  • --wasm-enforce-bounds-checks: reproduced;
  • --wasm-num-compilation-tasks=1: reproduced;
  • --no-wasm-tier-up: passed 100/100 waves in that version;
  • --no-liftoff: reproduced.

The Node 27 canary nevertheless reproduced with both --no-wasm-tier-up and --liftoff-only, so no single WASM tier is universally required.

I also built the exact PGlite 0.5.4 tag using its official pnpm build:all:debug workflow. That PostgreSQL WASM used -g,
-gsource-map, and --no-wasm-opt; LLVM verified 1,322 DWARF compilation units without errors. With Node 25.2.1, a sequential control passed 5/5 waves and eight concurrent children passed 5/5 waves, while 16 concurrent children
reproduced SIGSEGV in wave 2. The cgroup peak was 67,306,729,472 bytes (62.7 GiB), with oom=0 and oom_kill=0.

Runtime controls using the same PGlite version and concurrency:

  • Bun 1.3.11 / JavaScriptCore: 50/50 waves passed;
  • Deno 2.8.0 / V8 14.9.207.2, default: 50/50 passed;
  • Deno 2.8.0 / V8 14.9.207.2, --no-liftoff: 50/50 passed.

The evidence therefore establishes a Node-runtime-specific interaction in the tested configurations, but does not yet distinguish Node's WASM trap/embedding layer or build configuration from a V8 defect exposed by Node. d8 is V8's standalone JavaScript shell; reproducing there without Node or its APIs would prove that the bug exists independently in V8. I would appreciate guidance on whether this should be routed or cross-filed to V8.

On 2026-07-14 I searched open Node and V8 reports for PGlite, TrapWebAssemblyOrContinue, WasmCompileLazy, LiftoffAssembler, and concurrent WASM SIGSEGVs, and found no exact duplicate. The reproduction README records the related but materially different reports.

Lastly,

  • Co-authored with Codex gpt-5.6-sol, reasoning: high, with significant human input.
  • This issue was opened by me personally by hand and without any automation.

Edit 1: Added PGLite debug build info in "Additional Information" above.

Activity

  1. gadicc commented on Jul 30, 2026

    @gadicc
    Author

    Investigation update: the evidence points to platform-level address corruption on this machine, not to Node or V8. I am closing this issue.

    I spent a day doing native debugging on the affected machine (the same host as the original report). Short version: multiple independent crash captures show the kernel-reported fault address equal to the architecturally intended address with a single high bit set (2^42), while the intended address was validly mapped and the fault-time registers are clean. No software mechanism can produce that. A second machine is clean. I now believe this is a hardware/platform issue on this specific machine, with Node + PGlite acting as an unusually effective trigger.

    The signature

    Method: children run under gdb --batch with handle SIGSEGV stop nopass, capturing the pristine fault context before Node's trap handler runs.

    Capture Faulting instruction Intended address (from registers) si_addr (kernel) Intended address mapped?
    Node 25.2.1, gdb addl $1, 0x1c0(%r13), r13=0x6720080 0x6720240 0x40006720240 yes, inside [heap] (rw)
    Node 25.2.1, gdb mov %rbp, 0xb0(%r13), r13=0x6720080 0x6720130 0x40006720130 yes, inside [heap] (rw)
    Node 25.2.1, strace mov %rbp, 0xb0(%r13) (wasm entry) 0x44905130 0x40044905130 —

    Every sample: si_addr = intended_address + 2^42. strace -e %memory additionally shows the faulting address was never mapped, unmapped, or protected by the process at any point.

    Why this cannot be software: a software bug can only corrupt architectural state (registers or memory), which would be visible in the fault-time register dump. The dump is clean and self-consistent — r13 holds the intended value, which points into a live heap mapping. There is no x86-64 mechanism (segment bases, LAM, canonicalization) that adds 2^42 to a plain register-relative access, and page-table or TLB-shootdown bugs cannot change the linear address reported in CR2. The varying crash sites recorded earlier (Liftoff code emission, wasm entry, GC marking barrier) are explained by the anomaly striking whatever memory access is in flight.

    Cross-checks

    • Deno 2.9.3 (V8 14.9.207.2) crashes under the identical gdb harness, with a flipped-high-bit fault address (si_addr=0x167f20915dad versus live pointer rbx=0x127f20915c71, approximately bit 34). My earlier "Deno clean" control was insufficient sampling.
    • Second machine (i9-10885H, Comet Lake): 100/100 waves = 1,600 child-runs clean. On the affected machine a crash occurs roughly every 40–80 child-runs; the clean run is conclusive to better than 1e-14.
    • Kernels: reproduced on 6.18 (CachyOS) and 7.1.5 (Arch).
    • Flag/version matrix: all rows fail identically. Correction to my earlier report: --no-wasm-tier-up passing 100/100 waves did not reproduce — it was luck; intermittent clean runs are not evidence of mitigation.
    • Population evidence: the affected machine's journal shows other applications (Chromium, and Electron/Signal V8 int3 aborts) crashing with anomalous fault addresses over the same period.

    Conclusion

    Sporadic single-high-bit corruption of faulting linear addresses under heavy concurrent load on this machine (Core Ultra 9 285HX, stepping 2, microcode 0x122; the CPU reports 42-bit physical addressing, so the flipped bit is — intriguingly but unproven — exactly the MAXPHYADDR boundary). This is most consistent with a platform-level issue (CPU erratum or marginal platform behavior under load). Node, V8, and PGlite are the trigger, not the cause. The observed crashes are the subset where the corrupted address is unmapped; a corrupted store address landing on a mapped page would silently corrupt memory.

    I am closing this as not-a-Node-bug. Next steps on my side are a stock-BIOS/firmware retest and, if the behavior persists, an erratum report to Intel; I will add a comment here with the firmware outcome either way. The repro repository remains available as a trigger, and its README carries the full investigation write-up. If anyone with Core Ultra 200-series hardware wants to help confirm the platform angle, npm run repro in the repo is the fastest way. Thank you to everyone who spent time reading.

  2. rainder commented on Jul 31, 2026

    @rainder

    We are hitting what looks like a closely related crash in the same subsystem (V8 ThreadIsolation + WASM code allocation under PGlite churn), but with a different signature: a clean V8 CHECK failed abort rather than a SIGSEGV.

    Fatal error in ../../deps/v8/src/heap/code-range.cc
    Check failed: jit_page_->allocations_.erase(addr) == 1
    

    (ThreadIsolation::JitPageReference::UnregisterAllocation — the JIT page allocation registry is asked to unregister an address it doesn't have.)

    Environment:

    • Node 24.18.0 (V8 13.x), Linux x64, Google Cloud Build container (nproc=1)
    • @electric-sql/pglite 0.5.4 (in-memory)
    • CPU: Intel Xeon @ 2.20GHz, PKU absent — we verified in V8 source that this CHECK is not PKU-gated (UnregisterAllocation has no Enabled() guard), so it fires regardless of memory-protection-key support.

    Trigger shape — one data point that differs from this issue's repro: our crash happens inside a single Node process (a node:test suite creating and closing many in-memory PGlite instances sequentially, up to ~60 per test file), not across concurrent child processes. So single-process instantiate/free churn appears sufficient to corrupt/desync the ThreadIsolation allocation registry, not just cross-process concurrency.

    What we ruled out (details in our internal investigation, happy to share more):

    • Instance count alone is not the trigger: the abort has hit a file that creates only 8 instances, while 300 back-to-back create/close cycles in a minimal harness on the same CI hardware never crashed. Some co-factor (schema DDL load, app query mix, or node:test context) is required.
    • Intermittent and rare: ~5% of CI runs (3 of 64 builds over the observed window), never reproducible on demand.
    • Re-running only the failed tests immediately afterwards always passes.

    We have not yet tried --no-wasm-tier-up against this signature; will report back if we gather data on it.

    Mentioning here because the trigger (many PGlite WASM instances churning) matches this issue exactly and the failing component (V8 ThreadIsolation / JIT page tracking) is the same — if this gets escalated to the V8 team, the deterministic CHECK failed form above may be a more tractable lead than the SIGSEGV form.

    Cross-posted to electric-sql/pglite#1053.

  3. gadicc commented on Jul 31, 2026

    @gadicc
    Author

    Hey, thanks for your detailed post above. Re-opening for now due to common interest.

    From my side, I'm pretty set on the hardware angle now, with the above just being a reliable trigger. We (my LLMs and I :)) weren't yet able to reproduce in a single C file. Presumably just due to all the heavy allocations on this stack. However, some important updates:

    • I originally wrote "Concurrency is the stable trigger" No longer true
    • Can now reliably reproduce single core, but only on certain cores
    • Still SIGSEGV, still the exact same signature where si_addr = intended + 2^42

    In brief, used taskset to run the suite on groups of cpu-cores. The group of 4x P-cores all passed, 1/4 groups of 4x E-cores passed. Within those 3 faliing groups, each group had 1 failing core. Presumably during boost/turbo... initial testing shows that locking a failing core at a lower clock speed mitigates the issue.

    A bit more info in this update on the repo README but still doing further testing and still need to significantly clean up the README too. I'm likely heading for an RMA on my CPU though :D

  4. reopened this on Jul 31, 2026
  5. rainder commented on Aug 27, 2026

    @rainder

    Follow-up on --no-wasm-tier-up.

    We ran that flag on our CI server-test suite for 26 days.
    Node versions: 24.16.0, then 24.18.1, then 24.19.0.
    Host: Linux x64, Google Cloud Build, one process, @electric-sql/pglite 0.5.4.

    The V8 CHECK still fires:

    Fatal error in ../../deps/v8/src/heap/code-range.cc
    Check failed: jit_page_->allocations_.erase(addr) == 1
    

    We also saw Check failed: jit_page.has_value().

    Counts for 2026-08-01..27: 74 unique CI builds hit that fatal, out of 1669 treated builds (~4.4%).
    Pre-flag baseline was ~4.7% (3 of 64 builds).

    The flag did not stop this abort. We have dropped it.
    We still retry the failed tests once when those CHECK strings appear.

    This remains a single-process PGlite instantiate/free crash on Intel Xeon with no PKU.
    It does not match the SIGSEGV + 2^42 address-bit signature in this issue.

  6. gadicc commented on Sep 18, 2026

    @gadicc
    Author

    Final update on my original SIGSEGV report

    The original failure is best explained by a machine-specific hardware/platform fault, not a Node or V8 defect.

    The investigation changed several conclusions from the original report:

    • Concurrency was not required. A single process pinned to a susceptible logical CPU reproduced the fault; concurrency mainly increased scheduler exposure.
    • Reproduction was strongly CPU-affinity-dependent. In the completed exact-CPU tests, failures were observed only on particular E-cores, with CPU 19 producing the highest rate under load.
    • PGlite was not required. A dependency-free Node/V8 workload that repeatedly compiled, instantiated, executed, and discarded fresh WebAssembly modules reproduced reliably.
    • A mapping-heavy native C workload produced kernel oopses on CPU 19 with the same address anomaly: the reported fault address was the intended address plus 2^42. Native execution-only controls remained clean.
    • After the affected hardware was replaced under RMA, the same trigger workloads have not reproduced the fault in my post-RMA testing.

    Taken together, the evidence is most consistent with a fault in the original machine’s platform assembly. Node, V8, and PGlite were unusually effective ways to expose it, but I found no evidence that they caused it. Because the RMA replaced multiple possible contributors together, I cannot attribute the fault more narrowly to the CPU, motherboard, power delivery, or another individual component.

    The investigation, measurements, and evidence boundaries are documented in the [Fault Affinity case study](https://github.com/gadicc/fault-affinity/tree/dev/docs/case-study). That same repo also includes all my experiments, the most reliable and highest rate triggers, and a harness to check CPUs, execute experiments, and record results.

    I’m leaving this issue open briefly because @rainder’s later ThreadIsolation / JIT allocation-tracking CHECK failure is not explained by my diagnosis. It has a different signature, environment, and trigger shape, and deserves to remain visible.

    @rainder, would you be willing to open a separate Node issue for that failure and link it here? I’ll cross-link it from this report, after which I think it would be appropriate to close this original issue as a hardware-specific case.

    Thanks again to everyone who reviewed the evidence.

  7. owen-harborcoat commented on Sep 28, 2026

    @owen-harborcoat

    Opened #66366 for the jit_page_->allocations_.erase(addr) == 1 CHECK. It's the import wrapper double free fixed in V8 9b8ca54d5a, which isn't in 24.x. A v24.x-staging build with that commit (plus 68210d500a) stops it in my repro.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions