Skip to content

Experimental YJIT port to Windows x64 (mingw-ucrt) - #1

Open
Largo wants to merge 10 commits into
masterfrom
yjit-windows
Open

Largo wants to merge 10 commits into
masterfrom
yjit-windows

Conversation

@Largo

@Largo Largo commented Aug 9, 2026 •

Copy link
Copy Markdown
Owner

WARNING: Very experimental

Experimental YJIT port to Windows x86_64 (x86_64-w64-mingw-ucrt, MSYS2/UCRT64 toolchain, rustc from mingw-w64-ucrt-x86_64-rust).

Before this, rb_jit_reserve_addr_space in jit.c simply returned NULL under _WIN32, so --enable-yjit produced a build that could not allocate executable memory. With these commits ruby.exe --yjit JIT-compiles and runs correctly.

Benchmarks

All numbers are the same binary, YJIT vs. interpreter, best-of steady state, measured in the Windows VM.

ruby/ruby-bench, pure-Ruby CPU benchmarks

benchmark interp (ms) yjit (ms) speedup
30k_methods 1448 65 22.2×
30k_ifelse 2066 95 21.7×
ruby-xor 514 48 10.6×
respond_to 875 82 10.6×
nqueens 837 88 9.6×
fib 879 100 8.8×
keyword_args 807 139 5.8×
matmul 1935 441 4.4×
cfunc_itself 391 103 3.8×
getivar 451 453 1.0×

Geometric mean 7.3× over the 10 benchmarks. getivar is memory-bound, so ~1× is the expected result there.

Macro benchmarks (the vendored, gem-free subset of ruby-bench)

benchmark interp (ms) yjit (ms) speedup
optcarrot (NES emulator) 18108 3612 5.0×
nbody 348 75 4.6×
binarytrees 1057 344 3.1×
rubykon (Go AI) 2657 978 2.7×
fannkuchredux 1557 1561 1.0×

fannkuchredux is permutation/GC-bound and is flat under YJIT upstream too. lee and the Ractor benchmarks call Ractor.make_shareable and were skipped.

The ~25 gem-based macro benchmarks (lobsters, railsbench, rubocop, sequel, …) need bundle install of large native-gem trees, and Rails additionally needs the socket and strscan extensions, which do not build in this 4.1.0dev tree for reasons unrelated to YJIT. Not covered here.

miniruby microbenchmarks

workload interp yjit speedup
fib(31) (recursion) 515 ms 52 ms 9.8×
method + ivar dispatch 769 ms 317 ms 2.4×
predicate / branch 859 ms 463 ms 1.9×
string build/scan 138 ms 107 ms 1.3×
array sort/reduce 251 ms 236 ms 1.06×

Tests

suite result
bootstraptest/test_yjit.rb 369/369
bootstraptest/test_yjit_30k_ifelse.rb, ..._30k_methods.rb pass
test/ruby/test_yjit.rb 139 tests, 375 assertions, 0 failures, 0 errors, 1 skip

test/ruby/test_yjit.rb started at 65 assertions / 120 errors, almost all of them a harness limitation rather than YJIT (see below).

What the port required

The x86_64 backend hard-coded the System V AMD64 ABI; Windows uses the Microsoft x64 ABI. Three differences matter for codegen:

  1. Argument registers. SysV passes six integer args in RDI, RSI, RDX, RCX, R8, R9. Win64 passes four, in RCX, RDX, R8, R9, and the caller reserves 32 bytes of shadow space below the return address; args 5+ go on the stack. C_ARG_OPNDS and the CCall lowering/emit were rewritten.
  2. Callee-saved registers. RSI and RDI are callee-saved on Win64 but caller-saved on SysV, and YJIT used them as scratch. The port makes RSI/RDI the allocation registers, preserved once per JIT frame in FrameSetup/FrameTeardown, and uses RCX, RDX, R8, R9, R10 as temps — five temps, all caller-saved, matching SysV's arithmetic.
  3. The callback convention. c_callable! pins the Rust runtime callbacks to extern "sysv64" on all x86_64. On Windows that fights the native ABI, so the macro uses extern "C" (= MS x64) there and everything speaks one convention.

Platform plumbing:

  • jit.c — executable memory via VirtualAlloc(MEM_RESERVE) / VirtualAlloc(MEM_COMMIT) / VirtualProtect(PAGE_EXECUTE_READ) / FlushInstructionCache / VirtualFree(MEM_DECOMMIT), page size from GetSystemInfo. The reservation probes downward from the module address to stay within ±2 GiB of Ruby's C functions so 32-bit relative call/jmp can reach them.
  • configure.ac — a separate YJIT_TARGET_OK for x86_64-*mingw* (widening the shared JIT_TARGET_OK would implicitly enable ZJIT, which has not enabled itself on Windows), plus the Rust staticlib's system dependencies -lbcrypt -lntdll -luserenv -lsynchronization.
  • defs/jit.mk — binutils ld -r cannot partial-link the Rust libyjit.a on PE; it chokes on the TLS directory Rust std emits (unable to fill in DataDirectory[9]: _tls_used). Instead the port builds a localized copy of the archive: strip the .dwo members, then objcopy --localize-symbols on only the symbols that actually collide with $(MISSING) (the nm intersection — lgamma_r, tgamma, …). A blanket --keep-global-symbol=rb_* would sever Rust's inter-CGU symbols.

LLP64, the second and subtler mismatch

long is 32-bit on Windows even though VALUE is 64-bit, so RUBY_FIXNUM_MAX is 2^30 — a Win64 Fixnum holds 31 bits, not 63. Three separate bugs came out of that, each producing wrong-but-plausible values rather than an obvious crash:

  • Integer fast paths. opt_plus/opt_minus/opt_mult/opt_succ do 64-bit arithmetic and detect overflow with a 64-bit jo, which only trips near 2^63, so results between 2^31 and 2^63 stayed Fixnum-tagged and hit [BUG] Unnormalized Fixnum. Fixed with a Windows-only range guard that side-exits to the interpreter for Bignum promotion. No-op on LP64.
  • Array length. get_array_len reads the heap length as a C long (32-bit here) and csels it against the 64-bit-derived embedded length, tripping match_num_bits. Fixed by reading 64 bits and masking. The masking and clobbers flags, so it has to be emitted before the RARRAY_EMBED_FLAG test the csel consumes — getting that order wrong made heap arrays silently report length 0.
  • String length. String#bytesize/#getbyte read RSTRING_LEN the same way. Surfaced by the ruby-bench ruby-xor benchmark.

Also ID and st_data_t are uintptr_t in C but bindgen recorded them as c_ulong; they are now blocklisted in the generator and emitted as usize, so a regenerated cruby_bindings.inc.rs stays correct.

Verified byte-identical to the interpreter with a differential harness: arithmetic across the Fixnum boundary in both directions, deep recursion, blocks and accumulators past 2^32, factorials, strings, arrays, hashes, procs/lambdas, and RubyVM::YJIT.runtime_stats.

Test-harness changes

Two commits touch test/ruby/test_yjit.rb rather than YJIT:

  • assert_compiles has the child Marshal-dump its stats to fd 3, which Windows spawn rejects outright (process.c refuses fd >= 3; CreateProcess only inherits 0/1/2). That errored out 119 tests before any code ran. On Windows the child now writes to a temp file whose path arrives in YJIT_TEST_STATS_FILE; every other platform keeps the pipe.
  • test_tracing_str_uplus asserts an exact putspecialobject deopt count. On Windows YJIT compiles that instruction where other platforms side-exit; the result is identical, only the deopt profile differs, so the exit assertion is :any there.

Status

Experimental. Built and tested on one Windows VM with MSYS2/UCRT64; there is no CI for this target, and the MS-ABI paths do not get exercised by a Linux build.

Largo added 10 commits August 10, 2026 00:27
Bring up YJIT on x86_64-w64-mingw-ucrt, previously unsupported (rb_jit_reserve_addr_space returned NULL on _WIN32).

Changes:
- configure.ac: allow JIT on x86_64-*mingw*; link Rust std deps (bcrypt/ntdll/userenv/synchronization).
- jit.c: VirtualAlloc/VirtualProtect/VirtualFree for reserve/mark_writable/mark_executable/mark_unused; GetSystemInfo page size; FlushInstructionCache.
- defs/jit.mk: on mingw, binutils ld -r cannot partial-link the Rust staticlib (PE TLS DataDirectory). Instead build a symbol-localized copy of the archive (drop .dwo members, objcopy --localize-symbols on just the symbols that collide with Ruby's missing/*.o) and link it directly.
- yjit backend (x86_64): MS x64 ABI. C_ARG_OPNDS = RCX,RDX,R8,R9 with stack args + 32-byte shadow space for >4 args; TEMP_REGS = RCX,RDX,R8,R9,R10; alloc regs = RAX,RSI,RDI (RSI/RDI callee-saved, preserved once per frame in FrameSetup); Win64 caller-save set.
- yjit utils.rs: c_callable! uses extern "C" (MS x64) on Windows instead of extern "sysv64" so runtime callbacks match the generated calls.
- yjit cruby_bindings: ID/st_data_t are usize (uintptr_t), not c_ulong, which is 32-bit on LLP64.
- yjit fd handling (options/log/disasm): RawFd -> RawFileRef (RawHandle on Windows).
- assorted usize/u64/i32/i64 cast fixes for LLP64.

Status: scalar arithmetic, loops, method calls, deep recursion (fib), strings, arrays, procs, and blocks all run correctly and pass. Benchmarks vs interpreter: fib(31) 9.8x, method/ivar dispatch 2.4x, predicate/branch 1.9x.

Known bug: an accumulator built up over many JIT'd block iterations that crosses 2^32 triggers a spurious 'Unnormalized Fixnum' check (the computed value is numerically correct). RubyVM::YJIT.runtime_stats also asserts. Both are isolated follow-ups; core codegen is correct.
On Windows x64 (LLP64) long is 32-bit, so Ruby's Fixnum range is LONG_MAX/2 (+/-2^30) even though VALUE is 64-bit. YJIT's integer fast paths do 64-bit arithmetic and only detect overflow at 2^63, so a result between 2^31 and 2^63 stayed Fixnum-tagged when it should promote to Bignum -> [BUG] Unnormalized Fixnum.

Add guard_fixnum_in_long_range(): after opt_plus/opt_minus/opt_mult/opt_succ compute the tagged result, side-exit to the interpreter (which promotes to Bignum) when it leaves the signed-32-bit range. No-op on LP64 where the 64-bit overflow check already matches the boundary. Costs 2 cmp+jcc per op; benchmarks unchanged (fib 9.8x etc).

Also add VALUE::num_from_usize() (rb_ull2inum-backed) and route RubyVM::YJIT.runtime_stats counter packing through it, so counters exceeding the LLP64 Fixnum range become Bignums instead of asserting in fixnum_from_usize.

Verified byte-identical to the interpreter across boundary-crossing add/sub/mult, power-via-mult, succ, factorial/Bignum, and the previously-failing block accumulators >2^32. runtime_stats now returns a Hash.
RARRAY heap length is a C long (32-bit on Windows LLP64), but get_array_len read it at c_long::BITS and csel'd it against the embedded length derived from the 64-bit flags word -> match_num_bits panic (operands of incompatible sizes 64 vs 32), hit by the Array#empty?/length/size cfunc specializations and array aref/splat when compiled at runtime.

Read a full VALUE-width word at the heap-len offset (the 4 bytes above len belong to the adjacent aux field) and mask to the low  bits; array lengths are non-negative so this zero-extends to a 64-bit operand matching the embedded length and downstream Fixnum tagging. The masking  clobbers flags, so it is computed before the RARRAY_EMBED_FLAG test whose result the csel consumes.

Verified byte-identical to the interpreter (--yjit-call-threshold=2): length/size/empty on embedded and heap arrays, aref+length, runtime RubyVM::YJIT.enable. bootstraptest/test_yjit.rb 369/369 and the 30k ifelse/methods stress tests pass.
String#bytesize and String#getbyte read RSTRING_LEN, a C long that is 32-bit on Windows (LLP64), then used it in 64-bit contexts (Fixnum tagging in bytesize; cmp against a 64-bit index in getbyte). Mixing a 32-bit field with 64-bit operands tripped the backend operand-size assertion and crashed with [BUG] YJIT panicked (asm/x86_64/mod.rs) - hit by the ruby-bench ruby-xor benchmark.

Add load_long_len_field() which reads a full VALUE-width word at the length offset and masks to the low  bits (lengths are non-negative, so this zero-extends), matching the approach used for array length. Verified byte-identical to the interpreter for bytesize/getbyte on embedded and heap strings including negative/out-of-bounds indices; ruby-xor now runs (10.6x). bootstraptest 369/369.
The invokebuiltin / leaf-builtin fast paths bail out when the builtin needs
more arguments than C_ARG_OPNDS holds, which is 4 on the MS x64 ABI against 6
on SysV. ccall() already spills arguments past the fourth onto the stack, so
the register count is not the real limit: builtins with 3-4 arguments were
side-exiting on Windows for no reason, and the exit profile diverged from
every other platform.

Use the SysV limit of 6 on Windows too, so the same builtins compile
everywhere.
rb_jit_mark_executable is handed a page-aligned range that can contain pages
rb_jit_mark_writable never committed, or pages rb_jit_mark_unused decommitted
during code GC. Linux mprotect tolerates that; VirtualProtect fails the whole
call with ERROR_INVALID_ADDRESS if even one page in the range is uncommitted,
which showed up as "[BUG] Couldn't make JIT page executable" under code GC and
with a small --yjit-exec-mem-size.

Commit the range with VirtualAlloc(MEM_COMMIT) first, which is idempotent on
already-committed pages, then set the protection.
The Windows port makes YJIT work on x86_64-*mingw*, but JIT_TARGET_OK also
gates ZJIT, which has not enabled itself on Windows. Widening JIT_TARGET_OK
would implicitly turn on a JIT the platform does not support yet, so derive
YJIT_TARGET_OK from it and add mingw only there.
Both are uintptr_t in C, but bindgen run on LP64 records them as c_ulong,
which is 32-bit once the checked-in bindings are compiled on LLP64 (Windows).
Blocklist the two types and emit them as usize, so a regenerated
cruby_bindings.inc.rs stays correct instead of depending on a hand-edit
surviving the next bindgen run.
The harness has the child write its marshaled stats to fd 3. Windows spawn
rejects fd >= 3 as a redirect key (process.c guard; CreateProcess only wires
up fds 0/1/2), which errored out 119 tests with "wrong file descriptor (3)".

On Windows, hand the child a temp-file path in YJIT_TEST_STATS_FILE and read
the results back from there instead. Every other platform keeps the pipe.
… Windows

test_tracing_str_uplus asserts an exact putspecialobject side-exit count under
object-allocation tracing. On Windows YJIT compiles putspecialobject where
other platforms deopt; the result is identical (the frozen string's allocation
source line is correct and tracing invalidation works), only the exact deopt
profile differs. Assert :any there, as the suite does for other
platform-specific differences.

With this, test/ruby/test_yjit.rb on Windows is 139 tests, 375 assertions,
0 failures, 0 errors, 1 skip.
@github-actions github-actions Bot added the jit label Aug 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant