Repository navigation
JIT: Assertion failure for WITHIN_STACK_BOUNDS() in optimize_uops #141976
Description
Activity
- addedinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)type-crashA hard crash of the interpreter, possibly with a core dumpA hard crash of the interpreter, possibly with a core dump
on Nov 26, 2025 In the future, could you please run
PYTHON_LLTRACEtogether withPYTHON_OPT_DEBUG? It's hard to make sense of PYTHON_OPT_DEBUG if the execution of PYTHON_LLTRACE is not interleaved. Thank you!Sorry, will do.
Reacted by Ken JinReacted by Ken JinNo need to apologise :). Thank you for all the great fuzzing work you do.
Reacted by Chris EiblSo this is a very interesting bug. It's in the trace recorder but manifests in the optimizer. What I think is happening is the following:
CALL
specializes immediately to
CALL_ALLOC_AND_ENTER_INIT
which immediately deopts due toDEOPT_IF(!_PyThreadState_HasStackSpace(tstate, code->co_framesize + _Py_InitCleanup.co_framesize));or some other deopt.So it's very possible for CALL_ALLOC_AND_ENTER_INIT to make no progress due to stack space checks which are dynamic!
@markshannon this is an example of a specialization that makes no progress but isn't buggy. In this case, the easiest fix instead of counters it to just inspect the adaptive counter i think. If it's non-zero, that indicates that a deopt happened.
Fix up at #141989
I don't see how the behavior of the specializer matters. The generated trace should be correct regardless of what the specializer does.
I don't see how the behavior of the specializer matters. The generated trace should be correct regardless of what the specializer does.
No the problem is not the specializer. The problem is that the previous instruction is not actually what was actually executed. Please see the PR.
IIUC, the instruction recorded was
CALL_ALLOC_AND_ENTER_INIT, but the executed path was to deopt toCALL. This should be fine as the guard immediately after the call will fail at runtime, or be converted to an exit by the optimizer, which would be expecting execution in the__init__function, not immediately after the call.Are we missing a guard?
IIUC, the instruction recorded was CALL_ALLOC_AND_ENTER_INIT, but the executed path was to deopt to CALL.
Yes.
This should be fine as the guard immediately after the call will fail at runtime, or be converted to an exit by the optimizer
No the optimizer does not recognise that the sequence as invalid and tries to optimize across it, manifesting in the assertion failure.
We shouldn't fix this in the optimizer as there's nothing it's doing that is wrong there. The sequence of uops produced is just invalid due to the trace recorder.Note that the assertion failure is in the optimizer, not in the interpreter. So it's the optimizer complaining about the trace.
The previous instruction is what was executed, just not the fast path through it. Which isn't invalid, just unusual.
It looks to me like we are missing a guard after the
CALL_ALLOC_AND_ENTER_INITIt looks to me like we are missing a guard after the
CALL_ALLOC_AND_ENTER_INITA guard for what?
A guard for the control flow that we didn't expect.
The bug is that
CALL_ALLOC_AND_ENTER_INITbecomesCALLand that confuses the tracer.
We need a general solution for instructionX_SPECIALIZEDbecoming instructionX(or vice versa) during tracing.
This is probably only a problem whenXandX_SPECIALIZEDinvolve different control flow.I don't know what the solution is. Your PR may do that, but I'm worried that it mostly handles the problem, leaving an even harder to find bug behind.
What happens if we turn the optimizer off?
I assume it crashes when we try to run the trace.There are two cases:
- X becomes X_SPECIALIZED
- X_SPECIALIZED deopts to X but stays as X_SPECIALIZED in the bytecode
The first case is already safe, because X_SPECIALIZED is the actual instruction executed.
The second case is currently not safe, because X is the actual instruction executed not X_SPECIALIZED.
The fix is to reflect the actual executed instruction properly in the trace stream. This is why I had the deopt counters in the past, to handle such cases. We ended up removing that but now we still have to deal with the specializer not reflecting the actual instruction executed.
For case 2, we execute
X_SPECIALIZEDeven though that execution ends up inXso we should recordX_SPECIALIZED.
Since the the state of VM will be different than if we had executedX_SPECIALIZEDnormally, we need a guard to check we are executing in the correct place after the instruction has completed.
Don't we already emit that guard afterCALLand its specializations?As long as we insert guards after jumps, we can record any specialization of an instruction and the trace will be correct, just not very efficient.
So this does look like a bug in the optimizer after all. We should abandon the trace if we get stack out of bounds, not crash.
Having said that, we do want to get the right specialization. I've created #142183 for the specific deopt we are seeing here.
Can we do both? Ie fix the optimizer and fix the trace recorder. The problem is that if we continue recording these types of traces where a deopt actually happened, the chances of contradiction in the optimizer go up. Which in turn means the optimizer is less likely to do its job properly.
Yes let's do both. But we should fix the optimizer first, otherwise we won't know if it is fixed.
- added a commit that references this issue
on Dec 10, 2025 Thanks again for the bug report!
Crash report
What happened?
It's possible to cause an abort in a patched JIT build with the code below. Removing some more code still aborts, but usually starts causing ignored exceptions or making the reproduction take longer (to the point of becoming probabilistic).
Here's the patch used, it might be possible to reduce it:
Here's the MRE for quick, reliable reproduction:
Here's the output and backtrace, it shows that the issue happens on the first
f1call, with the last print not being called:Output from running with PYTHON_LLTRACE=4:
1989_abort_lltrace.txt
Output from running with PYTHON_DEBUG=4:
1989_abort_opt_debug.txt
Found using lafleur.
CPython versions tested on:
CPython main branch
Operating systems tested on:
Linux
Output from running 'python -VV' on the command line:
Python 3.15.0a2+ (heads/main-dirty:dc9d2eea587, Nov 24 2025, 06:26:38) [Clang 21.1.2 (2ubuntu6)]
Linked PRs