Repository navigation
JIT: invalid memory read in _PyEval_EvalFrameDefault #141861
Description
Activity
- addedinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)type-crashA hard crash of the interpreter, possibly with a core dumpA hard crash of the interpreter, possibly with a core dump
on Nov 22, 2025 I can reproduce this, but I'm completely stumped by exactly what is causing it. It seems to be a mixture of executor insertion and freeing, but I don't know. We seem to somehow grab a freed executor???
In any case, it's not a buffer overflow. Rather, it's a invalid memory read.
Anyone else is free to try debug this as well and provide a fix. I will try but I can't promise anything.
I would like to look into it.
Reacted by Savannah Ostrowski- changed the title
[-]JIT: Heap buffer overflow in _PyEval_EvalFrameDefault[/-][+]JIT: invalid memory read in _PyEval_EvalFrameDefault[/+]on Nov 23, 2025 This bug is causing the sphinx benchmark to crash.
Reacted by Savannah Ostrowski@sergey-miryanov I found a fix, could you please apply this patch and add the test case in a PR? Thank you.
diff --git a/Python/bytecodes.c b/Python/bytecodes.c index 12ee506e4f2..6129ea2e723 100644 --- a/Python/bytecodes.c +++ b/Python/bytecodes.c @@ -3018,7 +3018,7 @@ dummy_func( goto stop_tracing; } PyCodeObject *code = _PyFrame_GetCode(frame); - _PyExecutorObject *executor = code->co_executors->executors[oparg & 255]; + _PyExecutorObject *executor = code->co_executors->executors[this_instr->op.arg]; assert(executor->vm_data.index == INSTR_OFFSET() - 1); assert(executor->vm_data.code == code); assert(executor->vm_data.valid); diff --git a/Python/generated_cases.c.h b/Python/generated_cases.c.h index b83b7c528e9..47805c270f9 100644 --- a/Python/generated_cases.c.h +++ b/Python/generated_cases.c.h @@ -5476,7 +5476,7 @@ JUMP_TO_LABEL(stop_tracing); } PyCodeObject *code = _PyFrame_GetCode(frame); - _PyExecutorObject *executor = code->co_executors->executors[oparg & 255]; + _PyExecutorObject *executor = code->co_executors->executors[this_instr->op.arg]; assert(executor->vm_data.index == INSTR_OFFSET() - 1); assert(executor->vm_data.code == code); assert(executor->vm_data.valid);The problem is that it's possible for oparg to change under our feet while tracing right before an opcode switches to ENTER_EXECUTOR, thus causing the oparg and the index to be different.
Reacted by Savannah Ostrowski and Chris EiblReacted by Sergey Miryanov, Chris Eibl and SaculThis was a very tricky bug to find.
A slightly reduced repro:
import sys sys.setrecursionlimit(30) # reduce time of the run str_v1 = '' tuple_v2 = (None, None, None, None, None) small_int_v3 = 4 def f1(): for _ in range(10): abs(0) tuple_v2[small_int_v3] tuple_v2[small_int_v3] tuple_v2[small_int_v3] def recursive_wrapper_4569(): str_v1 > str_v1 str_v1 > str_v1 str_v1 > str_v1 recursive_wrapper_4569() recursive_wrapper_4569() for i_f1 in range(19000): try: f1() except RecursionError: passThank you for the bug report @devdanzin and thank you for the PR @sergey-miryanov !
Reacted by Sergey Miryanov and Mikhail EfimovI'm not convinced that the fix is right.
I think we need to know what's happening before closing the issue.Reacted by yihong@markshannon the fix is correct for the root cause. It's nothing to do with atomicity. This is what is happening (to my understanding):
- TRACE_RECORD contains the oparg of the next instruction to be executed.
- Sometimes, TRACE_RECORD terminates and inserts an executor.
- The executor may just so happen to be the next instruction.
- In which case, the oparg is still the old oparg (the one from before ENTER_EXECUTOR was inserted).
- This causes the ENTER_EXECUTOR oparg to be stale and thus lead to segfault.
There are two possible fixes: force TRACE_RECORD to reload the oparg of the next instruction, OR fix ENTER_EXECUTOR.
I'm not in favour of the first fix, as we need to reload it carefully in TRACE_RECORD, which slows it down even further and complicates it. So the easiest fix is to reload it at ENTER_EXECUTOR time.
"Fixing" this by modifying
ENTER_EXECUTORis adding undocumented coupling between trace recording andENTER_EXECUTOR. Please don't.Shouldn't the dispatch in
TRACE_RECORDafterstop_tracing_and_jit(tstate, frame)be a normalDISPATCHnot aDISPATCH_GOTO_NON_TRACING?
https://github.com/python/cpython/blob/main/Python/bytecodes.c#L5653Reacted by Ken Jin and Sergey MiryanovTRACE_RECORD
Good spot. That indeed fixes it too. Let's do that instead! @sergey-miryanov sorry to bother you, do you want to take this up? Otherwise I can do it. We don't need to revert your PR, just open another PR on top applying the one line fix in
TRACE_RECORDinstructionif (full) { LEAVE_TRACING(); _PyFrame_SetStackPointer(frame, stack_pointer); int err = stop_tracing_and_jit(tstate, frame); stack_pointer = _PyFrame_GetStackPointer(frame); if (err < 0) { JUMP_TO_LABEL(error); } DISPATCH_GOTO_NON_TRACING(); // CHANGE THIS TO NORMAL DISPATCH }The test can stay the same, and we remove the one line change to ENTER_EXECUTOR previously done.
@Fidget-Spinner Yes, of course.
Reacted by Ken JinThe root cause is addressed, hopefully to Mark's satisfaction. So I'm closing this again.
Crash report
What happened?
ASan will detect a heap buffer overflow from running the code below in a JIT build. In a patched build, the error comes at iteration 175, while on an unpatched build it happens at around 12.000 iterations.
The necessary patch to speed up the overflow:
The code that triggers the overflow:
ASan output:
Output from running with PYTHON_LLTRACE=4:
139_overflow_lltrace.txt
Output from running with PYTHON_DEBUG=4:
139_overflow_opt_debug.txt
Interestingly, the original fuzzing code would trigger the overflow at iteration 127 instead of 175, something I cut out when reducing increased the necessary number of iterations.
Found using lafleur.
CPython versions tested on:
CPython main branch
Operating systems tested on:
Linux
Output from running 'python -VV' on the command line:
Python 3.15.0a2+ (heads/main-dirty:dc9d2eea587, Nov 22 2025, 19:28:52) [Clang 21.1.2 (2ubuntu6)]
Linked PRs