Repository navigation
Move stackref buffer to per-eval loop to reduce interp stack usage #138115
Description
Activity
- addedtype-featureA feature request or enhancementA feature request or enhancement
on Aug 24, 2025 - addedinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Aug 24, 2025 Thanks to talking to @colesbury , there's a cheaper approach here.
- Allocate per-thread/per-interpreter buffer for stackref scratch, similar to the datastack chunk.
- At each
_PyThreadState_PushFrame, do the bounds check to see if we have enough space left on the stackref buffer froco->co_stacksize. If we don't, allocate a new chunk. - At each callsite, no bounds check is needed thanks to 2., just bump the pointer in the stackref chunk and use the scratch in
STACKREFS_TO_PYOBJECTS. - After each call, shrink the pointer in the stackref chunk.
I think it needs to be per-thread (even with the GIL) so that the pointer bumps/shrinks in (3) and (4) match up. I think with a single per-interpreter stack you could get incorrect interleavings.
Reacted by Ken JinThanks, yes that seems right. I think we need this now for correctness in 3.14 and 3.15 though, to un-crash Clang builds. So it's not just about perf anymore (though the perf should be around the same in most cases).
See #148284 for an actual segfault in the wild on Clang 21 builds.
This seems like a bit of a complex change to backport to 3.14.
I'm not sure if this would work, but a simpler 3.14 change might be:
- In 3.14, when building without the tail call interpreter, declare the temporary storage for STACKREFS_TO_PYOBJECTS once inside PyEval_FrameDefault
- Use that common PyObject** array in each
STACKREFS_TO_PYOBJECTSwhen not building with the tail call interpreter
Reacted by Ken Jin- added3.14bugs and security fixesbugs and security fixes3.15bugs and security fixesbugs and security fixes
on Apr 9, 2026 Just checking, I built CPython with
CFLAGS="-g" LDFLAGS="-fuse-ld=lld-22" RANLIB=llvm-ranlib-22 CC=clang-22 ./configure --enable-optimizations --with-lto --enable-shared && make clean && make -j18LD_LIBRARY_PATH=/home/ken/Documents/GitHub/cpython llvm-objdump-22 --disassemble-symbols=_PyEval_EvalFrameDefault --source ./libpython3.14.so > 1.txtgives me the following dissassembly:
with @colesbury suggested fix, I get
Stack usage went from
subq $0x6f8, %rsptosubq $0x268, %rsp, or roughly 1/3rd now. Wow!Reacted by Sam Gross- changed the title
[-]Move stackref buffer to thread state to reduce interp stack usage[/-][+]Move stackref buffer to per-eval loop to reduce interp stack usage[/+]on Apr 9, 2026 Is this still relevant, or can we close it?
Feature or enhancement
Proposal:
The interpreter main loop's stack usage is huge. We should try to reduce it a little. Currently, the stackref buffer takes up 10 words on 64-bit machines. We could lessen that by moving it to the heap (thread state).
This might mean slightly less perf due to worse locality and one memory indirection. So let's benchmark this to be sure.
Has this already been discussed elsewhere?
No response given
Links to previous discussion of this feature:
No response
Linked PRs