Skip to content

Move stackref buffer to per-eval loop to reduce interp stack usage #138115

Description

@Fidget-Spinner

Feature or enhancement

Proposal:

The interpreter main loop's stack usage is huge. We should try to reduce it a little. Currently, the stackref buffer takes up 10 words on 64-bit machines. We could lessen that by moving it to the heap (thread state).

This might mean slightly less perf due to worse locality and one memory indirection. So let's benchmark this to be sure.

Has this already been discussed elsewhere?

No response given

Links to previous discussion of this feature:

No response

Linked PRs

Activity

  1. Fidget-Spinner commented on Apr 9, 2026

    @Fidget-Spinner
    MemberAuthor

    Thanks to talking to @colesbury , there's a cheaper approach here.

    1. Allocate per-thread/per-interpreter buffer for stackref scratch, similar to the datastack chunk.
    2. At each _PyThreadState_PushFrame, do the bounds check to see if we have enough space left on the stackref buffer fro co->co_stacksize. If we don't, allocate a new chunk.
    3. At each callsite, no bounds check is needed thanks to 2., just bump the pointer in the stackref chunk and use the scratch in STACKREFS_TO_PYOBJECTS.
    4. After each call, shrink the pointer in the stackref chunk.
  2. colesbury commented on Apr 9, 2026

    @colesbury
    Contributor

    I think it needs to be per-thread (even with the GIL) so that the pointer bumps/shrinks in (3) and (4) match up. I think with a single per-interpreter stack you could get incorrect interleavings.

  3. Fidget-Spinner commented on Apr 9, 2026

    @Fidget-Spinner
    MemberAuthor

    Thanks, yes that seems right. I think we need this now for correctness in 3.14 and 3.15 though, to un-crash Clang builds. So it's not just about perf anymore (though the perf should be around the same in most cases).

    See #148284 for an actual segfault in the wild on Clang 21 builds.

  4. colesbury commented on Apr 9, 2026

    @colesbury
    Contributor

    This seems like a bit of a complex change to backport to 3.14.

    I'm not sure if this would work, but a simpler 3.14 change might be:

    1. In 3.14, when building without the tail call interpreter, declare the temporary storage for STACKREFS_TO_PYOBJECTS once inside PyEval_FrameDefault
    2. Use that common PyObject** array in each STACKREFS_TO_PYOBJECTS when not building with the tail call interpreter
  5. Fidget-Spinner commented on Apr 9, 2026

    @Fidget-Spinner
    MemberAuthor

    Just checking, I built CPython with CFLAGS="-g" LDFLAGS="-fuse-ld=lld-22" RANLIB=llvm-ranlib-22 CC=clang-22 ./configure --enable-optimizations --with-lto --enable-shared && make clean && make -j18

    LD_LIBRARY_PATH=/home/ken/Documents/GitHub/cpython llvm-objdump-22 --disassemble-symbols=_PyEval_EvalFrameDefault --source ./libpython3.14.so > 1.txt

    gives me the following dissassembly:

    1.txt

    with @colesbury suggested fix, I get

    2.txt

    Stack usage went from subq $0x6f8, %rsp to subq $0x268, %rsp, or roughly 1/3rd now. Wow!

  6. changed the title [-]Move stackref buffer to thread state to reduce interp stack usage[/-] [+]Move stackref buffer to per-eval loop to reduce interp stack usage[/+] on Apr 9, 2026
  7. sergey-miryanov commented on May 4, 2026

    @sergey-miryanov
    Contributor

    Is this still relevant, or can we close it?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    3.14bugs and security fixes3.15bugs and security fixesinterpreter-core(Objects, Python, Grammar, and Parser dirs)type-featureA feature request or enhancement

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions