Skip to content

Python 3.14 stack overflow detection is incompatible with C++ Boost make_fcontext() coroutines #139653

Description

@vstinner

Python 3.14 introduced a new stack overflow detection mecanism: InternalDocs/stack_protection.md (#130396).

The KiCad application uses C++ Boost make_fcontext() coroutines which runs coroutine in their own stack.

Code example from fcontext doc:

// context-function
void f(intptr);

// creates a new stack
std::size_t size = 8192;
void* sp(std::malloc(size));

// context fc uses f() as context function
// fcontext_t is placed on top of context stack
// a pointer to fcontext_t is returned
fcontext_t fc(make_fcontext(sp,size,f));

_Py_InitializeRecursionLimits() is called in the main thread, whereas _Py_CheckRecursiveCall() is called for the first time in a coroutine (make_fcontext()).

Problem: Python detects a stack overflow because it's not aware that the stack base address and size changed when make_fcontext() was called.

pthread functions such as pthread_attr_getguardsize() are incompatible with make_fcontext().

cc @markshannon

Linked PRs

Activity

  1. vstinner commented on Oct 6, 2025

    @vstinner
    MemberAuthor

    @hugovk: I used the release-blocker label to raise awareness of this issue, but I'm not sure that it should hold Python 3.14.0 final release. We might be able to develop a workaround in Python 3.14.1.

  2. vstinner commented on Oct 6, 2025

    @vstinner
    MemberAuthor
  3. encukou commented on Oct 6, 2025

    @encukou
    Member
  4. penguin42 commented on Oct 6, 2025

    @penguin42

    To me the simplest partial fix would be to reread the stack limits before throwing the abort; that would stop
    spurious aborts (and be low cost), however there's also the opposite case, where if the new stack is higher than the old one, the check won't trigger at all.

  5. zooba commented on Oct 6, 2025

    @zooba
    Member

    Do we need a new (internal) API that can be called after jump_fcontext? Presumably it only occurs when user code explicitly requests it, which should mean that it can also call a CPython API after switching if needed, and it looks like the members of fcontext_t are public, so we can accept them as arguments.

  6. encukou commented on Oct 6, 2025

    @encukou
    Member

    Do these coroutines each use their own Python thread state?

    reread the stack limits before throwing the abort

    Unfortunately on some platforms where we can't read stack limits, so we approximate stack top using the current address. In that case, stack limits will be different each time.

  7. markshannon commented on Oct 6, 2025

    @markshannon
    Member

    We can (mostly) do this without a new API, using a variant of @penguin42's suggestion.

    What we would do is:

    • Change the recursion check from here_addr < tstate->c_stack_soft_limit to here_addr - state->c_stack_window_base < WINDOW_SIZE
    • If the check fails, check if there is an actual stack overflow, then update the window if we are within the stack.

    The actual window size doesn't matter, but it should be a constant to keep the cost low. It will a bit slower, but probably not measurably so.

    Unfortunately on some platforms where we can't read stack limits, so we approximate stack top using the current address. In that case, stack limits will be different each time.

    True, but this bug has only surfaced on linux (so far).

  8. markshannon commented on Oct 6, 2025

    @markshannon
    Member

    Having said that, new unstable ABI to notify of stack changes is probably the best long term option.

  9. vstinner commented on Oct 6, 2025

    @vstinner
    MemberAuthor

    Do these coroutines each use their own Python thread state?

    Coroutines reuse the existing Python thread state.

    Do we need a new (internal) API that can be called after jump_fcontext?

    It sounds like a good idea, but this function would need a stack base address and stack size parameters, since pthread_attr_getguardsize() doesn't work in this case. make_fcontext() has a known stack base address and size: its the first and second arguments.

  10. added a commit that references this issue on Oct 6, 2025
  11. 60 remaining items

  12. alexreinking commented on May 31, 2026

    @alexreinking

    We're hitting this in Halide, too. We switch stacks internally while running compiler/lowering/codegen work (makecontext/swapcontext on POSIX, fibers on Windows). The crash happens only when code running on that alternate stack calls back into Python. In our reproducer, Halide emits a compile-time warning, the Python bindings route that through py::print, and Python 3.14 aborts in _Py_CheckRecursiveCall with a reported stack usage around 1.5GB.

    If we disable Halide's alternate compiler stack, the same Python warning test passes. If we explicitly re-enable the alternate stack, it fails again. So this seems to be specifically "callback into Python from a non-thread stack", not ordinary recursion or actual stack exhaustion.

    We can work around this in Halide's Python bindings by disabling stack switching by default on Python 3.14+. But that is a
    workaround, not ideal: Halide uses the alternate stack because some large/pathological compile cases can otherwise overflow the native stack.

    This also makes unstable/private APIs awkward for us long-term, since we want the Python bindings to move toward the Stable API once the needed buffer APIs are available.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    3.14bugs and security fixesinterpreter-core(Objects, Python, Grammar, and Parser dirs)

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions