Skip to content

Tuples should be immutable and safe in C, as well as in Python. #127058

Description

@markshannon

Bug report

Bug description:

[Apologies if this sounds a bit like a rant. I'm not blaming anyone. Just because something is the wrong choice now, doesn't mean it wasn't the right choice historically]

Tuples are immutable in Python, but we play all sorts of games in C with tuples, filling them will NULLs, mutating them and reusing them.

We do this in the mistaken belief that it improves performance.
But it doesn't. It makes the code base more complicated and fragile as we need to work around tuples that misbehave and do strange things. Any local performance gain is overwhelmed by slowdowns caused by the extra complexity in tuple code, the garbage collector and a few other places.

So let's fix this.

We need to:

  • Provide a new C API PyTuple_MakePair(). Pairs are by far the most common type of tuple that we play games with. By providing a fast way to create pairs, we can provide an upgrade path for C code that creates tuples in unsafe ways to do so safely and quickly.
  • Deprecate PyTuple_New. I don't know when we'll be able to remove it, but we should deprecate it ASAP.
  • Change PyTuple_New to fill the tuple with pointers to None instead of NULL. This doesn't fix the mutability issue, but it at least means the GC will only see valid objects. (This might break too much code, so we might just have to clearly document that tuples should be fully initialized in one go, before the tuple escapes the function it was created in)
  • Fix our own code to not use PyTuple_New() or perform tuple shenanigans. We can't reasonably expect third-party package authors to follow the rules if we don't.

CPython versions tested on:

CPython main branch

Operating systems tested on:

No response

Linked PRs

Activity

  1. added
    type-bugAn unexpected behavior, bug, or error
    interpreter-core(Objects, Python, Grammar, and Parser dirs)
    3.14bugs and security fixes
    on Nov 20, 2024
  2. markshannon commented on Nov 20, 2024

    @markshannon
    MemberAuthor

    I'm filing this as a bug rather than an enhancement as it does cause bugs. Most notably: https://github.com/python/cpython/blob/main/Lib/test/crashers/gc_inspection.py

  3. markshannon commented on Nov 20, 2024

    @markshannon
    MemberAuthor
  4. markshannon commented on Nov 20, 2024

    @markshannon
    MemberAuthor

    PyTuple_SetItem and PyTuple_SET_ITEM need to go too.

  5. colesbury commented on Nov 20, 2024

    @colesbury
    Contributor

    This requires a massive change to C extensions, but doesn't actually address the problem that "tuples should be immutable."

    Change PyTuple_New to fill the tuple with pointers to None instead of NULL

    This would break backwards compatibility with C extensions. See some examples: https://github.com/search?q=%2FPyTuple_GET_ITEM.*%3D%3D.*NULL%2F&type=code

    gc.get_referrers has a lot of issues, not just with tuples. As the linked files writes:

    Note that this is only an example. There are many ways to crash Python
    by using gc.get_referrers(), as well as many extension modules (even
    when they are using perfectly documented patterns to build objects).

  6. markshannon commented on Nov 20, 2024

    @markshannon
    MemberAuthor

    It might take a while for people to stop using PyTuple_New, but that doesn't mean we shouldn't deprecate it.
    I suspect we won't be removing it for a long time.

    I think you need to refine your search for PyTuple_GET_ITEM and NULL. A lot of the results look like foo(PyTuple_GET_ITEM(tuple, index)) == NULL which is fine.

    gc.get_referrers has a lot of issues, not just with tuples.

    Such as?

    Partially created tuples may not be the only culprit, but they are the main one, I think.

  7. colesbury commented on Nov 20, 2024

    @colesbury
    Contributor

    It might take a while for people to stop using PyTuple_New, but that doesn't mean we shouldn't deprecate it

    Deprecating commonly used APIs imposes a cost to C API extension authors even when we don't remove the existing API. Many extension authors will change their code in order to adopt the new best practices. If we don't have a foreseeable path to actually removing PyTuple_New() then it's unlikely that we'll ever reap the benefits of the deprecation. We'll have created churn for extension authors without delivering any benefits.

    I think you need to refine your search...

    My point is that it's a breaking change. The search does not capture all the uses that would break either.

    gc.get_referrers has a lot of issues, not just with tuples. Such as?

    As Armin wrote in #39117: "Expecting an object not to be seen before you first hand it out is extremely common, and get_referrers() breaks that assumption."

    PyType_GenericAlloc is the most common way to create extension objects, and it returns an object that is already tracked, but not yet initialized (other than zero initialization).


    • In addition to deprecating PyTuple_New(), won't this also require deprecating PyTuple_SET_ITEM and PyTuple_SetItem for initializing tuples?
    • Other than the proposed PyTuple_MakePair, what are the intended replacements for PyTuple_New()? PyTuple_Pack()? How is someone supposed to create a tuple with a dynamic number of arguments?
    • Where is the extra complexity that this would improve? There's not a whole lot of tuple-specific code in the GC, and I don't see any GC code that is dedicated to handling tuple mutability.
  8. markshannon commented on Nov 22, 2024

    @markshannon
    MemberAuthor

    Other than the proposed PyTuple_MakePair, what are the intended replacements for PyTuple_New()? PyTuple_Pack()? How is someone supposed to create a tuple with a dynamic number of arguments?

    We already have PyTuple_FromArray which is an efficient way to create a tuple. We could expose PyTuple_FromArraySteal which is even more efficient if you don't need the references to the values in the array.

    For creating tuples of unknown size, the best way is to create a list and then convert it to a tuple with PyList_AsTuple
    If the extra overhead of clearing the list is a concern we could add PyList_AsTupleAndClear which would clear the list at the same time as creating the tuple, saving the cost of modifying reference counts.

    Looking at the uses of PyTuple_New in CPython, prepending an object to a tuple is surprising common. So we could add a new function PyTuple_Prepend(PyObject *first_item, PyObject *tuple). PyTupleConcat is another possibility.

    Where is the extra complexity that this would improve?

    Not just the GC, but in the optimizer and in memory management. Knowing that tuples are genuinely immutable allows some useful savings and a few tricks.

    Mainly it allows us to make reasoned improvements. For example, when untracking tuples in the GC we should be able to assume that, thanks to immutability, a tuple must be younger than the objects it contains. Therefore a simple oldest-first scan should find all tuples that can be untracked. Sadly, this isn't the case. Likewise, we would like to know that tuples cannot contain themselves.

  9. added
    type-featureA feature request or enhancement
    type-bugAn unexpected behavior, bug, or error
    3.14bugs and security fixes
    triagedThe issue has been accepted as valid by a triager.
    and removed
    type-bugAn unexpected behavior, bug, or error
    3.14bugs and security fixes
    type-featureA feature request or enhancement
    on Dec 9, 2024
  10. picnixz commented on Dec 9, 2024

    @picnixz
    Member

    (sorry I missed the comment on triaged)

  11. added a commit that references this issue on Dec 11, 2024
  12. added a commit that references this issue on Jan 8, 2025
  13. davidhewitt commented on Apr 23, 2026

    @davidhewitt
    Contributor

    With PyTuple_FromArray now incoming in 3.15, should I consider switching PyO3 to it from PyTuple_New + PyTuple_SET_ITEM? A quick test with 3.15.0a7 suggests that crude benchmarks which convert Rust 2- and 12-tuples of integers to Python tuples gets ~25% slower using PyTuple_FromArray.

    using PyTuple_New + PyTuple_SET_ITEM:

    tuple_into_pyobject_2      time:   [7.5477 ns 7.5943 ns 7.6465 ns]
    tuple_into_pyobject_12     time:   [22.022 ns 22.183 ns 22.341 ns]
    

    using PyTuple_FromArray:

    tuple_into_pyobject_2      time:   [9.4254 ns 9.4982 ns 9.5704 ns]
                               change: [+23.217% +24.299% +25.284%] (p = 0.00 < 0.05)
                               Performance has regressed.
    tuple_into_pyobject_12     time:   [27.944 ns 28.062 ns 28.181 ns]
                               change: [+25.246% +26.118% +27.009%] (p = 0.00 < 0.05)
                               Performance has regressed.
    

    I am sad that PyTuple_FromArraySteal was voted against by the C API working group, I think particularly at the FFI boundary PyO3 is sensitive to this because there are many cases where Python objects are built from Rust values and immediately placed into Python tuples in a fashion similar to the benchmark above.

    Not only is the performance of PyTuple_FromArray worse than the status quo, but it's also worse for code size because the Rust compiler has to emit calls to Py_DecRef for all these new Python objects being placed directly into the tuples.

    If the concensus is that PyO3 should switch to PyTuple_FromArray now so that we can clean up the C API, I'm happy to eat the cost for the greater good. Hopefully we can regain that performance in the future 🤞

  14. colesbury commented on Apr 23, 2026

    @colesbury
    Contributor

    In my opinion, no. What do you gain? This is one of the most commonly used C API functions.

  15. davidhewitt commented on Apr 23, 2026

    @davidhewitt
    Contributor

    PyO3 indeed doesn't gain much from using PyTuple_FromArray.

    With the stealing variant it might be the opposite; we could delete some existing code which fills tuples created by PyTuple_New, and possibly even benefit from the interpreter being able to knowingly skip stages like zeroing the tuple memory before filling it. Now I think about it, we also don't have PyTuple_SET_ITEM in the stable ABI, we make repeated calls to PyTuple_SetItem instead, so extensions built for the stable ABI might be accidentally suffering here. A single FFI call would be optimal.

    I keep seeing discussions like this one and capi-workgroup/problems#56 which makes me believe that the long term desire from many is to deprecate PyTuple_New, provided suitable replacements exist. If that's the case, perhaps we want extensions to stop using it where possible already, as a gentle forcing function towards finding good replacement APIs.

  16. markshannon commented on May 1, 2026

    @markshannon
    MemberAuthor

    From a performance perspective, PyTuple_FromArray is likely to be faster only if you need to retain the references to the objects in the array.

    @davidhewitt
    OOI, when was PyTuple_FromArraySteal voted against by the C API working group?

  17. davidhewitt commented on May 1, 2026

    @davidhewitt
    Contributor

    There was the opinion that it was a micro-optimization not worth having. capi-workgroup/decisions#78

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    3.14bugs and security fixesinterpreter-core(Objects, Python, Grammar, and Parser dirs)triagedThe issue has been accepted as valid by a triager.type-bugAn unexpected behavior, bug, or error

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions