Repository navigation
[C API] Add PyTupleWriter API #139888
Description
Activity
- addedtype-featureA feature request or enhancementA feature request or enhancementinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Oct 10, 2025 ... but why? We're adding
PyTuple_FromArrayalready, and it's going to be just as easy for a user to fill out a native array and turn it into a tuple in one go. The majority of tuples aren't variable length, they're known at compile time (because that's what a tuple represents - a list is for dynamically sized objects).Where's the real world benchmarks showing that this API is worth any performance or correctness benefit? We can still deprecate the bad APIs in favour of "just use a native array until you're ready".
Reacted by Petr ViktorinThe majority of tuples aren't variable length
This API is mostly a replacement to
_PyTuple_Resize(): function when the input size is not known in advance. In that case, you should allocate a few items, fill these items, allocate more items, etc.Having to manage manually the buffer/array is non trivial and so
PyTupleWriteroffers a high-level API for that.When the input size is known, there are other existing safe functions:
PyTuple_FromArray()(new! I just added it),PyTuple_Pack(),Py_BuildValue(), etc.Where's the real world benchmarks showing that this API is worth any performance or correctness benefit?
I ran a micro-benchmark: #139891 (comment)
About correctness, this API should fix bugs when it is possible to discover an incomplete tuple through the GC:
- Incomplete tuple created by PyTuple_New() and accessed via the GC can trigged a crash #59313: the main issue about the GC tracking issue
- gc.get_referrers() is inherently dangerous #39117
- segfault due to null pointer in tuple #70998
- Correct reuse argument tuple in property descriptor #68464
PySequence_Tuple() was fixed recently (Python 3.14) against such bug: commit 5a23994.
What do you think about adding a
PyTuple_FromSingleandPyTuple_FromPair(orPyTuple_MakeSingleandPyTuple_MakePair) (according to my measurements for pyperformance 1-size and 2-size tuples have about 80% of occurrences)?Following is a data of occurrences of calls that can be replaced with
PyTuple_FromSingleandPyTuple_FromPair:file function count _asyncmodule.cPyTuple_New(2)2 _collectionsmodule.cPyTuple_Pack(1)2 _csv.cPyTuple_Pack(1)1 _datetimemodule.cPyTuple_Pack(2)4 PyTuple_Pack(1)3 _elementtree.cPyTuple_Pack(2)4 _functoolsmodule.cPyTuple_New(2)2 _interpretersmodule.cPyTuple_Pack(2)1 _json.cPyTuple_New(2)1 PyTuple_Pack(2)1 PyTuple_Pack(1)1 _operator.cPyTuple_Pack(2)1 _pickle.cPyTuple_Pack(2)3 PyTuple_New(2)2 PyTuple_New(1)2 _ssl.cPyTuple_New(2)6 PyTuple_Pack(2)1 _threadmodule.cPyTuple_New(2)1 _tkinter.cPyTuple_Pack(1)arraymodule.cPyTuple_New(2)2 itertoolsmodule.cPyTuple_Pack(2)2 PyTuple_New(2)1 main.cPyTuple_Pack(2)1 overlapped.cPyTuple_New(2)2 posixmodule.cPyTuple_Pack(2)1 pyexpat.cPyTuple_New(1)1 selectmodule.cPyTuple_Pack(2)1 PyTuple_New(2)1 signal_module.cPyTuple_New(2)1 socket_module.cPyTuple_Pack(2)3 termios.cPyTuple_New(2)2 _ctypes.cPyTuple_Pack(2)2 stgdict.cPyTuple_Pack(2)1 decimal.cPyTuple_Pack(2)7 PyTuple_Pack(1)2 microprotocol.cPyTuple_Pack(2)2 _sre.cPyTuple_New(2)1 datetime.cPyTuple_Pack(1)2 PyTuple_Pack(2)1 getargs.cPyTuple_Pack(1)1 heaptype.cPyTuple_Pack(2)1 PyTuple_Pack(1)2 PyTuple_New(2)1 vectorcall_limited.cPyTuple_New(1)2 multibytecodec.cPyTuple_New(2)1 codeobject.cPyTuple_Pack(2)7 dictobject.cPyTuple_Pack(2)2 PyTuple_New(2)4 enumobject.cPyTuple_Pack(2)1 PyTuple_New(2)2 exceptions.cPyTuple_Pack(2)7 floatobject.cPyTuple_Pack(2)1 frameobject.cPyTuple_Pack(2)2 PyTuple_Pack(1)1 genericaliasobject.cPyTuple_Pack(1)2 listobject.cPyTuple_Pack(2)1 longobject.cPyTuple_Pack(2)1 PyTuple_New(2)2 odictobject.cPyTuple_Pack(2)2 PyTuple_New(2)1 setobject.cPyTuple_Pack(1)1 typeobject.cPyTuple_Pack(2)2 PyTuple_Pack(1)5 typevarobject.cPyTuple_Pack(2)1 PyTuple_Pack(1)2 unicode_format.hPyTuple_Pack(2)2 pegen_errors.cPyTuple_Pack(2)2 _warnings.cPyTuple_Pack(2)1 bltnmodule.cPyTuple_Pack(2)1 ceval.cPyTuple_Pack(1)1 _codegen.cPyTuple_Pack(1)2 compile.cPyTuple_Pack(2)1 crossinterp.cPyTuple_Pack(1)1 errors.cPyTuple_Pack(1)1 hamt.cPyTuple_Pack(2)1 marshal.cPyTuple_Pack(2)1 pylifecycle.cPyTuple_Pack(2)1 Python-tokenize.cPyTuple_Pack(2)1 sysmodule.cPyTuple_Pack(1)1 tracemalloc.cPyTuple_New(2)1 Is it worth to implement those methods, replace in the codebase and benchmark it?
I believe that my question interleaves with #140009 a bit.What do you think about adding a
PyTuple_FromSingleandPyTuple_FromPair(orPyTuple_MakeSingleandPyTuple_MakePair) (according to my measurements for pyperformance 1-size and 2-size tuples have about 80% of occurrences)?This was suggested in #118222
Reacted by Sergey MiryanovWhat do you think about adding a PyTuple_FromSingle and PyTuple_FromPair (or PyTuple_MakeSingle and PyTuple_MakePair) (according to my measurements for pyperformance 1-size and 2-size tuples have about 80% of occurrences)?
Would you mind to open a separated issue for that?
@eendebakpt Thanks!
@vstinner Done - #140052.Reacted by Victor Stinner- added 6 commits that reference this issue
on Oct 14, 2025 I wrote this API to get rid of
_PyTuple_Resize()which treats an immutable tuple as mutable. I would like to avoid that: do not provide any function to mutate an immutable tuple.I failed to convince other core developers that this
PyTupleWriterAPI is useful. Moreover, it's 1.4x slower than the current API based on macros (which mutate an immutable tuple!).I prefer to give up on
PyTupleWriter. It seems like_PyTuple_Resize()has to stay for a few more years. In the meanwhile, one option is to create alistand then callPyList_AsTuple()to convert it to atuple.
Feature or enhancement
Hi,
Creating a tuple with the current C API has multiple issues:
PyTuple_SetItem()andPyTuple_SET_ITEM()modify an immutable tuple.PyTuple_New()creates an incomplete object: items are set toNULL. This is bad:PyTuple_New()tracks directly the tuple in the garbage collector. For example,gc.get_objects()gives access to the incomplete tuple. Using the tuple, like callingrepr(tuple), can crash Python._PyTuple_Resize()is private which is surprising for a documented API. It stays private because it has issues._PyTuple_Resize()modifies an immutable tuple._PyTuple_Resize()must not be used of the refcount is greater than1: the API is fragile.I propose adding a new efficient
PyTupleWriterAPI: work on a temporary "writer" object, and then callFinish()on it to get the tuple.I already proposed a similar API in 2023:
_PyTupleBuilder. Since that, Python C API got thePyUnicodeWriterAPI and thePyBytesWriterAPI which are efficient writers forstrandbytesobjects, and the C API Working Group was created. The proposed API is now public and allocates the structure on the heap memory to hide the implementation details (the structure).Mark Shannon asked if it would be possible to work on a list and then convert the list to a tuple, but it's less efficient. Tuples are commonly used in Python, and so creating a tuple should be efficient.
Mark Shannon also tried to initialize tuple items to
Noneinstead ofNULLinPyTuple_New(). His attempt failed because of implementation issues. Also, this change only fix some of the issues that I listed, not all of them.An alternative is to fill a C array of Python objects, call the new
PyTuple_FromArray(), and then callPy_DECREF()on the array items. It requires to allocate and deallocate an array, and callPy_DECREF()on items. It can be less efficient.Linked PRs