Repository navigation
[C API] Replace PyTuple_Pack(1,2) with PyTuple_Make[Single,Pair] to optimize creation of tuples #140052
Description
Activity
- addedtype-featureA feature request or enhancementA feature request or enhancement
on Oct 13, 2025 - addedperformancePerformance or resource usagePerformance or resource usageextension-modulesC modules in the Modules dirC modules in the Modules dir
on Oct 13, 2025 sergey-miryanov commented
on Oct 13, 2025 ContributorAuthorMore actionsFor Steal version I have following results:
Geometric mean - 1.01x slower
+--------------------------+----------+------------------------+------------------------+ | Benchmark | m | o | s | +==========================+==========+========================+========================+ | async_generators | 277 ms | 279 ms: 1.01x slower | 275 ms: 1.01x faster | +--------------------------+----------+------------------------+------------------------+ | asyncio_tcp_ssl | 812 ms | not significant | 803 ms: 1.01x faster | +--------------------------+----------+------------------------+------------------------+ | asyncio_websockets | 242 ms | 241 ms: 1.00x faster | not significant | +--------------------------+----------+------------------------+------------------------+ | chaos | 36.0 ms | 36.6 ms: 1.02x slower | 36.7 ms: 1.02x slower | +--------------------------+----------+------------------------+------------------------+ | comprehensions | 10.3 us | 10.1 us: 1.02x faster | 10.1 us: 1.01x faster | +--------------------------+----------+------------------------+------------------------+ | coroutines | 15.0 ms | 14.5 ms: 1.03x faster | 14.7 ms: 1.02x faster | +--------------------------+----------+------------------------+------------------------+ | coverage | 54.9 ms | 56.1 ms: 1.02x slower | 56.2 ms: 1.02x slower | +--------------------------+----------+------------------------+------------------------+ | crypto_pyaes | 45.3 ms | 46.1 ms: 1.02x slower | 45.6 ms: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | deepcopy | 171 us | 168 us: 1.02x faster | 168 us: 1.02x faster | +--------------------------+----------+------------------------+------------------------+ | deepcopy_reduce | 1.87 us | 1.84 us: 1.02x faster | 1.83 us: 1.02x faster | +--------------------------+----------+------------------------+------------------------+ | deepcopy_memo | 16.5 us | 17.5 us: 1.06x slower | 16.9 us: 1.02x slower | +--------------------------+----------+------------------------+------------------------+ | deltablue | 1.97 ms | 1.99 ms: 1.01x slower | not significant | +--------------------------+----------+------------------------+------------------------+ | django_template | 23.1 ms | 23.4 ms: 1.02x slower | 23.4 ms: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | docutils | 1.70 sec | 1.68 sec: 1.01x faster | not significant | +--------------------------+----------+------------------------+------------------------+ | dulwich_log | 43.3 ms | not significant | 43.7 ms: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | fannkuch | 245 ms | 243 ms: 1.01x faster | 238 ms: 1.03x faster | +--------------------------+----------+------------------------+------------------------+ | float | 42.4 ms | 43.4 ms: 1.02x slower | 43.3 ms: 1.02x slower | +--------------------------+----------+------------------------+------------------------+ | create_gc_cycles | 1.18 ms | not significant | 1.19 ms: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | gc_traversal | 2.93 ms | 2.82 ms: 1.04x faster | 3.03 ms: 1.03x slower | +--------------------------+----------+------------------------+------------------------+ | generators | 19.1 ms | 19.5 ms: 1.02x slower | 19.3 ms: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | genshi_text | 14.2 ms | 14.4 ms: 1.02x slower | 14.4 ms: 1.02x slower | +--------------------------+----------+------------------------+------------------------+ | genshi_xml | 32.8 ms | 33.1 ms: 1.01x slower | 33.3 ms: 1.02x slower | +--------------------------+----------+------------------------+------------------------+ | go | 68.7 ms | 70.2 ms: 1.02x slower | 69.3 ms: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | hexiom | 3.64 ms | 3.61 ms: 1.01x faster | 3.62 ms: 1.00x faster | +--------------------------+----------+------------------------+------------------------+ | json_dumps | 6.34 ms | 6.26 ms: 1.01x faster | 6.43 ms: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | json_loads | 15.7 us | 16.1 us: 1.02x slower | 16.6 us: 1.06x slower | +--------------------------+----------+------------------------+------------------------+ | logging_format | 3.88 us | not significant | 3.91 us: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | logging_silent | 63.6 ns | 59.4 ns: 1.07x faster | 59.0 ns: 1.08x faster | +--------------------------+----------+------------------------+------------------------+ | logging_simple | 3.49 us | 3.52 us: 1.01x slower | 3.56 us: 1.02x slower | +--------------------------+----------+------------------------+------------------------+ | mako | 7.00 ms | 6.94 ms: 1.01x faster | 6.90 ms: 1.01x faster | +--------------------------+----------+------------------------+------------------------+ | mdp | 788 ms | 780 ms: 1.01x faster | 781 ms: 1.01x faster | +--------------------------+----------+------------------------+------------------------+ | meteor_contest | 68.0 ms | 68.2 ms: 1.00x slower | 69.2 ms: 1.02x slower | +--------------------------+----------+------------------------+------------------------+ | nbody | 55.4 ms | 55.0 ms: 1.01x faster | 55.9 ms: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | nqueens | 55.8 ms | not significant | 57.0 ms: 1.02x slower | +--------------------------+----------+------------------------+------------------------+ | pathlib | 9.71 ms | not significant | 9.79 ms: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | pickle | 7.29 us | not significant | 7.37 us: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | pickle_dict | 18.5 us | 18.8 us: 1.02x slower | 18.7 us: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | pickle_list | 2.93 us | 2.98 us: 1.02x slower | 3.02 us: 1.03x slower | +--------------------------+----------+------------------------+------------------------+ | pickle_pure_python | 208 us | 205 us: 1.01x faster | not significant | +--------------------------+----------+------------------------+------------------------+ | pidigits | 143 ms | 143 ms: 1.00x slower | 144 ms: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | pprint_safe_repr | 489 ms | 496 ms: 1.01x slower | 502 ms: 1.03x slower | +--------------------------+----------+------------------------+------------------------+ | pprint_pformat | 997 ms | 1.01 sec: 1.01x slower | 1.03 sec: 1.03x slower | +--------------------------+----------+------------------------+------------------------+ | pyflate | 259 ms | 260 ms: 1.00x slower | 258 ms: 1.00x faster | +--------------------------+----------+------------------------+------------------------+ | python_startup | 7.74 ms | not significant | 7.74 ms: 1.00x slower | +--------------------------+----------+------------------------+------------------------+ | python_startup_no_site | 5.09 ms | not significant | 5.10 ms: 1.00x slower | +--------------------------+----------+------------------------+------------------------+ | raytrace | 171 ms | not significant | 170 ms: 1.00x faster | +--------------------------+----------+------------------------+------------------------+ | regex_compile | 82.7 ms | 83.5 ms: 1.01x slower | 83.5 ms: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | regex_dna | 115 ms | 113 ms: 1.01x faster | not significant | +--------------------------+----------+------------------------+------------------------+ | regex_effbot | 1.68 ms | not significant | 1.63 ms: 1.03x faster | +--------------------------+----------+------------------------+------------------------+ | regex_v8 | 14.7 ms | 14.3 ms: 1.03x faster | 14.6 ms: 1.00x faster | +--------------------------+----------+------------------------+------------------------+ | richards | 27.3 ms | 27.1 ms: 1.01x faster | not significant | +--------------------------+----------+------------------------+------------------------+ | richards_super | 31.2 ms | 31.1 ms: 1.01x faster | 30.7 ms: 1.02x faster | +--------------------------+----------+------------------------+------------------------+ | scimark_fft | 174 ms | 178 ms: 1.02x slower | 177 ms: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | scimark_lu | 71.5 ms | 69.3 ms: 1.03x faster | 70.2 ms: 1.02x faster | +--------------------------+----------+------------------------+------------------------+ | scimark_monte_carlo | 41.9 ms | 41.2 ms: 1.02x faster | 41.8 ms: 1.00x faster | +--------------------------+----------+------------------------+------------------------+ | scimark_sor | 70.4 ms | 72.0 ms: 1.02x slower | not significant | +--------------------------+----------+------------------------+------------------------+ | scimark_sparse_mat_mult | 2.67 ms | 2.71 ms: 1.02x slower | not significant | +--------------------------+----------+------------------------+------------------------+ | spectral_norm | 59.2 ms | 58.8 ms: 1.01x faster | 61.2 ms: 1.03x slower | +--------------------------+----------+------------------------+------------------------+ | sqlglot_normalize | 176 ms | 179 ms: 1.01x slower | 179 ms: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | sqlglot_optimize | 34.0 ms | 34.1 ms: 1.00x slower | 34.3 ms: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | sqlglot_parse | 801 us | 792 us: 1.01x faster | 796 us: 1.01x faster | +--------------------------+----------+------------------------+------------------------+ | sqlglot_transpile | 1.01 ms | 994 us: 1.01x faster | 1.01 ms: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | sympy_expand | 311 ms | 306 ms: 1.02x faster | 312 ms: 1.00x slower | +--------------------------+----------+------------------------+------------------------+ | sympy_integrate | 12.5 ms | not significant | 12.5 ms: 1.00x slower | +--------------------------+----------+------------------------+------------------------+ | sympy_sum | 92.8 ms | 92.4 ms: 1.00x faster | 93.1 ms: 1.00x slower | +--------------------------+----------+------------------------+------------------------+ | sympy_str | 177 ms | 176 ms: 1.01x faster | not significant | +--------------------------+----------+------------------------+------------------------+ | telco | 112 ms | 111 ms: 1.00x faster | 113 ms: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | tomli_loads | 1.21 sec | 1.23 sec: 1.02x slower | 1.23 sec: 1.02x slower | +--------------------------+----------+------------------------+------------------------+ | typing_runtime_protocols | 107 us | 109 us: 1.01x slower | not significant | +--------------------------+----------+------------------------+------------------------+ | unpack_sequence | 25.2 ns | 26.3 ns: 1.04x slower | 25.4 ns: 1.01x slower | +--------------------------+----------+------------------------+------------------------+ | unpickle_list | 2.93 us | 3.01 us: 1.03x slower | 3.18 us: 1.09x slower | +--------------------------+----------+------------------------+------------------------+ | unpickle_pure_python | 137 us | 138 us: 1.01x slower | 139 us: 1.02x slower | +--------------------------+----------+------------------------+------------------------+ | xml_etree_iterparse | 62.9 ms | 62.0 ms: 1.01x faster | 62.4 ms: 1.01x faster | +--------------------------+----------+------------------------+------------------------+ | xml_etree_generate | 54.7 ms | 55.4 ms: 1.01x slower | 59.2 ms: 1.08x slower | +--------------------------+----------+------------------------+------------------------+ | xml_etree_process | 39.0 ms | 40.3 ms: 1.03x slower | 42.0 ms: 1.08x slower | +--------------------------+----------+------------------------+------------------------+ | Geometric mean | (ref) | 1.00x slower | 1.01x slower | +--------------------------+----------+------------------------+------------------------+I exclude
bench_mp_poolbecause it is flaky and have issues last time - #139881+---------------+---------+-----------------------+-----------------------+ | Benchmark | m | o | s | +===============+=========+=======================+=======================+ | bench_mp_pool | 66.1 ms | 43.0 ms: 1.54x faster | 94.1 ms: 1.42x slower | +---------------+---------+-----------------------+-----------------------+sergey-miryanov commented
on Oct 13, 2025 ContributorAuthorMore actionsImplementation for
PyTuple_Make[Single,PairTriplet]Steal:PyObject * PyTuple_MakeSingle(PyObject *one) { assert (one != NULL); PyTupleObject *op = tuple_alloc(1); if (op == NULL) { return NULL; } op->ob_item[0] = Py_NewRef(one); _PyObject_GC_TRACK(op); return (PyObject *) op; } PyObject * PyTuple_MakeSingleSteal(PyObject *one) { assert (one != NULL); PyTupleObject *op = tuple_alloc(1); if (op == NULL) { Py_DECREF(one); return NULL; } op->ob_item[0] = one; _PyObject_GC_TRACK(op); return (PyObject *) op; } PyObject * PyTuple_MakePair(PyObject *one, PyObject *two) { assert (one != NULL); assert (two != NULL); PyTupleObject *op = tuple_alloc(2); if (op == NULL) { return NULL; } op->ob_item[0] = Py_NewRef(one); op->ob_item[1] = Py_NewRef(two); _PyObject_GC_TRACK(op); return (PyObject *) op; } PyObject * PyTuple_MakePairSteal(PyObject *one, PyObject *two) { assert (one != NULL); assert (two != NULL); PyTupleObject *op = tuple_alloc(2); if (op == NULL) { Py_DECREF(one); Py_DECREF(two); return NULL; } op->ob_item[0] = one; op->ob_item[1] = two; _PyObject_GC_TRACK(op); return (PyObject *) op; } PyObject * PyTuple_MakeTriplet(PyObject *one, PyObject *two, PyObject *three) { assert (one != NULL); assert (two != NULL); assert (three != NULL); PyTupleObject *op = tuple_alloc(3); if (op == NULL) { return NULL; } op->ob_item[0] = Py_NewRef(one); op->ob_item[1] = Py_NewRef(two); op->ob_item[2] = Py_NewRef(three); _PyObject_GC_TRACK(op); return (PyObject *) op; }Do we really need Steal variants? I would prefer to restrict to the minimum, PyTuple_MakeSingle() and PyTuple_MakePair().
Reacted by Peter Biermasergey-miryanov commented
on Oct 13, 2025 ContributorAuthorMore actionsI'm not sure, I did it just for test.
We have little or nothing to steal. Some parts where we can replacePyTuple_New+PyTuple_SET_ITEM.I agree that we don't need "steal" variants. I think it's pretty easy to decref an object after calling
PyTuple_MakePairor whatever.I think it's pretty easy to decref an object after calling PyTuple_MakePair or whatever.
Well, it's convenient. But I would prefer to limit the number of added functions.
I would prefer to have only steal versions because there are less redundant reference counting operations. decrefs are expensive on hot paths and users can easily incref before passing to steal function as needed.
Considering that reference counting is more expensive on free-threading, I think we should encourage the use of Steal functions to avoid redundant reference counting operations.
See also #140010.
sergey-miryanov commented
on Nov 17, 2025 ContributorAuthorMore actionsIIUC, we don't want to add a steal version too. Closing this.
Reacted by Victor Stinner
Feature or enhancement
Proposal:
I ran benchmarks on
pyperformanceand found that tuples with one or two elements account for about 80% of the total.I checked the code and filled the following table with the number of occurrences of one- and two-element tuples:
Table with number of occurrences
_asyncmodule.cPyTuple_New(2)_collectionsmodule.cPyTuple_Pack(1)_csv.cPyTuple_Pack(1)_datetimemodule.cPyTuple_Pack(2)PyTuple_Pack(1)_elementtree.cPyTuple_Pack(2)_functoolsmodule.cPyTuple_New(2)_interpretersmodule.cPyTuple_Pack(2)_json.cPyTuple_New(2)PyTuple_Pack(2)PyTuple_Pack(1)_operator.cPyTuple_Pack(2)_pickle.cPyTuple_Pack(2)PyTuple_New(2)PyTuple_New(1)_ssl.cPyTuple_New(2)PyTuple_Pack(2)_threadmodule.cPyTuple_New(2)_tkinter.cPyTuple_Pack(1)arraymodule.cPyTuple_New(2)itertoolsmodule.cPyTuple_Pack(2)PyTuple_New(2)main.cPyTuple_Pack(2)overlapped.cPyTuple_New(2)posixmodule.cPyTuple_Pack(2)pyexpat.cPyTuple_New(1)selectmodule.cPyTuple_Pack(2)PyTuple_New(2)signal_module.cPyTuple_New(2)socket_module.cPyTuple_Pack(2)termios.cPyTuple_New(2)_ctypes.cPyTuple_Pack(2)stgdict.cPyTuple_Pack(2)decimal.cPyTuple_Pack(2)PyTuple_Pack(1)microprotocol.cPyTuple_Pack(2)_sre.cPyTuple_New(2)datetime.cPyTuple_Pack(1)PyTuple_Pack(2)getargs.cPyTuple_Pack(1)heaptype.cPyTuple_Pack(2)PyTuple_Pack(1)PyTuple_New(2)vectorcall_limited.cPyTuple_New(1)multibytecodec.cPyTuple_New(2)codeobject.cPyTuple_Pack(2)dictobject.cPyTuple_Pack(2)PyTuple_New(2)enumobject.cPyTuple_Pack(2)PyTuple_New(2)exceptions.cPyTuple_Pack(2)floatobject.cPyTuple_Pack(2)frameobject.cPyTuple_Pack(2)PyTuple_Pack(1)genericaliasobject.cPyTuple_Pack(1)listobject.cPyTuple_Pack(2)longobject.cPyTuple_Pack(2)PyTuple_New(2)odictobject.cPyTuple_Pack(2)PyTuple_New(2)setobject.cPyTuple_Pack(1)typeobject.cPyTuple_Pack(2)PyTuple_Pack(1)typevarobject.cPyTuple_Pack(2)PyTuple_Pack(1)unicode_format.hPyTuple_Pack(2)pegen_errors.cPyTuple_Pack(2)_warnings.cPyTuple_Pack(2)bltnmodule.cPyTuple_Pack(2)ceval.cPyTuple_Pack(1)_codegen.cPyTuple_Pack(1)compile.cPyTuple_Pack(2)crossinterp.cPyTuple_Pack(1)errors.cPyTuple_Pack(1)hamt.cPyTuple_Pack(2)marshal.cPyTuple_Pack(2)pylifecycle.cPyTuple_Pack(2)Python-tokenize.cPyTuple_Pack(2)sysmodule.cPyTuple_Pack(1)tracemalloc.cPyTuple_New(2)I came up with the idea of adding
PyTuple_MakeSingleandPyTuple_MakePairfor such cases to improve performance.Afterwards, @eendebakpt sent me a link with a previous attempt at this (many thanks!) - #118222.
Anyway, I implemented these changes and ran benchmarks.
If we replace
PyTuple_Pack(1,...)withPyTuple_MakeSingleandPyTuple_Pack(2,...)withPyTuple_MakePairthen we get following results (ran on ubuntu 24.04 x64, compiled with lto):Geometric mean - 1.00x faster
I plan to implement
PyTuple_Make[Single,Pair]Stealand also replacePyTuple_New(1,2).Has this already been discussed elsewhere?
This is a minor feature, which does not need previous discussion elsewhere
Links to previous discussion of this feature:
No response
Linked PRs