Skip to content

[C API] Replace PyTuple_Pack(1,2) with PyTuple_Make[Single,Pair] to optimize creation of tuples #140052

Description

@sergey-miryanov

Feature or enhancement

Proposal:

I ran benchmarks on pyperformance and found that tuples with one or two elements account for about 80% of the total.

Image

I checked the code and filled the following table with the number of occurrences of one- and two-element tuples:

Table with number of occurrences
file function count
_asyncmodule.c PyTuple_New(2) 2
_collectionsmodule.c PyTuple_Pack(1) 2
_csv.c PyTuple_Pack(1) 1
_datetimemodule.c PyTuple_Pack(2) 4
PyTuple_Pack(1) 3
_elementtree.c PyTuple_Pack(2) 4
_functoolsmodule.c PyTuple_New(2) 2
_interpretersmodule.c PyTuple_Pack(2) 1
_json.c PyTuple_New(2) 1
PyTuple_Pack(2) 1
PyTuple_Pack(1) 1
_operator.c PyTuple_Pack(2) 1
_pickle.c PyTuple_Pack(2) 3
PyTuple_New(2) 2
PyTuple_New(1) 2
_ssl.c PyTuple_New(2) 6
PyTuple_Pack(2) 1
_threadmodule.c PyTuple_New(2) 1
_tkinter.c PyTuple_Pack(1)
arraymodule.c PyTuple_New(2) 2
itertoolsmodule.c PyTuple_Pack(2) 2
PyTuple_New(2) 1
main.c PyTuple_Pack(2) 1
overlapped.c PyTuple_New(2) 2
posixmodule.c PyTuple_Pack(2) 1
pyexpat.c PyTuple_New(1) 1
selectmodule.c PyTuple_Pack(2) 1
PyTuple_New(2) 1
signal_module.c PyTuple_New(2) 1
socket_module.c PyTuple_Pack(2) 3
termios.c PyTuple_New(2) 2
_ctypes.c PyTuple_Pack(2) 2
stgdict.c PyTuple_Pack(2) 1
decimal.c PyTuple_Pack(2) 7
PyTuple_Pack(1) 2
microprotocol.c PyTuple_Pack(2) 2
_sre.c PyTuple_New(2) 1
datetime.c PyTuple_Pack(1) 2
PyTuple_Pack(2) 1
getargs.c PyTuple_Pack(1) 1
heaptype.c PyTuple_Pack(2) 1
PyTuple_Pack(1) 2
PyTuple_New(2) 1
vectorcall_limited.c PyTuple_New(1) 2
multibytecodec.c PyTuple_New(2) 1
codeobject.c PyTuple_Pack(2) 7
dictobject.c PyTuple_Pack(2) 2
PyTuple_New(2) 4
enumobject.c PyTuple_Pack(2) 1
PyTuple_New(2) 2
exceptions.c PyTuple_Pack(2) 7
floatobject.c PyTuple_Pack(2) 1
frameobject.c PyTuple_Pack(2) 2
PyTuple_Pack(1) 1
genericaliasobject.c PyTuple_Pack(1) 2
listobject.c PyTuple_Pack(2) 1
longobject.c PyTuple_Pack(2) 1
PyTuple_New(2) 2
odictobject.c PyTuple_Pack(2) 2
PyTuple_New(2) 1
setobject.c PyTuple_Pack(1) 1
typeobject.c PyTuple_Pack(2) 2
PyTuple_Pack(1) 5
typevarobject.c PyTuple_Pack(2) 1
PyTuple_Pack(1) 2
unicode_format.h PyTuple_Pack(2) 2
pegen_errors.c PyTuple_Pack(2) 2
_warnings.c PyTuple_Pack(2) 1
bltnmodule.c PyTuple_Pack(2) 1
ceval.c PyTuple_Pack(1) 1
_codegen.c PyTuple_Pack(1) 2
compile.c PyTuple_Pack(2) 1
crossinterp.c PyTuple_Pack(1) 1
errors.c PyTuple_Pack(1) 1
hamt.c PyTuple_Pack(2) 1
marshal.c PyTuple_Pack(2) 1
pylifecycle.c PyTuple_Pack(2) 1
Python-tokenize.c PyTuple_Pack(2) 1
sysmodule.c PyTuple_Pack(1) 1
tracemalloc.c PyTuple_New(2) 1

I came up with the idea of adding PyTuple_MakeSingle and PyTuple_MakePair for such cases to improve performance.

Afterwards, @eendebakpt sent me a link with a previous attempt at this (many thanks!) - #118222.

Anyway, I implemented these changes and ran benchmarks.

If we replace PyTuple_Pack(1,...) with PyTuple_MakeSingle and PyTuple_Pack(2,...) with PyTuple_MakePair then we get following results (ran on ubuntu 24.04 x64, compiled with lto):

Geometric mean - 1.00x faster
+--------------------------+----------+------------------------+
| Benchmark                | main     | opt                    |
+==========================+==========+========================+
| async_generators         | 277 ms   | 279 ms: 1.01x slower   |
+--------------------------+----------+------------------------+
| asyncio_websockets       | 242 ms   | 241 ms: 1.00x faster   |
+--------------------------+----------+------------------------+
| chaos                    | 36.0 ms  | 36.6 ms: 1.02x slower  |
+--------------------------+----------+------------------------+
| comprehensions           | 10.3 us  | 10.1 us: 1.02x faster  |
+--------------------------+----------+------------------------+
| bench_mp_pool            | 66.1 ms  | 43.0 ms: 1.54x faster  |
+--------------------------+----------+------------------------+
| coroutines               | 15.0 ms  | 14.5 ms: 1.03x faster  |
+--------------------------+----------+------------------------+
| coverage                 | 54.9 ms  | 56.1 ms: 1.02x slower  |
+--------------------------+----------+------------------------+
| crypto_pyaes             | 45.3 ms  | 46.1 ms: 1.02x slower  |
+--------------------------+----------+------------------------+
| deepcopy                 | 171 us   | 168 us: 1.02x faster   |
+--------------------------+----------+------------------------+
| deepcopy_reduce          | 1.87 us  | 1.84 us: 1.02x faster  |
+--------------------------+----------+------------------------+
| deepcopy_memo            | 16.5 us  | 17.5 us: 1.06x slower  |
+--------------------------+----------+------------------------+
| deltablue                | 1.97 ms  | 1.99 ms: 1.01x slower  |
+--------------------------+----------+------------------------+
| django_template          | 23.1 ms  | 23.4 ms: 1.02x slower  |
+--------------------------+----------+------------------------+
| docutils                 | 1.70 sec | 1.68 sec: 1.01x faster |
+--------------------------+----------+------------------------+
| fannkuch                 | 245 ms   | 243 ms: 1.01x faster   |
+--------------------------+----------+------------------------+
| float                    | 42.4 ms  | 43.4 ms: 1.02x slower  |
+--------------------------+----------+------------------------+
| gc_traversal             | 2.93 ms  | 2.82 ms: 1.04x faster  |
+--------------------------+----------+------------------------+
| generators               | 19.1 ms  | 19.5 ms: 1.02x slower  |
+--------------------------+----------+------------------------+
| genshi_text              | 14.2 ms  | 14.4 ms: 1.02x slower  |
+--------------------------+----------+------------------------+
| genshi_xml               | 32.8 ms  | 33.1 ms: 1.01x slower  |
+--------------------------+----------+------------------------+
| go                       | 68.7 ms  | 70.2 ms: 1.02x slower  |
+--------------------------+----------+------------------------+
| hexiom                   | 3.64 ms  | 3.61 ms: 1.01x faster  |
+--------------------------+----------+------------------------+
| json_dumps               | 6.34 ms  | 6.26 ms: 1.01x faster  |
+--------------------------+----------+------------------------+
| json_loads               | 15.7 us  | 16.1 us: 1.02x slower  |
+--------------------------+----------+------------------------+
| logging_silent           | 63.6 ns  | 59.4 ns: 1.07x faster  |
+--------------------------+----------+------------------------+
| logging_simple           | 3.49 us  | 3.52 us: 1.01x slower  |
+--------------------------+----------+------------------------+
| mako                     | 7.00 ms  | 6.94 ms: 1.01x faster  |
+--------------------------+----------+------------------------+
| mdp                      | 788 ms   | 780 ms: 1.01x faster   |
+--------------------------+----------+------------------------+
| meteor_contest           | 68.0 ms  | 68.2 ms: 1.00x slower  |
+--------------------------+----------+------------------------+
| nbody                    | 55.4 ms  | 55.0 ms: 1.01x faster  |
+--------------------------+----------+------------------------+
| pickle_dict              | 18.5 us  | 18.8 us: 1.02x slower  |
+--------------------------+----------+------------------------+
| pickle_list              | 2.93 us  | 2.98 us: 1.02x slower  |
+--------------------------+----------+------------------------+
| pickle_pure_python       | 208 us   | 205 us: 1.01x faster   |
+--------------------------+----------+------------------------+
| pidigits                 | 143 ms   | 143 ms: 1.00x slower   |
+--------------------------+----------+------------------------+
| pprint_safe_repr         | 489 ms   | 496 ms: 1.01x slower   |
+--------------------------+----------+------------------------+
| pprint_pformat           | 997 ms   | 1.01 sec: 1.01x slower |
+--------------------------+----------+------------------------+
| pyflate                  | 259 ms   | 260 ms: 1.00x slower   |
+--------------------------+----------+------------------------+
| regex_compile            | 82.7 ms  | 83.5 ms: 1.01x slower  |
+--------------------------+----------+------------------------+
| regex_dna                | 115 ms   | 113 ms: 1.01x faster   |
+--------------------------+----------+------------------------+
| regex_v8                 | 14.7 ms  | 14.3 ms: 1.03x faster  |
+--------------------------+----------+------------------------+
| richards                 | 27.3 ms  | 27.1 ms: 1.01x faster  |
+--------------------------+----------+------------------------+
| richards_super           | 31.2 ms  | 31.1 ms: 1.01x faster  |
+--------------------------+----------+------------------------+
| scimark_fft              | 174 ms   | 178 ms: 1.02x slower   |
+--------------------------+----------+------------------------+
| scimark_lu               | 71.5 ms  | 69.3 ms: 1.03x faster  |
+--------------------------+----------+------------------------+
| scimark_monte_carlo      | 41.9 ms  | 41.2 ms: 1.02x faster  |
+--------------------------+----------+------------------------+
| scimark_sor              | 70.4 ms  | 72.0 ms: 1.02x slower  |
+--------------------------+----------+------------------------+
| scimark_sparse_mat_mult  | 2.67 ms  | 2.71 ms: 1.02x slower  |
+--------------------------+----------+------------------------+
| spectral_norm            | 59.2 ms  | 58.8 ms: 1.01x faster  |
+--------------------------+----------+------------------------+
| sqlglot_normalize        | 176 ms   | 179 ms: 1.01x slower   |
+--------------------------+----------+------------------------+
| sqlglot_optimize         | 34.0 ms  | 34.1 ms: 1.00x slower  |
+--------------------------+----------+------------------------+
| sqlglot_parse            | 801 us   | 792 us: 1.01x faster   |
+--------------------------+----------+------------------------+
| sqlglot_transpile        | 1.01 ms  | 994 us: 1.01x faster   |
+--------------------------+----------+------------------------+
| sympy_expand             | 311 ms   | 306 ms: 1.02x faster   |
+--------------------------+----------+------------------------+
| sympy_sum                | 92.8 ms  | 92.4 ms: 1.00x faster  |
+--------------------------+----------+------------------------+
| sympy_str                | 177 ms   | 176 ms: 1.01x faster   |
+--------------------------+----------+------------------------+
| telco                    | 112 ms   | 111 ms: 1.00x faster   |
+--------------------------+----------+------------------------+
| tomli_loads              | 1.21 sec | 1.23 sec: 1.02x slower |
+--------------------------+----------+------------------------+
| typing_runtime_protocols | 107 us   | 109 us: 1.01x slower   |
+--------------------------+----------+------------------------+
| unpack_sequence          | 25.2 ns  | 26.3 ns: 1.04x slower  |
+--------------------------+----------+------------------------+
| unpickle_list            | 2.93 us  | 3.01 us: 1.03x slower  |
+--------------------------+----------+------------------------+
| unpickle_pure_python     | 137 us   | 138 us: 1.01x slower   |
+--------------------------+----------+------------------------+
| xml_etree_iterparse      | 62.9 ms  | 62.0 ms: 1.01x faster  |
+--------------------------+----------+------------------------+
| xml_etree_generate       | 54.7 ms  | 55.4 ms: 1.01x slower  |
+--------------------------+----------+------------------------+
| xml_etree_process        | 39.0 ms  | 40.3 ms: 1.03x slower  |
+--------------------------+----------+------------------------+
| Geometric mean           | (ref)    | 1.00x faster           |
+--------------------------+----------+------------------------+

Benchmark hidden because not significant (19): 2to3, asyncio_tcp, asyncio_tcp_ssl, bench_thread_pool, dulwich_log, create_gc_cycles, html5lib, logging_format, nqueens, pathlib, pickle, python_startup, python_startup_no_site, raytrace, regex_effbot, sqlite_synth, sympy_integrate, unpickle, xml_etree_parse

I plan to implement PyTuple_Make[Single,Pair]Steal and also replace PyTuple_New(1,2).

Has this already been discussed elsewhere?

This is a minor feature, which does not need previous discussion elsewhere

Links to previous discussion of this feature:

No response

Linked PRs

Activity

  1. sergey-miryanov commented on Oct 13, 2025

    @sergey-miryanov
    ContributorAuthor

    For Steal version I have following results:

    Geometric mean - 1.01x slower
    +--------------------------+----------+------------------------+------------------------+
    | Benchmark                | m        | o                      | s                      |
    +==========================+==========+========================+========================+
    | async_generators         | 277 ms   | 279 ms: 1.01x slower   | 275 ms: 1.01x faster   |
    +--------------------------+----------+------------------------+------------------------+
    | asyncio_tcp_ssl          | 812 ms   | not significant        | 803 ms: 1.01x faster   |
    +--------------------------+----------+------------------------+------------------------+
    | asyncio_websockets       | 242 ms   | 241 ms: 1.00x faster   | not significant        |
    +--------------------------+----------+------------------------+------------------------+
    | chaos                    | 36.0 ms  | 36.6 ms: 1.02x slower  | 36.7 ms: 1.02x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | comprehensions           | 10.3 us  | 10.1 us: 1.02x faster  | 10.1 us: 1.01x faster  |
    +--------------------------+----------+------------------------+------------------------+
    | coroutines               | 15.0 ms  | 14.5 ms: 1.03x faster  | 14.7 ms: 1.02x faster  |
    +--------------------------+----------+------------------------+------------------------+
    | coverage                 | 54.9 ms  | 56.1 ms: 1.02x slower  | 56.2 ms: 1.02x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | crypto_pyaes             | 45.3 ms  | 46.1 ms: 1.02x slower  | 45.6 ms: 1.01x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | deepcopy                 | 171 us   | 168 us: 1.02x faster   | 168 us: 1.02x faster   |
    +--------------------------+----------+------------------------+------------------------+
    | deepcopy_reduce          | 1.87 us  | 1.84 us: 1.02x faster  | 1.83 us: 1.02x faster  |
    +--------------------------+----------+------------------------+------------------------+
    | deepcopy_memo            | 16.5 us  | 17.5 us: 1.06x slower  | 16.9 us: 1.02x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | deltablue                | 1.97 ms  | 1.99 ms: 1.01x slower  | not significant        |
    +--------------------------+----------+------------------------+------------------------+
    | django_template          | 23.1 ms  | 23.4 ms: 1.02x slower  | 23.4 ms: 1.01x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | docutils                 | 1.70 sec | 1.68 sec: 1.01x faster | not significant        |
    +--------------------------+----------+------------------------+------------------------+
    | dulwich_log              | 43.3 ms  | not significant        | 43.7 ms: 1.01x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | fannkuch                 | 245 ms   | 243 ms: 1.01x faster   | 238 ms: 1.03x faster   |
    +--------------------------+----------+------------------------+------------------------+
    | float                    | 42.4 ms  | 43.4 ms: 1.02x slower  | 43.3 ms: 1.02x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | create_gc_cycles         | 1.18 ms  | not significant        | 1.19 ms: 1.01x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | gc_traversal             | 2.93 ms  | 2.82 ms: 1.04x faster  | 3.03 ms: 1.03x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | generators               | 19.1 ms  | 19.5 ms: 1.02x slower  | 19.3 ms: 1.01x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | genshi_text              | 14.2 ms  | 14.4 ms: 1.02x slower  | 14.4 ms: 1.02x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | genshi_xml               | 32.8 ms  | 33.1 ms: 1.01x slower  | 33.3 ms: 1.02x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | go                       | 68.7 ms  | 70.2 ms: 1.02x slower  | 69.3 ms: 1.01x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | hexiom                   | 3.64 ms  | 3.61 ms: 1.01x faster  | 3.62 ms: 1.00x faster  |
    +--------------------------+----------+------------------------+------------------------+
    | json_dumps               | 6.34 ms  | 6.26 ms: 1.01x faster  | 6.43 ms: 1.01x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | json_loads               | 15.7 us  | 16.1 us: 1.02x slower  | 16.6 us: 1.06x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | logging_format           | 3.88 us  | not significant        | 3.91 us: 1.01x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | logging_silent           | 63.6 ns  | 59.4 ns: 1.07x faster  | 59.0 ns: 1.08x faster  |
    +--------------------------+----------+------------------------+------------------------+
    | logging_simple           | 3.49 us  | 3.52 us: 1.01x slower  | 3.56 us: 1.02x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | mako                     | 7.00 ms  | 6.94 ms: 1.01x faster  | 6.90 ms: 1.01x faster  |
    +--------------------------+----------+------------------------+------------------------+
    | mdp                      | 788 ms   | 780 ms: 1.01x faster   | 781 ms: 1.01x faster   |
    +--------------------------+----------+------------------------+------------------------+
    | meteor_contest           | 68.0 ms  | 68.2 ms: 1.00x slower  | 69.2 ms: 1.02x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | nbody                    | 55.4 ms  | 55.0 ms: 1.01x faster  | 55.9 ms: 1.01x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | nqueens                  | 55.8 ms  | not significant        | 57.0 ms: 1.02x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | pathlib                  | 9.71 ms  | not significant        | 9.79 ms: 1.01x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | pickle                   | 7.29 us  | not significant        | 7.37 us: 1.01x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | pickle_dict              | 18.5 us  | 18.8 us: 1.02x slower  | 18.7 us: 1.01x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | pickle_list              | 2.93 us  | 2.98 us: 1.02x slower  | 3.02 us: 1.03x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | pickle_pure_python       | 208 us   | 205 us: 1.01x faster   | not significant        |
    +--------------------------+----------+------------------------+------------------------+
    | pidigits                 | 143 ms   | 143 ms: 1.00x slower   | 144 ms: 1.01x slower   |
    +--------------------------+----------+------------------------+------------------------+
    | pprint_safe_repr         | 489 ms   | 496 ms: 1.01x slower   | 502 ms: 1.03x slower   |
    +--------------------------+----------+------------------------+------------------------+
    | pprint_pformat           | 997 ms   | 1.01 sec: 1.01x slower | 1.03 sec: 1.03x slower |
    +--------------------------+----------+------------------------+------------------------+
    | pyflate                  | 259 ms   | 260 ms: 1.00x slower   | 258 ms: 1.00x faster   |
    +--------------------------+----------+------------------------+------------------------+
    | python_startup           | 7.74 ms  | not significant        | 7.74 ms: 1.00x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | python_startup_no_site   | 5.09 ms  | not significant        | 5.10 ms: 1.00x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | raytrace                 | 171 ms   | not significant        | 170 ms: 1.00x faster   |
    +--------------------------+----------+------------------------+------------------------+
    | regex_compile            | 82.7 ms  | 83.5 ms: 1.01x slower  | 83.5 ms: 1.01x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | regex_dna                | 115 ms   | 113 ms: 1.01x faster   | not significant        |
    +--------------------------+----------+------------------------+------------------------+
    | regex_effbot             | 1.68 ms  | not significant        | 1.63 ms: 1.03x faster  |
    +--------------------------+----------+------------------------+------------------------+
    | regex_v8                 | 14.7 ms  | 14.3 ms: 1.03x faster  | 14.6 ms: 1.00x faster  |
    +--------------------------+----------+------------------------+------------------------+
    | richards                 | 27.3 ms  | 27.1 ms: 1.01x faster  | not significant        |
    +--------------------------+----------+------------------------+------------------------+
    | richards_super           | 31.2 ms  | 31.1 ms: 1.01x faster  | 30.7 ms: 1.02x faster  |
    +--------------------------+----------+------------------------+------------------------+
    | scimark_fft              | 174 ms   | 178 ms: 1.02x slower   | 177 ms: 1.01x slower   |
    +--------------------------+----------+------------------------+------------------------+
    | scimark_lu               | 71.5 ms  | 69.3 ms: 1.03x faster  | 70.2 ms: 1.02x faster  |
    +--------------------------+----------+------------------------+------------------------+
    | scimark_monte_carlo      | 41.9 ms  | 41.2 ms: 1.02x faster  | 41.8 ms: 1.00x faster  |
    +--------------------------+----------+------------------------+------------------------+
    | scimark_sor              | 70.4 ms  | 72.0 ms: 1.02x slower  | not significant        |
    +--------------------------+----------+------------------------+------------------------+
    | scimark_sparse_mat_mult  | 2.67 ms  | 2.71 ms: 1.02x slower  | not significant        |
    +--------------------------+----------+------------------------+------------------------+
    | spectral_norm            | 59.2 ms  | 58.8 ms: 1.01x faster  | 61.2 ms: 1.03x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | sqlglot_normalize        | 176 ms   | 179 ms: 1.01x slower   | 179 ms: 1.01x slower   |
    +--------------------------+----------+------------------------+------------------------+
    | sqlglot_optimize         | 34.0 ms  | 34.1 ms: 1.00x slower  | 34.3 ms: 1.01x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | sqlglot_parse            | 801 us   | 792 us: 1.01x faster   | 796 us: 1.01x faster   |
    +--------------------------+----------+------------------------+------------------------+
    | sqlglot_transpile        | 1.01 ms  | 994 us: 1.01x faster   | 1.01 ms: 1.01x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | sympy_expand             | 311 ms   | 306 ms: 1.02x faster   | 312 ms: 1.00x slower   |
    +--------------------------+----------+------------------------+------------------------+
    | sympy_integrate          | 12.5 ms  | not significant        | 12.5 ms: 1.00x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | sympy_sum                | 92.8 ms  | 92.4 ms: 1.00x faster  | 93.1 ms: 1.00x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | sympy_str                | 177 ms   | 176 ms: 1.01x faster   | not significant        |
    +--------------------------+----------+------------------------+------------------------+
    | telco                    | 112 ms   | 111 ms: 1.00x faster   | 113 ms: 1.01x slower   |
    +--------------------------+----------+------------------------+------------------------+
    | tomli_loads              | 1.21 sec | 1.23 sec: 1.02x slower | 1.23 sec: 1.02x slower |
    +--------------------------+----------+------------------------+------------------------+
    | typing_runtime_protocols | 107 us   | 109 us: 1.01x slower   | not significant        |
    +--------------------------+----------+------------------------+------------------------+
    | unpack_sequence          | 25.2 ns  | 26.3 ns: 1.04x slower  | 25.4 ns: 1.01x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | unpickle_list            | 2.93 us  | 3.01 us: 1.03x slower  | 3.18 us: 1.09x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | unpickle_pure_python     | 137 us   | 138 us: 1.01x slower   | 139 us: 1.02x slower   |
    +--------------------------+----------+------------------------+------------------------+
    | xml_etree_iterparse      | 62.9 ms  | 62.0 ms: 1.01x faster  | 62.4 ms: 1.01x faster  |
    +--------------------------+----------+------------------------+------------------------+
    | xml_etree_generate       | 54.7 ms  | 55.4 ms: 1.01x slower  | 59.2 ms: 1.08x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | xml_etree_process        | 39.0 ms  | 40.3 ms: 1.03x slower  | 42.0 ms: 1.08x slower  |
    +--------------------------+----------+------------------------+------------------------+
    | Geometric mean           | (ref)    | 1.00x slower           | 1.01x slower           |
    +--------------------------+----------+------------------------+------------------------+
    

    I exclude bench_mp_pool because it is flaky and have issues last time - #139881

    +---------------+---------+-----------------------+-----------------------+
    | Benchmark     | m       | o                     | s                     |
    +===============+=========+=======================+=======================+
    | bench_mp_pool | 66.1 ms | 43.0 ms: 1.54x faster | 94.1 ms: 1.42x slower |
    +---------------+---------+-----------------------+-----------------------+
    
  2. sergey-miryanov commented on Oct 13, 2025

    @sergey-miryanov
    ContributorAuthor

    Implementation for PyTuple_Make[Single,PairTriplet]Steal:

    
    PyObject *
    PyTuple_MakeSingle(PyObject *one)
    {
        assert (one != NULL);
    
        PyTupleObject *op = tuple_alloc(1);
        if (op == NULL) {
            return NULL;
        }
        op->ob_item[0] = Py_NewRef(one);
        _PyObject_GC_TRACK(op);
        return (PyObject *) op;
    }
    
    
    PyObject *
    PyTuple_MakeSingleSteal(PyObject *one)
    {
        assert (one != NULL);
    
        PyTupleObject *op = tuple_alloc(1);
        if (op == NULL) {
            Py_DECREF(one);
            return NULL;
        }
        op->ob_item[0] = one;
        _PyObject_GC_TRACK(op);
        return (PyObject *) op;
    }
    
    PyObject *
    PyTuple_MakePair(PyObject *one, PyObject *two)
    {
        assert (one != NULL);
        assert (two != NULL);
    
        PyTupleObject *op = tuple_alloc(2);
        if (op == NULL) {
            return NULL;
        }
        op->ob_item[0] = Py_NewRef(one);
        op->ob_item[1] = Py_NewRef(two);
        _PyObject_GC_TRACK(op);
        return (PyObject *) op;
    }
    
    PyObject *
    PyTuple_MakePairSteal(PyObject *one, PyObject *two)
    {
        assert (one != NULL);
        assert (two != NULL);
    
        PyTupleObject *op = tuple_alloc(2);
        if (op == NULL) {
            Py_DECREF(one);
            Py_DECREF(two);
            return NULL;
        }
        op->ob_item[0] = one;
        op->ob_item[1] = two;
        _PyObject_GC_TRACK(op);
        return (PyObject *) op;
    }
    
    PyObject *
    PyTuple_MakeTriplet(PyObject *one, PyObject *two, PyObject *three)
    {
        assert (one != NULL);
        assert (two != NULL);
        assert (three != NULL);
    
        PyTupleObject *op = tuple_alloc(3);
        if (op == NULL) {
            return NULL;
        }
        op->ob_item[0] = Py_NewRef(one);
        op->ob_item[1] = Py_NewRef(two);
        op->ob_item[2] = Py_NewRef(three);
        _PyObject_GC_TRACK(op);
        return (PyObject *) op;
    }
    
    
  3. vstinner commented on Oct 13, 2025

    @vstinner
    Member

    Do we really need Steal variants? I would prefer to restrict to the minimum, PyTuple_MakeSingle() and PyTuple_MakePair().

  4. sergey-miryanov commented on Oct 13, 2025

    @sergey-miryanov
    ContributorAuthor

    I'm not sure, I did it just for test.
    We have little or nothing to steal. Some parts where we can replace PyTuple_New + PyTuple_SET_ITEM.

  5. ZeroIntensity commented on Oct 14, 2025

    @ZeroIntensity
    Member

    I agree that we don't need "steal" variants. I think it's pretty easy to decref an object after calling PyTuple_MakePair or whatever.

  6. vstinner commented on Oct 14, 2025

    @vstinner
    Member

    I think it's pretty easy to decref an object after calling PyTuple_MakePair or whatever.

    Well, it's convenient. But I would prefer to limit the number of added functions.

  7. kumaraditya303 commented on Oct 14, 2025

    @kumaraditya303
    Contributor

    I would prefer to have only steal versions because there are less redundant reference counting operations. decrefs are expensive on hot paths and users can easily incref before passing to steal function as needed.

    Considering that reference counting is more expensive on free-threading, I think we should encourage the use of Steal functions to avoid redundant reference counting operations.

  8. vstinner commented on Oct 14, 2025

    @vstinner
    Member

    See also #140010.

  9. sergey-miryanov commented on Nov 17, 2025

    @sergey-miryanov
    ContributorAuthor

    IIUC, we don't want to add a steal version too. Closing this.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions