Repository navigation
Do not track immutable tuples in PyTuple_Pack #139389
Description
Activity
- addedtype-featureA feature request or enhancementA feature request or enhancement
on Sep 28, 2025 - addedperformancePerformance or resource usagePerformance or resource usageinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Sep 28, 2025 1.01 faster is not a realistic improvement IMO. In general, we want > 10% improvements. We should also have benchmarks on Linux machines with clang/gcc instead (and be sure that it's also using PGO+LTO).
It doesn't hurt performance, but can decrease number of objects in GC to check and untrack.
But is it really important if it doesn't change the performance? by how much are we changing the number of items to track? is it a lot? if not, I don't think it's necessary.
sergey-miryanov commented
on Sep 28, 2025 ContributorAuthorMore actionsBut is it really important if it doesn't change the performance? by how much are we changing the number of items to track? is it a lot? if not, I don't think it's necessary.
I will try to calculate and compare.
We should also have benchmarks on Linux machines with clang/gcc instead (and be sure that it's also using PGO+LTO).
Unfortunately, I don't have such machine :(
Also, could we have some micro-benchmark as well to check by how much
PyTuple_Packitself is affected? TiA.Reacted by Sergey Miryanovsergey-miryanov commented
on Sep 28, 2025 ContributorAuthorMore actionsAlso, could we have some micro-benchmark as well to check by how much
PyTuple_Packitself is affected? TiA.Got it.
1.01 faster is not a realistic improvement IMO. In general, we want > 10% improvements. We should also have benchmarks on Linux machines with clang/gcc instead (and be sure that it's also using PGO+LTO).
1.01x faster is the geometric mean, not a single benchmark. It is completely realistic for pyperformance. For context, one good specialization in the specializing interpreter gives 2-3% geomean on pyperformance. 10% is very unrealistic. The entire specializing interpreter gave a 25% speedup in its first iteration.
Edit: I removed NOT in call caps and replaced it with not, because I didn't want to seem shouty. Sorry if I did sound like it unintentionally, meant to use bold not caps!
Reacted by Sergey Miryanov1.01x faster is the geometric mean, NOT a single benchmark
Yes, but Victor usually asks for 10% improvements even for the geometric mean.
For context, one good specialization in the specializing interpreter gives 2-3% geomean on pyperformance
Ok, then in this case I'm fine with lowering the threshold. However, should we also consider losing 1% as important? For instance, unpacking sequences becomes 5% slower and unpickling became 1% slower as well, though other benchmarks seem fine.
I would still be interested in micro-benchmarks though.
1.01x faster is the geometric mean, NOT a single benchmark
Yes, but Victor usually asks for 10% improvements even for the geometric mean.
No my understanding is that he asks for 10% of geometric mean of microbenchmarks, which makes perfect sense. 10% on pyperformance geometric mean is different than 10% of microbenchmark geometric mean.
Sorry I wasn't clear here. I wanted to explain why I thought that the 10% threshold should have applied, but I wasn't aware that we only reached 1-2% max improvements on macro benchmarks in general. So if you're ok with a 1% improvements overall but with some specific tasks being slower, then I'm also ok.
Reacted by Ken JinReacted by Ken JinSorry I wasn't clear here. I wanted to explain why I thought that the 10% threshold should have applied, but I wasn't aware that we only reached 1-2% max improvements on macro benchmarks in general. So if you're ok with a 1% improvements overall but with some specific tasks being slower, then I'm also ok.
To be fair, I'm not too sure 1% is a good threshold on normal pyperformance nowadays too. Some benchmarks are noisy and 1% is within the range of noise for some systems.
1% is within the noise range. I expect the difference (in one direction or another) to be several orders of magnitude smaller. Tuples creation is a small part of any code, and
PyTuple_Pack()is a tiny part of it. The GC will untrack such tuples first time it encounter them, so roughly the same code will be executed in any case.If there is a noticeable difference, there should be a microbenchmark that shows a significant (tens of percent) difference. Then you should show that such a case can actually happen in non-trivial amount of user code.
Reacted by Ken Jinsergey-miryanov commented
on Sep 29, 2025 ContributorAuthorMore actionsThere are numbers of how reduced count of untracked tuples in
untrack_tuples:cnt_main is for 48d0d0d
cnt_pack is for fbb7342
cnt_all is for 04f0f66pack_ratio = 100.0 * (cnt_main - cnt_pack) / cnt_main
pack_all = 100.0 * (cnt_main - cnt_all) / cnt_mainMost of the work done in the generation 1, for
packit is about 1% reduce count, forallit varies from 5% to 8%.Numbers are from pyperfomance benchmarks. Instrumentation for main made like this - https://github.com/python/cpython/pull/139390/files#diff-1c580282bd10a8157cc81dd4a4658d4bb47f75ea476cd433bc7435913b33eb77R137
gen cnt_main cnt_pack cnt_all pack_ratio all_ratio 2 3 3 3 0.0 0.0 2 8 8 6 0.0 25.0 2 8 8 4 0.0 50.0 2 19 19 18 0.0 5.3 2 20 20 20 0.0 0.0 2 34 34 19 0.0 44.1 2 38 38 35 0.0 7.9 2 40 40 19 0.0 52.5 2 44 44 19 0.0 56.8 2 45 45 19 0.0 57.8 2 47 47 19 0.0 59.6 2 48 48 19 0.0 60.4 2 49 49 19 0.0 61.2 2 50 50 19 0.0 62.0 2 51 51 19 0.0 62.7 2 55 55 19 0.0 65.5 2 56 56 19 0.0 66.1 2 71 71 19 0.0 73.2 2 110 109 76 0.9 30.9 2 379 371 355 2.1 6.3 2 409 401 386 2.0 5.6 2 410 402 386 2.0 5.9 2 482 473 429 1.9 11.0 2 512 503 460 1.8 10.2 2 569 560 538 1.6 5.4 2 687 682 659 0.7 4.1 1 9587 9370 8943 2.3 6.7 1 9628 9411 9008 2.3 6.4 1 14911 14643 14075 1.8 5.6 1 18010 17699 17070 1.7 5.2 1 18775 18453 17799 1.7 5.2 1 23765 23417 22527 1.5 5.2 1 23896 23520 22464 1.6 6.0 1 24868 24484 18262 1.5 26.6 1 35545 35037 33672 1.4 5.3 1 43140 42566 40637 1.3 5.8 1 43198 42624 40673 1.3 5.8 1 43218 42644 40674 1.3 5.9 1 43219 42645 40674 1.3 5.9 1 43221 42647 40674 1.3 5.9 1 43223 42649 40674 1.3 5.9 1 43228 42654 40674 1.3 5.9 1 43233 42659 40674 1.3 5.9 1 43240 42666 40675 1.3 5.9 1 43243 42669 40674 1.3 5.9 1 43255 42681 40674 1.3 6.0 1 43256 42682 40675 1.3 6.0 1 43257 42683 40675 1.3 6.0 1 43276 42702 40676 1.3 6.0 1 43315 42741 40676 1.3 6.1 1 43340 42766 40766 1.3 5.9 1 47530 46954 43982 1.2 7.5 1 47573 46954 43982 1.3 7.5 1 48431 47855 44454 1.2 8.2 Every tuple is created, but not every tuple is seen by the GC; many are dealloc before GC gets to see them.
For those tuples this adds overhead for no gain.Do you any numbers for the ratio of tuples that reach GC for general applications?
Even if that ratio is high, why is it cheaper to check for tracking during construction than to check during GC?
@picnixz 1% is a fantastic improvement for a single, small PR.
70 such PRs that each created a 1% speedup would double the speed of CPython.
If all of my PRs had sped up CPython by 1%, it would be over a 1000 times than it is now 🙂The challenge is determining whether the speedup is real, and that requires a more sophisticated approach than just running the benchmarks once.
Reacted by Sergey MiryanovEvery tuple is created, but not every tuple is seen by the GC; many are dealloc before GC gets to see them.
For those tuples this adds overhead for no gain.Yeah, I agree. After careful consideration, I think that most of tuples die before garbage collection.
Do you any numbers for the ratio of tuples that reach GC for general applications?
I have plans to collect total count of tuples that allocated and deallocated before GC.
Even if that ratio is high, why is it cheaper to check for tracking during construction than to check during GC?
Also, I plan to add microbenchmarks to measure impact of this micro-optimisation on PyTuple_* methods. My main job takes too much time :)
sergey-miryanov commented
on Oct 19, 2025 ContributorAuthorMore actionsHere are some numbers for tracked and untracked tuples in GC. I previously collected the data and just processed it now (data for
pyperformance, code for counting total untracked tuples - ff81eb2, for tracked tuples - a79a283).Below is a plot of the total number of created, tracked, and untracked tuples (log scale):
As can be seen, the total number of tracked tuples is about 10-30% of the total number of created tuples. Total tracked tuples are tuples that have been seen by GC. Untracked tuples are those that were untracked by
untrack_tuples.The percentage of the tracked and untracked tuples to the total number:
About 2% of the total number of created tuples are kept in GC:
Feature or enhancement
Proposal:
When we use
PyTuple_Packall objects already well constructed. If we know that they immutable we can skip tracking it in GC, because GC will untrack them eventually.I have a PR ready and benchmark results:
Geometric mean: 1.01x faster (Win11 x64, 11th Gen Intel(R) Core(TM) i5-11600K @ 3.90GHz, 48d0d0d)
All benchmarks:
Benchmark hidden because not significant (20): 2to3, chaos, deepcopy_reduce, genshi_xml, html5lib, json_loads, nqueens, pathlib, pickle, pickle_dict, pickle_list, pidigits, regex_dna, sqlglot_normalize, sqlglot_parse, sqlglot_transpile, sqlite_synth, sympy_integrate, unpickle_list, xml_etree_generate
It doesn't hurt performance, but can decrease number of objects in GC to check and untrack.
Has this already been discussed elsewhere?
This is a minor feature, which does not need previous discussion elsewhere
Links to previous discussion of this feature:
No response
Linked PRs