"""dict_dict_merge() cases. ctl rows do not run the per-item merge loop:
dense copies take the key-table clone fast path."""
import os
import pyperf
runner = pyperf.Runner()
def S(n): return "{'key%%d' %% i: i for i in range(%d)}" % n
CASES = [
("ctl: a + b", "a + b", "a = 1000; b = 2000"),
("ctl: dct_100[key]", "d[k]", "d = " + S(100) + "; k = 'key50'"),
("ctl: dict(dct_5) (clone path)", "dict(d)", "d = dict.fromkeys('abcde', 1)"),
("ctl: dict(dct_100) (clone path)", "dict(d)", "d = " + S(100)),
("dct.update(dct_5)", "c.update(b)", "b = dict.fromkeys('abcde', 1); c = dict(b)"),
("dct.update(dct_100)", "c.update(b)", "b = " + S(100) + "; c = dict(b)"),
("dct.update(dct_1000)", "c.update(b)", "b = " + S(1000) + "; c = dict(b)"),
("dct.update(dct_100) (int keys)", "c.update(b)", "b = {i: i for i in range(1000, 1100)}; c = dict(b)"),
("dct_100 | other_100", "a | b", "a = " + S(100) + "; b = {'other%d' % i: i for i in range(100)}"),
("{**dct_5, **other_5}", "{**a, **b}", "a = dict.fromkeys('abcde', 1); b = dict.fromkeys('fghij', 1)"),
("sparse_1000.copy() (500 deleted)", "d.copy()",
"d = " + S(1000) + "\nfor i in range(0, 1000, 2): del d['key%d' % i]"),
("dict(vars(obj)) (4 attributes)", "dict(v)",
"class C:\n def __init__(self): self.a = 1; self.b = 2; self.c = 3; self.d = 4\nv = vars(C())"),
]
for name, stmt, setup in CASES:
runner.timeit(name, stmt, setup=setup)
Feature or enhancement
The per-item loop in
dict_dict_merge()takes an extra reference tokeyandvaluebefore dispatching onoverride. On theoverride == 1pathinsertdict()already receives its own references throughPy_NewRef()andnothing uses
keyorvalueafterwards, so the outer pair is wasted refcount operations per item (atomic ones in the free-threaded build). The pair is still needed on the other path, where_PyDict_Contains_KnownHash()can run__eq__and*dupkeyhands the reference to the caller.override == 1coversdict.update(d),d1 \| d2,{**a, **b}, any merge into an empty dict, anddict(d)/d.copy()when the key-table clone fast path does not apply (sparse, split-table or subclass source).Benchmarks
ctl: a + bctl: dct_100[key]ctl: dict(dct_100) (clone path)dct.update(dct_5)dct.update(dct_100)dct.update(dct_1000)dct.update(dct_100) (int keys)dct_100 | other_100{**dct_5, **other_5}sparse_1000.copy() (500 deleted)dict(vars(obj)) (4 attributes)Free-threaded build shows a slightly larger gain.
Benchmark script
Generated with help from Claude Code
Has this already been discussed elsewhere?
This is a minor feature, which does not need previous discussion elsewhere
Links to previous discussion of this feature:
No response
Linked PRs