Skip to content

gh-158790: Improve performance for list.insert() and del list[i] free-threaded builds - #158791

Open
hetaozdh wants to merge 2 commits into
python:mainfrom
hetaozdh:list-shift-memmove
Open

hetaozdh wants to merge 2 commits into
python:mainfrom
hetaozdh:list-shift-memmove

Conversation

@hetaozdh

@hetaozdh hetaozdh commented Oct 4, 2026 •

Copy link
Copy Markdown

Functions ins1()(used by list.insert()) and list_ass_item_lock_held() shift the list with one atomic release store per element in free-threaded builds. Those atomic stores prevent the compiler from vectorizing the loops, making them much slower than memmove() for large shifts. I propose using the existing ptr_wise_atomic_memmove() helper here. It uses memmove() as a fast path when the list is owned by the current thread and is not shared, and keeps the atomic element-wise stores for lists that other threads can observe.

The default (GIL) build keeps the current loops because the compiler can optimize them directly in the calling function, and they perform better in benchmarks.

pyperf, free-threaded release build (--disable-gil, -O3):

del l[0]                 2.26x faster
del l[len(l)//2]         2.14x faster
l.insert(0, x)           1.95x faster
l.insert(len(l)//2, x)   1.86x faster
insert(0)+del[0], n=1000 2.86x faster
sliding window, n=1000   2.46x faster

Geometric mean           2.23x faster

Compiling both versions in GIL mode produces the same code for ins1() and list_ass_item_lock_held(), so the default build is not affected.

The benchmark following is produced by AI and verified by me.

#!/usr/bin/env python3
"""pyperf benchmark for list.insert() / del list[i] element shifting.

Usgae:
    PYTHONPATH=... <interpreter> bench_list_shift.py -o result.json -p 1 -n 10
    python -m pyperf compare_to base.json patched.json --table --table-format md

Each benchmark function performs one workload iteration; pyperf calibrates
the number of loops per value.  Results are compared with
``pyperf compare_to`` in the accompanying report.
"""

import pyperf


def bench_del_front(n=100_000, k=2_000):
    lst = list(range(n))
    for _ in range(k):
        del lst[0]


def bench_del_mid(n=100_000, k=2_000):
    lst = list(range(n))
    for _ in range(k):
        del lst[len(lst) // 2]


def bench_insert_front(n=100_000, k=2_000):
    lst = list(range(n))
    for _ in range(k):
        lst.insert(0, None)


def bench_insert_mid(n=100_000, k=2_000):
    lst = list(range(n))
    for _ in range(k):
        lst.insert(len(lst) // 2, None)


def _small(n, k):
    # insert(0) and del[0] both shift ~n elements; the size stays n
    lst = list(range(n))
    for _ in range(k):
        lst.insert(0, None)
        del lst[0]


def bench_small_n8(n=8, k=200_000):
    _small(n, k)


def bench_small_n16(n=16, k=200_000):
    _small(n, k)


def bench_small_n32(n=32, k=200_000):
    _small(n, k)


def bench_small_n1000(n=1000, k=20_000):
    _small(n, k)


def bench_sliding_window(n=1000, k=50_000):
    # deque-like pattern: push to the front, pop from the back
    lst = list(range(n))
    for i in range(k):
        lst.insert(0, i)
        lst.pop()


def bench_control_del_last(n=100_000):
    # no shift: exercise the unchanged "delete last element" guard path
    lst = list(range(n))
    for _ in range(n):
        del lst[-1]


def bench_control_append(n=1_000_000):
    # no shift: unrelated fast path, should stay flat
    lst = []
    append = lst.append
    for i in range(n):
        append(i)


BENCHMARKS = [
    ("del_front_n100k_k2k", bench_del_front),
    ("del_mid_n100k_k2k", bench_del_mid),
    ("insert_front_n100k_k2k", bench_insert_front),
    ("insert_mid_n100k_k2k", bench_insert_mid),
    ("small_n8_k200k", bench_small_n8),
    ("small_n16_k200k", bench_small_n16),
    ("small_n32_k200k", bench_small_n32),
    ("small_n1000_k20k", bench_small_n1000),
    ("sliding_window_n1000_k50k", bench_sliding_window),
    ("control_del_last_n100k", bench_control_del_last),
    ("control_append_n1M", bench_control_append),
]


if __name__ == "__main__":
    runner = pyperf.Runner()
    for name, func in BENCHMARKS:
        runner.bench_func(name, func)

…] in free-threaded builds

`ins1()`(used by `list.insert()`) and `list_ass_item_lock_held()` shift the list with one atomic release store per element in free-threaded builds.  Those atomic stores prevent the compiler from vectorizing the loops, making them much slower than `memmove()` for large shifts.  I propose using the existing `ptr_wise_atomic_memmove()` helper here. It uses `memmove()` as a fast path when the list is owned by the current thread and is not shared, and keeps the atomic element-wise stores for lists that other threads can observe.

The default (GIL) build keeps the current loops because the compiler can optimize them directly in the calling function, and they perform better in benchmarks.

pyperf, free-threaded release build (--disable-gil, -O3):

    del l[0]                 2.26x faster
    del l[len(l)//2]         2.14x faster
    l.insert(0, x)           1.95x faster
    l.insert(len(l)//2, x)   1.86x faster
    insert(0)+del[0], n=1000 2.86x faster
    sliding window, n=1000   2.46x faster

Compiling both versions in GIL mode produces the same code for `ins1()` and `list_ass_item_lock_held()`, so the default build is not affected.

This continues pythongh-129069, which introduced `ptr_wise_atomic_memmove()`.
@corona10

corona10 commented Oct 4, 2026

Copy link
Copy Markdown
Member

You are now making the new fragmented logic for the free-threaded build. How worth to maintain this? How many performance portion from this logic in general workload.

Comment thread Misc/NEWS.d/next/Core_and_Builtins/2026-10-04-01-40-31.gh-issue-158790.F3kQ7z.rst Outdated
Comment thread Objects/listobject.c Outdated
@picnixz

picnixz commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

You are now making the new fragmented logic for the free-threaded build. How worth to maintain this? How many performance portion from this logic in general workload.

We could add a macro for that pattern, or a small function which does it but i'm not sure the compiler would inline the call AND the loop anymore.

@corona10

corona10 commented Oct 4, 2026

Copy link
Copy Markdown
Member

We could add a macro for that pattern, or a small function which does it but i'm not sure the compiler would inline the call AND the loop anymore.

Yeah in that case, would be fine, but I prefer to not adding free-threading only logic as possible unless it's really worth to do.

Use Sphinx roles in the NEWS entry (:meth:`list.insert`,
:keyword:`del`, :manpage:`memmove(3)`) and drop a comment in ins1()
that repeated the rationale already given in the commit message.
@hetaozdh

hetaozdh commented Oct 4, 2026

Copy link
Copy Markdown
Author

@corona10 I agree that the branching does introduce some maintenance cost, but I think it's worths it.

The point is that memmove() is already a better implementation for large shifts in general. In the GIL build, replacing the existing loop with memmove() makes these benchmarks only about 10% slower overall (geometric mean). In the free-threaded build, it makes them 2.23x faster overall.

I think the best trade-off is to keep the existing compiler-optimized loop for the GIL build, while using ptr_wise_atomic_memmove() for the free-threaded build. The helper also retains the atomic element-wise path for shared lists and uses memmove() only when the list is owned by the current thread and is not shared.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

awaiting review type-feature A feature request or enhancement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants