Conversation
…] in free-threaded builds
`ins1()`(used by `list.insert()`) and `list_ass_item_lock_held()` shift the list with one atomic release store per element in free-threaded builds. Those atomic stores prevent the compiler from vectorizing the loops, making them much slower than `memmove()` for large shifts. I propose using the existing `ptr_wise_atomic_memmove()` helper here. It uses `memmove()` as a fast path when the list is owned by the current thread and is not shared, and keeps the atomic element-wise stores for lists that other threads can observe.
The default (GIL) build keeps the current loops because the compiler can optimize them directly in the calling function, and they perform better in benchmarks.
pyperf, free-threaded release build (--disable-gil, -O3):
del l[0] 2.26x faster
del l[len(l)//2] 2.14x faster
l.insert(0, x) 1.95x faster
l.insert(len(l)//2, x) 1.86x faster
insert(0)+del[0], n=1000 2.86x faster
sliding window, n=1000 2.46x faster
Compiling both versions in GIL mode produces the same code for `ins1()` and `list_ass_item_lock_held()`, so the default build is not affected.
This continues pythongh-129069, which introduced `ptr_wise_atomic_memmove()`.
|
You are now making the new fragmented logic for the free-threaded build. How worth to maintain this? How many performance portion from this logic in general workload. |
We could add a macro for that pattern, or a small function which does it but i'm not sure the compiler would inline the call AND the loop anymore. |
Yeah in that case, would be fine, but I prefer to not adding free-threading only logic as possible unless it's really worth to do. |
Use Sphinx roles in the NEWS entry (:meth:`list.insert`, :keyword:`del`, :manpage:`memmove(3)`) and drop a comment in ins1() that repeated the rationale already given in the commit message.
|
@corona10 I agree that the branching does introduce some maintenance cost, but I think it's worths it. The point is that I think the best trade-off is to keep the existing compiler-optimized loop for the GIL build, while using |
Functions
ins1()(used bylist.insert()) andlist_ass_item_lock_held()shift the list with one atomic release store per element in free-threaded builds. Those atomic stores prevent the compiler from vectorizing the loops, making them much slower thanmemmove()for large shifts. I propose using the existingptr_wise_atomic_memmove()helper here. It usesmemmove()as a fast path when the list is owned by the current thread and is not shared, and keeps the atomic element-wise stores for lists that other threads can observe.The default (GIL) build keeps the current loops because the compiler can optimize them directly in the calling function, and they perform better in benchmarks.
pyperf, free-threaded release build (--disable-gil, -O3):
Compiling both versions in GIL mode produces the same code for
ins1()andlist_ass_item_lock_held(), so the default build is not affected.The benchmark following is produced by AI and verified by me.