gh-158677: Fix Pool.imap(buffersize=...) blocking other tasks on the pool - #158792
oyiakoumis wants to merge 1 commit into
Conversation
…n the pool With buffersize, the task handler thread waited on a semaphore when the buffer of an imap() or imap_unordered() iterator was full. The pool has only one such thread, so no other task was submitted until the iterator was consumed, which could deadlock. The task handler no longer waits. When the buffer is full, it sets the iterator's tasks aside and moves on. They are put back on the task queue when a result is consumed.
|
Most changes to Python require a NEWS entry. Add one using the blurb_it web app or the blurb command-line tool. If this change has little impact on Python users, wait for a maintainer to apply the |
|
This is breaking Linux and Windows CI.
In the issue you said:
But you did not wait for anything. Really, don't use automated agents for making fixes if they do not leave the time for maintainers to even reply. |
|
Sorry @picnixz. I opened this PR as a draft to collaborate with @PrakharAgarwal17 on a potential patch, and I didn't realize that a draft would still create noise. I am closing it for now and will continue on the issue. |
Problem
With
buffersize,imap()andimap_unordered()throttle task submission by blocking on a semaphore inside the pool's task handler thread. The pool has only one such thread, so while it waits for the iterator to be consumed, no other task on the pool is submitted. This deadlocks, for example,zip()over two bufferedimap()iterators.Fix
The task handler no longer waits. When the buffer of an iterator is full, it sets that iterator's tasks aside and moves on to the next job. When a result is consumed,
next()puts them back on the task queue.The input iterable is still consumed by the task handler thread, so nothing else changes:
imap()returns immediately andnext()never waits for the iterable.close()is unchanged: a partially consumed iterator still stops early. That is gh-158675 and is left out of this PR.Alternative considered
The issue suggested submitting the next task from
next(), asExecutor.map(buffersize=...)does. I tried that approach first but moved away from it because the iterable would then be consumed in the caller's thread. As a result,next()could hold a ready result while waiting for the next input item, causing a deadlock if producing that item depends on the result being handled.Tests
apply_async(), a second bufferedimap()) run while a buffered iterator is waiting to be consumed. This test fails without the fix.next(). This case was not previously covered withbuffersize.Tested on macOS. Not tested on Linux, Windows or the free-threaded build.
No NEWS entry, since
buffersizeis new in 3.16 and not released yet.I used Claude Code (Claude Opus 5.5, effort: Medium) to help investigate the bug and develop the fix. I reviewed, understood, modified, and tested the resulting changes myself.
Pool.imap(buffersize=...)blocks all other tasks on the pool until the iterator is consumed #158677