Repository navigation
Leaked semaphore objects in test_concurrent_futures #104090
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on May 2, 2023 The leakage is from
FailingInitializerMixin. Don't know how to fix though.Two failing tests:
test.test_concurrent_futures.ProcessPoolForkserverFailingInitializerTest.test_initializertest.test_concurrent_futures.ProcessPoolSpawnFailingInitializerTest.test_initializer
- Found out using
git bisectthat the first failed commit is 6883007, bpo-4080: unittest durations #12271 - I splited the commit to multiple new commits each per file, so I can use
git bisectagain to find the specific problematic file - The file is
unittest/result.pyand this is the git diff
diff --git a/Lib/unittest/result.py b/Lib/unittest/result.py index 5ca4c23238..fa9bea47c8 100644 --- a/Lib/unittest/result.py +++ b/Lib/unittest/result.py @@ -43,6 +43,7 @@ def __init__(self, stream=None, descriptions=None, verbosity=None): self.skipped = [] self.expectedFailures = [] self.unexpectedSuccesses = [] + self.collectedDurations = [] self.shouldStop = False self.buffer = False self.tb_locals = False @@ -157,6 +158,12 @@ def addUnexpectedSuccess(self, test): """Called when a test was expected to fail, but succeed.""" self.unexpectedSuccesses.append(test) + def addDuration(self, test, elapsed): + """Called when a test finished to run, regardless of its outcome.""" + # support for a TextTestRunner using an old TestResult class + if hasattr(self, "collectedDurations"): + self.collectedDurations.append((test, elapsed)) + def wasSuccessful(self): """Tells whether or not this result was a success.""" # The hasattr check is for test_result's OldResult test. That- Indeed when I put a return before
self.collectedDurations.append((test, elapsed)), all is OK
$ ./python -m unittest test.test_concurrent_futures.ProcessPoolSpawnFailingInitializerTest.test_initializer 2>&1 . ---------------------------------------------------------------------- Ran 1 test in 0.398s OK- I found out with some debugging and print that on the failed case (
collectedDurations.append) theresource_trackeris called before thesem_unlinkwhich will try to unlink them later and fail, since the resource tracker already did it. Only 3 semaphore objects are failing here, trying to unlink afteraddDurationfunction.
case.py(638)run() -> self._callTearDown() -> run sem_unlink(name) on /mp-4lu9pdgs case.py(638)run() -> self._callTearDown() -> run sem_unlink(name) on /mp-61ynpjfw case.py(638)run() -> self._callTearDown() -> run sem_unlink(name) on /mp-wrenmqfa case.py(638)run() -> self._callTearDown() -> run sem_unlink(name) on /mp-e01eptc5 case.py(638)run() -> self._callTearDown() -> run sem_unlink(name) on /mp-sssjxme9 0.27s run addDuration Lib/multiprocessing/resource_tracker.py:236: UserWarning: resource_tracker: There appear to be 3 leaked semaphore objects to clean up at shutdown call _CLEANUP_FUNCS semaphore function for /mp-u8k6h3v7 file call _CLEANUP_FUNCS semaphore function for /mp-ihqrt4f2 file call _CLEANUP_FUNCS semaphore function for /mp-n4h31gcc file suite.py(132)run() -> self._handleModuleTearDown(result) -> run sem_unlink(name) on /mp-n4h31gcc FileNotFoundError: [Errno 2] No such file or directory suite.py(132)run() -> self._handleModuleTearDown(result) -> run sem_unlink(name) on /mp-u8k6h3v7 FileNotFoundError: [Errno 2] No such file or directory suite.py(132)run() -> self._handleModuleTearDown(result) -> run sem_unlink(name) on /mp-ihqrt4f2 FileNotFoundError: [Errno 2] No such file or directory(reduced the output so it will be clearer)
5. In other hand, when I ignore this line, thesem_ulinkfor those 3 semaphore objects are called suceesfully after theaddDurationfunction but from another tracecase.py(638)run() -> self._callTearDown() -> run sem_unlink(name) on /mp-ch65p5qm case.py(638)run() -> self._callTearDown() -> run sem_unlink(name) on /mp-z2m0y_8x case.py(638)run() -> self._callTearDown() -> run sem_unlink(name) on /mp-mi_764q8 case.py(638)run() -> self._callTearDown() -> run sem_unlink(name) on /mp-ntvllres case.py(638)run() -> self._callTearDown() -> run sem_unlink(name) on /mp-f5553813 0.27s run addDuration .suite.py(122)test(result) -> suite.py(84)return self.run(*args, **kwds) -> run sem_unlink(name) on /mp-jvota4jy suite.py(122)test(result) -> suite.py(84)return self.run(*args, **kwds) -> run sem_unlink(name) on /mp-hxf6fzql suite.py(122)test(result) -> suite.py(84)return self.run(*args, **kwds) -> run sem_unlink(name) on /mp-fusub384- I'm trying now to understand what caused the
resource_trackerto trigger befure thesem_unlink. For now I think that for some reason thesem_unlinkfromTestSuiteline 122 is not reached, and thefore the resouce tracker is trying to release the resources. And later the call tosem_unlinkfromTestSuiteline 132 is trying to run and fails

- The key is to understand why when the
self.collectedDurationsis not empty, it blocks us from cleaning as usual those 3 semaphore objects
Reacted by shailshouryya- Found out using
Ok, found more -
The leaked semaphore objects are fromFailingInitializerMixin(as @sunmy2019 mentioned before, but now it's clearer for me).This line creates a new queue with 3 semaphore related and when the
collectedDurationsis not empty theself.log_queueis not released (still in used I guess)################ FailingInitializerMixin.setUp: Calling self.mp_context.Queue() Queue __init__: Set self._rlock = ctx.Lock() Add semlock to register and finalize self._semlock.name='/mp-lnmmngl0', self.__class__.__name__='Lock' Queue __init__: Set self._wlock = ctx.Lock() Add semlock to register and finalize self._semlock.name='/mp-5fhxbfsv', self.__class__.__name__='Lock' Queue __init__: Set ctx.BoundedSemaphore(maxsize) Add semlock to register and finalize self._semlock.name='/mp-8mzh94sd', self.__class__.__name__='BoundedSemaphore' ################Still investigating
For some additional context, I also ran
git bisectand landed on the same commit identified above. I did not get as far as the post above me did in terms of identifying the exact reasons why the 6 leaked semaphore objects were not cleaned up before shutdown, but I also narrowed down the problem and found the6 leaked semaphore objectswarning goes away when the following block (also mentioned above):def addDuration(self, test, elapsed): """Called when a test finished to run, regardless of its outcome.""" # support for a TextTestRunner using an old TestResult class if hasattr(self, "collectedDurations"): self.collectedDurations.append((test, elapsed))is completely commented out:
# def addDuration(self, test, elapsed): # """Called when a test finished to run, regardless of its outcome.""" # # support for a TextTestRunner using an old TestResult class # if hasattr(self, "collectedDurations"): # self.collectedDurations.append((test, elapsed))Another thing that I found interesting was the leaked semaphore objects look like
TemporaryFiles:/mp-ch65p5qm /mp-z2m0y_8x /mp-mi_764q8 /mp-ntvllres /mp-f5553813However, I'm not sure about this, since it does not look like
TemporaryFileis called anywhere intest_concurrent_futures.pyexcept forbut this does not look relevant for the problem here (but I'm adding this observation here in case it is relevant).cpython/Lib/test/test_concurrent_futures.py
Lines 1248 to 1249 in 89867d2
from tempfile import TemporaryFile with TemporaryFile(mode="w+") as f: Not related to TemporaryFile, those are virtual files used by each semaphore object as the system locking mechanisem
Accessing named semaphores via the file system
On Linux, named semaphores are created in a virtual file system, normally mounted under /dev/shmSource: https://linux.die.net/man/7/sem_overview
For example, if you stop the process in the middle
you will be able to see those virtual files here
which will be deleted on
_multiprocessing.sem_unlinkmethodReacted by shailshouryyaOk, found how to fix the issue! 🚀
Full details
The current code calls
self.collectedDurations.append((test, elapsed))onaddDurationmethod (as wrote in previous comment).This causes the 3 sempahors of
FailingInitializerMixin.log_queueto not being released before theresource trackeris stopped and clear by itself all left resources.Example 1 (without
.collectedDurations.append), theResourceTrackeris stopped after we clear the 3 semaphors -Example 2 (current code), the
ResourceTrackeris stopped before we clear the 3 semaphors -This happens since we call
_run_finalizers()afterresource_tracker._resource_tracker._stop()

So, optional solution No.1 will be to change the order of the calls, like this -
And it solves the issue perfectly

@vstinner -
I saw that you were the one who wrote this part,
Can we do such change?
If yes, I'll make a PR for it.@iritkatriel -
I'll happy to here what you think here, thanks.Still open question for me, is why the objects are not finalize normally (by the gc I assume) like it was before.
I'm still investigating on thisReacted by sunmy2019Hey,
Found better solution and also what's happening here.
The reason those 3 semaphores were still alive and not deleted, it's because there weren't garabage collected yet, since they have a reference.
How?
Those 3 semaphors are related to
FailingInitializerMixin.log_queueobject (they are created as part of this queue).This mixin is used to create the test class
ProcessPoolSpawnFailingInitializerTestwhich was used on the test result private objectcollectedDurations, and since the result object continue to live after the test object ended, the test object can't be collected and therefore the semaphors too...Why we need the test object there?
For displaying it on
unittest.runner._printDurationsmethod.

So how to fix it?
Very simple, instead of passing the test object, we just need here the test repr only!
if hasattr(self, "collectedDurations"): self.collectedDurations.append((str(test), elapsed))
And this works perfect! 🚀
I'll create a PR for this
@iritkatrielReacted by Irit KatrielGood find!
- added 4 commits that reference this issue
on Jul 15, 2023 16 remaining items
That PR was a helper to make failures like this easier to diagnose.
The issue itself is not fixed; it still appears in refleak tests (./python -m test test_concurrent_futures -R:).I get the same error in my case. I'm doing distributed training using spawn multiprocessing and DDP, also at the end of running my application I get a specific number of leaked semaphores.
This is the link of the issue.
I'm wondering if this issue is solved or not.
- added3.12only security fixesonly security fixes3.13only security fixesonly security fixes3.14bugs and security fixesbugs and security fixesand removedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Mar 22, 2025 - addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Mar 23, 2025 - added3.15bugs and security fixesbugs and security fixesand removed3.12only security fixesonly security fixes
on Aug 19, 2025
Metadata
Metadata
Assignees
Labels
Projects
- StatusShow more project fieldsNo status






Bug report
Configuration:
Your environment
Linked PRs