Repository navigation
Exceptions slow in 3.11, depending on location #109181
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or error
on Sep 9, 2023 - addedinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)performancePerformance or resource usagePerformance or resource usage
on Sep 9, 2023 They use an O(n) implementation internally.
Just for the sake of completeness, the length of unreachable code without a raise statement does not influence the computation time.
Benchmark script
from time import perf_counter as time from timeit import repeat for e in range(2, 6): n = 10 ** e exec(f'''def f(): if 0 == 1: {'unreached;' * n} return''') number = 10**6 // n t = min(repeat(f, number=number)) / number print(f'{n:6} {t * 1e6 :7.1f} μs')By comparing the performance record between the two we can see that Addr2Line is the culprit
And indeed, when commenting out the corresponding loops I see a reduction in the runtime
Same test script as initally provided, now with commented out loop
100 0.2 μs 1000 0.2 μs 10000 0.2 μs 100000 0.2 μs 1000000 0.3 μsvs
100 0.6 μs 1000 4.7 μs 10000 46.0 μs 100000 495.2 μs 1000000 5049.9 μsMore specifically, it seems most time is spent looking for 0 bytes here. This is about where my current understanding of what this code even does ends :)
- added3.11only security fixesonly security fixes3.12only security fixesonly security fixes3.13only security fixesonly security fixes
on Sep 11, 2023 @AlexWaygood I am also experiencing this slowdown in python 3.10.12 (since you added 3.11, 3.12 and 3.13)
@AlexWaygood I am also experiencing this slowdown in python 3.10.12 (since you added 3.11, 3.12 and 3.13)
Python 3.10 is old enough that it is now only accepting bugfixes if they relate to security issues, and this isn't a security issue. For us, therefore, this is only an issue for Python 3.11+, even if it can be reproduced on Python 3.10.
Reacted by Niels Mündler-Sasahara and sunmy2019@pablogsal and thoughts on this?
@pablogsal and thoughts on this?
I am currently on sick leave so I cannot look at this in detail but I feel this is the balance between the compression to save disk size and the speed to "uncompress" the line table.
On the other hand maybe we could defer the computation of the line table offsets only when exceptions buble up to top level, but it may be too late at that point so we may need to save extra information and maybe even the code object if is not reachable already in all cases.
In any case, I don't think this is a bug. It may qualify as a regression if we all agree but I am not even sure if I would call it that, since is just a consequence of the feature.
I agree it's not a bug. Extra cost for new features improving error messages is fine.
Someone may want a switch to disable this feature, but I would argue that you should rather keep the
try ... except ...block as short as possible.Reacted by Erlend E. AaslandI agree it's not a bug. Extra cost for new features improving error messages is fine.
Someone may want a switch to disable this feature, but I would argue that you should rather keep the
try ... except ...block as short as possible.I think we can make that work with
-X no_debug_rangesorPYTHONNODEBUGRANGES8 remaining items
- added 4 commits that reference this issue
on Oct 31, 2023 Backporting to 3.12 and 3.11 as technically this is a regression
These changes introduced references leaks in 3.11 and 3.12 branches: see issue gh-111947.
Reacted by Gregory P. Smith- added 2 commits that reference this issue
on Aug 8, 2024 - added a commit that references this issue
on Aug 12, 2024


Bug report
Bug description:
(From Discourse)
Consider these two functions:
The only difference is that
long()has 100 unreached statements instead of just one. But it takes much longer in Python 3.11 (and a bit longer in Python 3.10). Times from @jamestwebber (here):Why? Shouldn't it just jump over them all and be just as fast as
short()?Benchmark script
Attempt This Online!):
In fact it takes time linear in how many unreached statements there are. Times for 100 to 100000 unreached statements (on one line, before the
try):Benchmark script
Attempt This Online!
The slowness happens when the unreached statements are anywhere before the
raise, and not when they're anywhere after theraise(demo). So it seems what matters is location of theraisein the function. Long code before it somehow makes it slow.This has a noticeable impact on real code I wrote (assuming I pinpointed the issue correctly): two solutions for a task, and one was oddly slower (~760 vs ~660 ns) despite executing the exact same sequence of bytecode operations. Just one jump length differed, leading to a
raiseat a larger address.Benchmark script with those two solutions and the relevant test case:
The functions shall return the one item from the iterable, or raise an exception if there are fewer or more than one. Testing with an empty iterable, both get the iterator, iterate it (nothing, since it's empty), then raise. The relevant difference appears to be that the slower one has the
raisewritten at the bottom, whereas the faster one has it near the top.Sample times:
Code:
Attempt This Online!
CPython versions tested on:
3.10, 3.11
Operating systems tested on:
Linux, macOS
Linked PRs