Repository navigation
Explore using PyBytesWriter API for compression libraries output buffers #139877
Description
Activity
- addedtype-featureA feature request or enhancementA feature request or enhancementperformancePerformance or resource usagePerformance or resource usageextension-modulesC modules in the Modules dirC modules in the Modules dir3.15bugs and security fixesbugs and security fixes
on Oct 10, 2025 _BlocksOutputBuffer_Finishseems to be doing a"".join()onself->list. "stringlib" has a fast/fairly tuned implementation of that, wondering if the end join could just become something like:PyObject *new_buffer = PyObject_CallMethodOneArg( Py_GetConstant(Py_CONSTANT_EMPTY_BYTES), &_Py_ID(join), self->list);
Reacted by msmojtabafarGood idea, I will also test that. I expect
PyBytesWriterwill still be faster because of locality, but I can benchmark both.Reacted by Cody MaloneyI have more data, looks like not all of the (de)compressors rely on the output buffer code as much as zstd. That makes sense since the actual compression code is slower so the output buffer allocations are relatively less important.
I've started with zlib as that is the most used format most likely. It's also the (de)compressor for wheels so performance gains there have huge multipliers :)
My system:
AMD Ryzen 9 9950X
6000 MT/s DDR5 RAM
Running Fedora 42Here's
mainvs using"".join():zlib.compress(1M): Mean +- std dev: [main] 13.5 ms +- 0.1 ms -> [bytes_join] 13.5 ms +- 0.1 ms: 1.00x slower zlib.compress(1G): Mean +- std dev: [main] 11.4 sec +- 0.0 sec -> [bytes_join] 11.4 sec +- 0.0 sec: 1.00x slower zlib.decompress(1K): Mean +- std dev: [main] 1.42 us +- 0.01 us -> [bytes_join] 1.44 us +- 0.01 us: 1.01x slower zlib.decompress(1M): Mean +- std dev: [main] 1.29 ms +- 0.00 ms -> [bytes_join] 1.20 ms +- 0.00 ms: 1.07x faster zlib.decompress(1G): Mean +- std dev: [main] 1.36 sec +- 0.00 sec -> [bytes_join] 1.39 sec +- 0.01 sec: 1.02x slower Benchmark hidden because not significant (1): zlib.compress(1K) Geometric mean: 1.00x fasterNote that I had to add a
_PyBytes_Resizeimmediately after thejoin()because the output bytes object needs to be truncated to the length of the output data.Here's comparing
mainvsPyBytesWriter:zlib.compress(1M): Mean +- std dev: [main] 13.5 ms +- 0.1 ms -> [pybyteswriter] 13.4 ms +- 0.0 ms: 1.00x faster zlib.compress(1G): Mean +- std dev: [main] 11.4 sec +- 0.0 sec -> [pybyteswriter] 11.3 sec +- 0.0 sec: 1.00x faster zlib.decompress(1K): Mean +- std dev: [main] 1.42 us +- 0.01 us -> [pybyteswriter] 1.39 us +- 0.01 us: 1.02x faster zlib.decompress(1M): Mean +- std dev: [main] 1.29 ms +- 0.00 ms -> [pybyteswriter] 1.17 ms +- 0.00 ms: 1.10x faster zlib.decompress(1G): Mean +- std dev: [main] 1.36 sec +- 0.00 sec -> [pybyteswriter] 1.17 sec +- 0.00 sec: 1.17x faster Benchmark hidden because not significant (1): zlib.compress(1K) Geometric mean: 1.05x faster5% overall is not bad! 10% faster decompression for data >1MB is also great, that will be a big win for wheels and many other use cases I imagine.
I assume the numbers for xz and bz2 are pretty similar or a more minor win, since those (de)compressors are as slow or slower than zlib.
I'm going to collect numbers for zstd and verify my initial numbers next, then work on a patch.
Reacted by Cody MaloneyOkay! As expected, the improvements are a lot better for zstd (using default compression level 3):
zstd.compress(1K): Mean +- std dev: [main_zstd_3] 3.01 us +- 0.03 us -> [pybyteswriter_zstd_3] 3.00 us +- 0.03 us: 1.01x faster zstd.compress(1M): Mean +- std dev: [main_zstd_3] 2.92 ms +- 0.02 ms -> [pybyteswriter_zstd_3] 2.89 ms +- 0.02 ms: 1.01x faster zstd.compress(1G): Mean +- std dev: [main_zstd_3] 2.72 sec +- 0.01 sec -> [pybyteswriter_zstd_3] 2.67 sec +- 0.01 sec: 1.02x faster zstd.decompress(1K): Mean +- std dev: [main_zstd_3] 1.40 us +- 0.01 us -> [pybyteswriter_zstd_3] 1.38 us +- 0.01 us: 1.01x faster zstd.decompress(1M): Mean +- std dev: [main_zstd_3] 734 us +- 4 us -> [pybyteswriter_zstd_3] 546 us +- 3 us: 1.34x faster zstd.decompress(1G): Mean +- std dev: [main_zstd_3] 790 ms +- 4 ms -> [pybyteswriter_zstd_3] 634 ms +- 3 ms: 1.25x faster Geometric mean: 1.10x fasterReacted by Cody MaloneyCan we close this issue?
Yep!
Reacted by Victor Stinner
Metadata
Metadata
Assignees
Labels
Projects
- StatusShow more project fieldsDone
Feature or enhancement
The new
PyBytesWriter()API is fast and easy to use. I expect it will bring a nice improvement both to maintainability and speed for compression output buffer management.I have some perf recordings showing that a large portion (>50%!) of time in decompression for a mix of data sizes (1K, 1M, 1G) is in
_BlocksOutputBuffer_Finish, re-assembling the output buffer.I also made a very hacky modification to pycore_blocks_output_buffer.h to use
PyBytesWriter()and found it greatly sped up decompression time:The below two tests are operating on compressed enwiki content with zstd compression.
Those are 25-30% speedups!
I think this is enough to motivate a refactor of this code to use
PyBytesWriter()and benchmark against the current implementation across compression modules and data sizes.cc @vstinner for viz
Linked PRs