Skip to content

Explore using PyBytesWriter API for compression libraries output buffers #139877

Description

@emmatyping

Feature or enhancement

The new PyBytesWriter() API is fast and easy to use. I expect it will bring a nice improvement both to maintainability and speed for compression output buffer management.

I have some perf recordings showing that a large portion (>50%!) of time in decompression for a mix of data sizes (1K, 1M, 1G) is in _BlocksOutputBuffer_Finish, re-assembling the output buffer.

I also made a very hacky modification to pycore_blocks_output_buffer.h to use PyBytesWriter() and found it greatly sped up decompression time:

The below two tests are operating on compressed enwiki content with zstd compression.

test main PyBytesWriter()
decompress 1M 2.15ms 1.65ms
decompress 1G 2.2s 1.73s

Those are 25-30% speedups!

I think this is enough to motivate a refactor of this code to use PyBytesWriter() and benchmark against the current implementation across compression modules and data sizes.

cc @vstinner for viz

Linked PRs

Activity

  1. self-assigned this
    on Oct 10, 2025
  2. added
    type-featureA feature request or enhancement
    performancePerformance or resource usage
    3.15bugs and security fixes
    on Oct 10, 2025
  3. cmaloney commented on Oct 10, 2025

    @cmaloney
    Contributor

    _BlocksOutputBuffer_Finish seems to be doing a "".join() on self->list. "stringlib" has a fast/fairly tuned implementation of that, wondering if the end join could just become something like:

        PyObject *new_buffer = PyObject_CallMethodOneArg(
            Py_GetConstant(Py_CONSTANT_EMPTY_BYTES), &_Py_ID(join), self->list);
  4. emmatyping commented on Oct 10, 2025

    @emmatyping
    MemberAuthor

    Good idea, I will also test that. I expect PyBytesWriter will still be faster because of locality, but I can benchmark both.

  5. emmatyping commented on Oct 12, 2025

    @emmatyping
    MemberAuthor

    I have more data, looks like not all of the (de)compressors rely on the output buffer code as much as zstd. That makes sense since the actual compression code is slower so the output buffer allocations are relatively less important.

    I've started with zlib as that is the most used format most likely. It's also the (de)compressor for wheels so performance gains there have huge multipliers :)

    My system:
    AMD Ryzen 9 9950X
    6000 MT/s DDR5 RAM
    Running Fedora 42

    Here's main vs using "".join():

    zlib.compress(1M): Mean +- std dev: [main] 13.5 ms +- 0.1 ms -> [bytes_join] 13.5 ms +- 0.1 ms: 1.00x slower
    zlib.compress(1G): Mean +- std dev: [main] 11.4 sec +- 0.0 sec -> [bytes_join] 11.4 sec +- 0.0 sec: 1.00x slower
    zlib.decompress(1K): Mean +- std dev: [main] 1.42 us +- 0.01 us -> [bytes_join] 1.44 us +- 0.01 us: 1.01x slower
    zlib.decompress(1M): Mean +- std dev: [main] 1.29 ms +- 0.00 ms -> [bytes_join] 1.20 ms +- 0.00 ms: 1.07x faster
    zlib.decompress(1G): Mean +- std dev: [main] 1.36 sec +- 0.00 sec -> [bytes_join] 1.39 sec +- 0.01 sec: 1.02x slower
    
    Benchmark hidden because not significant (1): zlib.compress(1K)
    Geometric mean: 1.00x faster
    

    Note that I had to add a _PyBytes_Resize immediately after the join() because the output bytes object needs to be truncated to the length of the output data.

    Here's comparing main vs PyBytesWriter:

    zlib.compress(1M): Mean +- std dev: [main] 13.5 ms +- 0.1 ms -> [pybyteswriter] 13.4 ms +- 0.0 ms: 1.00x faster
    zlib.compress(1G): Mean +- std dev: [main] 11.4 sec +- 0.0 sec -> [pybyteswriter] 11.3 sec +- 0.0 sec: 1.00x faster
    zlib.decompress(1K): Mean +- std dev: [main] 1.42 us +- 0.01 us -> [pybyteswriter] 1.39 us +- 0.01 us: 1.02x faster
    zlib.decompress(1M): Mean +- std dev: [main] 1.29 ms +- 0.00 ms -> [pybyteswriter] 1.17 ms +- 0.00 ms: 1.10x faster
    zlib.decompress(1G): Mean +- std dev: [main] 1.36 sec +- 0.00 sec -> [pybyteswriter] 1.17 sec +- 0.00 sec: 1.17x faster
    
    Benchmark hidden because not significant (1): zlib.compress(1K)
    
    Geometric mean: 1.05x faster
    

    5% overall is not bad! 10% faster decompression for data >1MB is also great, that will be a big win for wheels and many other use cases I imagine.

    I assume the numbers for xz and bz2 are pretty similar or a more minor win, since those (de)compressors are as slow or slower than zlib.

    I'm going to collect numbers for zstd and verify my initial numbers next, then work on a patch.

  6. emmatyping commented on Oct 12, 2025

    @emmatyping
    MemberAuthor

    Okay! As expected, the improvements are a lot better for zstd (using default compression level 3):

    zstd.compress(1K): Mean +- std dev: [main_zstd_3] 3.01 us +- 0.03 us -> [pybyteswriter_zstd_3] 3.00 us +- 0.03 us: 1.01x faster
    zstd.compress(1M): Mean +- std dev: [main_zstd_3] 2.92 ms +- 0.02 ms -> [pybyteswriter_zstd_3] 2.89 ms +- 0.02 ms: 1.01x faster
    zstd.compress(1G): Mean +- std dev: [main_zstd_3] 2.72 sec +- 0.01 sec -> [pybyteswriter_zstd_3] 2.67 sec +- 0.01 sec: 1.02x faster
    zstd.decompress(1K): Mean +- std dev: [main_zstd_3] 1.40 us +- 0.01 us -> [pybyteswriter_zstd_3] 1.38 us +- 0.01 us: 1.01x faster
    zstd.decompress(1M): Mean +- std dev: [main_zstd_3] 734 us +- 4 us -> [pybyteswriter_zstd_3] 546 us +- 3 us: 1.34x faster
    zstd.decompress(1G): Mean +- std dev: [main_zstd_3] 790 ms +- 4 ms -> [pybyteswriter_zstd_3] 634 ms +- 3 ms: 1.25x faster
    
    Geometric mean: 1.10x faster
    
  7. added a commit that references this issue on Oct 14, 2025
  8. vstinner commented on Oct 14, 2025

    @vstinner
    Member

    Can we close this issue?

  9. emmatyping commented on Oct 14, 2025

    @emmatyping
    MemberAuthor

    Yep!

  10. added a commit that references this issue on Dec 6, 2025
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

extension-modulesC modules in the Modules dirperformancePerformance or resource usagetype-featureA feature request or enhancement

Projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions