Skip to content

Add bytes_per_line parameter to binascii.b2a_base64 #141966

Description

@cmaloney

Feature or enhancement

Proposal:

Currently base64, email.contentmanager, imaplib, and plistlib all have code which takes a contiguous bytes, splits it into at most bytes_per_line (or maxbinsize) length chunks. These chunks are passed to binascii.b2a_base64, the results collected, and then joined back together to create a final result. For example:

cpython/Lib/base64.py

Lines 565 to 573 in 33efd71

def encodebytes(s):
"""Encode a bytestring into a bytes object containing multiple lines
of base-64 data."""
_input_type_check(s)
pieces = []
for i in range(0, len(s), MAXBINSIZE):
chunk = s[i : i + MAXBINSIZE]
pieces.append(binascii.b2a_base64(chunk))
return b"".join(pieces)

Internally b2a_base64 is using PyBytesWriter to manage the buffer and that could hold the final joined together bytes. To do that, need to teach it to handle bytes_per_line terminating lines after that many bytes and inserting a newline if required.

That reduces the amount of code to call as well as increasing efficiency of these cases by reducing the number of Python objects involved as well as the number of times data is copied.

Proposed new signature:

b2a_base64(data, *, newline=True, bytes_per_line=None)

Sample implementation: cmaloney@705bd9b

Has this already been discussed elsewhere?

This is a minor feature, which does not need previous discussion elsewhere

Links to previous discussion of this feature:

No response

Activity

  1. AnshArya927 commented on Nov 26, 2025

    @AnshArya927

    Hi team,

    I would like to work on improving binascii.b2a_base64 by adding an optional bytes_per_line parameter. Currently, modules like base64, email, imaplib, and plistlib split input bytes in Python before calling b2a_base64, which creates multiple Python objects and increases memory copies.

    My proposal:

    • Add bytes_per_line: Optional[int] = None to b2a_base64.
    • If set, the output will wrap encoded bytes at that line length, inserting newlines as needed.
    • Default behavior remains unchanged for backward compatibility.

    This change will reduce overhead, simplify code, and improve efficiency for large inputs. I’m happy to provide a patch with tests and benchmarks.

    Please assign this issue to me if it’s acceptable.

    Thank you!

  2. cmaloney commented on Nov 26, 2025

    @cmaloney
    ContributorAuthor

    Hi @AnshArya927 great enthusiasm! In the CPython repository issues aren't typically assigned to individuals. In this case there's already an existing patch linked in the issue / code ready, the issue is here for discussion around what the API should look like and if CPython wants this change in general.

  3. added
    performancePerformance or resource usage
    stdlibStandard Library Python modules in the Lib/ directory
    on Nov 26, 2025
  4. picnixz commented on Nov 27, 2025

    @picnixz
    Member

    FTR, b2a_base64 must conform to RFC 3548. email & co follow MIME RFC which has the notion of line feeds and cutoff. Here is the relevant paragraph for RFC 3548:

    MIME [3] is often used as a reference for base 64 encoding. However,
    MIME does not define "base 64" per se, but rather a "base 64
    Content-Transfer-Encoding" for use within MIME. As such, MIME
    enforces a limit on line length of base 64 encoded data to 76
    characters. MIME inherits the encoding from PEM [2] stating it is
    "virtually identical", however PEM uses a line length of 64
    characters. The MIME and PEM limits are both due to limits within
    SMTP.

    Implementations MUST NOT not add line feeds to base encoded data
    unless the specification referring to this document explicitly
    directs base encoders to add line feeds after a specific number of
    characters.

    AFAIU, this means that we could have an interface that automatically adds whatever needs to be added (newlines and splits). So I would probably be in favor of this change. Some observations and questions:

    • Are performances affected? how much do we gain with the for-loop version vs this one? how much do we lose for the paths that don't use this feature?
    • newline is meant to be used exactly once, namely add a trailing newline. If I use newline=False and bytes_per_line=100, it doesn't always make sense. Sometimes it fits 100 chars, sometimes not. In particular, one parameter affects the other which is usually not a good design choice for Python itself (it's usually a pattern we want to avoid). Now, in this case, it looks like worth the change (it's not unheard of for people to wrap b64 data).
    • bytes_per_line seems overly verbose but I don't think ew have a better name. Are there precedents elsewhere (other languages)?
  5. cmaloney commented on Dec 10, 2025

    @cmaloney
    ContributorAuthor

    Re: Performance, it should be still working on gathering numbers (the list of buffers + join -> just build a bytes mirrors the optimization in gh-139877 and I suspect will be about the same magnitude)

    re: bytes_per_line, existing comments around these pieces suggested max_line_length but the code functions on number of input bytes not number of output bytes. Definitely open to suggestions, I tend to bias towards verbose names.

  6. serhiy-storchaka commented on Dec 27, 2025

    @serhiy-storchaka
    Member

    Oh, I missed this issue and opened almost identical one: #143214. Details are slightly different:

    • A newline at the end and newlines between lines are controlled by orthogonal options, because we often need multiline output without a trailing newline.
    • The option specifies the length of the output line, not the number of input bytes per line.

    bytes_per_line looks similar to bytes_per_sep in bytes.hex() which has a counterintuitive behavior -- you usually need to specify a negative value for it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    performancePerformance or resource usagestdlibStandard Library Python modules in the Lib/ directorytype-featureA feature request or enhancement

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions