Repository navigation
Add .take_bytes([n]) a zero-copy path from bytearray to bytes #139871
Description
Activity
- addedtype-featureA feature request or enhancementA feature request or enhancement
on Oct 9, 2025 - addedinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Oct 9, 2025 - added a commit that references this issue
on Oct 14, 2025 See my arguments against this idea: https://discuss.python.org/t/add-zero-copy-conversion-of-bytearray-to-bytes-by-providing-bytes/79164/15
It will not work because it is not compatible with the C API.
PyByteArray_AS_STRING()cannot return a reference to the bytes object returned bytake_bytes(), because it can be mutated.PyByteArray_AS_STRING()cannot allocate a new bytes object, because it never fails, but a memory allocation can fail. The only solution is iftake_bytes()allocates a new bytes object, but this kills its purpose.The implementation for this issue is related to that initial discussion thread but iterated on with your feedback incorporated. The final form is in: https://discuss.python.org/t/add-take-bytes-n-to-bytearray-providing-a-zero-copy-path-to-bytes/103804. The initial post there still talks about
PyBytesWriterbut after further discussion consensus reached usingPyBytesObjectdirectly.In particular with the current design the
bytesobject insidebytearrayis either the empty bytes constant or a uniquely referencedbytesobject. Thatbytesobject is only exposed if a user calls.take_bytes([n])and it returns successfully. There is no__bytes__function or similar which exposes the internal tobytearrayPyObject*.For the particular points:
- re:
PyByteArray_AS_STRING(): This is equivalent to.clear()and other operations that resize thebytearray. A pointer stashed fromPyByteArray_AS_STRING()before a length-mutating operation (ex..clear()) is not valid after. The implementation oftake_bytes([n])does the same checks as.clear(): Working inside a critical section / lock so only one resizing operation may be in progress at a time and checkingob_exports == 0. There is a free threading test added to help validate correctness. - re: Allocation failures: The current implementation changes
bytearrayto always containbytes, by defaultPy_GetConstant(Py_CONSTANT_EMPTY_BYTES);._PyBytes_Resize's existing implementation handles resizing from empty bytes to larger by allocating a new bytes, resizing to 0 by changing to the empty bytes constant, and resizing in other cases by ensuring thebytesis uniquely referenced before performing any in-place mutation. If a_PyBytes_Resizefails for any reason thebytearrayis kept in a good state by re-initializing itself with the empty bytes object. Forba.take_bytes([n])in particularbawill end in one of these states:
a. pointing to a newbytesobject with all the bytes aftern(successfully completed)
b. It errors before getting to the actual byte taking and is left unmodified (ex.ob_exports != 0so it can't be resized)
c. A mid-resizing error occurs andbais set to the empty bytes constant, an exception is raised, and NULL is returned.
- re:
PyByteArray_AS_STRING() cannot return a reference to the bytes object returned by take_bytes(), because it can be mutated.
I'm not sure that I understand your concern.
PyByteArray.ob_bytes_objectis private and cannot be accessed (especially in Python). Yes, theob_bytes_objectobject is mutated: bytes can be modified and its size can change via bytearray operation (bytearray += bytes,bytearray.resize(), etc.). But I don't see how this is an issue, mutating a bytes object is done by all functions using the soft deprecatedPyBytes_FromStringAndSize(NULL, size). PEP 782 implementation usesPyBytesWriterin many cases to avoid that, butPyBytesWriteris implemented internally by ... mutating abytesobject :-)bytearray.take_bytes()replaces itsob_bytes_objectobject with a new (empty) one. The returnedbytesobject should now be treated as immutable.Links to previous discussion of this feature:
https://discuss.python.org/t/add-take-bytes-n-to-bytearray-providing-a-zero-copy-path-to-bytes/103804/5That was in September. There is also the older discussion (February): https://discuss.python.org/t/add-zero-copy-conversion-of-bytearray-to-bytes-by-providing-bytes/79164
- added a commit that references this issue
on Nov 13, 2025 - added a commit that references this issue
on Dec 3, 2025
Feature or enhancement
Proposal:
Update
bytearrayto have its internal buffer always be aPyBytesObjectand add a new method,.take_bytes([n])which extracts that buffer..take_bytes([n])adds a way to go from a "mutable" buffer of bytes to an immutable one without requiring a memcpy.This creates a path to resolve gh-60107
When would I use this?
Any code which makes a
ba = bytearray(), modifies it, then callsbytes(ba). Other common patterns arebytes(ba); ba.clear()orbytes(ba[:n]); del ba[:n]. Note that if you want to discard data past a point, the most efficient pattern becomesba.resize(n); ba.take_bytes()so thattake_bytes([n])doesn’t need to keep around the soon to be discarded extra bytes.Has this already been discussed elsewhere?
I have already discussed this feature proposal on Discourse
Links to previous discussion of this feature:
https://discuss.python.org/t/add-take-bytes-n-to-bytearray-providing-a-zero-copy-path-to-bytes/103804/5
https://discuss.python.org/t/add-zero-copy-conversion-of-bytearray-to-bytes-by-providing-bytes/79164 -- edit: added
Linked PRs
bytearray.take_bytes([n])to efficiently extractbytes#140128