Skip to content

implement PEP 3118 struct changes #47382

Description

@benjaminp
BPO 3132
Nosy @warsaw, @mdickinson, @ncoghlan, @abalkin, @pitrou, @devdanzin, @benjaminp, @pv, @skrah, @meadori, @vadmium
Files
  • pep-3118.patch
  • struct-string.py3k.patch: Patch for 'T{}' syntax and multiple byte order specifiers.
  • struct-string.py3k.2.patch: Patch with fixed assertions
  • struct-string.py3k.3.patch: Patch for 'T{}' against py3k r87813
  • grammar.y
  • Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

    Show more details

    GitHub fields:

    assignee = None
    closed_at = None
    created_at = <Date 2008-06-17.22:30:31.496>
    labels = ['type-feature', 'library']
    title = 'implement PEP 3118 struct changes'
    updated_at = <Date 2016-04-13.10:20:21.199>
    user = 'https://github.com/benjaminp'

    bugs.python.org fields:

    activity = <Date 2016-04-13.10:20:21.199>
    actor = 'skrah'
    assignee = 'none'
    closed = False
    closed_date = None
    closer = None
    components = ['Library (Lib)']
    creation = <Date 2008-06-17.22:30:31.496>
    creator = 'benjamin.peterson'
    dependencies = []
    files = ['16242', '17386', '17416', '20298', '42451']
    hgrepos = []
    issue_num = 3132
    keywords = ['patch']
    message_count = 58.0
    messages = ['68347', '68507', '71313', '71316', '71338', '71342', '71882', '87921', '99296', '99297', '99309', '99312', '99313', '99460', '99472', '99474', '99551', '99655', '99656', '99677', '99711', '99771', '105952', '105955', '105970', '106087', '106088', '106089', '106090', '106091', '106153', '106155', '106157', '106164', '106168', '106173', '106175', '106177', '106180', '106181', '106188', '106416', '123093', '123204', '123205', '123226', '123366', '125617', '130694', '130695', '130696', '143505', '143509', '167963', '187583', '187589', '187591', '263321']
    nosy_count = 17.0
    nosy_names = ['barry', 'teoliphant', 'mark.dickinson', 'ncoghlan', 'belopolsky', 'pitrou', 'inducer', 'ajaksu2', 'MrJean1', 'benjamin.peterson', 'pv', 'Arfrever', 'noufal', 'skrah', 'meador.inge', 'martin.panter', 'paulehoffman']
    pr_nums = []
    priority = 'high'
    resolution = None
    stage = 'patch review'
    status = 'open'
    superseder = None
    type = 'enhancement'
    url = 'https://bugs.python.org/issue3132'
    versions = ['Python 3.6']

    Activity

    1. benjaminp commented on Jun 17, 2008

      @benjaminp
      ContributorAuthor

      It seems the new modifiers to the struct.unpack/pack module that were
      proposed in PEP-3118 haven't been implemented yet.

    2. MrJean1 commented on Jun 21, 2008

      MrJean1mannequin
      Mannequin

      If the struct changes are made, add also 2 formats for C types ssize_t and
      size_t, perhaps 'z' resp. 'Z'. In particular since on platforms
      sizeof(size_t) != sizeof(long).

    3. warsaw commented on Aug 18, 2008

      @warsaw
      Member

      It's looking pessimistic that this is going to make it by beta 3. If
      they can't get in by then, it's too late.

    4. pitrou commented on Aug 18, 2008

      @pitrou
      Member

      Let's retarget it to 3.1 then. It's a new feature, not a behaviour
      change or a deprecation, so adding it to 3.0 isn't a necessity.

    5. added
      stdlibStandard Library Python modules in the Lib/ directory
      and removed on Aug 18, 2008
    6. benjaminp commented on Aug 18, 2008

      @benjaminp
      ContributorAuthor

      Actually, this may be a requirement of bpo-2394; PEP-3118 states that
      memoryview.tolist would use the struct module to do the unpacking.

    7. pitrou commented on Aug 18, 2008

      @pitrou
      Member

      Actually, this may be a requirement of bpo-2394; PEP-3118 states that
      memoryview.tolist would use the struct module to do the unpacking.

      :-(
      However, we don't have any examples of the buffer API / memoryview
      object working with something else than 1-dimensional contiguous char
      arrays (e.g. bytearray). Therefore, I suggest that Python 3.0 provide
      official support only for 1-dimensional contiguous char arrays. Then
      tolist() will be easy to implement even without using the struct module
      (just a list of integers, if I understand the functionality).

    8. teoliphant commented on Aug 24, 2008

      teoliphantmannequin
      Mannequin

      This can be re-targeted to 3.1 as described.

    9. devdanzin commented on May 16, 2009

      devdanzinmannequin
      Mannequin

      Travis,
      Do you think you can contribute for this to actually land in 3.2? Having
      a critical issue slipping from 3.0 to 3.3 would be bad...

      Does this supersede bpo-2395 or is this a subset of that one.?

    10. meadori commented on Feb 13, 2010

      @meadori
      Member

      Is anyone working on implementing these new struct modifiers? If not, then I would love to take a shot at it.

    11. benjaminp commented on Feb 13, 2010

      @benjaminp
      ContributorAuthor

      2010/2/12 Meador Inge <report@bugs.python.org>:

      Meador Inge <meadori@gmail.com> added the comment:

      Is anyone working on implementing these new struct modifiers?  If not, then I would love to take a shot at it.

      Not to my knowledge.

    12. 46 remaining items

    13. meadori commented on Sep 5, 2011

      @meadori
      Member

      Is this work something that might be suitable for the features/pep-3118 repo (http://hg.python.org/features/pep-3118/) ?

    14. skrah commented on Sep 5, 2011

      skrahmannequin
      Mannequin

      Yes, definitely. I'm going to push a new memoryview implementation
      (complete for all 1D/native format cases) in a couple of days.

      Once that is done, perhaps we could create a memoryview-struct
      branch on top of that.

    15. ncoghlan commented on Aug 11, 2012

      @ncoghlan
      Contributor

      Following up here after rejecting bpo-15622 as invalid

      The "unicode" codes in PEP-3118 need to be seriously rethought before any related changes are made in the struct module.

      1. The 'c' and 's' codes are currently used for raw bytes data (represented as bytes objects at the Python layer). This means the 'c' code cannot be used as described in PEP-3118 in a world with strict binary/text separation.

      2. Any format codes for UCS1, UCS2 and UCS4 are more usefully modelled on 's' than they are on 'c' (so that repeat counts create longer strings rather than lists of strings that each contain a single code point)

      3. Given some of the other proposals in PEP-3118, it seems more useful to define an embedded text format as "S{<encoding>}".

      UCS1 would then be "S{latin-1}", UCS2 would be approximated as "S{utf-16}" and UCS4 would be "S{utf-32}" and arbitrary encodings would also be supported. struct packing would implicitly encode from text to bytes while unpacking would implicitly decode bytes to text. As with 's' a length mismatch in the encoded form would mean an error.

    16. paulehoffman commented on Apr 22, 2013

      paulehoffmanmannequin
      Mannequin

      Following up on http://mail.python.org/pipermail/python-ideas/2011-March/009656.html, I would like to request that struct also handle half-precision floats directly. It's a short change, and half-precision floats are becoming much more popular in applications.

      Adding this to struct would also maybe need to change math.isinf and math.isnan, but maybe not.

    17. mdickinson commented on Apr 22, 2013

      @mdickinson
      Member

      Paul: there's already an open issue for adding float16 to the struct module: see bpo-11734.

    18. paulehoffman commented on Apr 22, 2013

      paulehoffmanmannequin
      Mannequin

      Whoops, never mind. Thanks for the pointer to 11734.

    19. skrah commented on Apr 13, 2016

      skrahmannequin
      Mannequin

      Here's a grammar that roughly describes the subset that NumPy supports.

      As for implementing this in the struct module: There is a new data
      description language on the horizon:

      http://datashape.readthedocs.org/en/latest/

      It does not have all the low-level capabilities (e.g changing alignment
      on the fly), but it is far more readable. Example:

      PEP-3118: "(2,3)10f0fZdT{10B:x:(2,3)d:y:Q:z:}B"
      Datashape: "2 * 3 * (10 * float32, 0 * float32, complex128, {x: 10 * uint8, y: 2 * 3 * float64, z: int64}, uint8)"

      There are a lot of open questions still. Should "10f" be viewed as an
      array[10] of float, i.e. equivalent to (10)f?

      In the context of PEP-3118, I think so.

    20. transferred this issue fromon Apr 10, 2022
    21. JelleZijlstra commented on May 24, 2023

      @JelleZijlstra
      Member

      It's been more than ten years now and most of the additions to struct proposed by PEP-3118 (https://peps.python.org/pep-3118/#additions-to-the-struct-string-syntax) have not been implemented. The ? and c codes have been added, though.

      At this point, I would propose to close the issue. If there is interest in adding new codes, that should be discussed in a new feature with a more specific motivation. It doesn't make sense to implement them based on a PEP from more than a decade ago.

    22. pitrou commented on May 24, 2023

      @pitrou
      Member

      I'll note that Numpy does seem to support at least part of these additions:

      >>> dt = np.dtype([('x', np.float64), ('y', np.float64)])
      >>> a = np.array([(1, 2)], dtype=dt)
      >>> m = memoryview(a)
      >>> m.format
      'T{d:x:d:y:}'
    23. pitrou commented on May 24, 2023

      @pitrou
      Member

      If there is interest in adding new codes, that should be discussed in a new feature with a more specific motivation. It doesn't make sense to implement them based on a PEP from more than a decade ago.

      It would make sense to support at least what Numpy supports.

    24. encukou commented on Jan 22, 2025

      @encukou
      Member

      Also note that ctypes supports ?, g, O (...), T{...}, and :name: but chose:

      • c for char, not ucs-1 string
      • u for wchar_t, not ucs-2 string
      • C, E, F (not Zd, Zf, Zg) for complex types (cc @skirpichev -- just FYI)
      >>> import ctypes
      >>> class C(ctypes.Structure):
      ...     _fields_ = [(f'f{num}', type) for num, type in enumerate([
      ...         ctypes.c_bool,
      ...         ctypes.c_longdouble,
      ...         ctypes.py_object,
      ...         ctypes.c_double_complex,
      ...         ctypes.c_wchar * 2,
      ...     ])]
      ...     
      >>> memoryview(C()).format
      'T{<?:f0:15x<g:f1:<O:f2:<C:f3:(2)<u:f4:}'

      There are formats in the wild that struct can't parse. Making struct parse them in an incompatible way could lead to data corruption.

      IMO, at this point, standardizing needs a new PEP.

    25. skirpichev commented on Jan 22, 2025

      @skirpichev
      Member

      C, E, F (not Zd, Zf, Zg) for complex types

      That was discussed, see e.g. here. Unfortunately, the fielddesc struct seems to be supporting only single-character names.

    Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

    Metadata

    Metadata

    Assignees

    No one assigned

      Labels

      stdlibStandard Library Python modules in the Lib/ directorytype-featureA feature request or enhancement

      Projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions