Skip to content

Provide detection for SIMD features in autoconf and at runtime #125022

Description

@picnixz

Feature or enhancement

Proposal:

In #124951, there has been some initial discussion on improving the performances of base64 and possibly {bytearray,bytes,str}.translate using SIMD instructions.

More generally, if we want to use specific SIMD instructions, it'd be good if we at least know whether the processor supports them or not. Note that we already support SIMD in blake2 when possible. As such, I suggest an internal framework for detecting SIMD features for other part of the library as well as a compiler flag support detection.

Note that a single part of the code could benefit from some SIMD calls without having to link the entire library against the entire SIMD-128 or SIMD-256 instruction sets. Note that having a way to detect SIMD support should probably be independent of whether we would use them or not apart from the blake2 module because it could only benefit the standard library if we were to include them.


The blake2 module SIMD support is fairly... complicated due to the wide variety of platforms that need to be supported and due to the mixture of many SIMD instructions. So I don't think I want to touch that part and make it work under the new interface (at least, not for now). While I can say that I'm confident in detecting features on "widely used" systems, there are definitely systems that I don't know so I'd appreciate any help on this topic.

Has this already been discussed elsewhere?

I don't want to open a Discourse thread for now since it's mainly something that will be used internally and not to be exposed to the world.

Links to previous discussion of this feature:

There has been some discussion on Discourse already about SIMD in general and whether to include them (e.g., https://discuss.python.org/t/standard-library-support-for-simd/35138) but the number of results containing "SIMD" or "AVX" is very small. Either this is because the topic is too advanced (detecting CPU features is NOT funny and there is a lack of documentation, the best one being the Wikipedia page) or the feature request is too broad.

Linked PRs

Activity

  1. added
    type-featureA feature request or enhancement
    buildThe build process and cross-build
    on Oct 6, 2024
  2. self-assigned this
    on Oct 6, 2024
  3. corona10 commented on Oct 6, 2024

    @corona10
    Member

    My general opinion about managing SIMD logics in CPython side: #124951 (comment)
    And if we begin to depend on SIMD detection, do you have any concern that an unexpected illegal instruction error can occur from the unsupported machine side because of the difference between the build machine and the execution machine?

  4. picnixz commented on Oct 6, 2024

    @picnixz
    MemberAuthor

    I do have concerns and that's why I'd like to hear from people that 1) know about weird architectures 2) deal with real-life scenarios.

    What I have in mind:

    • The feature should be entirely opt-in. Possibly under some optimization flag as well.
    • We already have SIMD in blake2 so we may already possibly have those "unexpected illegal instruction" situation. I'm currently (well not today) trying to harden the detection of AVX instructions because even recognizing -mavx may not be sufficient (e.g., we also need to handle XSAVE and how it handles YMM registers).

    So, yes, I definitely have concerns on the differences. Using SIMD instructions could probably make local builds faster or builds managed by distributions themselves though we should be really careful. This is also the reason why I want to keep runtime detection to avoid issues.

    The idea was to open a wider discussion on SIMD support itself. If you want we can move to Discourse though I'm not sure whether it's better to keep it internal for now (the PR is just a PoC and it probably won't cover those cases we're worried about).


    I don't think we should add SIMD for every possible parts of the library, only those that are critical enough IMO. And they should be carefully worked out. However, in order to investigate them (and test them using the CI), I think having an internal detection framework would at least be the first step (or maybe I'm wrong here?).

  5. picnixz commented on Oct 7, 2024

    @picnixz
    MemberAuthor

    I've harden the detection of AVX instructions. I've also learned that macOS may not like AVX-512 at all (or at least some registers states won't be restored correctly upon context-switching). So there are real-life issues that we should address. What I'll maybe do is first try to make a PoC for str.translate and see how AVX could be used and how it could improve Python, then I'll come back (as Gregory said on the othere issue, we are targetting relatively simpler algorithms).

  6. Starbuck5 commented on Oct 14, 2024

    @Starbuck5

    Hello, I'm one of the maintainers of pygame-ce, a Python C extension library that uses SIMD extensively to speed up pixel processing operations. We've had various bits of SIMD for a long time and use runtime checks to manage it. I'd like to share some information about our approach, in the hope it is helpful.

    We SIMD accelerate at the SSE2 and AVX2 levels. SSE2 is part of the baseline of x86_64, but we've also had no problems with it on our 32 bit builds. AVX2 is where isolation and runtime checking is much more important.

    Each SIMD level of a module has its own file and is compiled into its own object. See https://github.com/pygame-community/pygame-ce/blob/6e0e0c67c799c7cc1fa9c96a71598a7751ae2fba/src_c/simd_transform_avx2.c for an example. Our build config for this looks like so: https://github.com/pygame-community/pygame-ce/blob/6e0e0c67c799c7cc1fa9c96a71598a7751ae2fba/src_c/meson.build#L215-L254. In this example, our transform module is not compiled with any special flags, but it is linked with objects that expose functions that can be called to get SIMD acceleration. An example of how the dispatch looks: https://github.com/pygame-community/pygame-ce/blob/6e0e0c67c799c7cc1fa9c96a71598a7751ae2fba/src_c/transform.c#L2158-L2181.

    The SIMD compilation itself is very conservative, it will only compile the backend if the computer doing the build supports that backend, using compile time macros to check that. I'm not sure if this is actually necessary.

    About our SIMD code itself, we use intrinsics rather than hardcoded assembly or frameworks like https://github.com/google/highway. https://www.intel.com/content/www/us/en/docs/intrinsics-guide/index.html is a great reference on this. Intrinsics are better for us than hardcoded assembly because they are more portable, both between compilers and even between architectures. For example we compile all of our "SSE2" code to NEON for ARM support using https://github.com/DLTcollab/sse2neon. Emscripten also allows compile time translation of SIMD intrinsics to Webassembly SIMD, https://emscripten.org/docs/porting/simd.html, although we do not take advantage of this currently.

    For runtime detection, we rely on https://github.com/libsdl-org/SDL, which is very easy for us because our entire library is built on top of the functionality provided by SDL. If you'd like to check your PR against their implementation of runtime checks, the source seems to be here: https://github.com/libsdl-org/SDL/tree/main/src/cpuinfo

    I think there could be value in exposing runtime SIMD support checking into the public C API for extension authors that aren't lucky enough to have an existing dependency to rely on for this. I've followed issues about Pillow and Pillow-SIMD where the authors are like these are impossible to merge because we don't have the resources to figure out runtime SIMD checks. I don't think it would get a ton of usage, but it would be extremely valuable functionality for any who need it.

    In terms of CPython and SIMD I'm not sure how much potential there is, but there may be cool things that could be done. Could 4 or 8 or 16 PyObjects get their reference count changed at once? Could unboxed integers do arithmetic in parallel?Could the JIT decide to use more efficient templates because it knows AVX2 is around? Knowing the runtime SIMD level is an advantage for a JIT over an AOT compiler. But these are just my musings.

  7. picnixz commented on Oct 14, 2024

    @picnixz
    MemberAuthor

    @Starbuck5 Thank you very much for all these insights!

    I'd like to share some information about our approach, in the hope it is helpful.

    It was definitely helpful.

    Each SIMD level of a module has its own file and is compiled into its own object

    Yup, that's what the blake2 authors did so we'll probably do something similar. One thing is that it could lead to blowing the code up if we have many different levels... or if we decide to split up files according to architecture itself (like it is done in https://github.com/aklomp/base64/tree/22a3e9d421ee25b25bc6af7a02d4076c49dd323f/lib/arch for instance). I personally think it's nicer to split them by architecture and folders but this leads to too many files and similar ones (which is not very nice for maintaining the whole thing).

    The SIMD compilation itself is very conservative, it will only compile the backend if the computer doing the build supports that backend, using compile time macros to check that. I'm not sure if this is actually necessary.

    I think it's always better to be safe than sorry, unless we're absolutely sure that we won't cause #UD at runtime.

    About our SIMD code itself, we use intrinsics rather than hardcoded assembly

    I also think it's better to use intrinsics for the same reasons as you cited but also because you don't need to know about ASM :') (and portability is key). The projects for translating intrinsics will definitely be helpful in the future if we were to eventually use SIMD instructions.

    For runtime detection, we rely on libsdl-org/SDL, which is very easy for us because our entire library is built on top of the functionality provided by SDL

    Thanks for this. I'll probably borrow some of their ideas but I don't think we can vendor this specific part of their library in CPython :( However it will definitely help in improving the detection algorithm (for now the algorithm is quite crude).

    I think there could be value in exposing runtime SIMD support checking into the public C API for extension authors that aren't lucky enough to have an existing dependency to rely on for this [...] In terms of CPython and SIMD I'm not sure how much potential there is, but there may be cool things that could be done

    That was my original intent, though limited to the CPython internals. Python is great but Python is sometimes slow on some aspects, and it'd be great if we could make it faster. We can always make Python faster by changing algorithms but if we have the possibility of making it faster using CPU features then we should probably try to benefit from them, at least in the important areas.

    Could 4 or 8 or 16 PyObjects get their reference count changed at once

    We could in some situations do it but this will probably need to be synchronized with the ongoing work on deferred reference counts (or so I think).

    Could unboxed integers do arithmetic in parallel

    @skirpichev do we have places where the arithmetic could be sped up using SIMD instructions? I think we either rely on mpdecimal or glibc directly for "advanced" arithmetic and I don't know whether we have a lot of places where we have additions in batches for instance.

    Could the JIT decide to use more efficient templates because it knows AVX2 is around? Knowing the runtime SIMD level is an advantage for a JIT over an AOT compiler

    I'm not JIT expert :') so let's ask someone who knows about it: @brandtbucher

  8. skirpichev commented on Oct 14, 2024

    @skirpichev
    Member

    do we have places where the arithmetic could be sped up using SIMD instructions?

    AFAIK, GMP isn't utilizes this too much so far.

  9. Zheaoli commented on Oct 14, 2024

    @Zheaoli
    Contributor

    Could the JIT decide to use more efficient templates because it knows AVX2 is around? Knowing the runtime SIMD level is an advantage for a JIT over an AOT compiler

    I'm not JIT expert :') so let's ask someone who knows about it: @brandtbucher

    For now, we pre compile the JIT code from template. So it depends on the meachine we use to release the official binary(But the JIT is not a default feature yet.). As I know, we dont have any instruction detect on the JIT build script, in another world, Use SIMD or not is depends on the compiler decision.

  10. Zheaoli commented on Oct 14, 2024

    @Zheaoli
    Contributor

    Could unboxed integers do arithmetic in parallel?

    I think arithmetic is not a common use case. The people may take care of there data layout to fit the parallel requirement. Otherwise, it may slower than normal operation.

    IMHO, I think some string operation it more suitable for SIMD, like JSON operation or pickle operation. FYI https://github.com/simdjson/simdjson

  11. Starbuck5 commented on Oct 16, 2024

    @Starbuck5

    For now, we pre compile the JIT code from template. So it depends on the meachine we use to release the official binary(But the JIT is not a default feature yet.). As I know, we dont have any instruction detect on the JIT build script, in another world, Use SIMD or not is depends on the compiler decision.

    On x86 the default compilation (at least on MSVC) goes up to SSE2. So there could be auto-vectorization opportunities. I actually investigated this quite a bit ago and got the JIT templates to compile with an explicit AVX2 flag, and it barely changed the templates at all. I think it would have to be more intentionally set up, like if unboxed integer arithmetic becomes a thing there could then be a uop that does 2 / 4 / 8 arithmetics at once and then the compiler would be able to do some auto vectorization there.

    But I'm fully aware all my ideas about this and the refcount thing are just ideas, not anywhere close to a concrete proposal.

    I think arithmetic is not a common use case. The people may take care of there data layout to fit the parallel requirement. Otherwise, it may slower than normal operation.

    I think the popularity of Numba showcases demand for higher performance number crunching.


    In terms of actionable SIMD items, an internal api for runtime detection seems like a great step to take. An external api could also be helpful to certain projects.

    I haven't looked in detail into the blake SIMD implementations, but if it just supports x86 right now it would be possible to bring those speedups to ARM using sse2neon.

    Personally I've never done any string operations with SIMD, but I agree with @Zheaoli that there is certainly potential to speed up things with it!

  12. diegorusso commented on Oct 21, 2024

    @diegorusso
    Contributor

    Hello, thanks for raising this. I think there is definitely some room for vector instructions in CPython. In the coming weeks I'll spend some time investigating and I'll be watching this space as well.

  13. cosmicexplorer commented on Feb 22, 2026

    @cosmicexplorer

    I was directed to this issue upon describing how I dove into URL quoting performance for pip and found that there were some very hard limits we were running into with the current string/byte search mechanisms available in cpython. I have a sketch in my head of an API using standard and portable SIMD techniques for matching simple sequences from low cardinality sets and/or ranges (intervals) of byte values.

    Vague Proposal: Byte Literal Matching API

    This is a rather lengthy reply, so the primary points I would like to make are:

    1. sparse (infrequently matching) string search operations are surprisingly common in real-world tasks.
      • URL quoting needs to match against a byte set--so does xmlcharrefreplace().
      • Generating a sequence of match offsets much more performant than splitting and joining strings.
      • Match offsets can be composed with other search logic, and are particularly useful to expose as coroutines.
    2. most real-world SIMD applications to match literal strings and/or sets of bytes are relatively simple and portable.
      • See existing usage of memchr() from libc in cpython.

    The provided primitives for string search/matching result in load-bearing calls to functions that are not designed to solve the given problem, but happen to achieve better performance (see rstrip() below). The re module is not designed for performance, so users are subtly incentivized to use string primitives instead of allowing re to actually parse the input. This results in less safe, less maintainable, and less performant python code.

    I think there is room to support a limited set of new primitives to:

    • compile a matcher for a (parallel) set of bytes or a (serial) sequence of bytes
    • execute the matcher against an immediate string and/or a stream (iterable) of strings
    • retrieve match offsets corresponding to positions in the input through a coroutine/iterable interface

    I particularly suspect this would be useful for interactive subprocess output.

    Finally,

    1. I think making a module wrapping hyperscan would be extremely cool.
      • (providing hyperscan as a dependency to autoconf)
      • Relying on hyperscan would sidestep any concerns about configuring processor instructions in cpython's autoconf script.
        • See their docs on Instruction Set Specialization1.

    The rest of this post is an attempt to further motivate the idea of SIMD for string matching as a lightweight and composable paradigm.

    Case Study: the load-bearing rstrip()

    Background: The most significant performance improvements I have achieved in pip are the result of separating and independently caching its distinct phases of execution (pypa/pip#12921). String parsing (in the form of parsing URL components) and especially URL quoting is the first time I've ever encountered a performance problem that cannot be solved in pure python (see my packagingcon 2023 talk2 which expounds at length on I/O performance in pip).

    JSON parsing the massive responses pypi sends back also represents a very significant performance bottleneck--the json module requires parsing the entire document at once and unconditionally allocates a new string then invokes a callback for every key and value in the document. There is no ability to tell the decoder to skip an object key, for example. The stdlib docs loudly acknowledge this3, but offer no solution or alternative, even though simdjson4 has demonstrated production-ready robustness along with their extremely innovative SIMD parsing approach (the paper is so cool and accessible!5).

    As a revealing case study, this precise line is by far the most significant contributor to the performance of urllib.parse.quote_plus():

    if not bs.rstrip(_ALWAYS_SAFE_BYTES + safe):

    This function is called thousands of times per cli invocation in current pip. This can be reduced by a constant factor with some in-memory caching, but this remains an intrinsic problem because most of what pip does is parse URLs. URL parsing is a code injection vector, so pip cannot risk doing it wrong. In that light, the reliance upon this single uncommented line of code and its implicit C-level semantics for performance seems suboptimal at best.

    The load-bearing line of code again:

         if not bs.rstrip(_ALWAYS_SAFE_BYTES + safe):

    To explain what this does: Calling .rstrip() here achieves early-exit semantics in the special case where we find no characters necessitating quote translation. This exits without ever engaging the _Quoter caching logic or the chunking by square root taking up the second half of the function below the load-bearing line. .rstrip() is more efficient than any alternative here because:

    • rstrip() may call the C function do_xstrip():
      return do_argstrip(self, RIGHTSTRIP, bytes);
    • do_xstrip() may call the standard libc memchr() function, which often internally is implemented with SIMD instructions
      while (i < len && memchr(sep, Py_CHARMASK(s[i]), seplen)) {

    One might then ask why the stdlib only invokes this once, instead of using it to iteratively identify the contiguous regions it needs to quote by splitting off non-matching chunks. Two answers to this:

    1. The .rstrip() trick that relies on memchr() only works when the set of bytes to match is very small. To split the string into matching and non-matching regions, it needs to switch off matching against the set of all bytes except the specific bytes in that small initial byte set. This means invoking memchr() in a loop around 250 times per contiguous region.
    2. The stdlib does perform that exact iterative splitting process.....for unquotes! 2e279e8
      • As per that commit description, .finditer() reduces memory consumption over the .split() approach which eagerly allocates a list of strings, inducing memory copies and further allocation pressure.
        • Should strings that reference other strings be able to avoid copying over that string data?
      • I see @gpshead contributed that change--I was super impressed to see this very effective and concise analysis of the engineering tradeoffs involved: gh-88500: Reduce memory use of urllib.unquote #96763 (comment)
      • The .split() approach was introduced 13 years ago in 8ea4616 -- as I describe below, this exactly the solution I landed on and found to be more performant for pip's purposes.
        • This indicates a potential conflict between memory pressure (important for long-running processes like the pypi warehouse server) and runtime performance (important for cli tools like pip). Can python serve both use cases without compromises?

    Alternatives: Splitting, Chunking, Bitsets

    The one improvement I have been able to achieve over the stdlib in pure python code was through using the re library's split() method: https://codeberg.org/cosmicexplorer/pip/src/commit/3bde75faebeae014e05b0c818b450a798a62a9a9/src/pip/_internal/utils/urls.py#L781

            split = self._unsafe_regions.split(path)

    I also found some benefit in a modified chunking approach: https://codeberg.org/cosmicexplorer/pip/src/commit/3bde75faebeae014e05b0c818b450a798a62a9a9/src/pip/_internal/utils/urls.py#L785-L792

                if len(prev) > self._split_chunks_length:
                    # NB: list comprehension is faster than generator here.
                    split[i] = "".join(
                        [
                            self[prev[j : j + self._split_chunks_length]]
                            for j in range(0, len(prev), self._split_chunks_length)
                        ]
                    )

    While the stdlib _Quoter maps single bytes to single strings, this approach instead splits the input into (non-overlapping) chunks <= a static maximum length parameter. This chunking reduces the number of lookups and consequently the cardinality of the resulting string components to join together after translating each chunk. This is vaguely reminiscent of an n-gram index, but without recursive hashing (overlapping chunk windows) it's not really analogous. I mention this because SIMD recursive hashing to lookup against a bloom filter may be an interesting future research direction.

    Two additional observations:

    • using an entire dict is wasteful for a fixed-cardinality input set, particular if the input is also consistently sized in memory.
    • instead of calculating contiguous regions to quote in advance, the stdlib instead performs a full dict membership check for every byte of the input string.

    Thinking in Intervals!

    By way of contrast, emacs lisp provides an opaque char-table data structure to lisp code, offering an interface to set a value for a contiguous range (aka "interval") of small nonnegative integer indices. This is internally used to encode unicode properties for efficient lookup, and consequently supports an optimize-char-table operation, which minimizes any internal fragmentation after a sequence of modifications.

    Analogy: SIMD for emacs string search

    The char-table data structure is also employed heavily to implement the regex-emacs matching logic, looking up each consecutive decoded codepoint within the active syntax table. This is remarkably reminiscent of cpython's current byte-by-byte iteration for url quoting. In fact, at emacsconf 2024 I discussed at length the prospect of introducing SIMD approaches for string matching operations in emacs lisp code: https://emacsconf.org/2024/talks/regex/.

    A summary of the consensus from emacs-devel discussions around that investigation (https://lists.gnu.org/archive/html/emacs-devel/2024-08/msg00108.html) mentions two points worth consideration for SIMD in cpython as well:

    1. Identifying "pathological" inputs for cpython's current string search/match techniques.
      • Unquoted URLs with multiple non-contiguous non-empty regions of bytes we need to encode.
      • Sequences of literals demarcated by regions of indefinite length (e.g. r'a.*b').
    2. Building in support for concurrent evaluation contexts.
      • Advancing the no-gil workstream.
      • Leveraging cpython's incredibly powerful built-in concurrency primitives to explicitly encode internal state transitions (compiling a search pattern, deallocating the compiled matcher, allocating intermediate results like match data).

    Finally, there has been a lot of exciting work to improve the mutable bytearray type recently (see e.g. #141864 and particularly the workstream in #139871), matched only by the very thoughtfully architected API of PyBytesWriter (see #129813). This means that any internal buffering used for search or matching operations should interface safely and robustly with other python code.

    Prior Art: rebar, zip, re2, hyperscan

    As a prerequisite to any of the below discussion of regex engines, I strongly recommend perusing the rebar benchmarks by @BurntSushi, which contains extremely thoughtful and effective comparisons of relative contributions to regex performance. It particularly identifies the prefilter6 as an exceptionally important contribution to performance, and specifically names the SIMD algorithms developed by hyperscan as the reason for its benchmark results.

    Regarding concrete implementation considerations, this page7 serves as a fantastic introductory text. It also cites one of the hyperscan authors.

    zip

    I introduced SIMD literal byte matching to the rust zip crate to fix yet another byte-by-byte iteration: zip-rs/zip2#93. To my knowledge this was the first implementation of zip file parsing to employ SIMD. It was subsequently vastly improved and cleaned up in zip-rs/zip2@33c71cc. It successfully encodes some extremely nontrivial stateful interactions: https://github.com/zip-rs/zip2/blob/5fcfad0bdfa587351a2c6f9b322c5dccbfc51aff/src/read/magic_finder.rs#L212

    /// A magic bytes finder with an optimistic guess that is tried before
    /// the inner finder begins searching from end. This enables much faster
    /// lookup in files without appended junk, because the magic bytes will be
    /// found directly.
    ///
    /// The guess can be marked as mandatory to produce an error. This is useful
    /// if the `ArchiveOffset` is known and auto-detection is not desired.
    pub struct OptimisticMagicFinder<Direction> {
        inner: MagicFinder<Direction>,
        initial_guess: Option<(u64, bool)>,
    }

    I believe we could incorporate these exact techniques into an alternate zipfile implementation in cpython, through a native module which can make use of any configured vector intrinsics.

    re2

    I developed a rust wrapper for the RE2 C++ regex engine in concert with the late maintainer a couple years back. One of the ideas we both found terribly interesting was how to incorporate literal search/match operations into regex matching, and characterizing the class of problems solvable with literal matching logic alone.

    My docs for it8 describe how RE2 compiles a regex pattern string into a sequence of literal match operations. The "atoms" (consecutive literals extracted from the compiled regex pattern) it generates can be directly consumed by a variety of standard approaches for string search, including e.g. bloom filters and n-gram indices as well as SIMD intrinsics7.

    hyperscan

    The hyperscan project deserves more love. I also created a rust wrapper for their work, and in doing so ended up producing an incredibly thorough analysis of what hyperscan does differently9. In particular I focused on their callback API10 as an important contribution to both performance and ergonomics.

    cpython has developed stable ABIs for coroutines, which are perfectly compatible with hyperscan's unique callback interface. hyperscan has also worked very hard to support streaming (i.e. non-contiguous) input11, which is completely unique among regex engines.

    Conclusion

    • I think there's room for literal byte set/sequence matchers using relatively basic SIMD techniques. I know it would speed up pip. The API generating a stream of match offsets would enable remaining in the C call stack until all matching offsets are generated instead of jumping back and forth from python to C.
    • hyperscan and simdjson both sound like really good proposals for optional modules to provide in the stdlib. I think would reduce the pressure to speed up the process of adding any SIMD instructions to cpython itself.
    • A string type that references another

    Footnotes

    1. https://intel.github.io/hyperscan/dev-reference/compilation.html#instruction-set-specialization ↩

    2. https://web.archive.org/web/20250218154403/https://cfp.packaging-con.org/2023/talk/hpuhu7/ ↩

    3. https://docs.python.org/3/library/json.html ↩

    4. https://simdjson.org/about/ ↩

    5. https://arxiv.org/pdf/1902.08318 ↩

    6. https://github.com/BurntSushi/rebar/blob/master/README.md#literal-alternate ↩

    7. http://0x80.pl/notesen/2018-10-18-simd-byte-lookup.html ↩ ↩2

    8. https://docs.rs/re2/latest/re2/filtered/struct.FilteredRE2.html ↩

    9. https://docs.rs/vectorscan-async ↩

    10. https://docs.rs/vectorscan-async/latest/vectorscan/matchers/index.html#vectorscan-callback-api ↩

    11. https://docs.rs/vectorscan-async/latest/vectorscan/stream/index.html#stream-parsing ↩

  14. gpshead commented on Feb 22, 2026

    @gpshead
    Member

    At a high level I think @cosmicexplorer 's comment may be over-focusing on hunting for use cases. Some of your ideas are probably best as feature requests on their own with example implementations and practical data. Don't be surprised to see a suggestion of proving those via a PyPI package first.

    The point of this issue as I see it is more enabling us to use the widely available hardware features where appropriate for what we already do offer, and in a more cohesive manner. We already have cases where SIMD is used (hashlib builtins) and areas where it can be useful, but most attempts so far have had to endure the pain of defining appropriate detection. I don't actually envision the mix of compilation and runtime detection pain going entirely away given the many shapes of hardware support and shapes of vector instruction intrinsics desired for a particular purpose, but it could at least be managed as a whole in one place rather than being scattered. Thus easing its use.


    • hyperscan [PyPI] and simdjson [GH] both sound like really good proposals for optional modules to provide in the stdlib.

    (linkified the quote) Those existing externally is good and working as intended (I did not linking to a PyPI simdjson wrapper as there are multiple and it isn't obvious to me which if any are deserving). We avoid "optional" in the stdlib, and the mantra "the stdlib is where libraries goes to die" or similar is pretty true - it isn't a place to actively iterate upon and improve designs of libraries to keep up with the latest and greatest thing if you want them in the hands of actual users. stdlib is long-term-stale. Anything in the stdlib by definition slows down its rate of change.


    • A string type that references another

    We sort of have this with memoryview but that is bytes level, a unicode py3-str variant of that was never created. Various libraries take pos and endpos style arguments as a result. Not ideal.

    Any such view like things are all PyObject's internally so there is overhead that usually means trying to use them everywhere is not worthwhile, but when you know you won't have a ton of them and the size of the things you are referencing are sufficient, using them can make sense once you know your data profile and how to track if they're useful vs a distraction.

    But a referencing view type is off topic for this issue.

  15. Reisande commented on Sep 20, 2026

    @Reisande

    @gpshead one of the earlier comments mentioned using SDL, I am relatively new to python/cpythoon development, but would we be able to simply take that as a dependency and check once at startup, perhaps caching into a temp file?

  16. picnixz commented on Sep 20, 2026

    @picnixz
    MemberAuthor

    No, we do not want another dependency. This is not the essence of Python as it means yet another requirement (and one that is nontrivial and mostly useless except for that)

  17. Reisande commented on Oct 5, 2026

    @Reisande

    Any update @picnixz ?

  18. picnixz commented on Oct 5, 2026

    @picnixz
    MemberAuthor

    There won't be an update until a long time. This is a nontrivial feature and I doubt we will accept non-core devs contributing to it. I plan to ask around at the next CPython sprint but don't expect this to be included before one year maybe.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

buildThe build process and cross-buildinterpreter-core(Objects, Python, Grammar, and Parser dirs)performancePerformance or resource usagetype-featureA feature request or enhancement

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions