Skip to content

Generate vectorcall code to parse arguments using Argument Clinic #87613

Description

@vstinner
BPO 43447
Nosy @vstinner, @serhiy-storchaka, @corona10, @erlend-aasland, @kumaraditya303

Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

Show more details

GitHub fields:

assignee = 'https://github.com/corona10'
closed_at = None
created_at = <Date 2021-03-09.11:34:23.944>
labels = ['expert-C-API', '3.10']
title = 'Generate vectorcall code to parse arguments using Argument Clinic'
updated_at = <Date 2022-03-06.09:38:05.594>
user = 'https://github.com/vstinner'

bugs.python.org fields:

activity = <Date 2022-03-06.09:38:05.594>
actor = 'kumaraditya'
assignee = 'corona10'
closed = False
closed_date = None
closer = None
components = ['C API']
creation = <Date 2021-03-09.11:34:23.944>
creator = 'vstinner'
dependencies = []
files = []
hgrepos = []
issue_num = 43447
keywords = []
message_count = 2.0
messages = ['388354', '389214']
nosy_count = 5.0
nosy_names = ['vstinner', 'serhiy.storchaka', 'corona10', 'erlendaasland', 'kumaraditya']
pr_nums = []
priority = 'normal'
resolution = None
stage = None
status = 'open'
superseder = None
type = None
url = 'https://bugs.python.org/issue43447'
versions = ['Python 3.10']

Linked PRs

Activity

  1. vstinner commented on Mar 9, 2021

    @vstinner
    MemberAuthor

    To optimize the creation of objects, a lot of "tp_new" methods are defined twice: once in the legacy way (tp_new slot), once with the new VECTORCALL calling convention (tp_vectorcall slot). My concern is that the VECTORCALL implementation copy/paste most of the code just to parse arguments, whereas the specific code is just a few lines.

    Example with the float type constructor:
    --------------

    static PyObject *
    float_new(PyTypeObject *type, PyObject *args, PyObject *kwargs)
    {
        PyObject *return_value = NULL;
        PyObject *x = NULL;
    
        if ((type == &PyFloat_Type) &&
            !_PyArg_NoKeywords("float", kwargs)) {
            goto exit;
        }
        if (!_PyArg_CheckPositional("float", PyTuple_GET_SIZE(args), 0, 1)) {
            goto exit;
        }
        if (PyTuple_GET_SIZE(args) < 1) {
            goto skip_optional;
        }
        x = PyTuple_GET_ITEM(args, 0);
    skip_optional:
        return_value = float_new_impl(type, x);
    
    exit:
        return return_value;
    }
    
    /*[clinic input]
    @classmethod
    float.\_\_new__ as float_new
        x: object(c_default="NULL") = 0
        /
    
    Convert a string or number to a floating point number, if possible.
    [clinic start generated code]*/
    
    static PyObject *
    float_new_impl(PyTypeObject *type, PyObject *x)
    /*[clinic end generated code: output=ccf1e8dc460ba6ba input=f43661b7de03e9d8]*/
    {
        if (type != &PyFloat_Type) {
            if (x == NULL) {
                x = _PyLong_GetZero();
            }
            return float_subtype_new(type, x); /* Wimp out */
        }
    
        if (x == NULL) {
            return PyFloat_FromDouble(0.0);
        }
        /* If it's a string, but not a string subclass, use
           PyFloat_FromString. */
        if (PyUnicode_CheckExact(x))
            return PyFloat_FromString(x);
        return PyNumber_Float(x);
    }
    
    static PyObject *
    float_vectorcall(PyObject *type, PyObject * const*args,
                     size_t nargsf, PyObject *kwnames)
    {
        if (!_PyArg_NoKwnames("float", kwnames)) {
            return NULL;
        }
    
        Py_ssize_t nargs = PyVectorcall_NARGS(nargsf);
        if (!_PyArg_CheckPositional("float", nargs, 0, 1)) {
            return NULL;
        }
    
        PyObject *x = nargs >= 1 ? args[0] : NULL;
        return float_new_impl((PyTypeObject *)type, x);
    }

    Here the float_new() function (tp_new slot) is implemented with Argument Clinic: float_new() C code is generated from the [clinic input] DSL: good!

    My concern is that float_vectorcall() code is hand written, it's boring to write and boring to maintain.

    Would it be possible to add a new [clinic input] DSL for vectorcall? I expect something like that:
    --------------

    static PyObject *
    float_vectorcall(PyObject *type, PyObject * const*args,
                     size_t nargsf, PyObject *kwnames)
    {
        if (!_PyArg_NoKwnames("float", kwnames)) {
            return NULL;
        }
    
        Py_ssize_t nargs = PyVectorcall_NARGS(nargsf);
        if (!_PyArg_CheckPositional("float", nargs, 0, 1)) {
            return NULL;
        }
    
        PyObject *x = nargs >= 1 ? args[0] : NULL;
        return float_vectorcall_impl(type, x);
    }
    
    static PyObject *
    float_vectorcall_impl(PyObject *type, PyObject *x)
    {
        return float_new_impl((PyTypeObject *)type, x);
    }

    where float_vectorcall() C code would be generated, and float_vectorcall_impl() would be the only part written manually. float_vectorcall_impl() gets a clean API and its body is way simpler to write and to maintain!

  2. corona10 commented on Mar 21, 2021

    @corona10
    Member

    I agree with we should update AC to generate vectorcall.
    I am going to investigate what we can :)

  3. transferred this issue fromon Apr 10, 2022
  4. vstinner commented on Nov 3, 2022

    @vstinner
    MemberAuthor

    I don't have the bandwidth to work on this issue, so I just close it.

  5. added
    type-featureA feature request or enhancement
    and removed on Aug 17, 2023
  6. 64 remaining items

  7. added 2 commits that reference this issue on Mar 1, 2026
  8. cmaloney commented on Mar 1, 2026

    @cmaloney
    Contributor

    Have a working implementation with performance matching or beating the currently handwritten implementations. Benchmark and full numbers below. Overall 1.03x faster.

    First PR: GH-145381. The changes I have split across a couple commits in my draft ac_vectorcall_v1 branch. Did a PR for just adding the support to AC; then not sure what would be best for the other changes (leaning: PR per type for easy reverts if needed?).

    • Adding @vectorcall decorator and codegen with test
    • Commit per type, replacing hand-coded vectorcall implementations with AC generated ones
    • Adding @vectorcall to additional types.
    • Microbenchmark for different constructions. Not planning to try and get into CPython; think it could be useful to track perf over time but not sure a good place for it to live.

    Performance numbers

    Benchmark code: cmaloney@ddcd3b6
    Configure options: ./configure --enable-optimizations --with-lto --with-static-libpython
    Host: 64 bit Arch Linux

    Benchmark bench_baseline2 bench_vectorcall2
    list(tuple) 303 ns 291 ns: 1.04x faster
    list_subclass 385 ns 395 ns: 1.02x slower
    bytes() 111 ns 89.5 ns: 1.24x faster
    bytes(int) 281 ns 224 ns: 1.26x faster
    bytearray() 210 ns 193 ns: 1.09x faster
    bytearray(bytes) 445 ns 394 ns: 1.13x faster
    bytearray(int) 190 ns 178 ns: 1.06x faster
    int() 73.5 ns 71.4 ns: 1.03x faster
    int(str) 175 ns 182 ns: 1.04x slower
    enumerate(list) 318 ns 329 ns: 1.04x slower
    Geometric mean (ref) 1.03x faster

    Benchmark hidden because not significant (13): list(), list(range), float(), float(int), float(str), str(), str(int), str(bytes,enc), tuple(), tuple(list), int(str,base), reversed(list), enumerate(list,start)

  9. added a commit that references this issue on Sep 10, 2026
  10. encukou commented on Sep 10, 2026

    @encukou
    Member

    Thank you for the improvement, @cmaloney!

    The docs for AC live in the devguide; would you mind making a PR there as well?

  11. added a commit that references this issue on Sep 12, 2026
  12. added a commit that references this issue on Sep 14, 2026
  13. cmaloney commented on Sep 22, 2026

    @cmaloney
    Contributor

    Working on the devguide updates + moving more types w/ pyperformance numbers. A number of benchmarks seemed to get slower on the day this was merged / September 10th (ex. https://speed.python.org/timeline/#/?exe=12&ben=async_tree_eager&env=1&revs=50&equid=off&quarts=on&extr=on). Investigating the root cause.

    Cleaned this up to no longer require static C type declarations and updated existing clinic tests. Adjusts the codegen type checks slightly but I think overall nicer: #157968

  14. StanFromIreland commented on Sep 25, 2026

    @StanFromIreland
    Member

    The async_tree_base benchmark on memory.python.org also shows a significant regression.

  15. cmaloney commented on Sep 25, 2026

    @cmaloney
    Contributor

    Bisected and memory + runtime increase seems to be from gh-157213, not this change. Left more details there.

  16. added 2 commits that reference this issue on Sep 26, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions