Repository navigation
Generate vectorcall code to parse arguments using Argument Clinic #87613
Description
Activity
To optimize the creation of objects, a lot of "tp_new" methods are defined twice: once in the legacy way (tp_new slot), once with the new VECTORCALL calling convention (tp_vectorcall slot). My concern is that the VECTORCALL implementation copy/paste most of the code just to parse arguments, whereas the specific code is just a few lines.
Example with the float type constructor:
--------------static PyObject * float_new(PyTypeObject *type, PyObject *args, PyObject *kwargs) { PyObject *return_value = NULL; PyObject *x = NULL; if ((type == &PyFloat_Type) && !_PyArg_NoKeywords("float", kwargs)) { goto exit; } if (!_PyArg_CheckPositional("float", PyTuple_GET_SIZE(args), 0, 1)) { goto exit; } if (PyTuple_GET_SIZE(args) < 1) { goto skip_optional; } x = PyTuple_GET_ITEM(args, 0); skip_optional: return_value = float_new_impl(type, x); exit: return return_value; } /*[clinic input] @classmethod float.\_\_new__ as float_new x: object(c_default="NULL") = 0 / Convert a string or number to a floating point number, if possible. [clinic start generated code]*/ static PyObject * float_new_impl(PyTypeObject *type, PyObject *x) /*[clinic end generated code: output=ccf1e8dc460ba6ba input=f43661b7de03e9d8]*/ { if (type != &PyFloat_Type) { if (x == NULL) { x = _PyLong_GetZero(); } return float_subtype_new(type, x); /* Wimp out */ } if (x == NULL) { return PyFloat_FromDouble(0.0); } /* If it's a string, but not a string subclass, use PyFloat_FromString. */ if (PyUnicode_CheckExact(x)) return PyFloat_FromString(x); return PyNumber_Float(x); } static PyObject * float_vectorcall(PyObject *type, PyObject * const*args, size_t nargsf, PyObject *kwnames) { if (!_PyArg_NoKwnames("float", kwnames)) { return NULL; } Py_ssize_t nargs = PyVectorcall_NARGS(nargsf); if (!_PyArg_CheckPositional("float", nargs, 0, 1)) { return NULL; } PyObject *x = nargs >= 1 ? args[0] : NULL; return float_new_impl((PyTypeObject *)type, x); }
Here the float_new() function (tp_new slot) is implemented with Argument Clinic: float_new() C code is generated from the [clinic input] DSL: good!
My concern is that float_vectorcall() code is hand written, it's boring to write and boring to maintain.
Would it be possible to add a new [clinic input] DSL for vectorcall? I expect something like that:
--------------static PyObject * float_vectorcall(PyObject *type, PyObject * const*args, size_t nargsf, PyObject *kwnames) { if (!_PyArg_NoKwnames("float", kwnames)) { return NULL; } Py_ssize_t nargs = PyVectorcall_NARGS(nargsf); if (!_PyArg_CheckPositional("float", nargs, 0, 1)) { return NULL; } PyObject *x = nargs >= 1 ? args[0] : NULL; return float_vectorcall_impl(type, x); } static PyObject * float_vectorcall_impl(PyObject *type, PyObject *x) { return float_new_impl((PyTypeObject *)type, x); }
where float_vectorcall() C code would be generated, and float_vectorcall_impl() would be the only part written manually. float_vectorcall_impl() gets a clean API and its body is way simpler to write and to maintain!
I agree with we should update AC to generate vectorcall.
I am going to investigate what we can :)I don't have the bandwidth to work on this issue, so I just close it.
- addedtype-featureA feature request or enhancementA feature request or enhancementand removed3.10 (EOL)end of lifeend of life
on Aug 17, 2023 64 remaining items
Have a working implementation with performance matching or beating the currently handwritten implementations. Benchmark and full numbers below. Overall 1.03x faster.
First PR: GH-145381. The changes I have split across a couple commits in my draft ac_vectorcall_v1 branch. Did a PR for just adding the support to AC; then not sure what would be best for the other changes (leaning: PR per type for easy reverts if needed?).
- Adding
@vectorcalldecorator and codegen with test - Commit per type, replacing hand-coded vectorcall implementations with AC generated ones
- Adding
@vectorcallto additional types. - Microbenchmark for different constructions. Not planning to try and get into CPython; think it could be useful to track perf over time but not sure a good place for it to live.
Performance numbers
Benchmark code: cmaloney@ddcd3b6
Configure options:./configure --enable-optimizations --with-lto --with-static-libpython
Host: 64 bit Arch LinuxBenchmark bench_baseline2 bench_vectorcall2 list(tuple) 303 ns 291 ns: 1.04x faster list_subclass 385 ns 395 ns: 1.02x slower bytes() 111 ns 89.5 ns: 1.24x faster bytes(int) 281 ns 224 ns: 1.26x faster bytearray() 210 ns 193 ns: 1.09x faster bytearray(bytes) 445 ns 394 ns: 1.13x faster bytearray(int) 190 ns 178 ns: 1.06x faster int() 73.5 ns 71.4 ns: 1.03x faster int(str) 175 ns 182 ns: 1.04x slower enumerate(list) 318 ns 329 ns: 1.04x slower Geometric mean (ref) 1.03x faster Benchmark hidden because not significant (13): list(), list(range), float(), float(int), float(str), str(), str(int), str(bytes,enc), tuple(), tuple(list), int(str,base), reversed(list), enumerate(list,start)
Reacted by Sergey Miryanov, Donghee Na, Chris Eibl and Erlend E. Aasland- Adding
Thank you for the improvement, @cmaloney!
The docs for AC live in the devguide; would you mind making a PR there as well?
Reacted by Cody MaloneyWorking on the devguide updates + moving more types w/ pyperformance numbers. A number of benchmarks seemed to get slower on the day this was merged / September 10th (ex. https://speed.python.org/timeline/#/?exe=12&ben=async_tree_eager&env=1&revs=50&equid=off&quarts=on&extr=on). Investigating the root cause.
Cleaned this up to no longer require static C type declarations and updated existing clinic tests. Adjusts the codegen type checks slightly but I think overall nicer: #157968
The
async_tree_basebenchmark onmemory.python.orgalso shows a significant regression.Bisected and memory + runtime increase seems to be from gh-157213, not this change. Left more details there.
- addedinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Sep 26, 2026
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields:
Linked PRs