Skip to content

Poor thread scaling when constructing instances or accessing attributes #139103

Description

@JukkaL

Bug report

Remaining scaling bugs

  • dataclass
  • namedtuple
  • enum

Bug description:

When constructing dataclass or NamedTuple instances on multiple threads (on a free threading build), or accessing enum class attributes, performance doesn't scale when using multiple threads.

Regular class example (scales well):

# b_regular_class.py
from threading import Thread
from time import time
import sys

class Foo:
    def __init__(self, x):
        self.x = x

niter = 5 * 1000 * 1000

def benchmark(n):
    for i in range(n):
        Foo(x=1)

for nth in (1, 4):
    t0 = time()
    threads = [Thread(target=benchmark, args=(niter,)) for _ in range(nth)]
    for t in threads:
        t.start()
    for t in threads:
        t.join()
    print(f"{nth=} {(time() - t0) / nth}")

Dataclass example (doesn't scale well):

# b_dataclass.py
from threading import Thread
from dataclasses import dataclass
from time import time
import sys

@dataclass
class Foo:
    x: int

niter = 5 * 1000 * 1000

def benchmark(n):
    for i in range(n):
        Foo(x=1)

for nth in (1, 4):
    t0 = time()
    threads = [Thread(target=benchmark, args=(niter,)) for _ in range(nth)]
    for t in threads:
        t.start()
    for t in threads:
        t.join()
    print(f"{nth=} {(time() - t0) / nth}")

Named tuple example (doesn't scale well):

# b_namedtuple.py
from threading import Thread
from typing import NamedTuple
from time import time
import sys

class Foo(NamedTuple):
    x: int

niter = 5 * 1000 * 1000

def benchmark(n):
    for i in range(n):
        Foo(x=1)

for nth in (1, 4):
    t0 = time()
    threads = [Thread(target=benchmark, args=(niter,)) for _ in range(nth)]
    for t in threads:
        t.start()
    for t in threads:
        t.join()
    print(f"{nth=} {(time() - t0) / nth}")

Enum example (doesn't scale well):

# b_enum.py
from threading import Thread
from time import time
from enum import Enum
import sys

class Foo(Enum):
    X = 1
    Y = 2

niter = 5 * 1000 * 1000

def benchmark(n):
    for i in range(n):
        Foo.X
        Foo.Y.value

for nth in (1, 4):
    t0 = time()
    threads = [Thread(target=benchmark, args=(niter,)) for _ in range(nth)]
    for t in threads:
        t.start()
    for t in threads:
        t.join()
    print(f"{nth=} {(time() - t0) / nth}")

Results on recent main branch (running on an EC2 instance):

(cpython-dev) jukka@jukka-coder-dbx free-threading-benchmarks $ py b_regular_class.py
nth=1 1.1085155010223389
nth=4 0.2796591520309448
(cpython-dev) jukka@jukka-coder-dbx free-threading-benchmarks $ py b_dataclass.py
nth=1 1.1910037994384766
nth=4 1.0931583642959595
(cpython-dev) jukka@jukka-coder-dbx free-threading-benchmarks $ py b_namedtuple.py
nth=1 1.5688557624816895
nth=4 2.0257126092910767
(cpython-dev) jukka@jukka-coder-dbx free-threading-benchmarks $ py b_enum.py
nth=1 0.9439797401428223
nth=4 2.272495985031128

The expected behavior is that when using 4 threads (nth=4), the elapsed time per benchmark iteration (the second printed value) goes down significantly compared to when using a single thread (nth=1), which happens with the first benchmark (b_regular_class.py) but not the others.

cc @colesbury (we discussed this at CPython Core Dev Sprint in person)

CPython versions tested on:

CPython main branch

Operating systems tested on:

Linux

Linked PRs

Activity

  1. added
    type-bugAn unexpected behavior, bug, or error
    on Sep 18, 2025
  2. changed the title [-]Poor thread scaling when constructing instances or accessings attributes[/-] [+]Poor thread scaling when constructing instances or accessing attributes[/+] on Sep 18, 2025
  3. colesbury commented on Sep 22, 2025

    @colesbury
    Contributor

    I've started looking into this. The following is mostly for my own future reference, but may be useful if anyone else wants to take a look at this as well.

    Here are the above examples added to Tools/ftscalingbench/ftscalingbench.py: https://gist.github.com/colesbury/65234f8e8d9afad361981255d8259d7f

    Dataclasses: there's reference count contention on the __init__ function. I think the function is created by calling exec() on a string, so the __init__ function doesn't have deferred reference counting enabled. I think we can fix this by enabling deferred reference counting when the function is set on the type object in type_setattro.

    typing.NamedTuple: Calling slot_tp_new doesn't scale well. PyObject_GetAttr() increments the refcount and the Py_DECREF decrements it. We want something like _PyObject_GetMethodStackRef() here that uses stackrefs so that we avoid reference count contention.

    cpython/Objects/typeobject.c

    Lines 10842 to 10856 in f0d8583

    static PyObject *
    slot_tp_new(PyTypeObject *type, PyObject *args, PyObject *kwds)
    {
    PyThreadState *tstate = _PyThreadState_GET();
    PyObject *func, *result;
    func = PyObject_GetAttr((PyObject *)type, &_Py_ID(__new__));
    if (func == NULL) {
    return NULL;
    }
    result = _PyObject_Call_Prepend(tstate, func, (PyObject *)type, args, kwds);
    Py_DECREF(func);
    return result;
    }

    Enum: I think there's reference count contention on instances of enum.property:

    class property(DynamicClassAttribute):

  4. LindaSummer commented on Oct 29, 2025

    @LindaSummer
    Contributor

    Hi @colesbury and @JukkaL ,

    I'm interested in this topic. Could I have a try? 😊

    Best Regards,
    Edward

  5. LindaSummer commented on Nov 6, 2025

    @LindaSummer
    Contributor

    Hi @colesbury ,

    Sorry to bother you for this issue. 😊

    I have read the dataclass code and find that it generated below code and runs with exec.

    def __create_fn__(__dataclass_type_x__,__dataclass_HAS_DEFAULT_FACTORY__,__dataclass_builtins_object__,__dataclass___init___return_type__,__dataclasses_recursive_repr):
     def __init__(self,x:__dataclass_type_x__)->__dataclass___init___return_type__:
      self.x=x
     @__dataclasses_recursive_repr()
     def __repr__(self):
      return f"{self.__class__.__qualname__}(x={self.x!r})"
     def __eq__(self,other):
      if self is other:
       return True
      if other.__class__ is self.__class__:
       return self.x==other.x
      return NotImplemented
     return (__init__,__repr__,__eq__,)

    Here is the reference code snippts.

    cpython/Lib/dataclasses.py

    Lines 492 to 524 in 2e5e6fd

    # txt is the entire function we're going to execute, including the
    # bodies of the functions we're defining. Here's a greatly simplified
    # version:
    # def __create_fn__():
    # def __init__(self, x, y):
    # self.x = x
    # self.y = y
    # @recursive_repr
    # def __repr__(self):
    # return f"cls(x={self.x!r},y={self.y!r})"
    # return __init__,__repr__
    txt = f"def __create_fn__({local_vars}):\n{fns_src}\n return {return_names}"
    ns = {}
    exec(txt, self.globals, ns)
    fns = ns['__create_fn__'](**self.locals)
    # Now that we've generated the functions, assign them into cls.
    for name, fn in zip(self.names, fns):
    fn.__qualname__ = f"{cls.__qualname__}.{fn.__name__}"
    if self.unconditional_adds.get(name, False):
    setattr(cls, name, fn)
    else:
    already_exists = _set_new_attribute(cls, name, fn)
    # See if it's an error to overwrite this particular function.
    if already_exists and (msg_extra := self.overwrite_errors.get(name)):
    error_msg = (f'Cannot overwrite attribute {fn.__name__} '
    f'in class {cls.__name__}')
    if not msg_extra is True:
    error_msg = f'{error_msg} {msg_extra}'
    raise TypeError(error_msg)

    So I create a case without dataclass below.

    from threading import Thread
    from time import time
    
    def test_1():
        class Foo:
            def __init__(self, x):
                self.x = x
    
        niter = 5 * 1000 * 1000
    
        def benchmark(n):
            for i in range(n):
                Foo(x=1)
    
        for nth in (1, 4):
            t0 = time()
            threads = [Thread(target=benchmark, args=(niter,)) for _ in range(nth)]
            for t in threads:
                t.start()
            for t in threads:
                t.join()
            print(f"{nth=} {(time() - t0) / nth}")
            
    def test_2():
        class Foo2:
            def __init__(self, x):
                    pass
            pass
        
        _Foo2_x = int
        
        create_str = """def create_init(_Foo2_x,):
            def __init__(self, x: _Foo2_x):
                self.x = x
            return (__init__,)
        """
        ns = {}
        exec(create_str, globals(), ns)
        fn = ns['create_init']({**locals()})
        setattr(Foo2, '__init__', fn[0])
        niter = 5 * 1000 * 1000
        def benchmark(n):
            for i in range(n):
                Foo2(x=1)
        
        for nth in (1, 4):
            t0 = time()
            threads = [Thread(target=benchmark, args=(niter,)) for _ in range(nth)]
            for t in threads:
                t.start()
            for t in threads:
                t.join()
            print(f"{nth=} {(time() - t0) / nth}")
            
    if __name__ == "__main__":
        print("------test_1-------")
        test_1()
        print("------test_2-------")
        test_2()

    But got below result.

    ------test_1-------
    nth=1 6.3677897453308105
    nth=4 2.0216556191444397
    ------test_2-------
    nth=1 6.5973052978515625
    nth=4 2.26946008205413
    

    It seems that this case follow the dataclass logic to exec string and generate an __init__.
    But the performance didn't change.

    I use below configure to build cpython on main branch 1697cb5 .

    ./configure --with-pydebug --disable-gil "CC=clang"

    Could you help correct me on the understanding of dataclass's case? 😊

    Wish you a good day!

    Best Regards,
    Edward

  6. colesbury commented on Nov 6, 2025

    @colesbury
    Contributor

    Don't use --with-pydebug when benchmarking:

    uv run -p 3.14t python example.py:

    ------test_1-------
    nth=1 1.1103649139404297
    nth=4 0.2983349561691284
    ------test_2-------
    nth=1 1.1990399360656738
    nth=4 1.652941882610321
    
  7. LindaSummer commented on Nov 6, 2025

    @LindaSummer
    Contributor

    Don't use --with-pydebug when benchmarking:

    uv run -p 3.14t python example.py:

    ------test_1-------
    nth=1 1.1103649139404297
    nth=4 0.2983349561691284
    ------test_2-------
    nth=1 1.1990399360656738
    nth=4 1.652941882610321
    

    Thanks very much for your kind help and patience!
    I will try to solve this issue. 😊

  8. eendebakpt commented on Nov 15, 2025

    @eendebakpt
    Contributor

    typing.NamedTuple: Calling slot_tp_new doesn't scale well. PyObject_GetAttr() increments the refcount and the Py_DECREF decrements it. We want something like _PyObject_GetMethodStackRef() here that uses stackrefs so that we avoid reference count contention.

    @colesbury In #141603 I used the _PyObject_GetMethodStackRef approach, but the scaling performance is unchanged. Not sure yet why that is.

  9. eendebakpt commented on Nov 16, 2025

    @eendebakpt
    Contributor

    A small update: for

    from collections import namedtuple        
    Bar= namedtuple('Bar', ('x'))
    

    the call B(1) is specialized to _CALL_NON_PY_GENERAL. Maybe that does not scale.

    Even with the modified slot_tp_new the scaling of getting an attribute is quite poor:

    Bar = namedtuple('Bar', ['x'])
    
    class Normal:
        pass
    
    @register_benchmark
    def attribute_normal_class():
        for i in range(1000 * WORK_SCALE):
            Normal.__new__
    
    @register_benchmark
    def attribute_namedtuple():
        for i in range(1000 * WORK_SCALE):
            Bar.__new__
    
    attribute_normal_class     3.6x faster
    attribute_namedtuple       3.9x slower
    
  10. added 3 commits that reference this issue on Nov 19, 2025
  11. 7 remaining items

  12. added 3 commits that reference this issue on Jan 29, 2026
  13. added a commit that references this issue on Feb 3, 2026
  14. added a commit that references this issue on Feb 3, 2026
  15. added 4 commits that reference this issue on Feb 3, 2026
  16. added 3 commits that reference this issue on Feb 15, 2026
  17. added a commit that references this issue on Feb 18, 2026
  18. added a commit that references this issue on Apr 25, 2026
  19. johng commented on Aug 16, 2026

    @johng
    Contributor

    Also see this on various class/instance attributes in addition to the above PR, shall I be putting these in a new issue?

    class MyClassWithInstanceAttr:
        def __init__(self):
            self.attr = object()
    
    # A single instance shared between threads. Reading its attribute (stored in
    # the inline __dict__ values, LOAD_ATTR_INSTANCE_VALUE) contends on the shared
    # value's reference count on every read.
    _shared_instance = MyClassWithInstanceAttr()
    
    @register_benchmark
    def shared_instance_attribute():
        obj = _shared_instance
        for _ in range(1000 * WORK_SCALE):
            obj.attr
            obj.attr
            obj.attr
    

    ./python_main.exe Tools/ftscalingbench/ftscalingbench.py shared_instance_attribute
    Running benchmarks with 18 threads
    shared_instance_attribute 25.0x slower

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions