Skip to content

JIT: Implement unique reference tracking in Tier 2 for reference count optimizations #143414

Description

@cocolato

Feature or enhancement

Proposal:

Motivation

We should implement unique reference tracking in Tier 2 to facilitate optimizations that reduce reference counting overhead. For example, when a tuple is known to be uniquely referenced, we can "steal" its element references during unpacking without performing any reference counting operations.

For reference: discussion in #142952

Technical Approach

  1. Reference Tracking Infrastructure
    • Add an REF_IS_UNIQUE bit (bit 1) to the JitOptRef union in pycore_optimizer.h (code reference).
    • Implement PyJitRef_MakeUnique() and PyJitRef_IsUnique() helper functions.
    • Update helper utilities including PyJitRef_StripReferenceInfo and JIT_BITS_TO_PTR_MASKED to support this unique reference bit.
  2. Apply unique reference tracking to UNPACK_SEQUENCE uops
  3. Expand support to more uops
    • After verifying performance and correctness, extend the use of unique reference tracking to additional uops and optimizations as identified.

Has this already been discussed elsewhere?

No response given

Links to previous discussion of this feature:

No response

Linked PRs

Activity

  1. cocolato commented on Jan 4, 2026

    @cocolato
    MemberAuthor

    @Fidget-Spinner Hi, Could you please review this issue and let me know your thoughts?

  2. Fidget-Spinner commented on Jan 4, 2026

    @Fidget-Spinner
    Member

    @cocolato seems correct to me. The key thing about the optimization is that

            op(_UNPACK_SEQUENCE_TWO_TUPLE, (seq -- val1, val0)) {
                assert(oparg == 2);
                PyObject *seq_o = PyStackRef_AsPyObjectBorrow(seq);
                assert(PyTuple_CheckExact(seq_o));
                DEOPT_IF(PyTuple_GET_SIZE(seq_o) != 2);
                STAT_INC(UNPACK_SEQUENCE, hit);
                val0 = PyStackRef_FromPyObjectNew(PyTuple_GET_ITEM(seq_o, 0));
                val1 = PyStackRef_FromPyObjectNew(PyTuple_GET_ITEM(seq_o, 1));
                PyStackRef_CLOSE(seq);
            }
    

    becomes

            op(_UNPACK_SEQUENCE_TWO_TUPLE_UNIQUE_STEAL, (seq -- val1, val0)) {
                assert(oparg == 2);
                PyObject *seq_o = PyStackRef_AsPyObjectBorrow(seq);
                assert(PyTuple_CheckExact(seq_o));
                DEOPT_IF(PyTuple_GET_SIZE(seq_o) != 2);
                STAT_INC(UNPACK_SEQUENCE, hit);
                val0 = PyStackRef_FromPyObjectSteal(PyTuple_GET_ITEM(seq_o, 0));
                PyTuple_SET_ITEM(seq_o, 0, NULL);
                val1 = PyStackRef_FromPyObjectSteal(PyTuple_GET_ITEM(seq_o, 1));
                PyTuple_SET_ITEM(seq_o, 1, NULL);
                PyStackRef_CLOSE_NO_ESCAPE(seq);
            }
    

    Which will allow the op to have no reference counts operations at all, and also be non-escaping. So we can stack cache over it.

  3. cocolato commented on Jan 4, 2026

    @cocolato
    Author
  4. Fidget-Spinner commented on Jan 4, 2026

    @Fidget-Spinner
    Member

    @cocolato now that I think about this, this is more complicated than I thought:

    The problem is you have to invalidate all unique references on any escaping uop, like see (_PyUop_Flags[opcode] & HAS_ESCAPES_FLAG). RETURN_VALUE is an escaping uop, so we cannot easily apply this optimization across RETURN_VALUE.

    However, I think there's one useful place you can still apply this optimization: object creation and initialization. See the CALL_ALLOC_AND_ENTER_INIT instruction for example.

    I think because of how complicated this is, I'll take over, sorry! However, I need your help on making this optimization better and can parallelize some work here: we need to specialize on more forms of object creation. Can you please add a specialization for CALL_SLOT_AND_ENTER_INIT?

    Basically the current code JITs:

    class A:
        def __init__(self, a):
            self.a = a
    
    
    def foo(n):
        for i in range(1, n + 1):
            x = A(1)
        return 1
    
    foo(4003)
    

    But once you add __slots__ to A, there's no more specialization, and the JIT cannot optimize and just fails. So we need a new specialization for calling __slots__. This turns up frequently in dataclasses and also the bm_float benchmark on pyperformance.

    Just by adding CALL_SLOT_AND_ENTER_INIT, you should see a speedup on that benchmark. This isn't an easy task, and I think you're a very capable/strong contributor, so I'm entrusting you with this. Do you mind taking it up? You can see an example of how to do add a specialization here #143389

  5. Fidget-Spinner commented on Jan 4, 2026

    @Fidget-Spinner
    Member

    Sorry for discouraging you from this optimization btw. I feel with the current state of the JIT, we can't use this (yet). If you manage to implement more specializations, that should make it possible though.

  6. Fidget-Spinner commented on Jan 4, 2026

    @Fidget-Spinner
    Member

    We can still do the optimizations, but we may have to insert extra guards. However, that's blocked by #143421 as well.

    Let me know which one you want to work on, and I can let you take it up! The latter got taken up by Donghee, so it will have to be the new specialization!

  7. cocolato commented on Jan 5, 2026

    @cocolato
    MemberAuthor

    add a specialization for CALL_SLOT_AND_ENTER_INIT

    Thanks for the explanation! I'm pleased to take on the task of adding a specialization for CALL_SLOT_AND_ENTER_INIT. I'm also very willing to help with other related development tasks in the future—Truly I need to start with simpler tasks to get more familiar with this area of optimization. Thanks for the trust!

  8. cocolato commented on Jan 6, 2026

    @cocolato
    Author
  9. Fidget-Spinner commented on Jan 6, 2026

    @Fidget-Spinner
    Member

    @cocolato I'm surprised, you're right. Really sorry, my mistake. I have another issue to work on. Can you please work on making the JIT optimization state per-thread?

    #143421 (comment)

    The main one is JitOptContext is stack allocated https://github.com/python/cpython/blob/main/Python/optimizer_analysis.c#L342

    We need it to be pushed to _PyThreadStateImpl https://github.com/python/cpython/blob/main/Include/internal/pycore_tstate.h#L149

    So that we don't stack overflow on the JIT. This time I've verified that the problem indeed exists. Do you want to take this one up?

  10. cocolato commented on Jan 6, 2026

    @cocolato
    MemberAuthor

    Ok, I will take some time to understand the issue and work on it!

  11. Fidget-Spinner commented on Jan 6, 2026

    @Fidget-Spinner
    Member

    Please don't work on the parent issue (Make the JIT optimizer buffer add to a new buffer, not in-place), instead it's the sub-issue, thanks!

  12. cocolato commented on Jan 8, 2026

    @cocolato
    MemberAuthor

    @Fidget-Spinner Hi, are there some other tasks we can do for this issue?

  13. 2 remaining items

  14. Fidget-Spinner commented on Jan 9, 2026

    @Fidget-Spinner
    Member

    @cocolato can I assign you this to do before that though? We first need the buffer to append instead of overwrite in place if we want to do this optimization properly.

    #143421

  15. cocolato commented on Jan 9, 2026

    @cocolato
    MemberAuthor

    @cocolato can I assign you this to do before that though? We first need the buffer to append instead of overwrite in place if we want to do this optimization properly.

    #143421

    Sure, I will work on this!

  16. Fidget-Spinner commented on Jan 17, 2026

    @Fidget-Spinner
    Member

    @cocolato We need to first implement a new symbolic type for objects. See optimizer_symbols.c and the things that are done there. I can let you take that up if you want.

    Alright, you can do this now, let me know if you face any issues. Thanks!

    I suspect we want to start simple. So just start with a symbolic type for objects with __slots__, then you just need to maintain an index to symbol mapping.

    Recall that we do not want allocations in the JIT optimizer if we can avoid it. So I think it's okay to just store it as a pair in an array and do a linear search for now.

  17. cocolato commented on Jan 20, 2026

    @cocolato
    MemberAuthor

    I suspect we want to start simple. So just start with a symbolic type for objects with __slots__, then you just need to maintain an index to symbol mapping.

    static_assert(sizeof(JitOptSymbol) <= 3 * sizeof(uint64_t), "JitOptSymbol has grown");

    Why do we have this limitation here? If we now need to extend JitOptSymbol to add the JitOptSlotsObject type, this might restrict the length of the slot array we're tracking.

  18. Fidget-Spinner commented on Jan 20, 2026

    @Fidget-Spinner
    Member

    @cocolato you can increase the size of the assert.

  19. Fidget-Spinner commented on Jan 20, 2026

    @Fidget-Spinner
    Member

    @cocolato on second thought, let's not change the size of that assert. That assert affects the size of all symbols. A possible solution without increasing the size is to contain a single pointer that points to an arena (see for example the t_arena and s_arena in the abstract optimizer struct). The arena itself will then be shared across all object types, and contain the slots/attributes for objects.

  20. Fidget-Spinner commented on Jan 24, 2026

    @Fidget-Spinner
    Member

    @cocolato sorry, do you mind if I let @reidenong do unique reference tracking while you're doing object property tracking? I think object property tracking will gain us huge immediate benefits. However, it's also quite large and complex. So I think you'll be occupied with it for awhile. In the meantime, we can parallelize some of this work.

  21. cocolato commented on Jan 24, 2026

    @cocolato
    MemberAuthor

    @cocolato sorry, do you mind if I let @reidenong do unique reference tracking while you're doing object property tracking? I think object property tracking will gain us huge immediate benefits. However, it's also quite large and complex. So I think you'll be occupied with it for awhile. In the meantime, we can parallelize some of this work.

    Thank you for letting me know! I don't mind at all. I will focus on object property tracking(but in the coming days, I may also have little time for development), and I'm sure I can learn a lot from other people's PRs in the meantime.

  22. eendebakpt commented on Mar 21, 2026

    @eendebakpt
    Contributor

    @reidenong @Fidget-Spinner I used the implementation of unique reference tracking from the PR to add inplace float operations to the jit. Results look promising on microbenchmarks:

    Expression              Speedup
    total += a*b + c:          2.1x
    total += a + b :           1.5x
    total += a*b + c*d:        1.6x
    

    On nbody I measured a 15% speedup. Once the PR is merged I will rebase my work and do some more benchmarking.

    Update: PR is at #146307

  23. added a commit that references this issue on Mar 22, 2026
  24. Fidget-Spinner commented on Mar 22, 2026

    @Fidget-Spinner
    Member

    @eendebakpt yeah that inplace op using refcnt==1 essentially was around in the older days (3.11-3.13) in the interpreter but removed due to it being not FT safe. It would be great to have that back in the JIT! I've merged the unique reference tracking PR, so please feel free to open a PR for float ops. Thanks!

  25. Fidget-Spinner commented on Mar 24, 2026

    @Fidget-Spinner
    Member

    I guess any further issues can be done as follow-up. Thanks Hai Zhu and Reiden!

  26. added a commit that references this issue on Apr 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    interpreter-core(Objects, Python, Grammar, and Parser dirs)topic-JITtype-featureA feature request or enhancement

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions