Skip to content

Add a module for monitoring GC statistics #146527

Description

@sergey-miryanov

Feature or enhancement

Proposal:

I propose adding a new internal module to read the gathered statistics from GC.
It is supposed to work out-of-process, so we can read and show statistics in non-intrusive way and without any pauses of the monitored process.

To achieve this goal, we need to follow these steps:

  1. Extend _PyDebugOffset with a pointer and size to the GC stats from _gc_runtime_state.
  2. Add more data to the GC stats(see below).
  3. Add a small module that connects to the process you're monitoring and displays statistics.

I have a working prototype.

I have added the following extra data to GC stats:

  • ts_start - the timestamp of the GC collect start
  • ts_stop - the timestamp of the GC collect stop
  • heap_size - the number of live objects.
  • work_to_do - an internal value from incremental GC. It controls the number of objects that are processed by one increment.
  • object_visits - the number of objects that were visited during GC work.
  • objects_transitively_reachable - the number of objects that were reachable from the global and local roots when incremental GC is called.
  • objects_not_transitively_reachable - the number of objects that were reachable while increment was constructed.

With this module we can build tools, that gathers stats in the "real-time" and output it to various formats. For example, we can format data to use in Perfetto UI (image from dpo post):

Image

I want to create two PRs, one for _PyDebugOffset and GC stats changes, and one for new module.

I want to start from GIL-enabled build.

cc @pablogsal @markshannon @nascheme @colesbury

Also link to the issue #131253

Has this already been discussed elsewhere?

No response given

Links to previous discussion of this feature:

https://discuss.python.org/t/add-a-module-for-monitoring-gc-statistics/106695

Linked PRs

Activity

  1. pablogsal commented on Mar 27, 2026

    @pablogsal
    Member

    I think this is fantastic! I am very excited this looks like a great idea. I think the only think I would prefer is not to introduce a new module but add support to this to the new Tachyon profiler as it already knows how to deal with all the debug offsets and it can already handle writing profiler-like outputs. We need to study either augmenting the existing profiles (as gecko format) or perhaps a separate mode where it just does this,

  2. sergey-miryanov commented on Mar 28, 2026

    @sergey-miryanov
    ContributorAuthor

    Thanks!

    I proposed a new module for the following reasons:

    1. We can build a lightweight sidecar to collect metrics from long-running process.
    2. We can build an application that gathers metrics from the list of long-running processes.

    I thought those cases are out of scope of Tachyon.

    I was only thinking about a simple module:

    static PyObject *
    _gc_monitor_handler_read(PyObject *op, PyObject *Py_UNUSED(ignored))
    {
        GCMonitorState *st = GCMonitor_GetStateFromType(Py_TYPE(op));
        GCMonitorHandler *h = GCMonitorHandler_CAST(op);
    
        PyThreadState *tstate = _PyThreadState_GET();
        struct _gc_runtime_state *gcstate = &tstate->interp->gc;
    
        uintptr_t interpreter_state_list_head =
            (uintptr_t)h->debug_offsets.runtime_state.interpreters_head;
    
        // TODO: all interpreters
        uintptr_t address_of_interpreter_state;
        if (_Py_RemoteDebug_ReadRemoteMemory(
                &h->handle,
                h->runtime_start_address + interpreter_state_list_head,
                sizeof(void*),
                &address_of_interpreter_state) < 0) {
            set_exception_cause(PyExc_RuntimeError, "Failed to read interpreter state address");
            return NULL;
        }
    
        if (address_of_interpreter_state == 0) {
            PyErr_SetString(PyExc_RuntimeError, "No interpreter state found");
            return NULL;
        }
    
        struct gc_stats stats;
        uintptr_t address = address_of_interpreter_state
            + h->debug_offsets.interpreter_state.gc
            + h->debug_offsets.gc.generation_stats;
        if (_Py_RemoteDebug_ReadRemoteMemory(&h->handle,
                                             address,
                                             h->debug_offsets.gc.generation_stats_size,
                                             &stats) < 0) {
            PyErr_SetString(PyExc_RuntimeError, "Failed to read GC state");
            return NULL;
        }
    
        PyObject *tuple = PyTuple_New(GC_YOUNG_STATS_SIZE + GC_OLD_STATS_SIZE * 2);
        if (tuple == NULL) {
            return NULL;
        }
    
        int index = 0;
        for(int gen = 0; gen < NUM_GENERATIONS; gen++) {
            struct gc_generation_stats **items;
            int size;
            if (gen == 0) {
                items = (struct gc_generation_stats **)&stats.young.items;
                size = GC_YOUNG_STATS_SIZE;
            }
            else {
                items = (struct gc_generation_stats **)&stats.old[gen-1].items;
                size = GC_OLD_STATS_SIZE;
            }
            for(int i = 0; i < size; i++, index++) {
                struct gc_generation_stats *stats_item = items[i];
                GCMonitorStatsItem *item = PyObject_New(GCMonitorStatsItem, st->GCMonitorStatsItem_Type);
                if (item == NULL) {
                    Py_DECREF(tuple);
                    return NULL;
                }
    
                item->ts_start = stats_item->ts_start;
                item->ts_stop = stats_item->ts_stop;
                item->gen = gen;
                item->collections = stats_item->collections;
                item->collected = stats_item->collected;
                item->uncollectable = stats_item->uncollectable;
                item->candidates = stats_item->candidates;
                item->object_visits = stats_item->object_visits;
                item->objects_transitively_reachable = stats_item->objects_transitively_reachable;
                item->objects_not_transitively_reachable = stats_item->objects_not_transitively_reachable;
                item->heap_size = stats_item->heap_size;
                item->work_to_do = stats_item->work_to_do;
    
                PyTuple_SET_ITEM(tuple, index, item);
            }
        }
    
        return tuple;
    }

    WDYT? Does it make sense?

  3. pablogsal commented on Mar 28, 2026

    @pablogsal
    Member

    A new module has a high bar. A top level module needs a PEP and a sub module needs a good story for future extension. If you really insist on this I think probably should be on the gc module itself ir perhaps some submodule of itself while the plumbing should probably exposed in the _remote_debugging.

    I don't think this gathers enough functionality by itself to be a sub module unless we start a bigger theme here but that required MUCH more thinking.

  4. pablogsal commented on Mar 28, 2026

    @pablogsal
    Member

    Thanks!

    I proposed a new module for the following reasons:

    1. We can build a lightweight sidecar to collect metrics from long-running process.
    2. We can build an application that gathers metrics from the list of long-running processes.

    I thought those cases are out of scope of Tachyon.

    I was only thinking about a simple module:

    static PyObject *
    _gc_monitor_handler_read(PyObject *op, PyObject *Py_UNUSED(ignored))
    {
        GCMonitorState *st = GCMonitor_GetStateFromType(Py_TYPE(op));
        GCMonitorHandler *h = GCMonitorHandler_CAST(op);
    
        PyThreadState *tstate = _PyThreadState_GET();
        struct _gc_runtime_state *gcstate = &tstate->interp->gc;
    
        uintptr_t interpreter_state_list_head =
            (uintptr_t)h->debug_offsets.runtime_state.interpreters_head;
    
        // TODO: all interpreters
        uintptr_t address_of_interpreter_state;
        if (_Py_RemoteDebug_ReadRemoteMemory(
                &h->handle,
                h->runtime_start_address + interpreter_state_list_head,
                sizeof(void*),
                &address_of_interpreter_state) < 0) {
            set_exception_cause(PyExc_RuntimeError, "Failed to read interpreter state address");
            return NULL;
        }
    
        if (address_of_interpreter_state == 0) {
            PyErr_SetString(PyExc_RuntimeError, "No interpreter state found");
            return NULL;
        }
    
        struct gc_stats stats;
        uintptr_t address = address_of_interpreter_state
            + h->debug_offsets.interpreter_state.gc
            + h->debug_offsets.gc.generation_stats;
        if (_Py_RemoteDebug_ReadRemoteMemory(&h->handle,
                                             address,
                                             h->debug_offsets.gc.generation_stats_size,
                                             &stats) < 0) {
            PyErr_SetString(PyExc_RuntimeError, "Failed to read GC state");
            return NULL;
        }
    
        PyObject *tuple = PyTuple_New(GC_YOUNG_STATS_SIZE + GC_OLD_STATS_SIZE * 2);
        if (tuple == NULL) {
            return NULL;
        }
    
        int index = 0;
        for(int gen = 0; gen < NUM_GENERATIONS; gen++) {
            struct gc_generation_stats **items;
            int size;
            if (gen == 0) {
                items = (struct gc_generation_stats **)&stats.young.items;
                size = GC_YOUNG_STATS_SIZE;
            }
            else {
                items = (struct gc_generation_stats **)&stats.old[gen-1].items;
                size = GC_OLD_STATS_SIZE;
            }
            for(int i = 0; i < size; i++, index++) {
                struct gc_generation_stats *stats_item = items[i];
                GCMonitorStatsItem *item = PyObject_New(GCMonitorStatsItem, st->GCMonitorStatsItem_Type);
                if (item == NULL) {
                    Py_DECREF(tuple);
                    return NULL;
                }
    
                item->ts_start = stats_item->ts_start;
                item->ts_stop = stats_item->ts_stop;
                item->gen = gen;
                item->collections = stats_item->collections;
                item->collected = stats_item->collected;
                item->uncollectable = stats_item->uncollectable;
                item->candidates = stats_item->candidates;
                item->object_visits = stats_item->object_visits;
                item->objects_transitively_reachable = stats_item->objects_transitively_reachable;
                item->objects_not_transitively_reachable = stats_item->objects_not_transitively_reachable;
                item->heap_size = stats_item->heap_size;
                item->work_to_do = stats_item->work_to_do;
    
                PyTuple_SET_ITEM(tuple, index, item);
            }
        }
    
        return tuple;
    }

    WDYT? Does it make sense?

    This is precisely why I want to reuse as much infrastructure from the remote debugging module: there are many important steps as you cannot just use your own offsets you need to copy the ones in the remote process and validate them to ensure all looks good, read the runtime, ...

    Also it's more complicated that that because you probably want some sort of iterator so you don't need to recalculate the addresses all the time without exposing the it internals....

  5. sergey-miryanov commented on Mar 28, 2026

    @sergey-miryanov
    ContributorAuthor

    If you really insist on this I think probably should be on the gc module

    It is a much better idea, than adding a module just for one or two functions. Now, I'm thinking that adding gc.get_stats_from_process(pid:int, all_interpreters:bool) -> list[dict[str, Any]] will be enough to build those applications that I mentioned above.

    the plumbing should probably exposed in the _remote_debugging.

    I'm glad to add needed functionality here and export it, to internal need from gcmodule for example.

    This is precisely why I want to reuse as much infrastructure from the remote debugging module

    I'm here on 100% with you. Doing this PoC I'm fully standing on your shoulders :)

    Also it's more complicated that that because you probably want some sort of iterator so you don't need to recalculate the addresses all the time without exposing the it internals

    C++ will be able us to make some type-safe visitor pattern for this :) But we have C. Maybe we can add some iterate_interpreters like `iterate_threads.

    If you don't mind, then instead of adding a new module I will add a get_stats_from_process function to the gcmodule and reuse maximum functionality from _remote_debugging.

  6. added a commit that references this issue on Mar 28, 2026
  7. added 2 commits that reference this issue on Apr 3, 2026
  8. added a commit that references this issue on Apr 16, 2026
  9. added 2 commits that reference this issue on Apr 25, 2026
  10. added a commit that references this issue on May 2, 2026
  11. 1 remaining item

  12. pablogsal commented on May 2, 2026

    @pablogsal
    Member

    We probably now want to re-export the function in the gc module as a public interface since the remote debugging module is currently private. @sergey-miryanov what do you think?

  13. added a commit that references this issue on May 2, 2026
  14. sergey-miryanov commented on May 2, 2026

    @sergey-miryanov
    ContributorAuthor

    @pablogsal I didn't intend to make this public in 3.15. Better to collect usage patterns now, finalize the API in 3.16 without triggering the public deprecation policy.

  15. pablogsal commented on May 2, 2026

    @pablogsal
    Member

    @pablogsal I didn't intend to make this public in 3.15. Better to collect usage patterns now, finalize the API in 3.16 without triggering the public deprecation policy.

    That works for me

  16. added 2 commits that reference this issue on Jun 5, 2026
  17. added a commit that references this issue on Jun 17, 2026
  18. added 6 commits that reference this issue on Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    interpreter-core(Objects, Python, Grammar, and Parser dirs)type-featureA feature request or enhancement

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions