Skip to content

test(runner-shared): benchmark finding module events in a memtrack artifact - #576

Open
not-matthias wants to merge 3 commits into
mainfrom
cod-3819-memtrack-investigate-8-minutes-spent-in-teardown
Open

not-matthias wants to merge 3 commits into
mainfrom
cod-3819-memtrack-investigate-8-minutes-spent-in-teardown

Conversation

@not-matthias

@not-matthias not-matthias commented Oct 9, 2026 •

Copy link
Copy Markdown
Member

TLDR: Since 5.4.0, the memory executor decodes the whole memtrack artifact after a run just to keep its few Mapping, Fork and Exec events. On artifacts with tens of millions of events, this adds minutes between memtrack's last log line and the upload. This PR only adds a benchmark for that search; the optimization is stacked on top in #578.

  • Move the filter from read_mappings_from_artifact into MemtrackArtifact::decode_module_events, so the bench measures the code the runner runs. No behavior change.
  • Add the memtrack_reader bench. Its artifact is laid out like memtrack writes one: allocations, forks and execs spread through the stream, and 64 mappings at the end. It runs without stacks and with an 8 KiB stack record every 400 events, and asserts that all module events are found.
  • It runs at 1M events. Larger sizes take minutes per case with the current decoder under simulation; perf(runner-shared): decode memtrack frames for module events in parallel #578 adds them back.
  • The artifact is generated lazily, written to a file under the target directory and mapped, so larger sizes never have to fit in memory.

Baseline, local walltime on a 32-thread machine: 258 ms without stacks, 267 ms with stacks.

@codspeed

codspeed Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Merging this PR will not alter performance

⚠️ Unknown Walltime execution environment detected

Using the Walltime instrument on standard Hosted Runners will lead to inconsistent data.

For the most accurate results, we recommend using CodSpeed Macro Runners: bare-metal machines fine-tuned for performance measurement consistency.

✅ 31 untouched benchmarks
🆕 6 new benchmarks
⏩ 6 skipped benchmarks1

Performance Changes

Mode Benchmark BASE HEAD Efficiency
🆕 Memory find_module_events_with_stacks[1000000] N/A 1.1 MB N/A
🆕 Memory find_module_events[1000000] N/A 1.1 MB N/A
🆕 Simulation find_module_events_with_stacks[1000000] N/A 2.5 s N/A
🆕 Simulation find_module_events[1000000] N/A 2.5 s N/A
🆕 WallTime find_module_events_with_stacks[1000000] N/A 1.3 s N/A
🆕 WallTime find_module_events[1000000] N/A 1.3 s N/A

Comparing cod-3819-memtrack-investigate-8-minutes-spent-in-teardown (2f1f983) with main (90cc8e9)

Open in CodSpeed

Footnotes

  1. 6 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩

@not-matthias

Copy link
Copy Markdown
Member Author

@codspeedbot suggest optimizations for find_module_events_with_stacks. sort by largest estimated speedup. estimate the diff size/complexity. no need to run anything. just explore the flamegraphs

@codspeed

This comment was marked as outdated.

@not-matthias
not-matthias force-pushed the cod-3819-memtrack-investigate-8-minutes-spent-in-teardown branch from 2bbb062 to 7044982 Compare October 9, 2026 11:20
@not-matthias
not-matthias added this pull request to stack #579 October 9, 2026 12:39
@not-matthias
not-matthias force-pushed the cod-3819-memtrack-investigate-8-minutes-spent-in-teardown branch from 7044982 to 4bd93bd Compare October 9, 2026 12:43
@not-matthias
not-matthias marked this pull request as ready for review October 9, 2026 12:57
@not-matthias
not-matthias force-pushed the cod-3819-memtrack-investigate-8-minutes-spent-in-teardown branch from 4bd93bd to 363ca54 Compare October 9, 2026 12:58
@greptile-apps

greptile-apps Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

[Medium impact] The PR appears safe to merge; no actionable issues were found.

Summary

Adds memtrack_reader benchmarks for finding module events in artifacts with and without captured stacks.

  • Both cases generate 1,048,576 events, write them to a temporary file, and check the number of module events.
  • Moves the existing filter into MemtrackArtifact::decode_module_events so the benchmark measures the runner’s code.
  • No actionable issues found. Tests and benchmarks were not executed.

Reviews (1) · Last reviewed commit: 4bd93bd · Reviewed by Greptile

@not-matthias

Copy link
Copy Markdown
Member Author

Note: SIZES currently only has 1 value, since larger values would completely timeout the CI. I'm adding larger values in the stacked PR.

@not-matthias
not-matthias force-pushed the cod-3819-memtrack-investigate-8-minutes-spent-in-teardown branch from 363ca54 to c85a5ef Compare October 9, 2026 13:59
…oder

The runner filtered the full memtrack event stream down to the Mapping,
Fork and Exec events inside read_mappings_from_artifact. Moving that filter
into MemtrackArtifact::decode_module_events puts it next to the decoder, so
it can be benchmarked and optimized in one place without changing the
runner's fork/exec replay.
…tifact

Since 5.4.0 the memory executor decodes every event of the memtrack
artifact after the run, only to keep the few Mapping, Fork and Exec events
it needs for module artifacts. On large artifacts this takes minutes
between memtrack's last log line and the upload.

The benchmark builds an artifact laid out like memtrack writes it:
allocation events with forks and execs spread through the stream and the
mapping suffix at the end. It runs without stacks and with an 8 KiB stack
record every 400 events, at 1M events, and checks the expected number of
module events is found. Larger artifacts take minutes per case with the
streamed decoder under the simulation instrument.

Events are generated as the encoder consumes them and the artifact is
written to a file under the target directory and mapped, so larger sizes
never have to fit in memory.
Generating and searching artifacts much larger than 1M events outlasts the
CI job under the simulation and memory instruments. Pick the sizes from
CODSPEED_RUNNER_MODE so these modes only search the smallest artifact,
leaving walltime and plain divan runs free to cover larger ones.
@not-matthias
not-matthias force-pushed the cod-3819-memtrack-investigate-8-minutes-spent-in-teardown branch from 9a334a0 to 2f1f983 Compare October 9, 2026 15:22
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant