Repository navigation
test(runner-shared): benchmark finding module events in a memtrack artifact - #576
not-matthias wants to merge 3 commits into
Conversation
Merging this PR will not alter performance
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| 🆕 | Memory | find_module_events_with_stacks[1000000] |
N/A | 1.1 MB | N/A |
| 🆕 | Memory | find_module_events[1000000] |
N/A | 1.1 MB | N/A |
| 🆕 | Simulation | find_module_events_with_stacks[1000000] |
N/A | 2.5 s | N/A |
| 🆕 | Simulation | find_module_events[1000000] |
N/A | 2.5 s | N/A |
| 🆕 | WallTime | find_module_events_with_stacks[1000000] |
N/A | 1.3 s | N/A |
| 🆕 | WallTime | find_module_events[1000000] |
N/A | 1.3 s | N/A |
Comparing cod-3819-memtrack-investigate-8-minutes-spent-in-teardown (2f1f983) with main (90cc8e9)
Footnotes
-
6 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
|
@codspeedbot suggest optimizations for find_module_events_with_stacks. sort by largest estimated speedup. estimate the diff size/complexity. no need to run anything. just explore the flamegraphs |
This comment was marked as outdated.
This comment was marked as outdated.
2bbb062 to
7044982
Compare
7044982 to
4bd93bd
Compare
4bd93bd to
363ca54
Compare
|
|
Note: |
363ca54 to
c85a5ef
Compare
…oder The runner filtered the full memtrack event stream down to the Mapping, Fork and Exec events inside read_mappings_from_artifact. Moving that filter into MemtrackArtifact::decode_module_events puts it next to the decoder, so it can be benchmarked and optimized in one place without changing the runner's fork/exec replay.
…tifact Since 5.4.0 the memory executor decodes every event of the memtrack artifact after the run, only to keep the few Mapping, Fork and Exec events it needs for module artifacts. On large artifacts this takes minutes between memtrack's last log line and the upload. The benchmark builds an artifact laid out like memtrack writes it: allocation events with forks and execs spread through the stream and the mapping suffix at the end. It runs without stacks and with an 8 KiB stack record every 400 events, at 1M events, and checks the expected number of module events is found. Larger artifacts take minutes per case with the streamed decoder under the simulation instrument. Events are generated as the encoder consumes them and the artifact is written to a file under the target directory and mapped, so larger sizes never have to fit in memory.
Generating and searching artifacts much larger than 1M events outlasts the CI job under the simulation and memory instruments. Pick the sizes from CODSPEED_RUNNER_MODE so these modes only search the smallest artifact, leaving walltime and plain divan runs free to cover larger ones.
9a334a0 to
2f1f983
Compare
TLDR: Since 5.4.0, the memory executor decodes the whole memtrack artifact after a run just to keep its few Mapping, Fork and Exec events. On artifacts with tens of millions of events, this adds minutes between memtrack's last log line and the upload. This PR only adds a benchmark for that search; the optimization is stacked on top in #578.
read_mappings_from_artifactintoMemtrackArtifact::decode_module_events, so the bench measures the code the runner runs. No behavior change.memtrack_readerbench. Its artifact is laid out like memtrack writes one: allocations, forks and execs spread through the stream, and 64 mappings at the end. It runs without stacks and with an 8 KiB stack record every 400 events, and asserts that all module events are found.Baseline, local walltime on a 32-thread machine: 258 ms without stacks, 267 ms with stacks.