Skip to content

feat: preserve virtual queue buffers across snapshots - #1891

Open
andreiltd wants to merge 2 commits into
mainfrom
virtq-lazy-pool-generations
Open

andreiltd wants to merge 2 commits into
mainfrom
virtq-lazy-pool-generations

Conversation

@andreiltd

@andreiltd andreiltd commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

This patch lets the guest keep referencing virtqueue data after a snapshot. Retained data is any guest value that keeps a transport pool slot alive:

  • H2G ByteChunks, which hold a host request
  • host replies, which become owner-backed guest Bytes

During normal execution:

  • the guest reads through the alias, which maps to the pool memory in scratch
  • dropping the last owner frees the slot for reuse

At checkpoint (prepare_snapshot):

  • the guest reclaims finished G2H work and resets both producers. After that, only retained slots are still live
  • prune_pool records the live slot ranges and unmaps every alias page they don't touch
  • H2G is refilled from free slots up to the queue size, including slots that share a page with retained ones

At snapshot:

  • the snapshot walker skips the scratch map but includes the aliases. Retained pages are therefore captured as ordinary memory

After checkpoint, the first transport entry (maybe_refresh):

  • maps missing alias pages to scratch
  • after restore, copies only the recorded retained bytes from snapshot alias pages into scratch, then remaps those pages to scratch. This preserves host writes in neighboring slots, such as the first request
  • copies nothing when no buffers are retained

Closes #1884

@andreiltd
andreiltd added this pull request to stack #1892 October 6, 2026 10:17
@andreiltd andreiltd added kind/enhancement For PRs adding features, improving functionality, docs, tests, etc. regen-goldens Regenerate snapshot golden fixtures labels Oct 6, 2026
@hyperlight-gh-bot

This comment has been minimized.

@andreiltd
andreiltd force-pushed the virtq-lazy-pool-generations branch from f718561 to 36eda9c Compare October 6, 2026 11:29
@andreiltd
andreiltd marked this pull request as ready for review October 6, 2026 11:45
Copilot AI balanced review requested due to automatic review settings October 6, 2026 11:45

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Alias page-table sizing and zero-capacity checkpoint cases can prevent supported sandboxes from initializing or continuing after restore.

Review effort: Balanced
Findings: 2 High severity · 1 Low severity

Open (3)
What changed in this PR

Preserves transport-backed guest buffers across snapshot capture and restore using stable aliases and pinned pages.

Changes:

  • Adds alias mapping, pruning, and mailbox checkpoint states.
  • Allows partially prefilled H2G rings in snapshots.
  • Adds restore tests, benchmarks, ABI 6, and documentation.
File Description
src/​tests/​rust_guests/​simpleguest/​src/​main.rs Adds retained buffer readers.
src/​tests/​rust_guests/​Cargo.lock Locks the new dependency.
src/​tests/​c_guests/​c_simpleguest/​main.c Adds C retention fixtures.
src/​hyperlight_host/​tests/​snapshot_goldens/​goldens_version.rs Advances goldens to v6.
src/​hyperlight_host/​tests/​sandbox_host_tests.rs Exercises repeated fragmented calls.
src/​hyperlight_host/​tests/​integration_test.rs Tests retained data restoration.
src/​hyperlight_host/​src/​sandbox/​uninitialized_evolve.rs Completes initial checkpoints.
src/​hyperlight_host/​src/​sandbox/​snapshot/​tripwires.rs Pins ABI 6.
src/​hyperlight_host/​src/​sandbox/​snapshot/​file/​transport.rs Updates malformed transport testing.
src/​hyperlight_host/​src/​sandbox/​snapshot/​file/​media_types.rs Bumps the snapshot ABI.
src/​hyperlight_host/​src/​sandbox/​snapshot/​file/​config.rs Updates pinned schemas.
src/​hyperlight_host/​src/​sandbox/​snapshot/​file_tests.rs Updates ABI rejection assertions.
src/​hyperlight_host/​src/​sandbox/​initialized.rs Enables retained-buffer snapshots.
src/​hyperlight_host/​src/​sandbox/​config.rs Clarifies scratch usage.
src/​hyperlight_host/​src/​mem/​virtq/​tests.rs Tests mailbox and partial prefill restore.
src/​hyperlight_host/​src/​mem/​virtq/​mod.rs Accepts bounded H2G prefixes.
src/​hyperlight_host/​src/​mem/​mgr.rs Implements mailbox checkpoint states.
src/​hyperlight_host/​benches/​benchmarks.rs Benchmarks retained-buffer restores.
src/​hyperlight_guest/​src/​transport/​mod.rs Exposes refresh and checkpoint entry points.
src/​hyperlight_guest/​src/​transport/​mem.rs Maps completed buffers through aliases.
src/​hyperlight_guest/​src/​transport/​context.rs Prunes and refreshes pool aliases.
src/​hyperlight_guest/​src/​transport/​backing.rs Implements alias paging and pinning.
src/​hyperlight_guest/​Cargo.toml Adds fixed bitset support.
src/​hyperlight_guest_bin/​src/​transport.rs Updates transport imports.
src/​hyperlight_guest_bin/​src/​lib.rs Checkpoints after guest initialization.
src/​hyperlight_guest_bin/​src/​guest_function/​call.rs Refreshes aliases before dispatch.
src/​hyperlight_common/​src/​virtq/​ring/​canonical.rs Adds prefix validation.
src/​hyperlight_common/​src/​virtq/​pool/​tests.rs Tests slot exclusion behavior.
src/​hyperlight_common/​src/​virtq/​pool/​slot.rs Adds persistent slot exclusions.
src/​hyperlight_common/​src/​virtq/​pool/​fuzz.rs Removes obsolete free-slot checks.
src/​hyperlight_common/​src/​transport.rs Defines mailbox wire values.
docs/​virtio-host-guest-communication.md Documents retained-buffer snapshots.
docs/​snapshot-versioning.md Documents ABI 6.
CHANGELOG.md Records the feature and ABI break.
Cargo.lock Locks the guest dependency.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread src/hyperlight_guest/src/transport/backing.rs Outdated
Comment thread src/hyperlight_guest/src/transport/backing.rs Outdated
Comment thread src/hyperlight_guest_bin/src/guest_function/call.rs Outdated
@hyperlight-gh-bot

This comment has been minimized.

@andreiltd
andreiltd marked this pull request as draft October 6, 2026 14:38
@andreiltd
andreiltd force-pushed the virtq-lazy-pool-generations branch from 36eda9c to 534cad5 Compare October 6, 2026 16:10
@andreiltd
andreiltd marked this pull request as ready for review October 6, 2026 16:22
@hyperlight-gh-bot

This comment has been minimized.

ludfjig
ludfjig previously approved these changes Oct 7, 2026

@ludfjig ludfjig left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm. Does this impact the performance of the first guest call after a restore? Do you happen to have any perf numers?

@andreiltd
andreiltd force-pushed the virtq-lazy-pool-generations branch from 534cad5 to 9395642 Compare October 7, 2026 09:12
@andreiltd
andreiltd removed this pull request from stack #1892 October 7, 2026 09:26
@andreiltd
andreiltd force-pushed the virtq-lazy-pool-generations branch from 9395642 to d2751f4 Compare October 7, 2026 09:29
@andreiltd
andreiltd changed the base branch from virtq-unmap-page to main October 7, 2026 09:29
@andreiltd
andreiltd force-pushed the virtq-lazy-pool-generations branch from d2751f4 to 255b9ab Compare October 7, 2026 09:46
@hyperlight-gh-bot

This comment has been minimized.

@andreiltd

Copy link
Copy Markdown
Member Author

lgtm. Does this impact the performance of the first guest call after a restore? Do you happen to have any perf numers?

I will get some numbers.

@andreiltd

andreiltd commented Oct 7, 2026 •

Copy link
Copy Markdown
Member Author

Hey @ludfjig, here are some numbers:

+---------------------------------------------------------------------------------------------------+
¦ Retained buffers      ¦ Parent 302175db ¦ Last commit 255b9ab1 ¦ Change [95% CI]                  ¦
+-----------------------+-----------------+----------------------+----------------------------------¦
¦ None,                 ¦        35.57 µs ¦             35.93 µs ¦ +1.03% [-0.48%, +2.40%]          ¦
+-----------------------+-----------------+----------------------+----------------------------------¦
¦ 8 KiB each direction  ¦     Unsupported ¦             37.48 µs ¦ +5.38% [+4.85%, +5.92%] vs empty ¦
+-----------------------+-----------------+----------------------+----------------------------------¦

That is going to scale with the size of the data that is retained because we are copying it back to scratch. Although I think we may do some tricks to avoid copying altogether.

ludfjig
ludfjig previously approved these changes Oct 7, 2026
@andreiltd
andreiltd added this pull request to stack #1897 October 7, 2026 18:22
@andreiltd
andreiltd force-pushed the virtq-lazy-pool-generations branch from 255b9ab to ad8fc12 Compare October 7, 2026 19:37
@hyperlight-gh-bot

This comment has been minimized.

@hyperlight-gh-bot

This comment has been minimized.

@andreiltd
andreiltd force-pushed the virtq-lazy-pool-generations branch from ad8fc12 to 9de91ff Compare October 8, 2026 08:39
@hyperlight-gh-bot

This comment has been minimized.

@syntactically syntactically left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This looks pretty good, but I think it could be simplified a lot by removing the invariant that the whole "pool" range is 1:1 mapped early, and instead doing the mappings when necessary. (There is a slight problem in ensuring that dropping a view/"returning" a buffer after snapshot does not actually return said buffer, since it is not scratch backed anymore---but this I think can be solved by keeping track of the generation of the view and no-op'ing it (or rather, changing it to just unmap the alias pages) if the generation is different). Then there is no need for the prune_pool/map_pool dance across restore.

I guess that lazily allocating the aliases (and only to guest-owned/used buffers) like I suggest could be a tiny bit worse for bytechunks performance, but because of the relatively light tlb invalidations needed, I would guess it would not really be super noticeable in the current workload, and it would trade off by needing a lot less work for initialisation. So i would expect that to be an overall win, probably.

Comment thread docs/virtio-host-guest-communication.md Outdated

### Retained virtual addresses

Each pool has a stable virtual alias range, mapped eagerly at initialization.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why map eagerly instead of when actually using a buffer? I think if doing the latter one can avoid the need to explicitly look for and unmap the buffers on checkpoint. And, it means we do not need to spend the memory on page tables (alias_table_len) unless references are actually created for a given buffer (which I think only happens in the bytechunks case right now iirc?)

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Perf being main reason, although please see the numbers below.

Comment thread src/hyperlight_common/src/layout.rs Outdated

/// Translate a scratch address into its guest physical address.
const fn scratch_gpa(addr: u64) -> u64 {
addr - (SCRATCH_TOP_GVA - SCRATCH_TOP_GPA) as u64

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we need to make this 1:1 assumption here (i.e. is walking the range too slow)? I don't think any other code makes that assumption, so maybe mention it in paging or layout docs if we are going to start making that an invariant.

@andreiltd andreiltd Oct 8, 2026 •

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I can try walking here, yes.


/// Prepare pool aliases after checkpointing or a generation change.
///
/// After restore, retained slots are copied from captured memory into

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Why is this copying back into the original scratch pool? The guest users of the page are already accessing it through their alias mappings.

I envision this functionality being useful mostly to allow the lifecycle of large data blobs to be independent from snapshot/restore (for example to allow replacing the init data section and as a more portable alternative to loading wasm binaries via mapping before the first snapshot). But, I think this large copy on restore would be quite bad for performance in that case?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pool slots that are less then page size would make neigbor slots unusable if they share a fixed alias page.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm still not clear why it is helpful to copy here. Everything accessing the retained data is doing so through the old alias which has now been disassociated from the buffer pool backing---why do we want to re-associate them?

Comment thread src/hyperlight_guest/src/transport/mem.rs Outdated
Comment thread src/hyperlight_host/src/sandbox/config.rs
Comment thread src/hyperlight_guest/src/transport/mem.rs Outdated
@andreiltd

Copy link
Copy Markdown
Member Author

Hey @syntactically ! I agree that the lazy model is a bit cleaner: GuestMapping has clear responsibility, own an alias while bytes exist, then release it. Restored value remains on snapshotted memory. Compared to eager approach the cost of mapping is per call rather than once on restore. It scales with buffer size. Here are some numbers:

+-------------------------------------------------+
¦ buffer size ¦ eager (us) ¦ lazy (us) ¦ overhead ¦
+-------------+------------+-----------+----------¦
¦ 64 B        ¦      14.19 ¦     14.51 ¦    +2.2% ¦
+-------------+------------+-----------+----------¦
¦ 4 KiB       ¦      15.11 ¦     15.59 ¦    +3.2% ¦
+-------------+------------+-----------+----------¦
¦ 12 KiB      ¦      16.81 ¦     17.50 ¦    +4.1% ¦
+-------------+------------+-----------+----------¦
¦ 64 KiB      ¦      26.39 ¦     28.96 ¦    +9.7% ¦
+-------------+------------+-----------+----------¦
¦ 1 MiB       ¦     201.53 ¦    241.06 ¦   +19.6% ¦
+-------------------------------------------------+

I think it could be simplified a lot by removing the invariant that the whole "pool" range is 1:1

I think I disagree. The complexity will move but I don't see how lazy mapping will simplify the implementation drastically:

  • we need alias allocator machinery for finding unused alias for new mapping because alias can already be occupied by retained buffer. The annoying part is that we cannot get away with simple bump pointer, because unmapping does not reclaim page tables so touching new virtual regions grows page tables and we will run out of scratch way before we exhaust virtual address space. The allocator need to maintain the list of freed regions, coalesce them and also maintain the next pointer. Nothing too crazy, I think we can even use vm-allocator crate for that but eager mapping avoid that all together.
  • pool and mapping need to gain epoch semantics. Again nothing too crazy but we need a mechanism to detach old leases from the pool free list so that retained buffers that already have backing from snapshot do not take space in the pool.

That being said, I we think the numbers are acceptable, I'm happy to give the lazy mapping a go?

@syntactically

Copy link
Copy Markdown
Member

Hey @syntactically ! I agree that the lazy model is a bit cleaner: GuestMapping has clear responsibility, own an alias while bytes exist, then release it. Restored value remains on snapshotted memory. Compared to eager approach the cost of mapping is per call rather than once on restore. It scales with buffer size. Here are some numbers:

+-------------------------------------------------+
¦ buffer size ¦ eager (us) ¦ lazy (us) ¦ overhead ¦
+-------------+------------+-----------+----------¦
¦ 64 B        ¦      14.19 ¦     14.51 ¦    +2.2% ¦
+-------------+------------+-----------+----------¦
¦ 4 KiB       ¦      15.11 ¦     15.59 ¦    +3.2% ¦
+-------------+------------+-----------+----------¦
¦ 12 KiB      ¦      16.81 ¦     17.50 ¦    +4.1% ¦
+-------------+------------+-----------+----------¦
¦ 64 KiB      ¦      26.39 ¦     28.96 ¦    +9.7% ¦
+-------------+------------+-----------+----------¦
¦ 1 MiB       ¦     201.53 ¦    241.06 ¦   +19.6% ¦
+-------------------------------------------------+

I am surprised the overhead is quite so large, but for a lot of the use cases we have that are just restore+call I think it is just shifting the overhead from restore to call?

I think it could be simplified a lot by removing the invariant that the whole "pool" range is 1:1

I think I disagree. The complexity will move but I don't see how lazy mapping will simplify the implementation drastically:

  • we need alias allocator machinery for finding unused alias for new mapping because alias can already be occupied by retained buffer. The annoying part is that we cannot get away with simple bump pointer, because unmapping does not reclaim page tables so touching new virtual regions grows page tables and we will run out of scratch way before we exhaust virtual address space. The allocator need to maintain the list of freed regions, coalesce them and also maintain the next pointer. Nothing too crazy, I think we can even use vm-allocator crate for that but eager mapping avoid that all together.

Good point about the lack of reclamation for the old page tables limiting the bump allocator semantics. They do get reclaimed on snapshot, but I could certainly see it being a problem in some workloads. We can't just try to switch to reclaiming them easily either, unless we change the whole physical page allocator to have a freelist, hm..

I see what you mean about how this basically makes the allocation for VA space used for data copies 1:1 with the buffer allocation itself, simplifying that problem, but I think it is quite undesirable to have the old buffers taking up space in the new pool. Apart from everything else, this means if you do send data + snapshot retaining data + send data you have to size the buffer pool appropriately for 2 rounds of data, which increases the size of the scratch region, which is quite unpleasant (it's important to minimize the size of the scratch region for a given workload, since the scratch region has to be zeroed on restore and directly contributes to restore latency).

  • pool and mapping need to gain epoch semantics. Again nothing too crazy but we need a mechanism to detach old leases from the pool free list so that retained buffers that already have backing from snapshot do not take space in the pool.

I think this is important to have anyway. In fact it seems like a big /dis/advantage to me that in the current PR we end up taking up space in the live descriptor pool for data that was read more than a snapshot ago (which I imagine is likely to survive for some time, a la the generational hypothesis). And, having that association means that we end up having the corresponding scratch pages which there is nothing obvious to do with, so we end up (I think unnecessarily?) copying the data into them just to keep them in sync. I do think it is pretty important for perf to avoid that extra copy that is happening on restore and I don't think should be necessary.

@andreiltd
andreiltd force-pushed the virtq-lazy-pool-generations branch 2 times, most recently from f1dd2a9 to 0d0576f Compare October 8, 2026 21:35
@andreiltd

Copy link
Copy Markdown
Member Author

I believe last commit addresses all the issues.

@hyperlight-gh-bot

This comment has been minimized.

ludfjig
ludfjig previously approved these changes Oct 9, 2026
Comment thread src/hyperlight_common/src/layout.rs Outdated
pub fn min_scratch_size(transport_len: usize) -> usize {
arch::min_scratch_size()
.and_then(|fixed| fixed.checked_add(transport_len))
.and_then(|size| size.checked_add(alias_table_len(transport_len)?))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is this still needed with lazy aliases?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Yeah, the transport length is no longer a reliable bound. I removed it.

Comment thread src/hyperlight_guest/Cargo.toml Outdated
anyhow = { version = "1.0.102", default-features = false }
serde_json = { version = "1.0", default-features = false, features = ["alloc"] }
hyperlight-common = { workspace = true, default-features = false }
fixedbitset = { version = "0.5.7", default-features = false }

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

i think this is unused

@andreiltd
andreiltd force-pushed the virtq-lazy-pool-generations branch 3 times, most recently from db7b79c to f9fb01b Compare October 9, 2026 06:30
@hyperlight-gh-bot

This comment has been minimized.

Keep transport pools and aliases stable. Pin pages backing live buffers
and exclude overlapping slots until a checkpoint finds those pages
empty.

Signed-off-by: Tomasz Andrzejak <andreiltd@gmail.com>
Signed-off-by: Tomasz Andrzejak <andreiltd@gmail.com>
@andreiltd
andreiltd force-pushed the virtq-lazy-pool-generations branch from f9fb01b to d1e7001 Compare October 9, 2026 22:25
@hyperlight-gh-bot

Copy link
Copy Markdown

Benchmark Results

Measured commit: d1e700136235
Baseline commit: dcb53c04a5a7

kvm / amd (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 796.73 ns (➖ 1.02x slower)
vec_bytes 588.91 ns (➖ 1.00x faster)
372.82 µs (➖ 1.00x faster)

payload_allocation

slot_pool_segmented
262144 523.96 ns (➖ 1.00x slower)
65536 141.70 ns (➖ 1.00x faster)

sandboxes

create_initialized_and_drop
medium 80.86 ms (➖ 1.00x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
8.22 ns (➖ 1.06x slower) 7.79 ns (➖ 1.01x slower) 7.79 ns (➖ 1.01x slower)

snapshot_files

load_snapshot_unverified
small 97.98 µs (➖ 1.05x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.09 µs (➖ 1.13x faster) 7.00 µs (➖ 1.22x faster)
65536 2.07 µs (➖ 1.08x slower) 1.94 µs (➖ 1.08x faster)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 6.28 µs (➖ 1.01x faster) 6.30 µs (➖ 1.13x faster)
8192 1.08 µs (➖ 1.04x slower) 1.09 µs (➖ 1.06x slower)
262144 27.46 µs (➖ 1.01x slower)
kvm / intel (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 707.45 ns (➖ 1.06x faster)
vec_bytes 539.27 ns (➖ 1.07x faster)
685.83 µs (➖ 1.08x faster)

payload_allocation

slot_pool_segmented
262144 595.63 ns (➖ 1.07x slower)
65536 149.46 ns (➖ 1.00x slower)

sandboxes

create_initialized_and_drop
medium 83.84 ms (➖ 1.03x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
7.98 ns (➖ 1.04x slower) 7.59 ns (➖ 1.01x faster) 7.54 ns (➖ 1.01x faster)

snapshot_files

load_snapshot_unverified
small 44.66 µs (➖ 1.12x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 8.38 µs (➖ 1.05x slower) 8.43 µs (➖ 1.05x slower)
65536 2.38 µs (➖ 1.06x slower) 2.26 µs (➖ 1.01x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 7.68 µs (➖ 1.01x slower) 7.77 µs (➖ 1.02x slower)
8192 838.12 ns (➖ 1.04x faster) 821.46 ns (➖ 1.00x slower)
262144 32.37 µs (➖ 1.04x slower)
mshv3 / amd (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 938.36 ns (➖ 1.04x faster)
vec_bytes 716.86 ns (➖ 1.01x slower)
318.67 µs (➖ 1.00x faster)

payload_allocation

slot_pool_segmented
262144 679.45 ns (➖ 1.02x faster)
65536 182.02 ns (➖ 1.01x faster)

sandboxes

create_initialized_and_drop
medium 58.12 ms (➖ 1.02x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
9.97 ns (➖ 1.00x slower) 9.93 ns (➖ 1.04x faster) 9.91 ns (➖ 1.00x faster)

snapshot_files

load_snapshot_unverified
small 86.36 µs (➖ 1.00x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 10.64 µs (➖ 1.11x faster) 9.48 µs (➖ 1.01x slower)
65536 2.63 µs (➖ 1.04x faster) 2.50 µs (➖ 1.02x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.09 µs (➖ 1.03x faster) 8.07 µs (➖ 1.00x slower)
8192 1.38 µs (➖ 1.02x slower) 1.34 µs (➖ 1.04x faster)
262144 34.86 µs (➖ 1.03x faster)
mshv3 / intel (Linux) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 927.44 ns (➖ 1.01x faster)
vec_bytes 661.45 ns (➖ 1.00x faster)
706.36 µs (➖ 1.08x slower)

payload_allocation

slot_pool_segmented
262144 640.82 ns (➖ 1.03x slower)
65536 170.97 ns (➖ 1.04x slower)

sandboxes

create_initialized_and_drop
medium 66.04 ms (➖ 1.05x slower)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
9.53 ns (➖ 1.02x faster) 9.16 ns (➖ 1.05x faster) 9.62 ns (➖ 1.10x slower)

snapshot_files

load_snapshot_unverified
small 45.19 µs (➖ 1.01x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 8.03 µs (➖ 1.01x slower) 7.95 µs (➖ 1.00x faster)
65536 2.28 µs (➖ 1.01x faster) 2.29 µs (➖ 1.02x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 7.60 µs (➖ 1.01x slower) 7.63 µs (➖ 1.00x slower)
8192 916.06 ns (➖ 1.01x faster) 940.90 ns (➖ 1.04x slower)
262144 39.13 µs (➖ 1.01x faster)
hyperv-ws2025 / amd (Windows) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 1.20 µs (➖ 1.07x faster)
vec_bytes 791.89 ns (➖ 1.00x slower)
1.95 ms (➖ 1.09x faster)

payload_allocation

slot_pool_segmented
262144 789.58 ns (➖ 1.03x faster)
65536 233.36 ns (➖ 1.01x slower)

sandboxes

create_initialized_and_drop
medium 80.80 ms (➖ 1.03x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
10.08 ns (➖ 1.18x faster) 10.43 ns (➖ 1.03x faster) 10.11 ns (➖ 1.03x faster)

snapshot_files

load_snapshot_unverified
small 642.80 µs (➖ 1.29x faster)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 9.00 µs (➖ 1.05x faster) 9.23 µs (➖ 1.02x slower)
65536 2.38 µs (➖ 1.01x faster) 2.38 µs (➖ 1.03x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 9.31 µs (➖ 1.08x slower) 8.84 µs (➖ 1.05x faster)
8192 1.31 µs (➖ 1.07x faster) 1.31 µs (➖ 1.02x faster)
262144 44.25 µs (➖ 1.06x slower)
hyperv-ws2025 / intel (Windows) (➖ stable)

No benchmark improved or regressed.

Benchmark Results

function_call_codec

encode_control decode_vec_bytes_copy
byte_chunks 1.17 µs (➖ 1.00x faster)
vec_bytes 770.63 ns (➖ 1.00x faster)
2.87 ms (➖ 1.06x faster)

payload_allocation

slot_pool_segmented
262144 731.18 ns (➖ 1.10x faster)
65536 214.08 ns (➖ 1.02x slower)

sandboxes

create_initialized_and_drop
medium 107.46 ms (➖ 1.09x faster)

slot_pool

alloc_dealloc_1500 alloc_dealloc_4096 alloc_dealloc_128
10.25 ns (➖ 1.02x faster) 10.09 ns (➖ 1.07x faster) 10.37 ns (➖ 1.02x faster)

snapshot_files

load_snapshot_unverified
small 608.57 µs (➖ 1.16x slower)

virtq_readonly

slot_pool_segmented_fragmented slot_pool_segmented
262144 7.80 µs (➖ 1.02x slower) 7.79 µs (➖ 1.01x slower)
65536 2.47 µs (➖ 1.07x slower) 2.34 µs (➖ 1.01x slower)

virtq_readwrite

slot_pool_segmented_fragmented slot_pool_segmented
65536 8.48 µs (➖ 1.00x faster) 8.44 µs (➖ 1.03x faster)
8192 1.11 µs (➖ 1.11x faster) 1.11 µs (➖ 1.03x slower)
262144 42.03 µs (➖ 1.18x faster)

Reported by cargo ci bench-report --candidate run:37999109555 --config-file bench_report.toml.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

kind/enhancement For PRs adding features, improving functionality, docs, tests, etc. regen-goldens Regenerate snapshot golden fixtures

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Guest retained virtual queue buffers should survive snapshotting

4 participants