Skip to content

MacBookPro15,2: t2bce VHCI dies and kernel faults after hibernate resume on linux-t2 7.1.8 #22

Description

@anildigital

Summary

Hibernating and resuming a 2018 13-inch Intel MacBook Pro causes the
t2bce_vhci controller to time out and die. The kernel subsequently emits two
general-protection faults and the machine becomes unusable, requiring a forced
power-off.

This uses the newer t2bce stack, not legacy apple-bce.

Hardware

  • Model: MacBookPro15,2
  • Board: Mac-827FB448E656EC26
  • Product: MacBook Pro 13-inch, 2018, four Thunderbolt 3 ports
  • Firmware: BIOS 2103.160.2.0.0
  • iBridge firmware: 23.16.16068.0.0

Software

  • Distribution: Omarchy 4.0.0 / Arch Linux
  • Kernel: 7.1.8-arch1-Watanare-T2-3-t2
  • Package: linux-t2 7.1.8.arch1-3
  • systemd: 261.2-1
  • Hyprland: 0.56.2-1
  • Root filesystem: Btrfs on dm-crypt
  • RAM: 15.5 GiB
  • Hibernation swapfile: 15.5 GiB Btrfs swapfile

Kernel command line included:

intel_iommu=on iommu=pt pm_async=off mem_sleep_default=deep
resume=/dev/mapper/root resume_offset=<configured>

Power states:

$ cat /sys/power/state
freeze mem disk

$ cat /sys/power/mem_sleep
s2idle [deep]

Hibernation was configured using Omarchy's standard hibernation setup. The
swapfile, initramfs resume hook, resume device and Btrfs resume offset were
already present.

Steps to reproduce

  1. Boot normally.
  2. Run systemctl hibernate from the graphical session.
  3. Power on/resume the system.
  4. The T2 virtual USB controller fails during restoration.
  5. The kernel faults and the system becomes unresponsive.
  6. Hold the power button to recover.

Relevant timeline

Hibernation was requested at 22:10:27:

systemd-logind: hibernate requested from client PID 1241948 ('systemctl')
systemd-logind: The system will hibernate now!
systemd-sleep: Performing sleep operation 'hibernate'...
kernel: PM: hibernation: hibernation entry

The hibernation image was successfully allocated:

kernel: PM: hibernation: Allocated 1693388 pages for snapshot
kernel: PM: hibernation: Allocated 6773552 kbytes in 20.28 seconds
kernel: ACPI: PM: Preparing to enter system sleep state S4
kernel: ACPI: PM: Waking up from system sleep state S4

The T2 VHCI restoration then failed with ETIMEDOUT:

kernel: t2bce_vhci: Possible desync, cmd cancel timed out
kernel: t2bce_vhci: tq resume set-active failed dev=1 port=5 ep=00 status=-110 ret_state=0 active=0 paused_by=4 stalled=0
kernel: t2bce_vhci: stateful resume queue failed: dev=1 ep=00 status=-110
kernel: t2bce_vhci: stateful resume exit status=-110
kernel: t2bce_vhci: bus_resume exit status=-110 no_state_resume=0
kernel: t2bce_vhci t2bce_vhci: HC died; cleaning up
kernel: usb usb7: PM: dpm_run_callback(): usb_dev_restore returns -110
kernel: usb usb7: PM: failed to restore: error -110
kernel: leds apple::kbd_backlight: Setting an LED's brightness failed (-19)

Immediately afterward, the first kernel general-protection fault occurred:

kernel: Oops: general protection fault, probably for non-canonical address 0x8000000100b00003: 0000 [#1] SMP PTI
kernel: CPU: 6 UID: 1000 PID: 1196382 Comm: quickshell
kernel: Tainted: G S       C
kernel: RIP: 0010:do_mprotect_pkey+0x242/0x5b0
kernel: RAX: 8000000100b00003

Its stack included:

__x64_sys_mprotect
do_syscall_64
wait_for_completion_io_timeout
__slab_free
zram_slot_free_notify
kmem_cache_free
zs_free
do_swap_page
handle_mm_fault

About eight seconds later, a second kernel fault occurred in a Chrome I/O
thread:

kernel: t2bce_dma: command queue timeout (slot 26)
kernel: t2bce_dma: SQ unregister failed
kernel: Oops: general protection fault, probably for non-canonical address 0xb8bcbc5bd5339ec5: 0000 [#2] SMP PTI
kernel: CPU: 1 UID: 1000 PID: 551417 Comm: Chrome_ChildIOT
kernel: Tainted: G S    D  C
kernel: RIP: 0010:__refill_objects_node+0x2db/0x660
kernel: RAX: b8bcbc5bd5339ec5
kernel: R10: dead000000000100

Its stack included:

refill_objects
__pcs_replace_empty_main
kmem_cache_alloc_noprof
vm_area_alloc
__mmap_region
mmap_region
do_mmap
vm_mmap_pgoff
ksys_mmap_pgoff

T2 communication continued timing out:

kernel: t2bce_dma: command queue timeout (slot 27)
kernel: t2bce_dma: CQ unregister failed
kernel: t2bce_vhci: Possible desync, cmd cancel timed out
kernel: t2bce_dma: command queue timeout (slot 28)
kernel: t2bce_dma: SQ unregister failed
kernel: t2bce_dma: command queue timeout (slot 29)
kernel: t2bce_dma: CQ unregister failed
kernel: t2bce_vhci: Possible desync, cmd cancel timed out

The journal then stopped without an orderly shutdown because the machine had to
be powered off forcibly.

Other observations

  • No OOM kill or memory-pressure event occurred.
  • No NVMe or filesystem I/O error preceded the failure.
  • No i915 GPU fault preceded the failure.
  • There was no relevant userspace coredump.
  • After the forced restart, Btrfs replayed its tree log and mounted successfully.
  • Normal suspend has previously resumed successfully on this installation.
  • The linux-t2 package was installed several days before this incident, rather
    than immediately before the failure.
  • The same boot contained a large number of malformed Bluetooth advertising
    packet messages from hci0.

Workaround

Hibernation, hybrid sleep and suspend-then-hibernate are now disabled:

[Sleep]
AllowHibernation=no
AllowSuspendThenHibernate=no
AllowHybridSleep=no

Normal suspend remains enabled and logind now reports:

CanSuspend=yes
CanHibernate=no
CanHybridSleep=no
CanSuspendThenHibernate=no

Related reports

  • #607 concerns hibernate resume with legacy apple-bce, but reports only
    keyboard and trackpad loss rather than kernel protection faults.
  • #729 contains similar -110 BCE/VHCI communication timeouts, but during
    Touch Bar probing at boot rather than hibernation resume.

Filed by GPT-5 via OpenAI Codex.

Activity

  1. ske5074 commented on Sep 30, 2026

    @ske5074

    Same failure on a MacBookAir8,1 (2018 Air, no Touch Bar), so it's not specific to the MacBookPro models.

    • Kernel: 7.2.7-arch1-Watanare-T2-2-t2 (linux-t2 7.2.7.arch1-2 from arch-mact2)
    • Firmware: 1554.140.20.0.0 (iBridge: 18.16.14759.0.1,0)
    • Distro: Omarchy 4.0.4 (Arch), LUKS root + btrfs swapfile, mem_sleep_default=deep pm_async=off
    • t2bce_core / t2bce_vhci / t2bce_dma / t2bce_audio are loaded in the initramfs (for the LUKS prompt)

    Default HibernateMode (platform): the machine reset about 3.5 minutes after PM: hibernation: hibernation entry. No signature was written, so the next boot logged PM: Image not found (code -22) and cold-booted. resume_offset is correct (it matches btrfs inspect-internal map-swapfile -r).

    pm_test freezer / devices / platform / processors / core: all pass. t2bce_vhci logs resume of a queue with pending submissions ×5 each time, but it recovers.

    /sys/power/disk=reboot (image written, then restored): the image restores, then the VHCI dies:

    t2bce_vhci: bus_resume entry
    t2bce_vhci: resume of a queue with pending submissions   (x5)
    t2bce_vhci: Possible desync, cmd cancel timed out
    t2bce_vhci: tq resume set-active failed dev=1 port=1 ep=00 status=-110 ret_state=0 active=0 paused_by=4 stalled=0
      ... (same for every endpoint on dev 1-5)
    t2bce_vhci: bus_resume exit status=-110
    t2bce_vhci t2bce_vhci: HC died; cleaning up
    usb usb5: PM: dpm_run_callback(): usb_dev_restore returns -110
    usb usb5: PM: failed to restore: error -110
    

    It then logs t2bce_dma: command queue timeout (slot N) / SQ unregister failed / CQ unregister failed in a loop, and the internal keyboard and trackpad are gone. I saw no GP faults, unlike the original report. It looks like the restored kernel resumes its old queue state against a BCE that the boot kernel already reinitialised.

    Unbind/rebind workaround doesn't work: unbinding t2bce_audio then t2bce_core from PCI and binding them again, while running normally with no hibernate involved, re-probes t2bce_core and t2bce_audio. t2bce_vhci never comes back, so the keyboard and trackpad are lost until reboot. The unbind also triggers:

    WARNING: drivers/base/core.c:1569 at device_links_driver_cleanup+0x110/0x120
    device_link_test(link, DL_FLAG_AUTOREMOVE_CONSUMER)
    

    Deep suspend (S3) works fine on this machine (resume path: stateful).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions