Skip to content

[BUG][MTL] DSP firmware boot -110 on resume — only when HDMI is connected AND audio streams via the internal analog path during system suspend (regression) #10955

Description

@edyfox

Describe the bug

On a ThinkPad T14s Gen 5 (Meteor Lake, sof-audio-pci-intel-mtl, PCI 0000:00:1f.3), the DSP fails to boot firmware with -110 (ETIMEDOUT) on resume from system suspend (S2idle), leaving audio permanently dead until a full reboot. Module reload, PCI remove/rescan, and driver unbind/rebind do not recover it — only a power-cycle does.

The trigger is very specific and fully reproducible: it requires an active playback stream on the internal analog path while an HDMI/DP display is connected at the moment of suspend.

Reproduction — requires ALL THREE conditions simultaneously

  1. An external display is connected via HDMI/DP (HDMI codec present in topology).
  2. Audio is actively streaming through the internal analog codec (in my case a background fluidsynth keeps the analog PCM open continuously).
  3. The system enters suspend (S2idle).

→ On resume, the DSP wedges with -110.

It is NOT reproducible if any one condition is removed (each verified):

HDMI connected Active stream path Suspend Result
Yes Analog Yes WEDGE (-110)
No Analog Yes OK
Yes HDMI Yes OK
Yes Analog No OK

The monitor model is irrelevant — reproduced with two different external monitors in two locations. Only the presence of a connected HDMI sink matters. This points to a cross-codec suspend/resume ordering issue: resuming the active analog path fails only when an idle-but-present HDMI codec also exists in the topology.

Regression

This worked correctly a few months ago — suspending with the same fluidsynth setup did not break audio. It has regressed since. (Happy to help bisect kernel / firmware versions to narrow the window.)

dmesg on the failed resume / forced re-probe

sof-audio-pci-intel-mtl 0000:00:1f.3: Loaded firmware library: ADSPFW, version: 2.12.0.1
sof-audio-pci-intel-mtl 0000:00:1f.3: timeout waiting for purge IPC done
sof-audio-pci-intel-mtl 0000:00:1f.3: ------------[ DSP dump start ]------------
sof-audio-pci-intel-mtl 0000:00:1f.3: Boot iteration failed: 3/3
sof-audio-pci-intel-mtl 0000:00:1f.3: fw_state: SOF_FW_BOOT_IN_PROGRESS (3)
sof-audio-pci-intel-mtl 0000:00:1f.3: 0xd000001c: module: ROM_EXT, state: REMOVE_ACCESS_CONTROL, not running
sof-audio-pci-intel-mtl 0000:00:1f.3: error code: 0x2328 (unknown)
sof-audio-pci-intel-mtl 0000:00:1f.3: ------------[ DSP dump end ]------------
sof-audio-pci-intel-mtl 0000:00:1f.3: error: dsp init failed after 3 attempts with err: -110
sof-audio-pci-intel-mtl 0000:00:1f.3: Failed to start DSP
sof-audio-pci-intel-mtl 0000:00:1f.3: error: failed to boot DSP firmware -110
sof-audio-pci-intel-mtl 0000:00:1f.3: error: sof_probe_work failed err: -110

Note: the failure is at the boot ROM stage (ROM_EXT / REMOVE_ACCESS_CONTROL), before firmware runs — so no PCI/driver/PM-level recovery works; only a full power cycle resets the DSP.

Environment

  • Hardware: Lenovo ThinkPad T14s Gen 5 (21LS0041SG), BIOS N46ET24W (1.14)
  • SoC / audio: Intel Meteor Lake, sof-hda-dsp, PCI 0000:00:1f.3
  • Kernel: 7.0.12+deb14.1-amd64 (Debian forky/sid)
  • SOF firmware: firmware-sof-signed 2025.01-1; loaded ADSPFW 2.12.0.1; sof-mtl.ri, IPC4
  • Topology: sof-ipc4-tplg/sof-hda-generic-2ch.tplg
  • alsa-ucm-conf: 1.2.16-1
  • Userspace: PipeWire + WirePlumber
  • Suspend type: S2idle (modern standby)

Possibly related

Same -110-on-suspend/resume IPC-timeout family: #1072, #2971, #4507, #4558 (different platforms/triggers, but similar resume failure signature).

Additional context

I can reliably reproduce on demand and am happy to collect sof-logger traces, test patches, or bisect. Workarounds in the meantime: disconnect HDMI before suspend, or keep the system on AC (no auto-suspend).

Activity

  1. lgirdwood commented on Jun 25, 2026

    @lgirdwood
    Member

    @edyfox thanks, this is goood data. Btw do you have a full kernel log around the suspend and resume operations. Is it also possible to disable pipewire and wireplumber (and reboot so they are not running) and just do an aplay -Dhw: command. i.e. to rule in/out any interactions for userspace discovery around the HDMI connection event. Thanks !

  2. edyfox commented on Jun 26, 2026

    @edyfox
    Author

    Thanks @lgirdwood! I hit it again this morning and grabbed the persistent journal (the ring buffer had already wrapped under the error spam, so the journal was essential). A few things that refine the original report:

    1. This time the failure is a DSP firmware panic/Oops, not the boot-ROM -110

    Importantly, the firmware booted successfully on resume (fw_state: SOF_FW_BOOT_COMPLETE (7)), then crashed with an Xtensa exception. The ROM_EXT … -110 from my original report is the secondary failure — it only appears afterwards when the driver tries to re-boot the panicked DSP. The primary event:

    sof-audio-pci-intel-mtl 0000:00:1f.3: ------------[ DSP dump start ]------------
    sof-audio-pci-intel-mtl 0000:00:1f.3: DSP panic!
    sof-audio-pci-intel-mtl 0000:00:1f.3: fw_state: SOF_FW_BOOT_COMPLETE (7)
    sof-audio-pci-intel-mtl 0000:00:1f.3: 0x50000005: module: ROM_EXT, state: FW_ENTERED, running
    sof-audio-pci-intel-mtl 0000:00:1f.3: Firmware state: 0x5, status/error code: 0x0
    sof-audio-pci-intel-mtl 0000:00:1f.3: error: DSP Firmware Oops
    sof-audio-pci-intel-mtl 0000:00:1f.3: error: Exception Cause: AllocaCause, MOVSP instruction, if caller's registers are not in the register file
    sof-audio-pci-intel-mtl 0000:00:1f.3: EXCCAUSE 0x00000005 EXCVADDR 0x00000000 PS 0x00060920 SAR 0x0000000c
    sof-audio-pci-intel-mtl 0000:00:1f.3: EPC1 0xa007a1d9 EPC2 0x00000000 EPC3 0x00000000 EPC4 0x00000000
    sof-audio-pci-intel-mtl 0000:00:1f.3: AR registers:
    sof-audio-pci-intel-mtl 0000:00:1f.3: 0x0: a0050c45 a0117a00 00000000 4015e880
    sof-audio-pci-intel-mtl 0000:00:1f.3: 0x10: a015e940 00000018 4014f590 a0117a00
    sof-audio-pci-intel-mtl 0000:00:1f.3: 0x20: a0062b8d a01179c0 4014f590 a00685a4
    sof-audio-pci-intel-mtl 0000:00:1f.3: 0x30: a0062b8d a01179c0 4014f590 a00685a4
    sof-audio-pci-intel-mtl 0000:00:1f.3: ------------[ DSP dump end ]------------
    

    Then the IPC times out and teardown fails:

    sof-audio-pci-intel-mtl 0000:00:1f.3: ipc timed out for 0xe030100|0x300
    sof-audio-pci-intel-mtl 0000:00:1f.3: IPC timeout
    sof-audio-pci-intel-mtl 0000:00:1f.3: ASoC error (-110): at soc_component_trigger() on 0000:00:1f.3
     HDMI1: ASoC error (-110): trigger FE cmd: 1 failed
    sof-audio-pci-intel-mtl 0000:00:1f.3: sof_ipc4_trigger_pipelines: pcm0 (HDA Analog), dir 0: failed to pause all pipelines
     HDA Analog: ASoC error (-19): trigger FE cmd: 0 failed
    sof-audio-pci-intel-mtl 0000:00:1f.3: failed to unbind modules mixin.1.1:0 -> mixout.2.1:0
    

    2. Failing cycle was an s2idle suspend/resume

    PM markers for the failing cycle (from the journal):

    6月 25 20:32:28  PM: suspend entry (s2idle)
    6月 26 10:15:18  Freezing user space processes completed (elapsed 0.276 seconds)
    6月 26 10:15:18  Restarting tasks: Done
    6月 26 10:15:18  PM: suspend exit
    6月 26 10:26:26  DSP panic! (first audio use after resume, ~11 min later)
    

    The crash-time teardown trace touches both the HDA Analog (pcm0) and HDMI1 pipelines (pcm0 (HDA Analog) … failed to pause all pipelines, HDMI1: trigger FE … failed), so both are present in the topology and entangled at the moment of failure. An external DisplayPort monitor (DP-1) was connected during this session.

    3. Full kernel log around suspend/resume

    Full journal (covers both suspend entries + the resume that panicked): https://gist.github.com/edyfox/15c84c63f4f1d110e91012d04b9820bf

    4. Your aplay -Dhw: test with PipeWire/WirePlumber disabled

    (a) Against the already-wedged device (PipeWire fully stopped): the wedge persists with zero userspace involvement. With pipewire/wireplumber killed, the card still enumerates but raw ALSA fails at device open:

    $ aplay -l   # device is present
    card 0: sofhdadsp [sof-hda-dsp], device 0: HDA Analog
    card 0: sofhdadsp [sof-hda-dsp], device 3: HDMI1 ...
    
    $ speaker-test -D hw:0,0 -c2 -t sine
    Playback open error: -22, Invalid argument
    
    $ aplay -D hw:0,0 -f S16_LE -r 48000 -c 2 < /dev/zero
    aplay: main:850: audio open error: Invalid argument
    
    # power/control=auto  runtime_status=error
    

    So PipeWire is not holding the device broken — once wedged, it rejects raw hw: opens with -22 (the DSP has parked in status=error by this point; the -110 panic was the originating event earlier).

    (b) Reproduce with PipeWire/WirePlumber MASKED (never running across the sleep cycle): I could not reproduce the wedge this way. With both pipewire and wireplumber masked (symlinked to /dev/null, confirmed no processes running), I held the analog path open with a raw aplay -D hw:0,0 /dev/zero and ran a full sleep cycle in both modes:

    Cycle PipeWire Analog stream held Result
    This morning (real) running fluidsynth WEDGED (firmware panic)
    Hibernate (test) masked raw aplay survived
    s2idle suspend (test) masked raw aplay survived

    In both masked cycles the held aplay logged the expected graceful recovery (Suspended. Trying resume. Failed. Restarting stream. Done.), the DSP stayed status=active, and a fresh aplay -D hw:0,0 opened and played normally afterward — no panic, no -110.

    Caveat: the bug is not 100%-per-cycle even with PipeWire running, so "survived 2/2 with PipeWire masked" is suggestive rather than conclusive. But it does point toward PipeWire/WirePlumber's resume-time activity (device re-discovery / re-routing) being involved in the trigger, rather than the raw active analog stream alone. Happy to run more masked cycles if you'd like a larger sample.

    Environment (unchanged from original)

    • ThinkPad T14s Gen 5 (21LS0041SG), BIOS N46ET24W (1.14)
    • Kernel 7.0.12+deb14.1-amd64 (Debian forky/sid)
    • firmware-sof-signed 2025.01-1, ADSPFW 2.12.0.1, IPC4, topology sof-hda-generic-2ch.tplg
  3. edyfox commented on Jun 29, 2026

    @edyfox
    Author

    Another fresh reproduction today (s2idle, PipeWire running normally, DP-1 connected). This one captured the full causal chain in a single clean sequence, and it adds a new clue I hadn't seen before: IMR restore failed.

    The whole chain, resume → wedge:

    12:09:21  PM: suspend entry (s2idle)
    13:45:05  PM: suspend exit
    13:45:25  sof_ipc4_trigger_pipelines: pcm0 (HDA Analog), dir 0: failed to pause all pipelines
    13:45:25  DSP panic!
    13:45:25  fw_state: SOF_FW_BOOT_COMPLETE (7)
    13:45:25  error: DSP Firmware Oops
    13:45:25  error: Exception Cause: AllocaCause, MOVSP instruction, ...
    13:45:25  EXCCAUSE 0x00000005 EXCVADDR 0x00000000 PS 0x00060d20 SAR 0x0000000c
    13:45:25  EPC1 0xa007a1d9 ...
    13:45:26  IMR restore failed, trying to cold boot
    13:45:27  timeout waiting for purge IPC done
    13:45:27  Boot iteration failed: 3/3
    13:45:27  0xd000001c: module: ROM_EXT, state: REMOVE_ACCESS_CONTROL, not running
    13:45:27  error code: 0x2328 (unknown)
    13:45:27  error: dsp init failed after 3 attempts with err: -110
    13:45:27  sof_pcm_prepare: pcm0 (HDA Analog), dir 0: failed to set hw_params after resume
    

    A few things that I think narrow this down:

    1. The panic is deterministic, not random corruption. The Xtensa exception is at the exact same address as the 2026-06-26 occurrence: EXCCAUSE 0x5 (AllocaCause), EPC1 0xa007a1d9. Same faulting instruction pointer across independent reproductions.

    2. IMR restore failed, trying to cold boot — on resume the driver attempts to restore the DSP from IMR, and that restore fails. It then falls back to a cold boot, which is what hits the ROM_EXT … 0x2328 / -110. So the -110 boot failure is downstream of an IMR-restore failure following the firmware panic.

    3. The driver explicitly attributes it to resume: failed to set hw_params after resume, and the teardown that immediately precedes the panic is on pcm0 (HDA Analog).

    So the sequence appears to be: resume → attempt to restore/re-trigger the active HDA Analog pipeline → DSP firmware Oops (AllocaCause at 0xa007a1d9) → IMR restore fails → cold boot fails at ROM with -110.

    Full kernel journal for this cycle: https://gist.github.com/edyfox/ffeaa0c8ae273d5faf150d3b7515b38f

    Environment unchanged from the report: kernel 7.0.12+deb14.1-amd64, firmware-sof-signed 2025.01-1, ADSPFW 2.12.0.1, IPC4, sof-hda-generic-2ch.tplg, ThinkPad T14s Gen 5 (21LS0041SG), BIOS N46ET24W (1.14).

    Happy to enable additional SOF trace/logging (e.g. dyndbg, sof-logger, higher trace levels) or test kernel-side debug patches if that would help pin down the Oops at 0xa007a1d9. I'd prefer to avoid flashing custom/debug DSP firmware on this machine, since the DSP has already shown it can wedge at the boot-ROM stage — but I can reproduce on demand and capture whatever traces are useful.

  4. added
    bugSomething isn't working as expected
    P2Critical bugs or normal features
    on Jun 30, 2026
  5. lyakh commented on Jul 1, 2026

    @lyakh
    Collaborator

    @edyfox hi, can you add a kernel module parameter to collect SOF driver debugging logs by adding a file, e.g. /etc/modprobe.d/sof.conf with the line options snd_sof sof_debug=0x41. Also, if you can collect firmware trace (you'll probably need to use mtrace-reader.py), that could help too!

  6. edyfox commented on Jul 2, 2026

    @edyfox
    Author

    Thanks @lyakh! I've set both up on my side:

    • Added /etc/modprobe.d/sof.conf with options snd_sof sof_debug=0x41.
    • Grabbed tools/mtrace/mtrace-reader.py and confirmed that with sof_debug=0x41 the /sys/kernel/debug/sof/mtrace/core0 interface appears (it's absent with the default sof_debug=0), so firmware trace capture is ready.

    Since the firmware trace has to be streaming before the panic, I've armed mtrace-reader.py to log continuously from boot, and I'll keep it running through my normal usage (HDMI connected + an audio stream holding the internal analog path open). The failure only shows up on resume from suspend, so I'll capture on the next reproduction and post:

    • the firmware trace (mtrace-reader.py) covering the crash,
    • the full kernel journal for that boot (the panic → IMR restore failed → ROM cold-boot -110 chain),
    • and the driver debugfs snapshots (fw_state, exception, fw_regs).

    I reproduce this fairly regularly, so it shouldn't be a long wait. Will report back with the traces attached.

  7. lyakh commented on Jul 2, 2026

    @lyakh
    Collaborator

    @edyfox thanks for setting things up. But just out of curiosity - it should be possible to trigger the bug artificially too, right? You can just close the lid of your laptop, which usually sends it to sleep, or send it to sleep from the menu, or use systemctl suspend or rtcwake -m mem -s 60 (both as root)?

  8. edyfox commented on Jul 3, 2026

    @edyfox
    Author

    @lyakh yes, I can trigger it artificially — systemctl suspend / closing the lid / rtcwake -m mem -s 60 all work, so I don't have to wait for it to happen on its own. Two caveats though:

    • It's not 100% per cycle. Some suspend/resume cycles come back clean. So it's "reproducible on demand" in the sense that I can keep trying, but not "one command and it always wedges."
    • The machine has to stay asleep for a while. A quick close-and-immediately-reopen doesn't do it — if I re-open the lid right away it resumes fine. It seems to need to actually settle into sleep for some time before the resume path wedges. rtcwake -s 60 sometimes isn't long enough; leaving it down for several minutes is more reliable.

    One practical constraint on my side: this is my work laptop and once the DSP wedges the only recovery is a full reboot, which is quite disruptive mid-workday. So realistically my best window to run an instrumented attempt is over lunch rather than "any time." I've got sof_debug=0x41 + mtrace-reader.py armed and I'm running attempts on that schedule — I'll post the firmware trace + kernel journal as soon as a cycle wedges with the capture live. Today's attempt came back clean; I'll keep at it over the next few days.

  9. lyakh commented on Jul 3, 2026

    @lyakh
    Collaborator

    @edyfox ah, sorry, I misinterpreted "fully reproducible" from the issue description, so it doesn't happen every single time! Sure, then we'll wait until it happens next time, thanks for your help

  10. edyfox commented on Jul 7, 2026

    @edyfox
    Author

    Quick status update — no failing capture yet, but an observation worth flagging.

    The bug appears to be timing-sensitive: enabling the firmware trace makes it much harder to reproduce. Before I set up sof_debug=0x41 + mtrace-reader.py, I could trigger the wedge fairly readily (that's how I got the three earlier reports). Since arming the firmware trace, I've run two deliberate suspend/resume attempts with every trigger condition confirmed live (analog PCM held open via a background stream → Headphones sink RUNNING, DP-1 connected, HDMI codecs present-but-suspended) and both resumed cleanly — one lunchtime s2idle, one ~14h overnight s2idle.

    That's a classic Heisenbug signature, and it fits the cross-codec suspend-ordering race hypothesis: the extra firmware-side trace activity (and the host reading the mtrace window) seems to perturb the resume-path timing just enough to lose the race that normally wedges the DSP. I don't think it's proof yet — the bug was never 100%/cycle — but two-for-two clean with the trace on, after it being readily reproducible with the trace off, is suggestive.

    To test that properly rather than assume it, I'm running a footprint-ramp over the next couple of days, holding the trigger conditions fixed and stepping the instrumentation down one notch at a time:

    1. Full — sof_debug=0x41 + live mtrace-reader streaming core0 (what I've been running).
    2. Reader off, FW trace still on — kill the host-side reader, keep sof_debug=0x41.
    3. Driver logs only — lower sof_debug so the firmware-trace bit is off (driver dmesg logging retained), reboot.
    4. Bare — revert sof.conf entirely, reboot.

    If the wedge reappears as I go down the ladder, that pins how much the observation itself is suppressing it — and tells us roughly where in the resume/teardown path the race lives. Whatever rung it fires on, I'll have the corresponding logs (FW trace if it's rung 1–2, kernel journal for all of them) and will post them.

    Two smaller notes:

    • A clean-resume firmware trace is here as a baseline: https://gist.github.com/edyfox/f8e6a82d4ee277e181c91f575990111e — it may be useful to diff against a failing capture once I get one. It's just steady ll_schedule heartbeat (overruns 0 throughout) with normal pipeline setup, spanning the full sleep (FW timestamp jumps ~14h across the s2idle and resumes cleanly).
    • On cadence: this is my daily work laptop and the only recovery from a wedge is a full reboot, so I'm running these attempts around lunch / end-of-day rather than on demand. Doing my best to get you a failing capture — will report back on each rung.
  11. edyfox commented on Jul 7, 2026

    @edyfox
    Author

    Got it — a live firmware trace captured across the panic. This is the first time I've had the DSP's own log for the crash itself (previous reports only had the kernel-side aftermath). s2idle resume, sof_debug=0x41 + mtrace-reader.py streaming core0, all trigger conditions live (HDA Analog held open → Headphones sink RUNNING, DP-1 connected, HDMI codecs present-but-suspended).

    Firmware trace: https://gist.github.com/edyfox/aa24e484eb5963e1142d9913dc6e30e2
    Kernel journal (same boot): https://gist.github.com/edyfox/d7993b950cdaff7e96468bc4a11fdfa0

    The precursor the kernel log never showed

    The firmware trace shows a chain_dma channel request failing on resume, after which the task starts anyway with a null DMA id, and the DSP immediately takes a null-pointer fatal exception:

    <inf> host_comp: host_get_copy_bytes_normal: comp:0 0x4 no bytes to copy, available samples: 384, free_samples: 0   (repeats — analog stream starved on resume)
    <inf> ipc:  ipc_cmd: rx : 0xe030100|0x300
    <inf> dma:  dma_get: dma_get() ID 0 sref = 2 busy channels 0
    <err> chain_dma: chain_init: comp:0 0x0 chain_init(): dma_request_channel() failed
    <inf> chain_dma: chain_task_start: comp:128 0x80 chain_task_start(), host_dma_id = 0x00000000
    <err> os: print_fatal_exception:  ** FATAL EXCEPTION
    <err> os: print_fatal_exception:  ** CPU 0 EXCCAUSE 13 (load/store PIF data error)
    <err> os: print_fatal_exception:  **  PC 0xa007a1d9 VADDR (nil)
    ...
    <err> os: z_fatal_error: >>> ZEPHYR FATAL ERROR 0: CPU exception on CPU 0
    <err> zephyr: k_sys_fatal_error_handler: Halting system
    

    Reading the chain firmware-side:

    1. On resume the active HDA Analog path re-triggers; host reports the stream starved (free_samples: 0, repeating).
    2. Host sends an IPC (0xe030100) that drives a chain_dma init.
    3. dma_request_channel() fails — no channel obtained.
    4. Execution continues past that failure: chain_task_start() runs with host_dma_id = 0x00000000.
    5. That null id is dereferenced → EXCCAUSE 13 (load/store PIF data error), VADDR (nil) → Zephyr fatal → DSP halts.

    The kernel side then plays out exactly as in the earlier reports — DSP panic! / fw_state: SOF_FW_BOOT_COMPLETE (7) → DSP Firmware Oops → fw_state: SOF_FW_CRASHED (8) → sof_pcm_prepare: pcm0 (HDA Analog): failed to set hw_params after resume → IMR restore failed, trying to cold boot → ROM REMOVE_ACCESS_CONTROL / error code: 0x2328 → dsp init failed after 3 attempts with err: -110. So the -110 is confirmed downstream of this firmware exception.

    Two notes / corrections

    • VADDR (nil) is a genuine null dereference, and it happens right after a logged dma_request_channel() failure that execution didn't bail on. From the trace this reads like a missing error-check on the chain_dma init failure path rather than only a timing artifact — the suspend/resume race seems to decide whether the channel request fails, but the crash once it fails looks like an unhandled error path. (I can't see the source line for 0xa007a1d9, so this is inference from the trace — you'll be able to map the PC/backtrace far better than I can.)
    • Exception cause differs from my earlier reports. The kernel-read register dumps in the previous comments said EXCCAUSE 0x5 (AllocaCause); the firmware's own trace here says EXCCAUSE 13 (load/store PIF data error) — same PC 0xa007a1d9. I'd trust the live firmware trace as the more authoritative crash-time source; flagging the discrepancy rather than hiding it.

    Backtrace from the trace (in case it helps map it): 0xa007a1d6 0xa0050c42 0xa0050612 0xa003f03a 0xa003daee 0xa003d7a3 0xa006a966 0xa006a3f5 0xa006dfa6 0xa005e501 0xa006215f.

    Firmware: firmware-sof-signed 2025.01-1 (ADSPFW as shipped), kernel 7.0.13+deb14-amd64. Debugfs exception/fw_regs reads at wedge time returned bad-address/garbage (DSP was already halted and couldn't service them), so the live mtrace is the reliable artifact here.

    This is my daily work laptop and the only recovery is a full reboot, so I need to reboot now rather than hold the DSP wedged — but I can reproduce this deliberately on demand. Happy to run more captures, grab additional debugfs, or test a patch on the chain_dma failure path on a future cycle whenever it'd help; just let me know what would be most useful to collect next.

  12. lyakh commented on Jul 8, 2026

    @lyakh
    Collaborator

    hi @edyfox thanks for the log. It confirms my previous analysis (not posted here before), that the exception happens in chain_host_start() and this is strange. Chain DMA should only be used with HDMI and you were playing audio to HDA. So it appears like a userspace application on your laptop is trying to use HDMI after the resume instead of HDA? That's why I asked you to enable kernel debugging the options snd_sof sof_debug=0x41 module parameter, but it looks like you haven't done that, the kernel log doesn't contain more information than before? I wanted to see a line in the kernel log like

    kernel: snd_sof:sof_pcm_open: sof-audio-pci-intel-ptl 0000:00:1f.3: pcm2 (Port2), dir 0: Entry: open
    

    Unfortunately this still wouldn't tell us which process does that but at least you could compare a working and a non-working case and confirm that in a working resume case this doesn't happen. You said that you're running Debian Sid on your work laptop, which is a rather bold thing to do IMHO... Not sure your IT department is particularly happy about that choice ;-) Maybe it would be better for you to downgrade to a more stable distribution.
    As for NULL dereference in chain_host_start() - we can add a guard there like #10979 but you'll only get it with the next SOF release

    UPDATE: sorry, that isn't the right parameter of course to enable kernel verbosity... You need to enable kernel dynamic debug logs. E.g. options snd_sof dyndbg=+pmf. You can also enable debugging for other SOF modules, but I think this is the most important one. Then you should get "open" messages as above and also individual IPC messages like

    kernel: snd_sof:sof_ipc4_log_header: sof-audio-pci-intel-ptl 0000:00:1f.3: ipc tx      : 0x44000000|0x3060004c: MOD_LARGE_CONFIG_SET [data size: 76]
    kernel: snd_sof:sof_ipc4_log_header: sof-audio-pci-intel-ptl 0000:00:1f.3: ipc tx reply: 0x64000000|0x3060004c: MOD_LARGE_CONFIG_SET
    kernel: snd_sof:sof_ipc4_log_header: sof-audio-pci-intel-ptl 0000:00:1f.3: ipc tx done : 0x44000000|0x3060004c: MOD_LARGE_CONFIG_SET [data size: 76]
    
  13. edyfox commented on Jul 9, 2026

    @edyfox
    Author

    Thanks @lyakh — and no worries about the parameter mix-up. Just to close the loop on that: I did have options snd_sof sof_debug=0x41 set (that's what gated the firmware mtrace interface and got us the crash-time trace), but you're right that it doesn't add kernel-side verbosity. I've now added dyndbg=+pmf on top of it:

    options snd_sof sof_debug=0x41 dyndbg=+pmf
    

    and confirmed it's producing exactly the lines you described — sof_pcm_open/prepare/trigger plus the ipc tx/reply/done sof_ipc4_log_header messages. (One note for anyone reproducing: on my kernel these land in dmesg at debug level rather than in journalctl -k's default view, so I'll pull the ring buffer directly on the next capture.)

    Your chain_host_start() / HDMI-vs-HDA read is really interesting and matches what bugged me about the trace too: the active stream was analog (HDA), so a chain_dma init firing at all on resume is the odd part — not just that it null-derefs once the channel request fails. So the diagnostic question is now "which PCM/port does userspace open on resume, and is it an HDMI one when it shouldn't be?" — which is precisely what dyndbg=+pmf will show.

    Plan for the next cycle:

    1. Reproduce the wedge with both sof_debug=0x41 (firmware mtrace) and dyndbg=+pmf (kernel PCM/IPC log) live.
    2. Capture a clean resume and a wedged resume under identical trigger conditions, and diff the sof_pcm_open sequence between them — specifically whether an HDMI/Port PCM opens in the failing case that doesn't open in the working case.
    3. Post both logs (kernel dmesg + firmware mtrace for the failing one).

    That won't name the offending process, as you say, but it should confirm whether an HDMI path is being opened post-resume against an analog-only stream — and if so, narrow it toward whatever userspace component is doing it.

    Thanks also for #10979 — good to know the NULL guard in chain_host_start() is queued regardless; I'll pick it up whenever it lands in a release. I'll report back with the working-vs-wedged comparison.

    (Re: the distro 😄 — small correction: it's actually Forky (testing), not Sid. It reports as forky/sid in os-release/debian_version (testing carries the /sid suffix), which is probably what you saw. And it is my corp machine — but IT is lenient: they were fine issuing me a Debian laptop with no official support attached. I tried a MacBook for my first five months here and bounced off it hard; I'm not a Mac person and won't be, so self-supported Debian it is. Which, silver lining, is how bugs like this end up getting reported.)

  14. 52 remaining items

  15. Kreinoee commented on Aug 28, 2026

    @Kreinoee

    Thanks @ujfalusi. My agent agrees with you conclusion. So now I just need to get the fix somehow.

    For reference, I have opened a request to backport this fix to ubuntus 7.0 kernel used in ubuntu 26.04: https://bugs.launchpad.net/ubuntu/+source/linux/+bug/2165740

  16. added
    P3Low-impact bugs or features
    and removed
    P2Critical bugs or normal features
    on Sep 1, 2026
  17. devis4wd commented on Sep 15, 2026

    @devis4wd

    I appear to be hitting the same or a very closely related resume failure on different Meteor Lake hardware.

    Hardware:

    • Huawei MateBook 14 2024
    • Intel Core Ultra 5 125H (Meteor Lake)
    • Audio: sof-hda-dsp / sof-audio-pci-intel-mtl
    • Two external ASUS monitors connected via HDMI/DisplayPort

    Distribution:

    • Linux Mint 22.1 (Ubuntu 24.04 base)
    • PipeWire 1.0.5

    Tested kernels:

    • 7.0.0-30-generic
    • 7.0.0-31-generic

    The failure has occurred with both kernels.

    I confirmed that the failure happens immediately after resume from system
    suspend. The system is using s2idle:

        $ cat /sys/power/mem_sleep
        [s2idle] deep
    

    At the failing resume:

        09:54:22 kernel: PM: suspend exit
        09:54:22 systemd-sleep: System returned from sleep operation 'suspend'
    

    Immediately afterwards:

        sof-audio-pci-intel-mtl 0000:00:1f.3:
        sof_pcm_prepare: pcm0 (HDA Analog), dir 0:
        failed to set hw_params after resume
    

    The kernel also reports:

         DSP panic!
        error: DSP Firmware Oops
        error: Exception Cause: AllocaCause
        IPC timeout
        IMR restore failed, trying to cold boot
        timeout waiting for purge IPC done
        Boot iteration failed: 3/3
        error code: 0x2328
        dsp init failed after 3 attempts with err: -110
        Failed to start DSP
    

    After this, playback and capture both stop working, including:

    • internal speakers
    • internal digital microphone
    • HDMI/DisplayPort audio

    ALSA still enumerates the devices and PipeWire/WirePlumber remain running, but PipeWire receives ALSA timeouts. Restarting PipeWire, WirePlumber or ALSA does not recover the DSP.

    A reboot is the only way to restore audio.

    One difference from the reproduction described in this issue is that I have not yet confirmed that an active stream was using the internal analog path at the exact moment of suspend. External HDMI/DP displays were connected.

    I can collect additional logs or test specific reproduction conditions if useful.

  18. ujfalusi commented on Sep 15, 2026

    @ujfalusi
    Contributor

    @devis4wd, linux-stable has not picked 0c0e418dbcf0 for 7.0.y as it is non LTS kernel, 7.1 have it and 7.2 as well.
    I think ubuntu should pick this patch for the 7.0.0-32 kernel.

  19. sqvist commented on Sep 15, 2026

    @sqvist

    Same problem here, I can reproduce the trigger conditions.

    OS: Linux Mint 22.1, based on Ubuntu 24.04.

    Hardware:

    * Huawei MateBook 14 2024
    
    * Intel Core Ultra 5 125H (Meteor Lake)
    
    * Audio driver: sof-audio-pci-intel-mtl / sof-hda-dsp
    

    The issue occurs intermittently after suspend/resume while external monitors are connected.

    When it happens, all audio stops working:

    * internal speakers
    
    * internal digital microphone
    
    * HDMI/DisplayPort audio
    

    PipeWire, pipewire-pulse and WirePlumber remain active. Restarting them does not recover audio.

    ALSA still enumerates the devices, but attempts to use them time out.

    Kernel log contains:

    sof-audio-pci-intel-mtl 0000:00:1f.3: DSP panic!
    sof-audio-pci-intel-mtl 0000:00:1f.3: error: DSP Firmware Oops
    sof-audio-pci-intel-mtl 0000:00:1f.3: error: Exception Cause: AllocaCause
    sof-audio-pci-intel-mtl 0000:00:1f.3: IPC timeout
    sof-audio-pci-intel-mtl 0000:00:1f.3: IMR restore failed, trying to cold boot
    sof-audio-pci-intel-mtl 0000:00:1f.3: error code: 0x2328
    sof-audio-pci-intel-mtl 0000:00:1f.3: dsp init failed after 3 attempts with err: -110
    sof-audio-pci-intel-mtl 0000:00:1f.3: Failed to start DSP
    

    PipeWire subsequently reports ALSA timeouts for both playback and DMIC, for example: spa.alsa: set_hw_params: Connection timed out

    The problem has occurred with both: 7.0.0-30-generic 7.0.0-31-generic

    A full reboot restores audio.

    This appears very similar to Ubuntu bug #2165740 and upstream SOF issue #10955. The kernel log signature is essentially the same.

    Hello!

    Its the same kernel as i have in my Kubuntu 26.04, but i installed kernel 7.1.5 with mainline kernel, and i applied the above suggested fix for this bug (you can see my struggles in this thread above), and i have not yet had another problem with this bug. So if you want to explore the world of applying kernel fixes, try out my stuff above. Or maybe try the latest mainline kernel 7.2.6, it might have the fix already applied.

  20. devis4wd commented on Sep 15, 2026

    @devis4wd

    @ujfalusi Thank you for the info!

  21. devis4wd commented on Sep 15, 2026

    @devis4wd

    @sqvist Hey, thank you very much for your help. I'll give it a try and see if I can finally get rid of this (very) annoying bug.

  22. ujfalusi commented on Sep 15, 2026

    @ujfalusi
    Contributor

    @devis4wd, I'm eager to hear from users on this as well. So far I did not got feedback from affected users, it is fixing the issue locally.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P3Low-impact bugs or featuresbugSomething isn't working as expected

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions