aboutsummaryrefslogtreecommitdiff
path: root/kernel/cgroup
diff options
context:
space:
mode:
authorLukas Wunner <lukas@wunner.de>2026-10-08 14:26:00 +0200
committerBjorn Helgaas <bhelgaas@google.com>2026-10-08 14:33:04 -0500
commit79c168e2aa957ed0054c93db17caba653b95e377 (patch)
tree254c799e92bc304ba52721cf1d85d90d2020626e /kernel/cgroup
parentcee9395acd8043be0644b25c34bfa86623f2b935 (diff)
PCI/AER: Skip error recovery on false alarms
Alex is seeing a probe failure of the amdgpu driver after the Root Port above an AMD Navi10 GPU has been reset. The reset was performed to recover from a Firmware First reported Fatal Error. However all status registers in the Root Port's AER Extended Capability are blank, so apparently the platform firmware raised a false alarm. The issue is only occurring since commit eddba19b8b5f ("PCI/AER: Support Advisory Non-Fatal Errors"). It looks like enabling Advisory Non-Fatal Errors causes code paths to be exercised in platform firmware which were never validated before. Skip error recovery on false alarms, i.e. if no unmasked errors were actually signaled. Note that this will also skip recovery if both the Status and Mask registers are "all ones", as would be the case for inaccessible devices. However that seems justified because it would imply either a hot-unplug event or a Surprise Down Error further up in the hierarchy. Interfering with recovery from that seems uncalled for. Fixes: eddba19b8b5f ("PCI/AER: Support Advisory Non-Fatal Errors") Reported-by: Alex Deucher <alexander.deucher@amd.com> Closes: https://bugzilla.kernel.org/show_bug.cgi?id=222095 Signed-off-by: Lukas Wunner <lukas@wunner.de> Signed-off-by: Bjorn Helgaas <bhelgaas@google.com> Tested-by: Alex Deucher <alexander.deucher@amd.com> Acked-by: Alex Deucher <alexander.deucher@amd.com> Link: https://patch.msgid.link/0552ed277e40a288e0157af799257ee6ec722534.1791460615.git.lukas@wunner.de
Diffstat (limited to 'kernel/cgroup')
0 files changed, 0 insertions, 0 deletions