Re: [PATCH v11 11/16] iommu/arm-smmu-v3: Add CMDQ_PROD_STOP_FLAG to gate CMDQ submissions
From: Pranjal Shrivastava <praan@google.com>
Date: 2026-10-02 21:11:50
Also in:
driver-core, linux-iommu, linux-pci, lkml
On Fri, Oct 02, 2026 at 01:47:48PM -0300, Jason Gunthorpe wrote:
On Thu, Oct 01, 2026 at 11:03:15AM -0700, Nicolin Chen wrote:quoted
quoted
So we can't issue ATC_INVs during suspend (the EP is already down), nor during resume (the SMMU resumes *before* the EP is made active). The EP can't use its ATC while suspended, and if it loses power/resets on the way back to D0 (from D3cold, or D3hot with No_Soft_Reset=0), it comes back with an empty ATC.. same assumption the PCI reset path makes today (pci_dev_reset_iommu_prepare()).In that case, would the STOP flag be too late? It's only set in the middle of the SMMU suspend. So, an ATC command (via doamin invalidation) might be issued prior to the Point of Commitment, which will be timed out due to the unresponding EP?How can you ever fix that?
The unfortunate reality is that this gap exists in the kernel even today.. upstream SMMUv3 has no RPM, so it's always on, while the EPs can runtime suspend independently. So an ATC_INV can already be issued to an EP that has suspended. I'd argue RPM improves this slightly, since once the STOP flag is set everything is elided, so the window closes at SMMU suspend instead of never.
How does power management really work, is it expected that the end device is already quieted by its driver?
Yes, power management would topo-sort all dependencies and invoke suspend callbacks accordingly, i.e. in our case the suspend callbacks of all SMMU clients would be called before the SMMU's suspend callback.
Could the first step in power management install a blocked STE? Then we don't have to worry about ATC desync and that automatically stops generating new ATC invalidations if we go and detact the domains too
Partially.. at SMMU suspend we set GBPA to abort and clear SMMUEN, so nothing gets through while the SMMU is off. But that's global and only happens after all EPs are down, it doesn't stop ATC_INVs in the window Nicolin pointed out. One way to ensure the ATC state is relying on the PCIe spec to lose ATC content during D0 entry from D3cold, or D3hot with No_Soft_Reset=0). Another way to enforce this, is to *somehow* ask the endpoint drivers disable ATS during *their* suspend, i.e. in the EP's driver's suspend they could call pci_disable_ats or a better suited helper from pci core and the in the pm_resume / rpm_resume they could call it's equivalent pci_enable_ats, counterpart ensuring a clean ATS state. Or maybe the pci_dev_reset_iommu_prepare/done() pair (with slight refactoring) in EP's suspend/resume? I could mention this explicitly in some comments or dev_warn if any of the masters have ATS state as ON during suspend? LMK what you guys think of that?
Maybe I'm wondering if power management should involve the core code so it detaches all the domains from the device, setups up blocking and then the iommu itself could power ofF?
I'm slightly against the blocking domain attach because it's a reasonable ask for the client drivers to be able to dma_map / unmap when they're suspended, given that most of the modern IOMMU state is in-memory and the only HW state is some sort of TLB/ATC maintenance. We can map/unmap when the IOMMU is off and just ensure a clean cache state. Drivers often want to pre-map everything, power ON just to run their workload, power off and then unmap. That said, I agree it would be nice to have the core code handle power management, which can be one of the next steps. (it would be complicated to see how or what each IOMMU might have to handle for power mangement in a generic way).
Jason
Thanks, Praan