Thread (41 messages) 41 messages, 3 authors, 2026-08-25

Re: [PATCH v4 10/17] KVM: arm64: Add a shrinker for pKVM

From: Fuad Tabba <fuad.tabba@linux.dev>
Date: 2026-08-24 18:11:17
Also in: kvmarm

On Mon, 24 Aug 2026 at 17:38, Vincent Donnefort [off-list ref] wrote:
On Tue, Aug 18, 2026 at 04:28:38PM +0100, Fuad Tabba wrote:
quoted
Hi Vincent,

On Fri, 31 Jul 2026 at 15:36, 'Vincent Donnefort' via kernel-team
[off-list ref] wrote:
quoted
Integrate the pKVM memory reclaim interface with the host's memory
management subsystem.

This allows the host to automatically recover unused memory fom the
hypervisor's heap allocator when the host is under memory pressure.

Tested-by: Fuad Tabba <fuad.tabba@linux.dev>
Signed-off-by: Vincent Donnefort <redacted>
diff --git a/arch/arm64/kvm/pkvm.c b/arch/arm64/kvm/pkvm.c
index d28422f5c3d6..bfbb1266491d 100644
--- a/arch/arm64/kvm/pkvm.c
+++ b/arch/arm64/kvm/pkvm.c
@@ -115,7 +115,7 @@ static int pkvm_hyp_topup(enum pkvm_topup_id id, unsigned long nr_pages)
        return ret;
 }

-static __maybe_unused unsigned long pkvm_hyp_reclaim(enum pkvm_topup_id id, unsigned long target)
+static unsigned long pkvm_hyp_reclaim(enum pkvm_topup_id id, unsigned long target)
This is the first caller of these, so it is where reclaim starts
running against a concurrent top-up.

Nothing marks the pages a top-up just put in allocator->mc as spoken
for, and hyp_allocator_reclaim() ends with an unbounded drain of it,
so a shrink with target 1 hands back the lot. Land that between a
top-up and the retry it was for, and the retry asks again, and
pkvm_call_hyp_req() goes round.
Sorry, I am not sure I follow here.

IIRC, the shrinker will only reclaim half of what is available. So the pressure
should be proportional to what is available and limit races with topup!
It's the ordering, not the amount.

topup and its retry are separate hypercalls, lock dropped between
them, so a shrink on another CPU can slip in, right? .
hyp_allocator_reclaim() drains allocator->mc, where the topup pages
sit, before any chunk, so half still comes out of them first: a target
of 1 fails the retry.
However now looking at it. I wonder if I don't want to ratelimit here the number
of pages reclaimed in one go to limit the time spent at EL2. Especially we do
all that with the allocator lock taken...
Ratelimiting would cap the time under the lock, but it wouldn't stop a
concurrent shrink from taking the topup pages, would it?

Cheers,
/fuad

quoted
Both are driven by memory pressure, so they are busiest together.
Worth holding back what a pending request asked for?
quoted
 {
        struct kvm_hyp_memcache mc;
        struct arm_smccc_res res;
@@ -133,7 +133,7 @@ static __maybe_unused unsigned long pkvm_hyp_reclaim(enum pkvm_topup_id id, unsi
        return reclaimed;
 }

-static __maybe_unused unsigned long pkvm_hyp_reclaimable(enum pkvm_topup_id id)
+static unsigned long pkvm_hyp_reclaimable(enum pkvm_topup_id id)
 {
        return kvm_call_hyp_nvhe(__pkvm_hyp_reclaimable, id);
 }
@@ -342,8 +342,19 @@ void __init pkvm_selftests(void)
 #endif
 }

+static unsigned long pkvm_shrinker_count(struct shrinker *shrink, struct shrink_control *sc)
+{
+       return pkvm_hyp_reclaimable(PKVM_TOPUP_HYP_ALLOC) ?: SHRINK_EMPTY;
+}
+
+static unsigned long pkvm_shrinker_scan(struct shrinker *shrink, struct shrink_control *sc)
+{
+       return pkvm_hyp_reclaim(PKVM_TOPUP_HYP_ALLOC, sc->nr_to_scan);
+}
Returning 0 rather than SHRINK_STOP when reclaim comes back empty
leaves do_shrink_slab() calling this until total_scan runs out, and
each call is an HVC. count_objects() covers the case where there is
nothing at all, but not the one where it goes stale between the two
calls.

Cheers,
/fuad
Ha right, that SHRINK_STOP seems indeed to be the convention for other users.

--
Vincent
quoted
quoted
+
 static int __init finalize_pkvm(void)
 {
+       struct shrinker *pkvm_shrinker;
        int ret;

        if (!is_protected_kvm_enabled() || !is_kvm_arm_initialised())
@@ -359,10 +370,21 @@ static int __init finalize_pkvm(void)
        kmemleak_free_part_phys(hyp_mem_base, hyp_mem_size);

        ret = pkvm_drop_host_privileges();
-       if (ret)
+       if (ret) {
                pr_err("Failed to finalize Hyp protection: %d\n", ret);
+               return ret;
+       }

-       return ret;
+       pkvm_shrinker = shrinker_alloc(0, "pkvm");
+       if (pkvm_shrinker) {
+               pkvm_shrinker->count_objects = pkvm_shrinker_count;
+               pkvm_shrinker->scan_objects = pkvm_shrinker_scan;
+               shrinker_register(pkvm_shrinker);
+       } else {
+               kvm_err("Failed to register shrinker for pKVM\n");
+       }
+
+       return 0;
 }
 device_initcall_sync(finalize_pkvm);

--
2.55.0.508.g3f0d502094-goog

To unsubscribe from this group and stop receiving emails from it, send an email to kernel-team+unsubscribe@android.com.
To unsubscribe from this group and stop receiving emails from it, send an email to kernel-team+unsubscribe@android.com.
  
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help