On 3/10/26 02:37, Baolin Wang wrote:
On 3/7/26 4:02 PM, Barry Song wrote:
quoted
On Sat, Mar 7, 2026 at 10:22 AM Baolin Wang
[off-list ref] wrote:
quoted
Thanks.
Yes. In addition, this will involve many architectures’ implementations
and their differing TLB flush mechanisms, so it’s difficult to make a
reasonable per-architecture measurement. If any architecture has a more
efficient flush method, I’d prefer to implement an architecture‑specific
clear_flush_young_ptes().
Right! Since TLBI is usually quite expensive, I wonder if a generic
implementation for architectures lacking clear_flush_young_ptes()
might benefit from something like the below (just a very rough idea):
int clear_flush_young_ptes(struct vm_area_struct *vma,
unsigned long addr, pte_t *ptep, unsigned int nr)
{
unsigned long curr_addr = addr;
int young = 0;
while (nr--) {
young |= ptep_test_and_clear_young(vma, curr_addr,
ptep);
ptep++;
curr_addr += PAGE_SIZE;
}
if (young)
flush_tlb_range(vma, addr, curr_addr);
return young;
}
I understand your point. I’m concerned that I can’t test this patch on
every architecture to validate the benefits. Anyway, let me try this on
my X86 machine first.
In any case, please make that a follow-up patch :)
--
Cheers,
David