From: Zi Yan <ziy@nvidia.com> Date: 2026-08-31 19:26:02
Hi all,
This patchset removes PG_private to make space for upcoming PG_folio
(reserved as __PG_folio) for identifying pages from a folio (more details
in Note below). Instead of checking PG_private, all code is changed to
check page/folio->private != NULL instead.
MM people are cc'd on all patches and subsystem people are cc'd on the
cover letter and corresponding patches.
Overview
===
Most code uses folio_attach/detach/change_private() functions, so folio
refcount is increased and decreased when folio->private is set and reset,
respectively. There is no need to change them.
Changes are needed for exceptional users:
1. zsmalloc uses PG_private to indicate first component zpdesc page and
page->private is used to store zspage in zpdesc. To remove PG_private,
is_first_zpdesc() is replaced by pointer comparison.
2. kernel/events/ring_buffer.c stores page order in page->private.
Replacing PG_private with page->private != NULL works.
3. drivers/xen/grant-table.c stores xen_page_foreign in page->private,
where on 32-bit, a pointer to xen_page_foreign is stored; on 64-bit,
page->private is used as xen_page_foreign. PG_private check is replaced
by page->private != NULL on 32-bit for xen_page_foreign deallocation.
On 64-bit, page->private is cleared unconditionally since {domid=0,
gref=0} (xen_page_foreign can be 0) is valid.
4. fs/crypto/crypto.c stores a folio pointer in page->private, PG_private
checks are replaced by page->private != NULL.
5. fs/erofs has two different uses:
5a. folio->private is used to form a reversed list of
the outputs of readahead_folio(). readahead_folio_last() is added to
output folios in reversed order, so that ->private is no longer needed.
5b. folio->private is used as an in-flight I/O counter. Convert the
code to use folio_attach/detach/get_private() and add bias==1 to the
counter to avoid folio->private being zero.
6. fs/nfs/write.c: folio refcount maintenance is in a bigger scope than
folio->private. So folio_attach/detach/get_private() is not used.
Nothing to change.
7. fs/f2fs uses attach_page_private() to first reset folio->private then
immediately sets PAGE_PRIVATE_NOT_POINTER bit on it. Change it to use
attach_page_private() to set PAGE_PRIVATE_NOT_POINTER bit directly to
avoid folio->private == NULL gap inside set_page_private_##name().
8. hugetlb uses folio_change_private(folio, NULL) without folio refcount
maintenance. Change it to folio->private = NULL.
After the above changes, PG_private ops are converted to
page/folio->private ops.
folio_test_fs_private() is added to check filesystem-only private data by
excluding swapcache and hugetlb folios, because swapcache folios overlap
swp_entry_t swap with ->private and hugetlb sets its own flags in
->private.
Note
===
1. KPF_PRIVATE is removed after PG_private is removed.
2. Documentation/mm/hugetlbfs_reserv.rst is outdated, so I did not remove
PG_private related text. It should be rewritten.
3. PG_folio is planned to be set on every page from a folio in
page_rmappable_folio(), so folios with any order (currently
PG_large_rmappable is used to identify >0 order folios) can be
identified, vm_insert_*() can correctly reject all folios, and rmap code
can accept only folios. Eventually, page_folio() will return NULL for
non-folio pages by checking PG_folio, but before that all existing users
that treat compound pages as folios will be converted.
Tests
===
1. allmodconfig build passed.
2. zsmalloc is tested using ext4 on a 1GB lz4 zram:
2a. zram load + zsmalloc compaction;
2b. concurrent zspage migration via memory compaction;
2c. confirmed that multi-page zspages actually formed.
Details: https://github.com/x-y-z/linux-dev/blob/b4/remove-pg_private/test_zsmalloc.md
3. erofs is tested on images created with -C4096 and lz4hc, lzma,
deflate, and zstd algorithms:
3a. cold read of all files, verify checksums match source;
3b. readahead + reclaim/migration race.
Details: https://github.com/x-y-z/linux-dev/blob/b4/remove-pg_private/test_erofs.md
4. fscrypt is tested on software-encrypted ext4 with writes to exercise
bounce pages.
Details: https://github.com/x-y-z/linux-dev/blob/b4/remove-pg_private/test_fscrypt.md
5. f2fs is tested on an image with inline_data,compress_algorithm=lz4:
5a. INLINE_INODE — lots of tiny files;
5b. REF_RESOURCE + general writeback — buffered write churn with fsync;
5c. ONGOING_MIGRATION — force GC / page migration;
5d. ATOMIC_WRITE — atomic-write ioctl path.
Details: https://github.com/x-y-z/linux-dev/blob/b4/remove-pg_private/test_f2fs.md
(I did not run xfstests)
6. MM selftests passed.
LLM use
===
Claude was used to form a concrete plan on what code needs to be changed
and how to change them. The plan was reviewed by Codex until no issue was
spotted.
Plan is at: https://github.com/x-y-z/linux-dev/blob/b4/remove-pg_private/plan.md
I then followed the plan to make code changes. I did bounce ideas with
Claude how to change fs/erofs, since I did not like the original idea.
After each change, I asked Claude to review my code and git commit message.
I also asked Claude to give me test plans (see above).
At last, Codex was used to review all patches.
Comments and suggestions are welcome. Thanks.
Assisted-by: Claude:claude-opus-4-8
Assisted-by: Codex:gpt-5
Signed-off-by: Zi Yan <ziy@nvidia.com>
---
Changes in v2:
1. removed is_first_zpdesc() in patch 1 and open coded the checks.
2. fixed wording in patch 2's commit message and clarified page_private()
also works when ring buffer's AUX page order is 0.
3. removed the empty loop in 64-bit gnttab_pages_set_private().
4. clarified folio->private will be reset to NULL by
fscrypt_free_bounce_page() in the commit message.
5. clarified why hugetlb needs to restore hugetlb_vmemmap_optimized.
6. renamed readahead_folio_reverse() readahead_folio_last() and
reimplemented readahead_folio_last() by adding a new readahead_control
private member, _forward, and a new helper __readahead_advance().
7. added a bias, 1, to erofs I/O counter, so that folio->private stays non
NULL between folio_attach_private() and folio_detach_private().
8. converted more call sites to use folio_test_fs_private().
- Link to v1: https://lore.kernel.org/r/20260731-remove-pg_private-v1-0-142c97ba3562@nvidia.com
---
Zi Yan (14):
mm/zsmalloc: replace PG_private with pointer comparison
perf/ring_buffer: stop using PG_private as AUX page high-order marker
xen/grant-table: stop setting PG_private on pages for grant mapping
fscrypt: stop setting PG_private on bounce page
mm/hugetlb: use direct assignment instead of folio_change_private()
f2fs: stop using PG_private
erofs: mm/pagemap: add readahead_folio_last() to avoid folio->private
erofs: use folio_attach/detach_private() instead of direct assignment
mm/page-flags: check page/folio->private instead of PG_private
mm/page-flags: introduce folio_test_fs_private()
treewide: remove folio_set/clear_private()
treewide: replace PagePrivate() with page_private()
treewide: adjust comments on PagePrivate and PG_private
mm/page-flags: remove PG_private
Documentation/admin-guide/kdump/vmcoreinfo.rst | 2 +-
Documentation/filesystems/vfs.rst | 6 +--
arch/x86/events/intel/bts.c | 3 --
arch/x86/events/intel/pt.c | 6 +--
drivers/md/md-bitmap.c | 6 +--
drivers/xen/balloon.c | 5 +++
drivers/xen/grant-table.c | 11 +++--
fs/ceph/addr.c | 8 ++--
fs/crypto/crypto.c | 2 -
fs/erofs/data.c | 16 ++++---
fs/erofs/zdata.c | 13 ++----
fs/f2fs/f2fs.h | 8 ++--
fs/nfs/file.c | 4 +-
fs/nfs/write.c | 2 -
fs/proc/page.c | 1 -
fs/ubifs/file.c | 8 ++--
include/linux/buffer_head.h | 6 ---
include/linux/kernel-page-flags.h | 1 -
include/linux/mm.h | 35 +++++++++------
include/linux/mm_types.h | 4 +-
include/linux/page-flags.h | 38 +++++++++++-----
include/linux/pagemap.h | 60 +++++++++++++++++++++-----
include/trace/events/mmflags.h | 2 +-
include/trace/events/pagemap.h | 2 +-
kernel/events/ring_buffer.c | 7 ++-
kernel/vmcore_info.c | 1 -
mm/huge_memory.c | 2 +-
mm/hugetlb.c | 6 +--
mm/migrate.c | 3 +-
mm/page-writeback.c | 2 +-
mm/vmscan.c | 2 +-
mm/zpdesc.h | 2 +-
mm/zsmalloc.c | 24 +++--------
tools/mm/page-types.c | 2 -
34 files changed, 163 insertions(+), 137 deletions(-)
---
base-commit: 443451c85ca8d6389d34b1299decada62128f1fe
change-id: 20260728-remove-pg_private-cfe926c7f83c
Best regards,
--
Yan, Zi
From: Zi Yan <ziy@nvidia.com> Date: 2026-08-31 19:26:13
folio->private != NULL indicates a folio carries private data, replacing
PG_private. All PG_private users are converted. Remove PG_private and
reserve the space as __PG_folio for future use.
Also update files in Documentation. hugetlbfs_reserv.rst is outdated and
left unchanged. It should be rewritten.
Assisted-by: Claude:claude-opus-4-8
Assisted-by: Codex:gpt-5
Signed-off-by: Zi Yan <ziy@nvidia.com>
To: Andrew Morton <akpm@linux-foundation.org>
To: Baoquan He <baoquan.he@linux.dev>
To: Mike Rapoport <rppt@kernel.org>
To: Pasha Tatashin <pasha.tatashin@soleen.com>
To: Pratyush Yadav <pratyush@kernel.org>
To: Jonathan Corbet <corbet@lwn.net>
To: "Matthew Wilcox (Oracle)" <willy@infradead.org>
To: Jan Kara <jack@suse.cz>
To: David Hildenbrand <david@kernel.org>
To: Steven Rostedt <rostedt@goodmis.org>
To: Masami Hiramatsu <mhiramat@kernel.org>
Cc: Dave Young <ruirui.yang@linux.dev>
Cc: Shuah Khan <skhan@linuxfoundation.org>
Cc: Lorenzo Stoakes <ljs@kernel.org>
Cc: "Liam R. Howlett" <liam@infradead.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: kexec@lists.infradead.org
Cc: linux-doc@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
Cc: linux-fsdevel@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: linux-trace-kernel@vger.kernel.org
---
Documentation/admin-guide/kdump/vmcoreinfo.rst | 2 +-
Documentation/filesystems/vfs.rst | 6 +++---
include/linux/page-flags.h | 19 ++-----------------
include/trace/events/mmflags.h | 2 +-
kernel/vmcore_info.c | 1 -
5 files changed, 7 insertions(+), 23 deletions(-)
@@ -325,7 +325,7 @@ NR_FREE_PAGES On linux-2.6.21 or later, the number of free pages is in vm_stat[NR_FREE_PAGES]. Used to get the number of free pages.-PG_lru|PG_private|PG_swapcache|PG_swapbacked|PG_hwpoison|PG_head_mask+PG_lru|PG_swapcache|PG_swapbacked|PG_hwpoison|PG_head_mask -------------------------------------------------------------------------- Page attributes. These flags are used to filter various unnecessary for
@@ -649,8 +649,8 @@ Writeback. The first can be used independently to the others. The VM can try to release clean pages in order to reuse them. To do this it can call-->release_folio on clean folios with the private-flag set. Clean pages without PagePrivate and with no external references+->release_folio on clean folios with folio->private set. Clean pages+without folio->private set and with no external references will be released without notice being given to the address_space. To achieve this functionality, pages need to be placed on an LRU with
@@ -674,7 +674,7 @@ filemap_fdatawait_range, to wait for all writeback to complete. An address_space handler may attach extra information to a page, typically using the 'private' field in the 'struct page'. If such-information is attached, the PG_Private flag should be set. This will+information is attached, non-NULL 'private' field will cause various VM routines to make extra calls into the address_space handler to deal with that data.
@@ -105,7 +101,7 @@ enum pageflags {PG_owner_2,/* Owner use. If pagecache, fs may use */PG_arch_1,PG_reserved,-PG_private,/* If pagecache, has fs-private data */+__PG_folio,/* Do not use: reserved for folio identification */PG_private_2,/* If pagecache, has fs aux data */PG_reclaim,/* To be reclaimed asap */PG_swapbacked,/* Page is backed by RAM/swap */
@@ -584,17 +580,6 @@ static __always_inline bool folio_test_private(const struct folio *folio)returnfolio->private;}-static__always_inlineintPagePrivate(conststructpage*page)-{-return!!page->private;-}--/* no-ops during transition */-static__always_inlinevoidfolio_set_private(structfolio*folio){}-static__always_inlinevoidfolio_clear_private(structfolio*folio){}-static__always_inlinevoidSetPagePrivate(structpage*page){}-static__always_inlinevoidClearPagePrivate(structpage*page){}-FOLIO_FLAG(private_2,FOLIO_HEAD_PAGE)/* owner_2 can be set on tail pages for anon memory */
From: Zi Yan <ziy@nvidia.com> Date: 2026-08-31 19:26:16
folio_test_fs_private() wraps folio->private != NULL check and excludes
swapcache and hugetlb folios, since swapcache uses swp_entry_t overlapping
with folio->private and hugetlb sets its own flags in folio->private.
Replace open code with the helper, since core MM does this check
frequently.
No functional change intended.
Assisted-by: Claude:claude-opus-4-8
Assisted-by: Codex:gpt-5
Signed-off-by: Zi Yan <ziy@nvidia.com>
To: Andrew Morton <akpm@linux-foundation.org>
To: David Hildenbrand <david@kernel.org>
To: Steven Rostedt <rostedt@goodmis.org>
To: Masami Hiramatsu <mhiramat@kernel.org>
To: Lorenzo Stoakes <ljs@kernel.org>
Cc: "Liam R. Howlett" <liam@infradead.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: Zi Yan <ziy@nvidia.com>
Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: Nico Pache <nico.pache@linux.dev>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: Dev Jain <dev.jain@arm.com>
Cc: Barry Song <baohua@kernel.org>
Cc: Lance Yang <lance.yang@linux.dev>
Cc: Usama Arif <usama.arif@linux.dev>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Joshua Hahn <joshua.hahnjy@gmail.com>
Cc: Rakie Kim <rakie.kim@sk.com>
Cc: Byungchul Park <byungchul@sk.com>
Cc: Gregory Price <gourry@gourry.net>
Cc: Ying Huang <ying.huang@linux.alibaba.com>
Cc: Alistair Popple <apopple@nvidia.com>
Cc: linux-kernel@vger.kernel.org
Cc: linux-fsdevel@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: linux-trace-kernel@vger.kernel.org
---
include/linux/mm.h | 4 +---
include/linux/page-flags.h | 21 ++++++++++++++++++---
include/trace/events/pagemap.h | 4 +---
mm/huge_memory.c | 4 +---
mm/migrate.c | 3 +--
mm/page-writeback.c | 5 +----
mm/vmscan.c | 3 +--
7 files changed, 24 insertions(+), 20 deletions(-)
@@ -955,8 +955,7 @@ static void folio_check_dirty_writeback(struct folio *folio,*writeback=folio_test_writeback(folio);/* Verify dirty/writeback state if the filesystem supports it */-if(!(folio_test_private(folio)&&!folio_test_swapcache(folio)&&-!folio_test_hugetlb(folio)))+if(!folio_test_fs_private(folio))return;mapping=folio_mapping(folio);
From: Zi Yan <ziy@nvidia.com> Date: 2026-08-31 19:26:18
After the changes of the prior commits, page/folio->private != NULL is now
equivalent to checking PG_private.
Stop checking PG_private on pages and folios and use page/folio->private
instead, except swapcache and hugetlb folios, because the former uses a
field (swp_entry_t swap) overlapping with ->private and the latter sets its
flags in ->private. Exclude swapcache and hugetlb when the code is meant to
check PG_private only.
folio_expected_ref_count() can be called without folio lock, so annotate
folio_test_private() with data_race() to avoid triggering race condition
checks. While at it, annotate folio->mapping too.
folio_set/clear_private() and Set/ClearPagePrivate() become no-ops.
PG_private is no longer checked at page free time.
Remove KPF_PRIVATE since PG_private is no longer used.
Assisted-by: Claude:claude-opus-4-8
Assisted-by: Codex:gpt-5
Signed-off-by: Zi Yan <ziy@nvidia.com>
To: Andrew Morton <akpm@linux-foundation.org>
To: David Hildenbrand <david@kernel.org>
To: Steven Rostedt <rostedt@goodmis.org>
To: Masami Hiramatsu <mhiramat@kernel.org>
To: Lorenzo Stoakes <ljs@kernel.org>
To: "Matthew Wilcox (Oracle)" <willy@infradead.org>
To: Jan Kara <jack@suse.cz>
To: Johannes Weiner <hannes@cmpxchg.org>
Cc: "Liam R. Howlett" <liam@infradead.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Mike Rapoport <rppt@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: Zi Yan <ziy@nvidia.com>
Cc: Baolin Wang <baolin.wang@linux.alibaba.com>
Cc: Nico Pache <nico.pache@linux.dev>
Cc: Ryan Roberts <ryan.roberts@arm.com>
Cc: Dev Jain <dev.jain@arm.com>
Cc: Barry Song <baohua@kernel.org>
Cc: Lance Yang <lance.yang@linux.dev>
Cc: Usama Arif <usama.arif@linux.dev>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Joshua Hahn <joshua.hahnjy@gmail.com>
Cc: Rakie Kim <rakie.kim@sk.com>
Cc: Byungchul Park <byungchul@sk.com>
Cc: Gregory Price <gourry@gourry.net>
Cc: Ying Huang <ying.huang@linux.alibaba.com>
Cc: Alistair Popple <apopple@nvidia.com>
Cc: Qi Zheng <qi.zheng@linux.dev>
Cc: Shakeel Butt <shakeel.butt@linux.dev>
Cc: Kairui Song <kasong@tencent.com>
Cc: Axel Rasmussen <axelrasmussen@google.com>
Cc: Yuanchu Xie <yuanchu@google.com>
Cc: Wei Xu <weixugc@google.com>
Cc: linux-kernel@vger.kernel.org
Cc: linux-fsdevel@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: linux-trace-kernel@vger.kernel.org
---
fs/proc/page.c | 1 -
include/linux/kernel-page-flags.h | 1 -
include/linux/mm.h | 22 +++++++++++++++-------
include/linux/page-flags.h | 26 +++++++++++++++++++++-----
include/trace/events/pagemap.h | 4 +++-
mm/huge_memory.c | 4 +++-
mm/migrate.c | 3 ++-
mm/page-writeback.c | 5 ++++-
mm/vmscan.c | 3 ++-
tools/mm/page-types.c | 2 --
10 files changed, 50 insertions(+), 21 deletions(-)
@@ -3062,10 +3062,18 @@ static inline int folio_expected_ref_count(const struct folio *folio)ref_count+=folio_test_swapcache(folio)<<order;if(!folio_test_anon(folio)){-/* One reference per page from the pagecache. */-ref_count+=!!folio->mapping<<order;-/* One reference from PG_private. */-ref_count+=folio_test_private(folio);+/*+*Onereferenceperpagefromthepagecache.+*Usedata_race()sincefoliomightnotbelocked.+*/+ref_count+=!!data_race(folio->mapping)<<order;+/*+*Onereferencefromfilesystemprivatedata.+*Usedata_race()sincefoliomightnotbelocked.+*/+ref_count+=data_race(folio_test_private(folio))&&+!folio_test_hugetlb(folio)&&+!folio_test_swapcache(folio);}/* One reference per page table mapping. */
@@ -578,7 +578,23 @@ FOLIO_FLAG(swapbacked, FOLIO_HEAD_PAGE)*foritsownpurposes.*-PG_privateandPG_private_2causerelease_folio()andcotobeinvoked*/-PAGEFLAG(Private,private,PF_ANY)++static__always_inlineboolfolio_test_private(conststructfolio*folio)+{+returnfolio->private;+}++static__always_inlineintPagePrivate(conststructpage*page)+{+return!!page->private;+}++/* no-ops during transition */+static__always_inlinevoidfolio_set_private(structfolio*folio){}+static__always_inlinevoidfolio_clear_private(structfolio*folio){}+static__always_inlinevoidSetPagePrivate(structpage*page){}+static__always_inlinevoidClearPagePrivate(structpage*page){}+FOLIO_FLAG(private_2,FOLIO_HEAD_PAGE)/* owner_2 can be set on tail pages for anon memory */
@@ -955,7 +955,8 @@ static void folio_check_dirty_writeback(struct folio *folio,*writeback=folio_test_writeback(folio);/* Verify dirty/writeback state if the filesystem supports it */-if(!folio_test_private(folio))+if(!(folio_test_private(folio)&&!folio_test_swapcache(folio)&&+!folio_test_hugetlb(folio)))return;mapping=folio_mapping(folio);
Thank you for your contribution! Sashiko AI review found 3 potential issue(s) to consider:
New issues:
- [High] Lockless TOCTOU race in `folio_expected_ref_count()` miscalculates expected references for shmem folios during swap cache removal, allowing `F_SEAL_WRITE` to bypass GUP pins.
- [Low] Lockless reads of `folio->private` via `folio_test_fs_private()` lack `data_race()` or `READ_ONCE()` annotations in several paths, triggering KCSAN warnings.
Pre-existing issues:
- [High] In `migrate_folio_unmap()`, an orphaned folio with private data is blindly assumed to have `buffer_head`s, causing a type confusion crash for filesystems that use other structures.
--
commit cfacca810fa1b9878c94f6054258d5e3c3b6824f
Author: Zi Yan [off-list ref]
mm/page-flags: check page/folio->private instead of PG_private
This patch changes the kernel to check page/folio->private instead of
PG_private, except for swapcache and hugetlb folios. It annotates
folio_expected_ref_count to prevent data races and removes KPF_PRIVATE.
@@ -3050,10 +3050,18 @@ static inline int folio_expected_ref_count(const struct folio *folio) ref_count += folio_test_swapcache(folio) << order; if (!folio_test_anon(folio)) {- /* One reference per page from the pagecache. */- ref_count += !!folio->mapping << order;- /* One reference from PG_private. */- ref_count += folio_test_private(folio);+ /*+ * One reference per page from the pagecache.+ * Use data_race() since folio might not be locked.+ */+ ref_count += !!data_race(folio->mapping) << order;+ /*+ * One reference from filesystem private data.+ * Use data_race() since folio might not be locked.+ */+ ref_count += data_race(folio_test_private(folio)) &&+ !folio_test_hugetlb(folio) &&+ !folio_test_swapcache(folio);
[Severity: High]
Could this lockless evaluation of folio->private and PG_swapcache lead to a
TOCTOU race for shmem folios during swap cache removal?
During __delete_from_swap_cache, folio->swap.val (which aliases
folio->private) is cleared before PG_swapcache. Without memory barriers, a
lockless reader like memfd_tag_pins calling folio_expected_ref_count could
observe the stale non-zero folio->private and the newly cleared PG_swapcache.
This would evaluate the condition above as true, falsely inflating the
expected refcount by 1. If the folio has exactly one extra GUP pin, the
inflated expected refcount would match the actual refcount, bypassing the
F_SEAL_WRITE protections.
[Severity: Low]
Should this lockless read of folio->private be annotated with data_race() or
READ_ONCE()?
Since tracepoints like TRACE_EVENT(mm_lru_insertion) are often executed
without the folio lock, concurrent modifications to folio->private could
trigger KCSAN data race warnings here.
@@ -4858,7 +4858,9 @@ static int split_huge_pages_pid(int pid, unsigned long vaddr_start,*willtrytodropitbeforesplitandthencheckifthefolio*canbesplitornot.Soskipthecheckhere.*/-if(!folio_test_private(folio)&&+if(!(folio_test_private(folio)&&+!folio_test_swapcache(folio)&&+!folio_test_hugetlb(folio))&&
[Severity: Low]
Can this similarly trigger a KCSAN data race warning?
This block in split_huge_pages_pid checks the private flags locklessly before
folio_trylock is called, so reading folio->private here without data_race()
could cause issues.
@@ -1327,7 +1327,8 @@ static int migrate_folio_unmap(new_folio_t get_new_folio,*freethemetadata,sothepagecanbefreed.*/if(!src->mapping){-if(folio_test_private(src)){+if(folio_test_private(src)&&!folio_test_swapcache(src)&&+!folio_test_hugetlb(src)){try_to_free_buffers(src);
[Severity: High]
This is a pre-existing issue, but does this generic migration code blindly
assume all private data on an orphaned folio corresponds to buffer heads?
If a folio is orphaned after truncation (src->mapping is NULL) but retains
private data because an earlier release failed, try_to_free_buffers is
unconditionally called. For filesystems like iomap, folio->private contains
custom structures (like iomap_folio_state), not a linked list of buffer
heads. Calling try_to_free_buffers here could lead to type confusion and a
kernel panic when accessing what it expects to be bh->b_this_page.
[Severity: Low]
Is a data_race() annotation missing here as well?
Functions like filemap_dirty_folio can be called locklessly (for example,
from zap_pte_range with only the pte lock held), resulting in an unannotated
read of folio->private that might race with concurrent modifications.
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260831-remove-pg_private-v2-0-3668159cd9e8@nvidia.com?part=9
Thank you for your contribution! Sashiko AI review found 1 potential issue(s) to consider:
- [High] Removal of PG_private safety check causes leaked filesystem private data to go undetected.
--
commit 9b27f18c6460a8f0a47a836096a558c0ddbb879d
Author: Zi Yan [off-list ref]
Subject: mm/page-flags: remove PG_private
This patch removes the PG_private flag and repurposes its space as
__PG_folio. It replaces the flag checks with checking if
folio->private != NULL to determine if a folio carries private data. It
also updates relevant documentation to reflect these changes.
[Severity: High]
Does removing PG_private from this mask cause leaked filesystem private data
to go undetected?
By removing the PG_private check without adding a replacement check for
folio->private != NULL on the head page, the memory allocator's validation
routines no longer catch this leak.
Instead, when a buggy filesystem or subsystem frees a page with a non-NULL
private field, free_pages_prepare() in mm/page_alloc.c will unconditionally
zero it without warning:
mm/page_alloc.c:free_pages_prepare() {
...
page->flags.f &= ~PAGE_FLAGS_CHECK_AT_PREP;
page->private = 0;
...
}
Could this silently erase the leaked private data (such as buffer heads) and
mask the underlying memory leak?
--
Sashiko AI review · https://sashiko.dev/#/patchset/20260831-remove-pg_private-v2-0-3668159cd9e8@nvidia.com?part=14
After the changes of the prior commits, page/folio->private != NULL is now
equivalent to checking PG_private.
Stop checking PG_private on pages and folios and use page/folio->private
instead, except swapcache and hugetlb folios, because the former uses a
field (swp_entry_t swap) overlapping with ->private and the latter sets its
flags in ->private. Exclude swapcache and hugetlb when the code is meant to
check PG_private only.
folio_expected_ref_count() can be called without folio lock, so annotate
folio_test_private() with data_race() to avoid triggering race condition
checks. While at it, annotate folio->mapping too.
folio_set/clear_private() and Set/ClearPagePrivate() become no-ops.
PG_private is no longer checked at page free time.
Remove KPF_PRIVATE since PG_private is no longer used.
@@ -3062,10 +3062,18 @@ static inline int folio_expected_ref_count(const struct folio *folio)ref_count+=folio_test_swapcache(folio)<<order;if(!folio_test_anon(folio)){-/* One reference per page from the pagecache. */-ref_count+=!!folio->mapping<<order;-/* One reference from PG_private. */-ref_count+=folio_test_private(folio);+/*+*Onereferenceperpagefromthepagecache.+*Usedata_race()sincefoliomightnotbelocked.+*/+ref_count+=!!data_race(folio->mapping)<<order;+/*+*Onereferencefromfilesystemprivatedata.+*Usedata_race()sincefoliomightnotbelocked.+*/+ref_count+=data_race(folio_test_private(folio))&&+!folio_test_hugetlb(folio)&&+!folio_test_swapcache(folio);
Sashiko said [Severity: High]:
Could this lockless evaluation of folio->private and PG_swapcache lead to a
TOCTOU race for shmem folios during swap cache removal?
During __delete_from_swap_cache, folio->swap.val (which aliases
folio->private) is cleared before PG_swapcache. Without memory barriers, a
lockless reader like memfd_tag_pins calling folio_expected_ref_count could
observe the stale non-zero folio->private and the newly cleared PG_swapcache.
This would evaluate the condition above as true, falsely inflating the
expected refcount by 1. If the folio has exactly one extra GUP pin, the
inflated expected refcount would match the actual refcount, bypassing the
F_SEAL_WRITE protections.
Answer:
Yes, it is a problem, since folio->swap.val and PG_swapcache cannot be
read as a whole, when __swap_cache_do_del_folio() (was
__delete_from_swap_cache()) clears folio->swap.val first then
PG_swapcache.
Fortunately, PG_swapbacked is stable during the process. So the code can
exclude swapcache folios by checking PG_swapbacked instead.
In the next patch, folio_test_fs_private() will be changed to replace the
new check: folio_test_private() && !folio_test_hugetlb(folio) &&
!folio_test_swapbacked().
<snip>
Sashiko said:
hould this lockless read of folio->private be annotated with data_race() or
READ_ONCE()?
Since tracepoints like TRACE_EVENT(mm_lru_insertion) are often executed
without the folio lock, concurrent modifications to folio->private could
trigger KCSAN data race warnings here.
Answer:
Yes, will annotate with data_race() here.
@@ -1327,7 +1327,8 @@ static int migrate_folio_unmap(new_folio_t get_new_folio,*freethemetadata,sothepagecanbefreed.*/if(!src->mapping){-if(folio_test_private(src)){+if(folio_test_private(src)&&!folio_test_swapcache(src)&&+!folio_test_hugetlb(src)){try_to_free_buffers(src);gotoout;}
Sashiko said [Severity: High]:
This is a pre-existing issue, but does this generic migration code blindly
assume all private data on an orphaned folio corresponds to buffer heads?
If a folio is orphaned after truncation (src->mapping is NULL) but retains
private data because an earlier release failed, try_to_free_buffers is
unconditionally called. For filesystems like iomap, folio->private contains
custom structures (like iomap_folio_state), not a linked list of buffer
heads. Calling try_to_free_buffers here could lead to type confusion and a
kernel panic when accessing what it expects to be bh->b_this_page.
Answer:
Not an issue. With the help of gpt-5.6-sol, this issue only affects
buffer heads. Folios using iomap always clears folio->private before
folio->mapping is cleared.
folio->private != NULL indicates a folio carries private data, replacing
PG_private. All PG_private users are converted. Remove PG_private and
reserve the space as __PG_folio for future use.
Also update files in Documentation. hugetlbfs_reserv.rst is outdated and
left unchanged. It should be rewritten.
@@ -105,7 +101,7 @@ enum pageflags {PG_owner_2,/* Owner use. If pagecache, fs may use */PG_arch_1,PG_reserved,-PG_private,/* If pagecache, has fs-private data */+__PG_folio,/* Do not use: reserved for folio identification */
Sashiko asked how to detect leaked filesystem private data during page
free time after PG_private is removed.
Answer copied from RFC:
After the conversion, folio/page should have an elevated refcount whenever
->private is set. That would help detect leaked private data. I tried to
enforce ->private needs to be NULL at page free time[1], but that might
cause trouble for certain use cases.
[1] https://lore.kernel.org/all/20260223032641.1859381-1-ziy@nvidia.com/
--
Best Regards,
Yan, Zi
Hmm, just for consistency sake, can we create a:
#define __DEF_PAGEFLAG_NAME(_name) { 1UL << __PG_##_name, __stringify(_name) }
Which is similar to:
#define DEF_PAGEFLAG_NAME(_name) { 1UL << PG_##_name, __stringify(_name) }
But adds the "__" to the name. Then the above would look like:
DEF_PAGEFLAG_NAME(reserved), \
__DEF_PAGEFLAG_NAME(folio), \
DEF_PAGEFLAG_NAME(private_2), \
DEF_PAGEFLAG_NAME(writeback), \
Where the __DEF_PAGEFLAG_NAME() with the "__" still stands out, but the
code looks better than open coding it in the middle and making one wonder
why it was open coded. (It took me a bit to noticed the "__" difference.)
-- Steve
Hmm, just for consistency sake, can we create a:
#define __DEF_PAGEFLAG_NAME(_name) { 1UL << __PG_##_name, __stringify(_name) }
Which is similar to:
#define DEF_PAGEFLAG_NAME(_name) { 1UL << PG_##_name, __stringify(_name) }
But adds the "__" to the name. Then the above would look like:
DEF_PAGEFLAG_NAME(reserved), \
__DEF_PAGEFLAG_NAME(folio), \
DEF_PAGEFLAG_NAME(private_2), \
DEF_PAGEFLAG_NAME(writeback), \
Where the __DEF_PAGEFLAG_NAME() with the "__" still stands out, but the
code looks better than open coding it in the middle and making one wonder
why it was open coded. (It took me a bit to noticed the "__" difference.)
No problem.
BTW, the plan is to add PG_folio after this series is picked up by Andrew/David.
So the new __DEF_PAGEFLAG_NAME() will have no user then and can be deleted.
It should be fine, right?
Best Regards,
Yan, Zi
From: Steven Rostedt <rostedt@goodmis.org> Date: 2026-09-01 17:49:52
On Tue, 01 Sep 2026 12:01:24 -0400
Zi Yan [off-list ref] wrote:
No problem.
BTW, the plan is to add PG_folio after this series is picked up by Andrew/David.
So the new __DEF_PAGEFLAG_NAME() will have no user then and can be deleted.
It should be fine, right?
Yeah, then we just remove that macro and change the one user of
__DEF_PAGEFLAG_NAME() to DEF_PAGEFLAG_NAME().
It will make that patch even easier ;-)
-- Steve
From: Usama Arif <usama.arif@linux.dev> Date: 2026-09-02 17:09:47
On Mon, 31 Aug 2026 15:25:37 -0400 Zi Yan [off-list ref] wrote:
quoted hunk
folio->private != NULL indicates a folio carries private data, replacing
PG_private. All PG_private users are converted. Remove PG_private and
reserve the space as __PG_folio for future use.
Also update files in Documentation. hugetlbfs_reserv.rst is outdated and
left unchanged. It should be rewritten.
Assisted-by: Claude:claude-opus-4-8
Assisted-by: Codex:gpt-5
Signed-off-by: Zi Yan <ziy@nvidia.com>
To: Andrew Morton <akpm@linux-foundation.org>
To: Baoquan He <baoquan.he@linux.dev>
To: Mike Rapoport <rppt@kernel.org>
To: Pasha Tatashin <pasha.tatashin@soleen.com>
To: Pratyush Yadav <pratyush@kernel.org>
To: Jonathan Corbet <corbet@lwn.net>
To: "Matthew Wilcox (Oracle)" <willy@infradead.org>
To: Jan Kara <jack@suse.cz>
To: David Hildenbrand <david@kernel.org>
To: Steven Rostedt <rostedt@goodmis.org>
To: Masami Hiramatsu <mhiramat@kernel.org>
Cc: Dave Young <ruirui.yang@linux.dev>
Cc: Shuah Khan <skhan@linuxfoundation.org>
Cc: Lorenzo Stoakes <ljs@kernel.org>
Cc: "Liam R. Howlett" <liam@infradead.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: kexec@lists.infradead.org
Cc: linux-doc@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
Cc: linux-fsdevel@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: linux-trace-kernel@vger.kernel.org
---
Documentation/admin-guide/kdump/vmcoreinfo.rst | 2 +-
Documentation/filesystems/vfs.rst | 6 +++---
include/linux/page-flags.h | 19 ++-----------------
include/trace/events/mmflags.h | 2 +-
kernel/vmcore_info.c | 1 -
5 files changed, 7 insertions(+), 23 deletions(-)
@@ -325,7 +325,7 @@ NR_FREE_PAGES On linux-2.6.21 or later, the number of free pages is in vm_stat[NR_FREE_PAGES]. Used to get the number of free pages.-PG_lru|PG_private|PG_swapcache|PG_swapbacked|PG_hwpoison|PG_head_mask+PG_lru|PG_swapcache|PG_swapbacked|PG_hwpoison|PG_head_mask -------------------------------------------------------------------------- Page attributes. These flags are used to filter various unnecessary for
@@ -649,8 +649,8 @@ Writeback. The first can be used independently to the others. The VM can try to release clean pages in order to reuse them. To do this it can call-->release_folio on clean folios with the private-flag set. Clean pages without PagePrivate and with no external references+->release_folio on clean folios with folio->private set. Clean pages+without folio->private set and with no external references will be released without notice being given to the address_space. To achieve this functionality, pages need to be placed on an LRU with
@@ -674,7 +674,7 @@ filemap_fdatawait_range, to wait for all writeback to complete. An address_space handler may attach extra information to a page, typically using the 'private' field in the 'struct page'. If such-information is attached, the PG_Private flag should be set. This will+information is attached, non-NULL 'private' field will cause various VM routines to make extra calls into the address_space handler to deal with that data.
@@ -105,7 +101,7 @@ enum pageflags {PG_owner_2,/* Owner use. If pagecache, fs may use */PG_arch_1,PG_reserved,-PG_private,/* If pagecache, has fs-private data */+__PG_folio,/* Do not use: reserved for folio identification */PG_private_2,/* If pagecache, has fs aux data */PG_reclaim,/* To be reclaimed asap */PG_swapbacked,/* Page is backed by RAM/swap */
@@ -584,17 +580,6 @@ static __always_inline bool folio_test_private(const struct folio *folio)returnfolio->private;}-static__always_inlineintPagePrivate(conststructpage*page)-{-return!!page->private;-}--/* no-ops during transition */-static__always_inlinevoidfolio_set_private(structfolio*folio){}-static__always_inlinevoidfolio_clear_private(structfolio*folio){}-static__always_inlinevoidSetPagePrivate(structpage*page){}-static__always_inlinevoidClearPagePrivate(structpage*page){}-FOLIO_FLAG(private_2,FOLIO_HEAD_PAGE)/* owner_2 can be set on tail pages for anon memory */
+ kdump maintainers and reviewers.
I believe makedumpfile reads VMCOREINFO. Removing it here, might cause issues
for older makedumpfile versions at crashdump?
Hopefully kdump folks will be able to comment better.
From: Zi Yan <ziy@nvidia.com> Date: 2026-09-02 17:57:26
On 2 Sep 2026, at 13:09, Usama Arif wrote:
On Mon, 31 Aug 2026 15:25:37 -0400 Zi Yan [off-list ref] wrote:
quoted
folio->private != NULL indicates a folio carries private data, replacing
PG_private. All PG_private users are converted. Remove PG_private and
reserve the space as __PG_folio for future use.
Also update files in Documentation. hugetlbfs_reserv.rst is outdated and
left unchanged. It should be rewritten.
Assisted-by: Claude:claude-opus-4-8
Assisted-by: Codex:gpt-5
Signed-off-by: Zi Yan <ziy@nvidia.com>
To: Andrew Morton <akpm@linux-foundation.org>
To: Baoquan He <baoquan.he@linux.dev>
To: Mike Rapoport <rppt@kernel.org>
To: Pasha Tatashin <pasha.tatashin@soleen.com>
To: Pratyush Yadav <pratyush@kernel.org>
To: Jonathan Corbet <corbet@lwn.net>
To: "Matthew Wilcox (Oracle)" <willy@infradead.org>
To: Jan Kara <jack@suse.cz>
To: David Hildenbrand <david@kernel.org>
To: Steven Rostedt <rostedt@goodmis.org>
To: Masami Hiramatsu <mhiramat@kernel.org>
Cc: Dave Young <ruirui.yang@linux.dev>
Cc: Shuah Khan <skhan@linuxfoundation.org>
Cc: Lorenzo Stoakes <ljs@kernel.org>
Cc: "Liam R. Howlett" <liam@infradead.org>
Cc: Vlastimil Babka <vbabka@kernel.org>
Cc: Suren Baghdasaryan <surenb@google.com>
Cc: Michal Hocko <mhocko@suse.com>
Cc: Mathieu Desnoyers <mathieu.desnoyers@efficios.com>
Cc: kexec@lists.infradead.org
Cc: linux-doc@vger.kernel.org
Cc: linux-kernel@vger.kernel.org
Cc: linux-fsdevel@vger.kernel.org
Cc: linux-mm@kvack.org
Cc: linux-trace-kernel@vger.kernel.org
---
Documentation/admin-guide/kdump/vmcoreinfo.rst | 2 +-
Documentation/filesystems/vfs.rst | 6 +++---
include/linux/page-flags.h | 19 ++-----------------
include/trace/events/mmflags.h | 2 +-
kernel/vmcore_info.c | 1 -
5 files changed, 7 insertions(+), 23 deletions(-)
@@ -325,7 +325,7 @@ NR_FREE_PAGES On linux-2.6.21 or later, the number of free pages is in vm_stat[NR_FREE_PAGES]. Used to get the number of free pages.-PG_lru|PG_private|PG_swapcache|PG_swapbacked|PG_hwpoison|PG_head_mask+PG_lru|PG_swapcache|PG_swapbacked|PG_hwpoison|PG_head_mask -------------------------------------------------------------------------- Page attributes. These flags are used to filter various unnecessary for
@@ -649,8 +649,8 @@ Writeback. The first can be used independently to the others. The VM can try to release clean pages in order to reuse them. To do this it can call-->release_folio on clean folios with the private-flag set. Clean pages without PagePrivate and with no external references+->release_folio on clean folios with folio->private set. Clean pages+without folio->private set and with no external references will be released without notice being given to the address_space. To achieve this functionality, pages need to be placed on an LRU with
@@ -674,7 +674,7 @@ filemap_fdatawait_range, to wait for all writeback to complete. An address_space handler may attach extra information to a page, typically using the 'private' field in the 'struct page'. If such-information is attached, the PG_Private flag should be set. This will+information is attached, non-NULL 'private' field will cause various VM routines to make extra calls into the address_space handler to deal with that data.
@@ -105,7 +101,7 @@ enum pageflags {PG_owner_2,/* Owner use. If pagecache, fs may use */PG_arch_1,PG_reserved,-PG_private,/* If pagecache, has fs-private data */+__PG_folio,/* Do not use: reserved for folio identification */PG_private_2,/* If pagecache, has fs aux data */PG_reclaim,/* To be reclaimed asap */PG_swapbacked,/* Page is backed by RAM/swap */
@@ -584,17 +580,6 @@ static __always_inline bool folio_test_private(const struct folio *folio)returnfolio->private;}-static__always_inlineintPagePrivate(conststructpage*page)-{-return!!page->private;-}--/* no-ops during transition */-static__always_inlinevoidfolio_set_private(structfolio*folio){}-static__always_inlinevoidfolio_clear_private(structfolio*folio){}-static__always_inlinevoidSetPagePrivate(structpage*page){}-static__always_inlinevoidClearPagePrivate(structpage*page){}-FOLIO_FLAG(private_2,FOLIO_HEAD_PAGE)/* owner_2 can be set on tail pages for anon memory */
I believe makedumpfile reads VMCOREINFO. Removing it here, might cause issues
for older makedumpfile versions at crashdump?
The expectation is that kdump userspace tools will adapt to this change.
Later, PG_folio will be added to identify folios and has the same value
of PG_private, so preserving PG_private will not work then.
Hopefully kdump folks will be able to comment better.
Hello:
This patch was applied to jaegeuk/f2fs.git (dev)
by Jaegeuk Kim [off-list ref]:
On Mon, 31 Aug 2026 15:25:23 -0400 you wrote:
Hi all,
This patchset removes PG_private to make space for upcoming PG_folio
(reserved as __PG_folio) for identifying pages from a folio (more details
in Note below). Instead of checking PG_private, all code is changed to
check page/folio->private != NULL instead.
[...]