Re: [PATCH v2 22/25] powerpc32: move xxxxx_dcache_range() functions inline
From: Joakim Tjernlund <hidden>
Date: 2015-09-22 19:34:39
Also in:
lkml
On Tue, 2015-09-22 at 13:58 -0500, Scott Wood wrote:
On Tue, 2015-09-22 at 18:12 +0000, Joakim Tjernlund wrote:quoted
On Tue, 2015-09-22 at 18:51 +0200, Christophe Leroy wrote:quoted
flush/clean/invalidate _dcache_range() functions are all very similar and are quite short. They are mainly used in __dma_sync() perf_event locate them in the top 3 consumming functions during heavy ethernet activity =20 They are good candidate for inlining, as __dma_sync() does almost nothing but calling them =20 Signed-off-by: Christophe Leroy <redacted> --- New in v2 =20 arch/powerpc/include/asm/cacheflush.h | 55 +++++++++++++++++++++++++=
++--
quoted
quoted
arch/powerpc/kernel/misc_32.S | 65 -------------------------=
-----
quoted
quoted
----- arch/powerpc/kernel/ppc_ksyms.c | 2 ++ 3 files changed, 54 insertions(+), 68 deletions(-) =20diff --git a/arch/powerpc/include/asm/cacheflush.h=20b/arch/powerpc/include/asm/cacheflush.h index 6229e6b..6169604 100644--- a/arch/powerpc/include/asm/cacheflush.h +++ b/arch/powerpc/include/asm/cacheflush.h@@ -47,12 +47,61 @@ static inline void=20__flush_dcache_icache_phys(unsigned long physaddr) } #endif =20 -extern void flush_dcache_range(unsigned long start, unsigned long st=
op);
quoted
quoted
#ifdef CONFIG_PPC32 -extern void clean_dcache_range(unsigned long start, unsigned long st=
op);
quoted
quoted
-extern void invalidate_dcache_range(unsigned long start, unsigned lo=
ng=20
quoted
quoted
stop); +/* + * Write any modified data cache blocks out to memory and invalidate=
=20
quoted
quoted
them. + * Does not invalidate the corresponding instruction cache blocks. + */ +static inline void flush_dcache_range(unsigned long start, unsigned =
long=20
quoted
quoted
stop) +{ + void *addr =3D (void *)(start & ~(L1_CACHE_BYTES - 1)); + unsigned int size =3D stop - (unsigned long)addr + (L1_CACHE_BYTE=
S - 1);
quoted
quoted
+ unsigned int i; + + for (i =3D 0; i < size >> L1_CACHE_SHIFT; i++, addr +=3D L1_CACHE=
_BYTES)
quoted
quoted
+ dcbf(addr); + if (i) + mb(); /* sync */ +}=20 This feels optimized for the uncommon case when there is no invalidatio=
n.
=20 If you mean the "if (i)", yes, that looks odd.
Yes.
=20quoted
I THINK it would be better to bail early=20=20 Bail under what conditions?
test for "i =3D 0" and return.=20
=20quoted
and use do { .. } while(--i); instead.=20 GCC knows how to optimize loops. Please don't make them less readable.
Been a while since I checked but it used to be bad att transforming post in= c to pre inc/dec I remain unconvinced until I have seen it. Jocke