From: Anton Blanchard <hidden> Date: 2010-01-12 02:21:51
commit ac4c2a3bbe5db5fc570b1d0ee1e474db7cb22585 (zlib: optimize inffast when
copying direct from output) referenced include/linux/autoconf.h which
is now called include/generated/autoconf.h.
Signed-off-by: Anton Blanchard <redacted>
---
Index: linux-cpumask/arch/powerpc/boot/Makefile
===================================================================
From: Stephen Rothwell <hidden> Date: 2010-01-12 03:01:42
Hi Anton,
On Tue, 12 Jan 2010 13:21:51 +1100 Anton Blanchard [off-list ref] wrote:
commit ac4c2a3bbe5db5fc570b1d0ee1e474db7cb22585 (zlib: optimize inffast when
copying direct from output) referenced include/linux/autoconf.h which
is now called include/generated/autoconf.h.
Even with this fix, you cannot build with a separate object directory.
See my other posting "linux-next: origin tree build failure" ...
--
Cheers,
Stephen Rothwell sfr@canb.auug.org.au
http://www.canb.auug.org.au/~sfr/
Anton Blanchard [off-list ref] wrote on 12/01/2010 03:21:51:
quoted hunk
commit ac4c2a3bbe5db5fc570b1d0ee1e474db7cb22585 (zlib: optimize inffast when
copying direct from output) referenced include/linux/autoconf.h which
is now called include/generated/autoconf.h.
Signed-off-by: Anton Blanchard <redacted>
---
Index: linux-cpumask/arch/powerpc/boot/Makefile
===================================================================
Hi Anton,
On Tue, 12 Jan 2010 13:21:51 +1100 Anton Blanchard [off-list ref] wrote:
quoted
commit ac4c2a3bbe5db5fc570b1d0ee1e474db7cb22585 (zlib: optimize inffast when
copying direct from output) referenced include/linux/autoconf.h which
is now called include/generated/autoconf.h.
Even with this fix, you cannot build with a separate object directory.
See my other posting "linux-next: origin tree build failure" ...
--
How does this work for you?
From 044f40d169bf5fe189d5cb058f56b7cd72675ca4 Mon Sep 17 00:00:00 2001
From: Joakim Tjernlund <redacted>
Date: Tue, 12 Jan 2010 11:20:36 +0100
Subject: [PATCH] powerpc: Fix build breakage due to incorrect location of autoconf.h
commit ac4c2a3bbe5db5fc570b1d0ee1e474db7cb22585 (zlib: optimize inffast when
copying direct from output) referenced include/linux/autoconf.h which
is now called include/generated/autoconf.h.
Also, -I include paths needs to be prefixed with $(srctree)
Signed-off-by: Joakim Tjernlund <redacted>
---
arch/powerpc/boot/Makefile | 4 ++--
1 files changed, 2 insertions(+), 2 deletions(-)
copying direct from output) referenced include/linux/autoconf.h which
is now called include/generated/autoconf.h.
Also, -I include paths needs to be prefixed with $(srctree)
=20
Signed-off-by: Joakim Tjernlund <redacted>
---
arch/powerpc/boot/Makefile | 4 ++--
1 files changed, 2 insertions(+), 2 deletions(-)
=20
Ack, this works for me (seeing as -rc4 doesn't generate uImages w/o it :)
I sent a different patch to Linus yesterday for that.
Seen it now as it is in Linus tree:
1) IMHO it would have been nicer to use #ifdef __KERNEL__
instead of CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS
as then arches that don't define CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS
at all will never use the new optimization or was that what you intended?
2) You really should add an comment in the Makefile about not using
autoconf.h/-D__KERNEL__
From: Benjamin Herrenschmidt <benh@kernel.crashing.org> Date: 2010-01-14 08:58:40
Seen it now as it is in Linus tree:
1) IMHO it would have been nicer to use #ifdef __KERNEL__
instead of CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS
as then arches that don't define CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS
at all will never use the new optimization or was that what you intended?
No, that was on purpose. If an arch doesn't have efficient unaligned
accesses, then they should not use the optimization since it will result
in a lot of unaligned accesses :-) In which case they are better off
falling back to the old byte-by-byte method.
The advantage also of doing it this way is that x86 will benefit from
the optimisation at boot time since it does include autoconf.h in its
boot wrapper (and deals with unaligned accesses just fine at any time)
though something tells me that it won't make much of a difference in
performances on any recent x86 (it might on some of the newer low power
embedded ones, I don't know for sure).
2) You really should add an comment in the Makefile about not using
autoconf.h/-D__KERNEL__
Benjamin Herrenschmidt [off-list ref] wrote on 14/01/2010 09:57:11:
Subject: Re: [PATCH]: powerpc: Fix build breakage due to incorrect location ofautoconf.h
quoted
Seen it now as it is in Linus tree:
1) IMHO it would have been nicer to use #ifdef __KERNEL__
instead of CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS
as then arches that don't define CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS
at all will never use the new optimization or was that what you intended?
No, that was on purpose. If an arch doesn't have efficient unaligned
accesses, then they should not use the optimization since it will result
in a lot of unaligned accesses :-) In which case they are better off
falling back to the old byte-by-byte method.
Not quite, it is only 1 of 4 accesses that uses unaligned and
that accesses is only unaligned 50% in average, it might still
be faster. We will never know now.
The advantage also of doing it this way is that x86 will benefit from
the optimisation at boot time since it does include autoconf.h in its
But so does my suggestion, doesn't it?
boot wrapper (and deals with unaligned accesses just fine at any time)
though something tells me that it won't make much of a difference in
performances on any recent x86 (it might on some of the newer low power
embedded ones, I don't know for sure).
From: Benjamin Herrenschmidt <benh@kernel.crashing.org> Date: 2010-01-14 09:45:07
On Thu, 2010-01-14 at 10:12 +0100, Joakim Tjernlund wrote:
quoted
No, that was on purpose. If an arch doesn't have efficient unaligned
accesses, then they should not use the optimization since it will
result
quoted
in a lot of unaligned accesses :-) In which case they are better off
falling back to the old byte-by-byte method.
Not quite, it is only 1 of 4 accesses that uses unaligned and
that accesses is only unaligned 50% in average, it might still
be faster. We will never know now.
Why ? If you think it's a win, then it's easy to make a patch to turn
it to __KERNEL__ and ask some people from ARM and MIPS or even sparc
land for example to give it a spin. If it's indeed a win, then submit it
to Linus and/or Andrew and there's no reason for it not to go in.
I simply took a more conservative approach for post -rc4
Cheers,
Ben.
Benjamin Herrenschmidt [off-list ref] wrote on 14/01/2010 10:43:44:
On Thu, 2010-01-14 at 10:12 +0100, Joakim Tjernlund wrote:
quoted
quoted
No, that was on purpose. If an arch doesn't have efficient unaligned
accesses, then they should not use the optimization since it will
result
quoted
in a lot of unaligned accesses :-) In which case they are better off
falling back to the old byte-by-byte method.
Not quite, it is only 1 of 4 accesses that uses unaligned and
that accesses is only unaligned 50% in average, it might still
be faster. We will never know now.
Why ? If you think it's a win, then it's easy to make a patch to turn
it to __KERNEL__ and ask some people from ARM and MIPS or even sparc
land for example to give it a spin. If it's indeed a win, then submit it
to Linus and/or Andrew and there's no reason for it not to go in.
I simply took a more conservative approach for post -rc4
Perhaps for the best this late. I will just leave it as is. If ARM/MIPS et. all
wants, they can test it whenever they want.
Jocke
Benjamin Herrenschmidt [off-list ref] wrote on 14/01/2010 10:43:44:
On Thu, 2010-01-14 at 10:12 +0100, Joakim Tjernlund wrote:
quoted
quoted
No, that was on purpose. If an arch doesn't have efficient unaligned
accesses, then they should not use the optimization since it will
result
quoted
in a lot of unaligned accesses :-) In which case they are better off
falling back to the old byte-by-byte method.
Not quite, it is only 1 of 4 accesses that uses unaligned and
that accesses is only unaligned 50% in average, it might still
be faster. We will never know now.
Why ? If you think it's a win, then it's easy to make a patch to turn
it to __KERNEL__ and ask some people from ARM and MIPS or even sparc
land for example to give it a spin. If it's indeed a win, then submit it
to Linus and/or Andrew and there's no reason for it not to go in.
I simply took a more conservative approach for post -rc4
It would probably be a good idea to redefine UP_UNALIGNED macro to
do 2 byte accesses in the non CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS case.
I don't think I will revisit this any time soon so I figured I should
mention it in case someone else wants to try it.
Jocke
From: Benjamin Herrenschmidt <benh@kernel.crashing.org> Date: 2010-01-14 20:01:12
On Thu, 2010-01-14 at 14:27 +0100, Joakim Tjernlund wrote:
It would probably be a good idea to redefine UP_UNALIGNED macro to
do 2 byte accesses in the non CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS
case.
I don't think I will revisit this any time soon so I figured I should
mention it in case someone else wants to try it.
Benjamin Herrenschmidt [off-list ref] wrote on 2010/01/14 09:57:11:
quoted
Seen it now as it is in Linus tree:
1) IMHO it would have been nicer to use #ifdef __KERNEL__
instead of CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS
as then arches that don't define CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS
at all will never use the new optimization or was that what you intended?
No, that was on purpose. If an arch doesn't have efficient unaligned
accesses, then they should not use the optimization since it will result
in a lot of unaligned accesses :-) In which case they are better off
falling back to the old byte-by-byte method.
The advantage also of doing it this way is that x86 will benefit from
the optimisation at boot time since it does include autoconf.h in its
boot wrapper (and deals with unaligned accesses just fine at any time)
though something tells me that it won't make much of a difference in
performances on any recent x86 (it might on some of the newer low power
embedded ones, I don't know for sure).
quoted
2) You really should add an comment in the Makefile about not using
autoconf.h/-D__KERNEL__
That's true :-)
So I fixed the new inflate to be endian independent and now
all arches can use the optimized version. Here goes:
From 4e769486e2520e34f532a2d4bf13ab13f05e3e76 Mon Sep 17 00:00:00 2001
From: Joakim Tjernlund <redacted>
Date: Sun, 24 Jan 2010 11:12:56 +0100
Subject: [PATCH] zlib: Make new optimized inflate endian independent
Commit 6846ee5ca68d81e6baccf0d56221d7a00c1be18b made the
new optimized inflate only available on arch's that
define CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS. This
fixes it by defining our own endian independent versions
of unaligned access.
Signed-off-by: Joakim Tjernlund <redacted>
---
lib/zlib_inflate/inffast.c | 70 +++++++++++++++++++-------------------------
1 files changed, 30 insertions(+), 40 deletions(-)
@@ -8,21 +8,6 @@#include"inflate.h"#include"inffast.h"-/* Only do the unaligned "Faster" variant when-*CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESSisset-*-*Onpowerpc,itwon'tbeaswedon'tincludeautoconf.h-*automaticallyforthebootwrapper,whichisintendedas-*weruninanenvironmentwherewemaynotbeabletodeal-*with(evenrare)alignmentfaults.Inaddition,wedonot-*define__KERNEL__forarch/powerpc/bootunlikex86-*/--#ifdef CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS-#include<asm/unaligned.h>-#include<asm/byteorder.h>-#endif-#ifndef ASMINF/* Allow machine dependent optimization for post-increment or pre-increment.
@@ -282,14 +287,13 @@ void inflate_fast(z_streamp strm, unsigned start)unsignedshortpat16;pat16=*(sout-2+2*OFF);-if(dist==1)-#if defined(__BIG_ENDIAN)-pat16=(pat16&0xff)|((pat16&0xff)<<8);-#elif defined(__LITTLE_ENDIAN)-pat16=(pat16&0xff00)|((pat16&0xff00)>>8);-#else-#error __BIG_ENDIAN nor __LITTLE_ENDIAN is defined-#endif+if(dist==1){+unionuumm;+/* copy one char pattern to both bytes */+mm.us=pat16;+mm.b[0]=mm.b[1];+pat16=mm.us;+}loops=len>>1;doPUP(sout)=pat16;
@@ -298,20 +302,6 @@ void inflate_fast(z_streamp strm, unsigned start)}if(len&1)PUP(out)=PUP(from);-#else /* CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS */-from=out-dist;/* copy direct from output */-do{/* minimum length is three */-PUP(out)=PUP(from);-PUP(out)=PUP(from);-PUP(out)=PUP(from);-len-=3;-}while(len>2);-if(len){-PUP(out)=PUP(from);-if(len>1)-PUP(out)=PUP(from);-}-#endif /* !CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS */}}elseif((op&64)==0){/* 2nd level distance code */--
Benjamin Herrenschmidt [off-list ref] wrote on 2010/01/14 09:57:11:
quoted
quoted
Seen it now as it is in Linus tree:
1) IMHO it would have been nicer to use #ifdef __KERNEL__
instead of CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS
as then arches that don't define CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS
at all will never use the new optimization or was that what you intended?
No, that was on purpose. If an arch doesn't have efficient unaligned
accesses, then they should not use the optimization since it will result
in a lot of unaligned accesses :-) In which case they are better off
falling back to the old byte-by-byte method.
The advantage also of doing it this way is that x86 will benefit from
the optimisation at boot time since it does include autoconf.h in its
boot wrapper (and deals with unaligned accesses just fine at any time)
though something tells me that it won't make much of a difference in
performances on any recent x86 (it might on some of the newer low power
embedded ones, I don't know for sure).
quoted
2) You really should add an comment in the Makefile about not using
autoconf.h/-D__KERNEL__
That's true :-)
So I fixed the new inflate to be endian independent and now
all arches can use the optimized version. Here goes:
No comments so far, I guess that is a good thing :)
Ben, Andrew could either of you carry this patch for me?
Jocke
quoted hunk
From 4e769486e2520e34f532a2d4bf13ab13f05e3e76 Mon Sep 17 00:00:00 2001
From: Joakim Tjernlund <redacted>
Date: Sun, 24 Jan 2010 11:12:56 +0100
Subject: [PATCH] zlib: Make new optimized inflate endian independent
Commit 6846ee5ca68d81e6baccf0d56221d7a00c1be18b made the
new optimized inflate only available on arch's that
define CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS. This
fixes it by defining our own endian independent versions
of unaligned access.
Signed-off-by: Joakim Tjernlund <redacted>
---
lib/zlib_inflate/inffast.c | 70 +++++++++++++++++++-------------------------
1 files changed, 30 insertions(+), 40 deletions(-)
@@ -8,21 +8,6 @@#include"inflate.h"#include"inffast.h"-/* Only do the unaligned "Faster" variant when-*CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESSisset-*-*Onpowerpc,itwon'tbeaswedon'tincludeautoconf.h-*automaticallyforthebootwrapper,whichisintendedas-*weruninanenvironmentwherewemaynotbeabletodeal-*with(evenrare)alignmentfaults.Inaddition,wedonot-*define__KERNEL__forarch/powerpc/bootunlikex86-*/--#ifdef CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS-#include<asm/unaligned.h>-#include<asm/byteorder.h>-#endif-#ifndef ASMINF/* Allow machine dependent optimization for post-increment or pre-increment.
@@ -282,14 +287,13 @@ void inflate_fast(z_streamp strm, unsigned start)unsignedshortpat16;pat16=*(sout-2+2*OFF);-if(dist==1)-#if defined(__BIG_ENDIAN)-pat16=(pat16&0xff)|((pat16&0xff)<<8);-#elif defined(__LITTLE_ENDIAN)-pat16=(pat16&0xff00)|((pat16&0xff00)>>8);-#else-#error __BIG_ENDIAN nor __LITTLE_ENDIAN is defined-#endif+if(dist==1){+unionuumm;+/* copy one char pattern to both bytes */+mm.us=pat16;+mm.b[0]=mm.b[1];+pat16=mm.us;+}loops=len>>1;doPUP(sout)=pat16;
@@ -298,20 +302,6 @@ void inflate_fast(z_streamp strm, unsigned start)}if(len&1)PUP(out)=PUP(from);-#else /* CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS */-from=out-dist;/* copy direct from output */-do{/* minimum length is three */-PUP(out)=PUP(from);-PUP(out)=PUP(from);-PUP(out)=PUP(from);-len-=3;-}while(len>2);-if(len){-PUP(out)=PUP(from);-if(len>1)-PUP(out)=PUP(from);-}-#endif /* !CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS */}}elseif((op&64)==0){/* 2nd level distance code */--
From: Andrew Morton <akpm@linux-foundation.org> Date: 2010-01-28 01:05:57
On Mon, 25 Jan 2010 09:19:59 +0100
Joakim Tjernlund [off-list ref] wrote:
Commit 6846ee5ca68d81e6baccf0d56221d7a00c1be18b made the
new optimized inflate only available on arch's that
define CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS. This
fixes it by defining our own endian independent versions
of unaligned access.
(I hope I picked up the right version of
whatever-it-was-i-was-supposed-to-pick-up - I wasn't paying attention).
The changelog sucks. You say the patch fixes "it", but what is "it"?
Is "it" a build error? If so, what? Or is "it" a
make-optimized-inflate-available-on-ppc patch, in which case it's a
"feature"?
Confuzzzzzed.
Andrew Morton [off-list ref] wrote on 2010/01/28 02:05:36:
On Mon, 25 Jan 2010 09:19:59 +0100
Joakim Tjernlund [off-list ref] wrote:
quoted
Commit 6846ee5ca68d81e6baccf0d56221d7a00c1be18b made the
new optimized inflate only available on arch's that
define CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS. This
fixes it by defining our own endian independent versions
of unaligned access.
(I hope I picked up the right version of
whatever-it-was-i-was-supposed-to-pick-up - I wasn't paying attention).
It was the right one.
The changelog sucks. You say the patch fixes "it", but what is "it"?
Is "it" a build error? If so, what? Or is "it" a
make-optimized-inflate-available-on-ppc patch, in which case it's a
"feature"?
Here is a new version with a somewhat better commit msg and
checkpatch fixes.
Jocke
BTW, I get duplicate mails from your patch handling tools,
one to Joakim.Tjernlund@transmode.se and one to joakim.tjernlund@transmode.se
From 612bfb4cc6cc55243c2a4f536ae2ae71065b57b0 Mon Sep 17 00:00:00 2001
From: Joakim Tjernlund <redacted>
Date: Sun, 24 Jan 2010 11:12:56 +0100
Subject: [PATCH] zlib: Make new optimized inflate endian independent
Commit 6846ee5ca68d81e6baccf0d56221d7a00c1be18b made the
new optimized inflate only available on arch's that
define CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS. This
will again enable the optimization for all arch's by
by defining our own endian independent version
of unaligned access. As an added bonus, arch's that
define CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS do a
plain load instead.
Signed-off-by: Joakim Tjernlund <redacted>
---
v2:
- fix checkpatch complaints.
- Improve commit msg.
lib/zlib_inflate/inffast.c | 70 +++++++++++++++++++-------------------------
1 files changed, 30 insertions(+), 40 deletions(-)
@@ -8,21 +8,6 @@#include"inflate.h"#include"inffast.h"-/* Only do the unaligned "Faster" variant when-*CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESSisset-*-*Onpowerpc,itwon'tbeaswedon'tincludeautoconf.h-*automaticallyforthebootwrapper,whichisintendedas-*weruninanenvironmentwherewemaynotbeabletodeal-*with(evenrare)alignmentfaults.Inaddition,wedonot-*define__KERNEL__forarch/powerpc/bootunlikex86-*/--#ifdef CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS-#include<asm/unaligned.h>-#include<asm/byteorder.h>-#endif-#ifndef ASMINF/* Allow machine dependent optimization for post-increment or pre-increment.
@@ -282,14 +287,13 @@ void inflate_fast(z_streamp strm, unsigned start)unsignedshortpat16;pat16=*(sout-2+2*OFF);-if(dist==1)-#if defined(__BIG_ENDIAN)-pat16=(pat16&0xff)|((pat16&0xff)<<8);-#elif defined(__LITTLE_ENDIAN)-pat16=(pat16&0xff00)|((pat16&0xff00)>>8);-#else-#error __BIG_ENDIAN nor __LITTLE_ENDIAN is defined-#endif+if(dist==1){+unionuumm;+/* copy one char pattern to both bytes */+mm.us=pat16;+mm.b[0]=mm.b[1];+pat16=mm.us;+}loops=len>>1;doPUP(sout)=pat16;
@@ -298,20 +302,6 @@ void inflate_fast(z_streamp strm, unsigned start)}if(len&1)PUP(out)=PUP(from);-#else /* CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS */-from=out-dist;/* copy direct from output */-do{/* minimum length is three */-PUP(out)=PUP(from);-PUP(out)=PUP(from);-PUP(out)=PUP(from);-len-=3;-}while(len>2);-if(len){-PUP(out)=PUP(from);-if(len>1)-PUP(out)=PUP(from);-}-#endif /* !CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS */}}elseif((op&64)==0){/* 2nd level distance code */--
From: Andrew Morton <akpm@linux-foundation.org> Date: 2010-01-29 23:49:58
On Thu, 28 Jan 2010 09:52:41 +0100
Joakim Tjernlund [off-list ref] wrote:
Commit 6846ee5ca68d81e6baccf0d56221d7a00c1be18b made the
new optimized inflate only available on arch's that
define CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS. This
will again enable the optimization for all arch's by
by defining our own endian independent version
of unaligned access. As an added bonus, arch's that
define CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS do a
plain load instead.
Given that we're at -rc5, that changelog says to me "this is a 2.6.34
patch".
If that is wrong, please tell me why.
Andrew Morton [off-list ref] wrote on 2010/01/30 00:49:24:
On Thu, 28 Jan 2010 09:52:41 +0100
Joakim Tjernlund [off-list ref] wrote:
quoted
Commit 6846ee5ca68d81e6baccf0d56221d7a00c1be18b made the
new optimized inflate only available on arch's that
define CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS. This
will again enable the optimization for all arch's by
by defining our own endian independent version
of unaligned access. As an added bonus, arch's that
define CONFIG_HAVE_EFFICIENT_UNALIGNED_ACCESS do a
plain load instead.
Given that we're at -rc5, that changelog says to me "this is a 2.6.34
patch".
If that is wrong, please tell me why.