Jeff,
Please apply. The rather long patch description is from the submitter,
Norbert Eicker, I don't know if that's alright, or if I should ask to
have it trimmed.
Thanks,
--linas
From: Norbert Eicker <redacted>
I found out that the spidernet-driver is unable to send fragmented IP
frames.
Let me just recall the basic structure of "normal" UDP/IP/Ethernet
frames (that actually work):
- It starts with the Ethernet header (dest MAC, src MAC, etc.)
- The next part is occupied by the IP header (version info, length of
packet, id=0, fragment offset=0, checksum, from / to address, etc.)
- Then comes the UDP header (src / dest port, length, checksum)
- Actual payload
- Ethernet checksum
Now what's different for IP fragment:
- The IP header has id set to some value (same for all fragments),
offset is set appropriately (i.e. 0 for first fragment, following
according to size of other fragments), size is the length of the frame.
- UDP header is unchanged. I.e. length is according to full UDP
datagram, not just the part within the actual frame! But this is only
true within the first frame: all following frames don't have a valid
UDP-header at all.
The spidernet silicon seems to be quite intelligent: It's able to
compute (IP / UDP / Ethernet) checksums on the fly and tests if frames
are conforming to RFC -- at least conforming to RFC on complete frames.
But IP fragments are different as explained above:
I.e. for IP fragments containing part of a UDP datagram it sees
incompatible length in the headers for IP and UDP in the first frame
and, thus, skips this frame. But the content *is* correct for IP
fragments. For all following frames it finds (most probably) no valid
UDP header at all. But this *is* also correct for IP fragments.
The Linux IP-stack seems to be clever in this point. It expects the
spidernet to calculate the checksum (since the module claims to be able
to do so) and marks the skb's for "normal" frames accordingly
(ip_summed set to CHECKSUM_HW).
But for the IP fragments it does not expect the driver to be capable to
handle the frames appropriately. Thus all checksums are allready
computed. This is also flaged within the skb (ip_summed set to
CHECKSUM_NONE).
Unfortunately the spidernet driver ignores that hints. It tries to send
the IP fragments of UDP datagrams as normal UDP/IP frames. Since they
have different structure the silicon detects them the be not
"well-formed" and skips them.
The following one-liner against 2.6.21-rc2 changes this behavior. If the
IP-stack claims to have done the checksumming, the driver should not
try to checksum (and analyze) the frame but send it as is.
Signed-off-by: Norbert Eicker <redacted>
Signed-off-by: Linas Vepstas <redacted>
----
From: Jeff Garzik <hidden> Date: 2007-03-09 16:53:42
Linas Vepstas wrote:
Jeff,
Please apply. The rather long patch description is from the submitter,
Norbert Eicker, I don't know if that's alright, or if I should ask to
have it trimmed.
Thanks,
--linas
From: Norbert Eicker <redacted>
I found out that the spidernet-driver is unable to send fragmented IP
frames.
Let me just recall the basic structure of "normal" UDP/IP/Ethernet
frames (that actually work):
- It starts with the Ethernet header (dest MAC, src MAC, etc.)
- The next part is occupied by the IP header (version info, length of
packet, id=0, fragment offset=0, checksum, from / to address, etc.)
- Then comes the UDP header (src / dest port, length, checksum)
- Actual payload
- Ethernet checksum
Now what's different for IP fragment:
- The IP header has id set to some value (same for all fragments),
offset is set appropriately (i.e. 0 for first fragment, following
according to size of other fragments), size is the length of the frame.
- UDP header is unchanged. I.e. length is according to full UDP
datagram, not just the part within the actual frame! But this is only
true within the first frame: all following frames don't have a valid
UDP-header at all.
The spidernet silicon seems to be quite intelligent: It's able to
compute (IP / UDP / Ethernet) checksums on the fly and tests if frames
are conforming to RFC -- at least conforming to RFC on complete frames.
But IP fragments are different as explained above:
I.e. for IP fragments containing part of a UDP datagram it sees
incompatible length in the headers for IP and UDP in the first frame
and, thus, skips this frame. But the content *is* correct for IP
fragments. For all following frames it finds (most probably) no valid
UDP header at all. But this *is* also correct for IP fragments.
The Linux IP-stack seems to be clever in this point. It expects the
spidernet to calculate the checksum (since the module claims to be able
to do so) and marks the skb's for "normal" frames accordingly
(ip_summed set to CHECKSUM_HW).
But for the IP fragments it does not expect the driver to be capable to
handle the frames appropriately. Thus all checksums are allready
computed. This is also flaged within the skb (ip_summed set to
CHECKSUM_NONE).
Unfortunately the spidernet driver ignores that hints. It tries to send
the IP fragments of UDP datagrams as normal UDP/IP frames. Since they
have different structure the silicon detects them the be not
"well-formed" and skips them.
The following one-liner against 2.6.21-rc2 changes this behavior. If the
IP-stack claims to have done the checksumming, the driver should not
try to checksum (and analyze) the frame but send it as is.
Signed-off-by: Norbert Eicker <redacted>
Signed-off-by: Linas Vepstas <redacted>
are you sure it can't send out fragmented IP frames? what's really
going on here?
patch was corrupted anyway, and could not be applied...
On Friday, 9. March 2007 17:53, Jeff Garzik wrote:
Linas Vepstas wrote:
quoted
Jeff,
Please apply. The rather long patch description is from the submitter,
Norbert Eicker, I don't know if that's alright, or if I should ask to
have it trimmed.
Thanks,
--linas
From: Norbert Eicker <redacted>
Signed-off-by: Norbert Eicker <redacted>
Signed-off-by: Linas Vepstas <redacted>
are you sure it can't send out fragmented IP frames? what's really
going on here?
Pretty sure that fragmented IP frames are not send out. Here's a small test
using ttcp and tcpdump reproducing the problem (I assume a MTU of 1500):
You need two node, lets call them src and dest.
On dest run 'tcpdump -nvvv host src'.
Furthermore start on dest: 'ttcp -u -r -s'
and on src: 'ttcp -u -t -s -l 1472 -n 4 dest
tcpdump will show 10 frames received from src, one of size 32 (UDP length 4),
four of size 1500 (UDP length 1472) (the payload), five more of size 32 (UDP
length 4). The size 32 frames are some startup/closing frames of ttcp.
Start on dest again: 'ttcp -u -r -s'
and on src: 'ttcp -u -t -s -l 1473 -n 4 dest
Now tcpdump only sees the six startup/closing frames. The payload frames
(which are fragmented by the IP-stack) disappear somehow.
With the patch applied you see in the second case the startup/closing frames
plus 4 pairs of frames in between. Each pair consists of one frame of size
1500 followed by one of size 21 (with an offset of 1480), i.e. the two
fragments of the too large UDP-datagram.
On src tcpdump will show the fragmented IP-frames in both cases (with and
without patch applied).
patch was corrupted anyway, and could not be applied...
On Friday, 9. March 2007 17:53, Jeff Garzik wrote:
quoted
Linas Vepstas wrote:
quoted
Please apply. The rather long patch description is from the submitter,
Norbert Eicker, I don't know if that's alright, or if I should ask to
have it trimmed.
Thanks,
--linas
From: Norbert Eicker <redacted>
Signed-off-by: Norbert Eicker <redacted>
Signed-off-by: Linas Vepstas <redacted>
are you sure it can't send out fragmented IP frames? what's really
going on here?
Pretty sure that fragmented IP frames are not send out. Here's a small test
using ttcp and tcpdump reproducing the problem (I assume a MTU of 1500):
Hence if I understand that correctly, NFS over UDP doesn't work at all as NFS
uses 8 KiB UDP packets by default?
Gr{oetje,eeting}s,
Geert
--
Geert Uytterhoeven -- Sony Network and Software Technology Center Europe (NSCE)
Geert.Uytterhoeven@sonycom.com ------- The Corporate Village, Da Vincilaan 7-D1
Voice +32-2-7008453 Fax +32-2-7008622 ---------------- B-1935 Zaventem, Belgium
On Monday, 12. March 2007 09:28, Geert Uytterhoeven wrote:
On Sat, 10 Mar 2007, Norbert Eicker wrote:
quoted
On Friday, 9. March 2007 17:53, Jeff Garzik wrote:
quoted
Linas Vepstas wrote:
quoted
Please apply. The rather long patch description is from the
submitter, Norbert Eicker, I don't know if that's alright, or
if I should ask to have it trimmed.
Thanks,
--linas
From: Norbert Eicker <redacted>
Signed-off-by: Norbert Eicker <redacted>
Signed-off-by: Linas Vepstas <redacted>
are you sure it can't send out fragmented IP frames? what's
really going on here?
Pretty sure that fragmented IP frames are not send out. Here's a
small test using ttcp and tcpdump reproducing the problem (I assume
a MTU of 1500):
Hence if I understand that correctly, NFS over UDP doesn't work at
all as NFS uses 8 KiB UDP packets by default?
NFS seems to work but considering what I saw in tcpdump it uses TCP
instead of UDP. I have not yet found out why NFS seems to default to
TCP on Cell (the man-pages claims that UDP is the default).
If I explicitely demand to use UDP it indeed does not work at all.
Again if I use tcpdump on both sides, I see that the NFS-server claims
to send out IP fragments containing the UDP datagrams (larger than
MTU). The NFS-client does not see any of these fragments.
Using the patched spidernet module makes NFS work like a charm.
Norbert
--
Fon ++49-(0)2461/61-1492
http://www.fz-juelich.de
On Monday, 12. March 2007 09:28, Geert Uytterhoeven wrote:
quoted
On Sat, 10 Mar 2007, Norbert Eicker wrote:
quoted
On Friday, 9. March 2007 17:53, Jeff Garzik wrote:
quoted
Linas Vepstas wrote:
quoted
Please apply. The rather long patch description is from the
submitter, Norbert Eicker, I don't know if that's alright, or
if I should ask to have it trimmed.
Thanks,
--linas
From: Norbert Eicker <redacted>
Signed-off-by: Norbert Eicker <redacted>
Signed-off-by: Linas Vepstas <redacted>
are you sure it can't send out fragmented IP frames? what's
really going on here?
Pretty sure that fragmented IP frames are not send out. Here's a
small test using ttcp and tcpdump reproducing the problem (I assume
a MTU of 1500):
Hence if I understand that correctly, NFS over UDP doesn't work at
all as NFS uses 8 KiB UDP packets by default?
NFS seems to work but considering what I saw in tcpdump it uses TCP
instead of UDP. I have not yet found out why NFS seems to default to
TCP on Cell (the man-pages claims that UDP is the default).
Hmm, I just checked and for me it defaults to UDP.
If I explicitely demand to use UDP it indeed does not work at all.
And NFS over UDP works fine on my PS3 with the gelic Ethernet driver.
Ethereal (on the server side) did show fragmented packets.
However, I remember having NFS problems when using a different NFS server
before. I just retried using that server, and got these error messages from
the gelic driver:
| error in received descriptor found, data_status=x70002300, data_error=x06100000
| ERROR DESTROY:6100000
| error in received descriptor found, data_status=x70002900, data_error=x06100000
| ERROR DESTROY:6100000
Ethereal (on the server side) did show fragmented packets with retransmissions.
So sometimes it works, sometimes it doesn't?
The differences between the 2 NFS servers are:
- working one runs 2.6.11, non-working one runs 2.4.17
- working one is on the same subnet as the PS3, non-working one isn't.
- both work fine with e.g. my laptop as an NFS client
- both have nfs-kernel-server 1:1.0.6-3.1
- both have 3c905 Tornado Ethernet
For completeness, sometimes (1-2 times a week) I do see a similar gelic error
message when using the working NFS server, but it recovers fine and never is a
problem.
Gr{oetje,eeting}s,
Geert
--
Geert Uytterhoeven -- Sony Network and Software Technology Center Europe (NSCE)
Geert.Uytterhoeven@sonycom.com ------- The Corporate Village, Da Vincilaan 7-D1
Voice +32-2-7008453 Fax +32-2-7008622 ---------------- B-1935 Zaventem, Belgium
On Monday, 12. March 2007 11:44, Geert Uytterhoeven wrote:
On Mon, 12 Mar 2007, Norbert Eicker wrote:
quoted
On Monday, 12. March 2007 09:28, Geert Uytterhoeven wrote:
quoted
On Sat, 10 Mar 2007, Norbert Eicker wrote:
quoted
On Friday, 9. March 2007 17:53, Jeff Garzik wrote:
quoted
Linas Vepstas wrote:
quoted
Please apply. The rather long patch description is from the
submitter, Norbert Eicker, I don't know if that's alright,
or if I should ask to have it trimmed.
Thanks,
--linas
From: Norbert Eicker <redacted>
Signed-off-by: Norbert Eicker <redacted>
Signed-off-by: Linas Vepstas <redacted>
are you sure it can't send out fragmented IP frames? what's
really going on here?
Pretty sure that fragmented IP frames are not send out. Here's
a small test using ttcp and tcpdump reproducing the problem (I
assume a MTU of 1500):
Hence if I understand that correctly, NFS over UDP doesn't work
at all as NFS uses 8 KiB UDP packets by default?
NFS seems to work but considering what I saw in tcpdump it uses TCP
instead of UDP. I have not yet found out why NFS seems to default
to TCP on Cell (the man-pages claims that UDP is the default).
Hmm, I just checked and for me it defaults to UDP.
quoted
If I explicitely demand to use UDP it indeed does not work at all.
And NFS over UDP works fine on my PS3 with the gelic Ethernet driver.
Ethereal (on the server side) did show fragmented packets.
Well, the gelic is a different driver, so for me it's no surprise it
shows different behavior.
In fact from a quick inspection I would guess that it has not the bug
that's inside the spidernet driver: The section I patched in
spider_net.c shows up similarly around line 839 in gelic_net.c. But an
important difference is in the lines before: It only appears in the
else-branch of an 'if'. This one tests if ip_summed is set to
CHECKSUM_PARTIAL.
I.e. gelic_net.c has the test that is missing in spider_net.c.
However, I remember having NFS problems when using a different NFS
server before. I just retried using that server, and got these error
messages from
the gelic driver:
| error in received descriptor found, data_status=x70002300,
| data_error=x06100000 ERROR DESTROY:6100000
| error in received descriptor found, data_status=x70002900,
| data_error=x06100000 ERROR DESTROY:6100000
Ethereal (on the server side) did show fragmented packets with
retransmissions.
So sometimes it works, sometimes it doesn't?
The differences between the 2 NFS servers are:
- working one runs 2.6.11, non-working one runs 2.4.17
- working one is on the same subnet as the PS3, non-working one
isn't. - both work fine with e.g. my laptop as an NFS client
- both have nfs-kernel-server 1:1.0.6-3.1
- both have 3c905 Tornado Ethernet
For completeness, sometimes (1-2 times a week) I do see a similar
gelic error message when using the working NFS server, but it
recovers fine and never is a problem.
IMHO the problem I reported is not an NFS problem at all (there might be
more problems in this arena;-)). As I have pointed out you can see the
problem even with ttcp (or some other tool sending UDP frames larger
than MTU - some_header_length).
Norbert
--
Fon ++49-(0)2461/61-1492
http://www.fz-juelich.de