I am unable to get networking to work with 2.6.18-mm1 on my system.
But 2.6.18 kernel on same system works fine. Here is some info about
the system/debug attempts. Attached are the lspci output and config.
Appreciate any help. Please let me know if you need more info.
Suka
System info:
x326, 2 CPU (AMD Opteron Processor 250)
Kernel info:
$ uname -a
Linux elm3b166 2.6.18-mm1 #4 SMP PREEMPT Tue Sep 26 18:11:58 PDT 2006
x86_64 GNU/Linux
Config tokens differing between the 2.6.18 kernel that works and
the 2.6.18-mm1 that does not are:
Tokens in 2.6.18 but not in 2.6.18-mm1 config
CONFIG_SCSI_FC_ATTRS=y
CONFIG_SCSI_SATA_SIL=y
CONFIG_SCSI_SATA=y
Tokens in 2.6.18-mm1 but not in 2.6.18 config
CONFIG_PROC_SYSCTL=y
CONFIG_SATA_SIL=y
CONFIG_ATA=y
CONFIG_ARCH_POPULATES_NODE_MAP=y
CONFIG_CRYPTO_ALGAPI=y
CONFIG_MICROCODE_OLD_INTERFACE=y
CONFIG_BLOCK=y
CONFIG_VIDEO_V4L1_COMPAT=y
CONFIG_ZONE_DMA=y
CONFIG_FB_DDC=y
All drivers compiled into kernel in both cases.
Debug info:
Checked hardware connections :-)
(Rebooting on 2.6.18 kernel works - consistently)
$ ethtool -i eth0
driver: e1000
version: 7.2.7-k2
firmware-version: N/A
$ ip addr
seems fine (up, broadcasting etc)
$ ip -s link
shows no errors/drops/overruns
$ ip route
shows the correct gw
$ ethtool -S eth0
shows non-zero tx/rx packets/bytes but *rx_missed_errors*
quite large (~138K) and increasing over time
$ ping <own-ip-addr>
works fine
$ ping <gateway>
no response.
$ tcpdump -i eth0 host <broken-host>
while pinging gateway, tcpdump shows messages like:
18:03:45.936161 arp who-has <gateway> tell <broken-host>
(Config file and lspci output are attached)
I am unable to get networking to work with 2.6.18-mm1 on my system.
But 2.6.18 kernel on same system works fine. Here is some info about
the system/debug attempts. Attached are the lspci output and config.
Appreciate any help. Please let me know if you need more info.
Suka
System info:
x326, 2 CPU (AMD Opteron Processor 250)
Kernel info:
$ uname -a
Linux elm3b166 2.6.18-mm1 #4 SMP PREEMPT Tue Sep 26 18:11:58 PDT 2006
x86_64 GNU/Linux
Config tokens differing between the 2.6.18 kernel that works and
the 2.6.18-mm1 that does not are:
Tokens in 2.6.18 but not in 2.6.18-mm1 config
CONFIG_SCSI_FC_ATTRS=y
CONFIG_SCSI_SATA_SIL=y
CONFIG_SCSI_SATA=y
Tokens in 2.6.18-mm1 but not in 2.6.18 config
CONFIG_PROC_SYSCTL=y
CONFIG_SATA_SIL=y
CONFIG_ATA=y
CONFIG_ARCH_POPULATES_NODE_MAP=y
CONFIG_CRYPTO_ALGAPI=y
CONFIG_MICROCODE_OLD_INTERFACE=y
CONFIG_BLOCK=y
CONFIG_VIDEO_V4L1_COMPAT=y
CONFIG_ZONE_DMA=y
CONFIG_FB_DDC=y
All drivers compiled into kernel in both cases.
Debug info:
Checked hardware connections :-)
(Rebooting on 2.6.18 kernel works - consistently)
$ ethtool -i eth0
driver: e1000
version: 7.2.7-k2
firmware-version: N/A
$ ip addr
seems fine (up, broadcasting etc)
$ ip -s link
shows no errors/drops/overruns
$ ip route
shows the correct gw
$ ethtool -S eth0
shows non-zero tx/rx packets/bytes but *rx_missed_errors*
quite large (~138K) and increasing over time
$ ping <own-ip-addr>
works fine
$ ping <gateway>
no response.
$ tcpdump -i eth0 host <broken-host>
while pinging gateway, tcpdump shows messages like:
18:03:45.936161 arp who-has <gateway> tell <broken-host>
(Config file and lspci output are attached)
how about dmesg? Perhaps it shows some valuable information.
also, since this is a networking problem, please include `ifconfig eth0` and the full
output of `ethtool eth0` and `ethtool -S eth0`
Cheers,
Auke
On 9/28/06, Sukadev Bhattiprolu [off-list ref] wrote:
Thanks. See below for additional info
Auke Kok [auke-jan.h.kok@intel.com] wrote:
| Sukadev Bhattiprolu wrote:
| >
| >I am unable to get networking to work with 2.6.18-mm1 on my system.
| >
| >But 2.6.18 kernel on same system works fine. Here is some info about
| >the system/debug attempts. Attached are the lspci output and config.
| >
| >Appreciate any help. Please let me know if you need more info.
It seems you're having interrupt delivery problems or interrupts are
getting lost.
rx_missed_errors indicates frames that were dropped due to the e1000
adapter's fifo getting full and over flowing.
rx_no_buffer_count: 310
rx_missed_errors: 5865
rx_no_buffer_count indicates that the driver didn't return buffers to
the hardware soon enough, but the hardware was able to store the
packet (at the time of reception) in the fifo to try again.
Both these indicate to me that there is something wrong with
interrupts. Maybe interrupt sharing
can you possibly try a back to back connection with another linux box
and run tcpdump on both ends then ping? it will tell us if traffic is
truely getting out and coming in okay.
also please send output of lspci -vv and cat /proc/interrupts
Jesse
Jesse Brandeburg [jesse.brandeburg@gmail.com] wrote:
| On 9/28/06, Sukadev Bhattiprolu [off-list ref] wrote:
| >Thanks. See below for additional info
| >
| >Auke Kok [auke-jan.h.kok@intel.com] wrote:
| >| Sukadev Bhattiprolu wrote:
| >| >
| >| >I am unable to get networking to work with 2.6.18-mm1 on my system.
| >| >
| >| >But 2.6.18 kernel on same system works fine. Here is some info about
| >| >the system/debug attempts. Attached are the lspci output and config.
| >| >
| >| >Appreciate any help. Please let me know if you need more info.
|
| It seems you're having interrupt delivery problems or interrupts are
| getting lost.
| rx_missed_errors indicates frames that were dropped due to the e1000
| adapter's fifo getting full and over flowing.
| >rx_no_buffer_count: 310
| >rx_missed_errors: 5865
| rx_no_buffer_count indicates that the driver didn't return buffers to
| the hardware soon enough, but the hardware was able to store the
| packet (at the time of reception) in the fifo to try again.
|
| Both these indicate to me that there is something wrong with
| interrupts. Maybe interrupt sharing
|
| can you possibly try a back to back connection with another linux box
| and run tcpdump on both ends then ping? it will tell us if traffic is
| truely getting out and coming in okay.
Unfortunately, I can't try this week, but can try it early next week.
|
| also please send output of lspci -vv and cat /proc/interrupts
lspci-vv.out is attached. Here is the /proc/interrupts:
$ cat /proc/interrupts
CPU0 CPU1
0: 18316 0 IO-APIC-edge timer
2: 0 0 XT-PIC-level cascade
4: 1023 0 IO-APIC-edge serial
8: 0 0 IO-APIC-edge rtc
17: 3380 0 IO-APIC-fasteoi libata
19: 174 0 IO-APIC-fasteoi ohci_hcd:usb1, ohci_hcd:usb2
28: 0 0 IO-APIC-fasteoi eth0
NMI: 96 35
LOC: 18251 18524
ERR: 0
you should be getting an interrupt every two seconds from the eth0
(e1000) driver. You are having interrupt delivery problems probably
due to something screwing up interrupt routing in the kernel.
Normally these issues are associated with MSI interrupts but your
adapter doesn't support those and is using generic IRQ
I'm guessing that if you somehow enable interrupts on your vga card on
the same bus as e1000 (bus 3) it will have interrupt delivery problems
as well. Maybe try xorg?
Jesse