Re: Odd Networking problem.

Al Hopper <al-Gniphm1Q+ng6rkq5uMl77VaTQe2KTcn/@public.gmane.org> Wed, 22 Jun 2011 03:42:32 -0500
Newsgroups gmane.org.operators.internet-access
Message-ID <[email protected]>
On Tue, Jun 21, 2011 at 6:14 PM, Keith <[email protected]> wrote:
>
> On Tue, 21 Jun 2011, Dan Gullick wrote:
>
> |->I have known this issue to happen with certain hardware vendors (Nortel) where the ARP table is not accurate.
> |->Out of curiosity, can you list the network equipment in the middle?
>
> It is just an IBM server directly connected to a 24 port Cisco 3560 with
> multiple Vlans.
>
> The 3560 has a connection to a 7206 and two other 3560's.
>
> Nothing fancy, and only this one machine it happens with. I am waiting for
> it to happen again now so I can play with the ARP table a bit to try and
> narrow this down.
>
> Thanks,
> Keith
> --
> Eat sushi frequently. - Avi
> [email protected] is the human contact address.
> [email protected] is the list posting address.
> See below URL for subscribe/unsubscribe and list options:
> http://inet-access.net/mailman/listinfo/list

I know you're going to think I'm nuts when you read this - but - I've
seen some bizaar behavior like this before
from 3 identical machines I built.  Same motherboard, same components
etc. etc.   Two of the machines were
fine - the 3rd machine displayed quirky network behavior before I
shipped it to the colo.  Thinking that I had typoed
something - I just dismissed it.  It occurred during network
configuration.  Turns out that there was some
'impossible-to-define' issue with the motherboard in the 3rd machine.
And it was returned to me to "fix" after living
in the colo for about 3 weeks.  Swapping out the motherboard resolved
all the weird behavior.

Lessons learned:  Modern motherboards are so complex, with so much
functionality embedded within complex
LSI (large scale integration) devices, that the possible failure modes
are *impossible* to define.
And probably, impossible to test for, in terms of an
end-of-the-production-line test.

I'm willing to bet, that if you swap out the motherboard in the wacko
machine,  all your issues will disappear.

In a simpler world, where many different chips were associated with
different functions, failure modes were pretty
easy to identify and define.  In today's world of complex, large scale
embedded chipsets that are part of the
modern motherboard - failure modes are often very, very bizaar and can
easily lead one to believe that it's a "neck
up" config issue, when, in fact, its some bloody transistor that has
failed in a chip (set) with 10's or 100's of millions
of transistors.

Swap out the motherboard in the "funky" server.   Don't even try to
RMA it.  Why?  Because the folks working on it
will simply not be able to replicate the fault condition and will just
return it to you.

And yes - I adopt stringent static discharge protection measures while
working on server level hardware.  Been
doing this for a *long* time now.

Think about this for a Second:  if one or two transistors die in a
North or South chip, can anyone define what effect
it'll have on the operation of the motherboard?  Answer: No.  It'll
simply manifest itself in some bizaar, impossible
to understand behavior and the thinking persons gut response is to
ask: "where did I mess up?".

Regards,

--
Al Hopper
-- 
Eat sushi frequently. - Avi
[email protected] is the human contact address.
[email protected] is the list posting address.
See below URL for subscribe/unsubscribe and list options:
http://inet-access.net/mailman/listinfo/list