[Bug 297337] nvme: timeouts and resets with Intel DC P3700

[email protected]
Newsgroups gmane.os.freebsd.bugs
Message-ID <[email protected]/bugzilla/>
https://bugs.freebsd.org/bugzilla/show_bug.cgi?id=297337

Warner Losh <[email protected]> changed:

           What    |Removed                     |Added
----------------------------------------------------------------------------
                 CC|                            |[email protected]

--- Comment #2 from Warner Losh <[email protected]> ---
(In reply to Hans Rosenfeld from comment #1)
Let's start with 'were there timeouts that it survived' in 15? It looks like
only one is having the trouble. and does this start after working for a while,
or at boot? Do different types NVMe drives in exactly the same slots have this
problem? Was the server they worked in the same server or one that's identical
or a different one?

Timeout come from a lot of different angles and we need to narrow down where.
Usually it's one of two things. Either we're sending a command to the drive
that makes it cranky (and the firmware resets, takes a while, we timeout and
then fail to reset it because the firmware hasn't finished resetting... but the
logs included so far don't show any exotics). Or there's some problem that's
persistent or that develops between the card and the CPU such that PCIe packets
don't get down to the card.

Since this is with older cards, I have a nickle on some command we're sending a
little different in -current than in 15, but can't rule out that this is, for
example, a hotplug problem (try with hw.pci.enable_pcie_hp=0) or some other
interrupt routing issue (though unlikely it's interrupt routing since that
usually recovers in the timeout watchdog).

-- 
You are receiving this mail because:
You are the assignee for the bug.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.