Re: s390x builder outage
Neal Gompa <[email protected]>
| Newsgroups | gmane.linux.redhat.fedora.devel |
|---|---|
| Message-ID | <CAEg-Je8NqZ4V8nj=xCS6WS8Oxn_Mb4JeECPszKmOPUGwFQbOEQ@mail.gmail.com> |
On Sat, Aug 15, 2026 at 6:30 PM Stephen J Smoogen <[email protected]> wrote: > > > > On Sat, Aug 15, 2026, at 18:13, Chris Adams wrote: > > Once upon a time, Fabio Valentini <[email protected]> said: > >> It might also be worth noting that s390x is *already* kind of a > >> second-tier architecture: > >> > >> - noarch builds are never scheduled to run on s390x > >> - SRPM builds are never scheduled to run on s390x > >> - koschei scratch builds always exclude s390x > > > > I admit I know minimal about how builds are handled. When I submit a > > build (which, I don't do lots of builds), the top-level task quite often > > seems to be on an s390x host (and ppc64le when it's not)... maybe that > > part uses minimal resources, but it's still there. SRPM build is often > > on ppc64le. > > > > And yes, if the hardware is getting overloaded, reducing simultaneous > > resource usage can improve throughput, because you reduce thrashing. I > > obviously have no idea if that's the case here, but a blanket "doing > > less slows things down" is not true. > > The issue has always been that the power and s390x have been spread thin. For the longest time the servers were loaners from IBM and when ones were purchased they were usually the bottom of the line because the cost difference cut into how many x86 systems were available. (It was something like 4:1) The s390x is spread between multiple Red Hat teams and if testing or other IO intensive work is going on our usage would be highly impacted. > > If I had to keep the architectures running, I would look at the problem differently… maybe there is an oversupply of resources for x86 and aarch64 plus an over demand of things wanting to be done. Cut those down to what the slowest arch and most easily overloaded arch can handle. > Honestly, likely the mistake was merging all the architectures into one Koji instance. The simplification of the infrastructure wasn't worth the pain it caused over time with ARM, POWER, and Z. Koji doesn't have the feature that OBS has where projects are chained in the same instance but don't block each other (ie openSUSE:Factory is x86 arches, openSUSE:Factory:ARM for ARM arches, openSUSE:Factory:RISCV for RISC-V, openSUSE:Factory:PowerPC for POWER, and openSUSE:Factory:zSystems for IBM Z contain linked project data to build the ports in a non-blocking way). It's true now that all ports remain in sync now, but that is only worth it if the performance is there. And if it never will be because it's too costly for someone to care enough to do it, then there's a problem. We know there isn't an oversupply on x86 and ARM because if a couple of boxes go down, then things slow down a lot there. So I would argue we're probably treading water on those architectures today. And that wasn't always the case. -- 真実はいつも一つ!/ Always, there's only one truth! -- _______________________________________________ devel mailing list -- [email protected] To unsubscribe send an email to [email protected] Fedora Code of Conduct: https://docs.fedoraproject.org/en-US/project/code-of-conduct/ List Guidelines: https://fedoraproject.org/wiki/Mailing_list_guidelines List Archives: https://lists.fedoraproject.org/archives/list/[email protected] Do not reply to spam, report it: https://forge.fedoraproject.org/infra/tickets/issues/new