CI Outages
Ben Cooksley <[email protected]> Fri, 1 May 2026 22:02:20 +1200
| Newsgroups | gmane.comp.kde.releases,gmane.comp.kde.devel.core,gmane.comp.kde.devel.general,gmane.comp.kde.devel.plasma,gmane.comp.kde.devel.frameworks |
|---|---|
| Message-ID | <CA+XidOG-1ZoXuAboB54sPOffLakbksuqu1SbytbbMNy2PrAcbg@mail.gmail.com> |
--00000000000013dd450650beae0e Content-Type: text/plain; charset="UTF-8" Hi all, Over the past 48 hours or so we've had a series of two unfortunate incidents that have significantly degraded the CI system. The first part of this took place approximately 36 hours ago, when an admin, in response to the announcement of the https://copy.fail/ Linux kernel exploit, installed updates on one or more of our VM Runners. The unfortunate side effect of this is that it also installed updates for gitlab-runner and upgraded it to a newer version. As part of it's work, Gitlab Runner requires the assistance of a helper binary within the VM, and this helper should ideally be the same version as is deployed on the VM runner hosts, or at the very least be a newer version. In the case of this update, there were incompatible changes as part of changes to how artifacts are captured, which is why we are seeing breakages related to an unrecognised timeout parameter which is causing a complete fatal failure of the CI jobs. The images that support the majority of our CI builds (Linux - Qt 6.11, Qt 6.12 and Qt 5.15, Android, Flatpak, Snap and Appimages) have been rebuilt to include the newer Gitlab Runner helper and those builds should now be functional again. Custom VM images utilised by Yocto, Buildstream, Neon and KDE Linux have also been rebuilt and should also be functional again. Windows builds require a replacement base image as the Gitlab Runner helper is burned into the base image - and that is in the process of being uploaded currently. Once uploaded, i'll rebuild the image that supports both general Windows CI and Craft builds which will restore those builds to working order as well. For FreeBSD, we will need our custom package repository updated to include the newer Gitlab Runner helper. This has been requested and should be completed in the next few days so those builds will remain broken for a bit longer i'm afraid. The second incident involved a service outage of the builder that supports Docker based jobs. This was caused by hoster maintenance related to the SAN that supports those hosts, and also caused service disruptions to all Notary Service operations, WebSVN and Sentry. This outage impacted us for approximately 12 hours and has now been corrected with all services fully returned to normal. Apologies for the disruption caused by these incidents, it is most regrettable - and in the case of the issue affecting VM builds - completely avoidable. Please let me know if you have any questions on the above. Many thanks, Ben --00000000000013dd450650beae0e Content-Type: text/html; charset="UTF-8" Content-Transfer-Encoding: quoted-printable <div dir=3D"ltr">Hi all,<div><br></div><div>Over the past 48 hours or so we= 've had a series of two unfortunate incidents that have significantly d= egraded the CI system.</div><div><br></div><div>The first part of this took= place approximately 36 hours ago, when an admin, in response to the announ= cement of the=C2=A0<a href=3D"https://copy.fail/">https://copy.fail/</a> Li= nux kernel exploit, installed updates on one or more of our VM Runners. The= unfortunate side effect of this is that it also installed updates for gitl= ab-runner and upgraded it to a newer version. As part of it's work, Git= lab Runner requires the assistance of a helper binary within the VM, and th= is helper should ideally be the same version as is deployed on the VM runne= r hosts, or at the very least be a newer version.=C2=A0</div><div><br></div= ><div>In the case of this update, there were incompatible changes as part o= f changes to how artifacts are captured, which is why we are seeing breakag= es related to an unrecognised timeout parameter which is causing a complete= fatal failure of the CI jobs.</div><div><br></div><div>The images that sup= port the majority=C2=A0of our CI builds (Linux - Qt 6.11, Qt 6.12 and Qt 5.= 15, Android, Flatpak, Snap and Appimages) have been rebuilt to include the = newer Gitlab Runner helper and those builds should now be functional again.= =C2=A0</div><div>Custom VM images utilised by Yocto, Buildstream, Neon and = KDE Linux have also been rebuilt and should also be functional again.</div>= <div><br></div><div>Windows builds require a replacement base image as the = Gitlab Runner helper is burned into the base image - and that is in the pro= cess of being uploaded currently.=C2=A0</div><div>Once uploaded, i'll r= ebuild the image that supports both general Windows CI and Craft builds whi= ch will restore those builds to working order as well.</div><div><br></div>= <div>For FreeBSD, we will need our custom package repository updated to inc= lude the newer Gitlab Runner helper.=C2=A0</div><div>This has been requeste= d and should be completed in the next few days so those builds will remain = broken for a bit longer i'm afraid.</div><div><br></div><div>The second= incident involved a service outage of the builder that supports Docker bas= ed jobs. This was caused by hoster maintenance related to the SAN that supp= orts those hosts, and also caused service disruptions to all Notary Service= operations, WebSVN and Sentry.</div><div>This outage impacted us for appro= ximately 12 hours and has now been corrected with all services fully return= ed to normal.</div><div><br></div><div>Apologies for the disruption caused = by these incidents, it is most regrettable - and in the case of the issue a= ffecting VM builds - completely avoidable.</div><div><br></div><div>Please = let me know if you have any questions on the above.</div><div><br></div><di= v>Many thanks,</div><div>Ben</div></div> --00000000000013dd450650beae0e--