Thread (1 message) 1 message, 1 author, 19d ago

Re: kerneltests.org master run: 23 slot-hours, 3h52 wall; questions

From: Guenter Roeck <linux@roeck-us.net>
Date: 2026-09-21 22:02:59

Hi,

On Mon, Sep 21, 2026 at 11:22:06PM +0400, Semyon Fedyukovich wrote:
Hi Guenter,

I read the public build pages of your master run for v7.3-rc4 to see
how long a full build/boot pass takes and what bounds it. I would like
to check my reading with you before I do anything with it.

What I read: 57 builders on 6 machines, 158 builds plus 608 qemu
boots (about 696k KUnit passes). The buildcommand steps sum to 23.0
hours; the run took 3h52 from first start to last end. Each machine
ran at most two builders at once, so the farm has 12 slots, yet on
average 5.9 steps were in flight (peak 10). Busy time per machine
ranged from 138 min (jupiter) to 231 min (mars). The longest groups:
powerpc 80 min, x86_64 80, qemu-arm-v7 78, arm 77, i386 and xtensa
about 61 each with a single build reported.

Replaying those step durations as independent tasks:

  one slot                                1379 min
  your run (12 slots, pinned builders)     232 min   5.9x
  12 slots, any task on any free slot      118 min  11.7x
  24 slots, same groups                     80 min  17.2x
  24 slots, one task per config             61 min  22.5x
That is pure theory: The machines do not have an endless amount of
CPU capacity. Builds don't get faster by adding more builds (or
slots, as you call it) to a machine. "mars" is still "only" a
system with an AMD Ryzen 9 5900X CPU. The systems are for all
practical purposes fully or almost fully utilized with two parallel
builds or three parallel qemu emulation runs.
  more slots                                no gain (longest build)

So with the hardware you have, free placement alone would roughly
halve the run; doubling the hardware gets it to 60-80 minutes.

Three things I cannot tell from the pages:
1. Do step times include waiting inside the scripts? i386 and xtensa
   both land just over an hour, which looks like a cap or a lock.
2. Is the builder-to-machine pinning deliberate (toolchains, machine
   speed, qemu versions)? If so the 118 figure is not reachable.
3. The run started at 00:00 your time, about ten hours after the tag.
   If that schedule is deliberate, faster compute may not matter to
   you at all.

Why I ask: I maintain dispat (https://github.com/yohimik/dispat), a
release tool whose next version places build and test tasks on worker
machines over plain git: no daemon, signed task branches, and build
outputs verified by digest before a dependent task may use them. Your
matrix is the cleanest real workload I know for it. Nothing of mine
has run against the kernel yet. If a shorter turnaround is useful to
you, I would prototype it on my own hardware with your public
linux-build-test scripts and send you measured numbers, not
projections. If the limit is people reading results rather than
machines, tell me and I will drop it.
Let's see ... first of all, as you noted, pretty much everything is public
at github.com:groeck/linux-build-test.git. I had not updated the repository
for a while, but I did that now, so it is current. You should be able
to find answers for most if not all your questions in there and in the
buildbot code.

There are no intentional "pinned builders". The system uses an older
version of buildbot which has its limitations. Maybe the configuration
is less than perfect, but either case "pinning" is not intentional.
There are some restrictions - for example, a single machine can
only test a single qemu architecture at any given time because there
is only a single checked out qemu tree per build machine. That should
only affect bulk stable release builds, though. I suspect that the
buildbot algorithm selecting the next build is less than perfect,
but I never spent time trying to track it down.

Either case, I never spent time trying to update the system to a more
recent version of buildbot, simply because its architecture was changed
completely, with limited guidance how to convert older configurations
to new ones. I am not sure I understand what you are suggesting here.
Replace buildbot with your system ? I don't immediately see how that
could work because result analysis is deeply built into the system
(see schedulers.py and shellcommands.py in the code). I most definitely
don't have the time to update that or I would have converted it to a
more recent version of buildbot or to something else a long time ago.
Feel free to try, though - I am not married to buildbot and I'd be happy
to replace it.

Build time start and coverage is limited to start at midnight due to the
nowadays outrageous electricity cost in California. The build systems
consumes ~300-400W of electricity while idle and up to 1.6kW of electricity
while active. I used to run the system more actively, but at 1.6kW it
consumes more than 1,000kW of energy per month if running all the time.
Imagine how much that would cost me when paying $.36 ker kwH at night
and $.59 in the late afternoon for electricity, even more so when adding
AC cost into the picture. We do now have a fancy new solar system which
covers most of our electricity cost, but running the builders more actively
would immediate get us back into the world of outrageous electricity
prices.

Hope that helps,
Guenter
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help