Re: kerneltests.org master run: 23 slot-hours, 3h52 wall; questions
From: Guenter Roeck <linux@roeck-us.net>
Date: 2026-09-21 22:02:59
Hi, On Mon, Sep 21, 2026 at 11:22:06PM +0400, Semyon Fedyukovich wrote:
Hi Guenter, I read the public build pages of your master run for v7.3-rc4 to see how long a full build/boot pass takes and what bounds it. I would like to check my reading with you before I do anything with it. What I read: 57 builders on 6 machines, 158 builds plus 608 qemu boots (about 696k KUnit passes). The buildcommand steps sum to 23.0 hours; the run took 3h52 from first start to last end. Each machine ran at most two builders at once, so the farm has 12 slots, yet on average 5.9 steps were in flight (peak 10). Busy time per machine ranged from 138 min (jupiter) to 231 min (mars). The longest groups: powerpc 80 min, x86_64 80, qemu-arm-v7 78, arm 77, i386 and xtensa about 61 each with a single build reported. Replaying those step durations as independent tasks: one slot 1379 min your run (12 slots, pinned builders) 232 min 5.9x 12 slots, any task on any free slot 118 min 11.7x 24 slots, same groups 80 min 17.2x 24 slots, one task per config 61 min 22.5x
That is pure theory: The machines do not have an endless amount of CPU capacity. Builds don't get faster by adding more builds (or slots, as you call it) to a machine. "mars" is still "only" a system with an AMD Ryzen 9 5900X CPU. The systems are for all practical purposes fully or almost fully utilized with two parallel builds or three parallel qemu emulation runs.
more slots no gain (longest build) So with the hardware you have, free placement alone would roughly halve the run; doubling the hardware gets it to 60-80 minutes. Three things I cannot tell from the pages: 1. Do step times include waiting inside the scripts? i386 and xtensa both land just over an hour, which looks like a cap or a lock. 2. Is the builder-to-machine pinning deliberate (toolchains, machine speed, qemu versions)? If so the 118 figure is not reachable. 3. The run started at 00:00 your time, about ten hours after the tag. If that schedule is deliberate, faster compute may not matter to you at all. Why I ask: I maintain dispat (https://github.com/yohimik/dispat), a release tool whose next version places build and test tasks on worker machines over plain git: no daemon, signed task branches, and build outputs verified by digest before a dependent task may use them. Your matrix is the cleanest real workload I know for it. Nothing of mine has run against the kernel yet. If a shorter turnaround is useful to you, I would prototype it on my own hardware with your public linux-build-test scripts and send you measured numbers, not projections. If the limit is people reading results rather than machines, tell me and I will drop it.
Let's see ... first of all, as you noted, pretty much everything is public at github.com:groeck/linux-build-test.git. I had not updated the repository for a while, but I did that now, so it is current. You should be able to find answers for most if not all your questions in there and in the buildbot code. There are no intentional "pinned builders". The system uses an older version of buildbot which has its limitations. Maybe the configuration is less than perfect, but either case "pinning" is not intentional. There are some restrictions - for example, a single machine can only test a single qemu architecture at any given time because there is only a single checked out qemu tree per build machine. That should only affect bulk stable release builds, though. I suspect that the buildbot algorithm selecting the next build is less than perfect, but I never spent time trying to track it down. Either case, I never spent time trying to update the system to a more recent version of buildbot, simply because its architecture was changed completely, with limited guidance how to convert older configurations to new ones. I am not sure I understand what you are suggesting here. Replace buildbot with your system ? I don't immediately see how that could work because result analysis is deeply built into the system (see schedulers.py and shellcommands.py in the code). I most definitely don't have the time to update that or I would have converted it to a more recent version of buildbot or to something else a long time ago. Feel free to try, though - I am not married to buildbot and I'd be happy to replace it. Build time start and coverage is limited to start at midnight due to the nowadays outrageous electricity cost in California. The build systems consumes ~300-400W of electricity while idle and up to 1.6kW of electricity while active. I used to run the system more actively, but at 1.6kW it consumes more than 1,000kW of energy per month if running all the time. Imagine how much that would cost me when paying $.36 ker kwH at night and $.59 in the late afternoon for electricity, even more so when adding AC cost into the picture. We do now have a fancy new solar system which covers most of our electricity cost, but running the builders more actively would immediate get us back into the world of outrageous electricity prices. Hope that helps, Guenter