[PATCH] accel/rocket: search every core slot when looking up a scheduler
From: Igor Paunovic <hidden>
Date: 2026-09-05 15:05:10
Also in:
dri-devel, linux-rockchip, lkml, stable
Subsystem:
drm accel driver for rockchip npu, drm compute accelerators drivers and framework, the rest · Maintainers:
Tomeu Vizoso, Oded Gabbay, Linus Torvalds
sched_to_core() walks rdev->cores[] up to rdev->num_cores, and rocket_remove() decrements num_cores for every core it removes. Unbind a core that is not the last one and the cores behind it fall outside the search, so sched_to_core() returns NULL for a core that is still bound and still running jobs. Neither caller checks the result: rocket_job_run(): rocket_fence_create(core), core->dev rocket_job_timedout(): dev_err(core->dev, "NPU job timed out") Unbinding the middle core of the three on an RK3588 while three clients are submitting to all of them faults twice, once from the surviving core's job queue and once from its reset work: KASAN: null-ptr-deref in range [0x0000000000000220-0x0000000000000227] Workqueue: fdad0000.npu drm_sched_run_job_work [gpu_sched] pc : rocket_job_run+0x234/0x838 [rocket] Call trace: rocket_job_run+0x234/0x838 [rocket] drm_sched_run_job_work+0x2cc/0xad8 [gpu_sched] process_one_work+0x640/0x14f0 KASAN: null-ptr-deref in range [0x0000000000000000-0x0000000000000007] Workqueue: rocket-reset-2 drm_sched_job_timedout [gpu_sched] pc : rocket_job_timedout+0xf0/0x1e0 [rocket] Call trace: rocket_job_timedout+0xf0/0x1e0 [rocket] drm_sched_job_timedout+0x188/0x6a0 [gpu_sched] Both are the third core: the workqueue names are its device and its core->index, and it was left at slot 2 while num_cores had dropped to 2. Search all the slots that were allocated, the way find_core_for_dev() now does. A core that is still bound is then found, and the two callers get the pointer they already assume they have. This does not make unbinding one core out of several safe. An open client keeps an entity pointing at the scheduler of the core that went away: drm_sched reports it as not ready for every job that lands on it, and the client waits in dma_fence_default_wait for a fence that will never signal. Stopping the NULL dereference is what belongs in a fix; the rest wants more thought. Reported-by: Sidong Yang <redacted> Closes: https://lore.kernel.org/dri-devel/apwUewaRnoTNXHCt@rock-5b-plus/ (local) Fixes: 0810d5ad88a1 ("accel/rocket: Add job submission IOCTL") Cc: stable@vger.kernel.org Signed-off-by: Igor Paunovic <redacted> Assisted-by: LLM sparse checkpatch --- drivers/accel/rocket/rocket_job.c | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-)
diff --git a/drivers/accel/rocket/rocket_job.c b/drivers/accel/rocket/rocket_job.c
index 3141f210fcd1b..a6c24dfe0563a 100644
--- a/drivers/accel/rocket/rocket_job.c
+++ b/drivers/accel/rocket/rocket_job.c@@ -283,7 +283,7 @@ static struct rocket_core *sched_to_core(struct rocket_device *rdev, { unsigned int core; - for (core = 0; core < rdev->num_cores; core++) { + for (core = 0; core < rdev->max_cores; core++) { if (&rdev->cores[core].sched == sched) return &rdev->cores[core]; }
base-commit: a9f09b5ea0c3db1e2d4c0f8d3ebdd612d8aa0366 prerequisite-patch-id: 519bcdfdde80d902309c8346f749ebc4bb6b29c0 -- 2.43.0