This patch set tries to fix that client might oops if nbd server send
unexpected message to client, for example, our syzkaller report a uaf
in nbd_read_stat():
Call trace:
dump_backtrace+0x0/0x310 arch/arm64/kernel/time.c:78
show_stack+0x28/0x38 arch/arm64/kernel/traps.c:158
__dump_stack lib/dump_stack.c:77 [inline]
dump_stack+0x144/0x1b4 lib/dump_stack.c:118
print_address_description+0x68/0x2d0 mm/kasan/report.c:253
kasan_report_error mm/kasan/report.c:351 [inline]
kasan_report+0x134/0x2f0 mm/kasan/report.c:409
check_memory_region_inline mm/kasan/kasan.c:260 [inline]
__asan_load4+0x88/0xb0 mm/kasan/kasan.c:699
__read_once_size include/linux/compiler.h:193 [inline]
blk_mq_rq_state block/blk-mq.h:106 [inline]
blk_mq_request_started+0x24/0x40 block/blk-mq.c:644
nbd_read_stat drivers/block/nbd.c:670 [inline]
recv_work+0x1bc/0x890 drivers/block/nbd.c:749
process_one_work+0x3ec/0x9e0 kernel/workqueue.c:2147
worker_thread+0x80/0x9d0 kernel/workqueue.c:2302
kthread+0x1d8/0x1e0 kernel/kthread.c:255
ret_from_fork+0x10/0x18 arch/arm64/kernel/entry.S:1174
1) At first, a normal io is submitted and completed with scheduler:
internel_tag = blk_mq_get_tag -> get tag from sched_tags
blk_mq_rq_ctx_init
sched_tags->rq[internel_tag] = sched_tag->static_rq[internel_tag]
...
blk_mq_get_driver_tag
__blk_mq_get_driver_tag -> get tag from tags
tags->rq[tag] = sched_tag->static_rq[internel_tag]
So, both tags->rq[tag] and sched_tags->rq[internel_tag] are pointing
to the request: sched_tags->static_rq[internal_tag]. Even if the
io is finished.
2) nbd server send a reply with random tag directly:
recv_work
nbd_read_stat
blk_mq_tag_to_rq(tags, tag)
rq = tags->rq[tag]
3) if the sched_tags->static_rq is freed:
blk_mq_sched_free_requests
blk_mq_free_rqs(q->tag_set, hctx->sched_tags, i)
-> step 2) access rq before clearing rq mapping
blk_mq_clear_rq_mapping(set, tags, hctx_idx);
__free_pages() -> rq is freed here
4) Then, nbd continue to use the freed request in nbd_read_stat()
Changes in v5:
- move patch 1 & 2 in v4 (patch 4 & 5 in v5) behind
- add some comment in patch 5
Changes in v4:
- change the name of the patchset, since uaf is not the only problem
if server send unexpected reply message.
- instead of adding new interface, use blk_mq_find_and_get_req().
- add patch 5 to this series
Changes in v3:
- v2 can't fix the problem thoroughly, add patch 3-4 to this series.
- modify descriptions.
- patch 5 is just a cleanup
Changes in v2:
- as Bart suggested, add a new helper function for drivers to get
request by tag.
Yu Kuai (6):
nbd: don't handle response without a corresponding request message
nbd: make sure request completion won't concurrent
nbd: check sock index in nbd_read_stat()
blk-mq: export two symbols to get request by tag
nbd: convert to use blk_mq_find_and_get_req()
nbd: don't start request if nbd_queue_rq() failed
block/blk-mq-tag.c | 5 +++--
block/blk-mq.c | 1 +
drivers/block/nbd.c | 51 ++++++++++++++++++++++++++++++++++++------
include/linux/blk-mq.h | 3 +++
4 files changed, 51 insertions(+), 9 deletions(-)
--
2.31.1
nbd has a defect that blk_mq_tag_to_rq() might return a freed
request in nbd_read_stat(). We need a new mechanism if we want to
fix this in nbd driver, which is rather complicated.
Thus use blk_mq_find_and_get_req() to replace blk_mq_tag_to_rq(),
which can make sure the returned request is not freed, and then we
can do more checking while 'cmd->lock' is hold.
Signed-off-by: Yu Kuai <redacted>
---
block/blk-mq-tag.c | 5 +++--
block/blk-mq.c | 1 +
include/linux/blk-mq.h | 3 +++
3 files changed, 7 insertions(+), 2 deletions(-)
blk_mq_tag_to_rq() can only ensure to return valid request in
following situation:
1) client send request message to server first
submit_bio
...
blk_mq_get_tag
...
blk_mq_get_driver_tag
...
nbd_queue_rq
nbd_handle_cmd
nbd_send_cmd
2) client receive respond message from server
recv_work
nbd_read_stat
blk_mq_tag_to_rq
If step 1) is missing, blk_mq_tag_to_rq() will return a stale
request, which might be freed. Thus convert to use
blk_mq_find_and_get_req() to make sure the returned request is not
freed.
Signed-off-by: Yu Kuai <redacted>
---
drivers/block/nbd.c | 15 ++++++++++++---
1 file changed, 12 insertions(+), 3 deletions(-)
Currently, blk_mq_end_request() will be called if nbd_queue_rq()
failed, thus start request in such situation is useless.
Signed-off-by: Yu Kuai <redacted>
Reviewed-by: Christoph Hellwig <hch@lst.de>
---
drivers/block/nbd.c | 3 ---
1 file changed, 3 deletions(-)
@@ -943,7 +943,6 @@ static int nbd_handle_cmd(struct nbd_cmd *cmd, int index)if(!refcount_inc_not_zero(&nbd->config_refs)){dev_err_ratelimited(disk_to_dev(nbd->disk),"Socks array is empty\n");-blk_mq_start_request(req);return-EINVAL;}config=nbd->config;
@@ -952,7 +951,6 @@ static int nbd_handle_cmd(struct nbd_cmd *cmd, int index)dev_err_ratelimited(disk_to_dev(nbd->disk),"Attempted send on invalid socket\n");nbd_config_put(nbd);-blk_mq_start_request(req);return-EINVAL;}cmd->status=BLK_STS_OK;
@@ -976,7 +974,6 @@ static int nbd_handle_cmd(struct nbd_cmd *cmd, int index)*/sock_shutdown(nbd);nbd_config_put(nbd);-blk_mq_start_request(req);return-EIO;}gotoagain;
commit cddce0116058 ("nbd: Aovid double completion of a request")
try to fix that nbd_clear_que() and recv_work() can complete a
request concurrently. However, the problem still exists:
t1 t2 t3
nbd_disconnect_and_put
flush_workqueue
recv_work
blk_mq_complete_request
blk_mq_complete_request_remote -> this is true
WRITE_ONCE(rq->state, MQ_RQ_COMPLETE)
blk_mq_raise_softirq
blk_done_softirq
blk_complete_reqs
nbd_complete_rq
blk_mq_end_request
blk_mq_free_request
WRITE_ONCE(rq->state, MQ_RQ_IDLE)
nbd_clear_que
blk_mq_tagset_busy_iter
nbd_clear_req
__blk_mq_free_request
blk_mq_put_tag
blk_mq_complete_request -> complete again
There are three places where request can be completed in nbd:
recv_work(), nbd_clear_que() and nbd_xmit_timeout(). Since they
all hold cmd->lock before completing the request, it's easy to
avoid the problem by setting and checking a cmd flag.
Signed-off-by: Yu Kuai <redacted>
---
drivers/block/nbd.c | 11 +++++++++--
1 file changed, 9 insertions(+), 2 deletions(-)
While handling a response message from server, nbd_read_stat() will
try to get request by tag, and then complete the request. However,
this is problematic if nbd haven't sent a corresponding request
message:
t1 t2
submit_bio
nbd_queue_rq
blk_mq_start_request
recv_work
nbd_read_stat
blk_mq_tag_to_rq
blk_mq_complete_request
nbd_send_cmd
Thus add a new cmd flag 'NBD_CMD_INFLIGHT', it will be set in
nbd_send_cmd() and checked in nbd_read_stat().
Noted that this patch can't fix that blk_mq_tag_to_rq() might
return a freed request, and this will be fixed in following
patches.
Signed-off-by: Yu Kuai <redacted>
---
drivers/block/nbd.c | 22 +++++++++++++++++++++-
1 file changed, 21 insertions(+), 1 deletion(-)
The sock that clent send request in nbd_send_cmd() and receive reply
in nbd_read_stat() should be the same.
Signed-off-by: Yu Kuai <redacted>
---
drivers/block/nbd.c | 4 ++++
1 file changed, 4 insertions(+)
On Thu, Sep 09, 2021 at 10:12:51PM +0800, Yu Kuai wrote:
While handling a response message from server, nbd_read_stat() will
try to get request by tag, and then complete the request. However,
this is problematic if nbd haven't sent a corresponding request
message:
t1 t2
submit_bio
nbd_queue_rq
blk_mq_start_request
recv_work
nbd_read_stat
blk_mq_tag_to_rq
blk_mq_complete_request
nbd_send_cmd
Thus add a new cmd flag 'NBD_CMD_INFLIGHT', it will be set in
nbd_send_cmd() and checked in nbd_read_stat().
Noted that this patch can't fix that blk_mq_tag_to_rq() might
return a freed request, and this will be fixed in following
patches.
Signed-off-by: Yu Kuai <redacted>
Looks fine:
Reviewed-by: Ming Lei <redacted>
--
Ming
On Thu, Sep 09, 2021 at 10:12:52PM +0800, Yu Kuai wrote:
quoted hunk
commit cddce0116058 ("nbd: Aovid double completion of a request")
try to fix that nbd_clear_que() and recv_work() can complete a
request concurrently. However, the problem still exists:
t1 t2 t3
nbd_disconnect_and_put
flush_workqueue
recv_work
blk_mq_complete_request
blk_mq_complete_request_remote -> this is true
WRITE_ONCE(rq->state, MQ_RQ_COMPLETE)
blk_mq_raise_softirq
blk_done_softirq
blk_complete_reqs
nbd_complete_rq
blk_mq_end_request
blk_mq_free_request
WRITE_ONCE(rq->state, MQ_RQ_IDLE)
nbd_clear_que
blk_mq_tagset_busy_iter
nbd_clear_req
__blk_mq_free_request
blk_mq_put_tag
blk_mq_complete_request -> complete again
There are three places where request can be completed in nbd:
recv_work(), nbd_clear_que() and nbd_xmit_timeout(). Since they
all hold cmd->lock before completing the request, it's easy to
avoid the problem by setting and checking a cmd flag.
Signed-off-by: Yu Kuai <redacted>
---
drivers/block/nbd.c | 11 +++++++++--
1 file changed, 9 insertions(+), 2 deletions(-)
On Thu, Sep 09, 2021 at 10:12:55PM +0800, Yu Kuai wrote:
blk_mq_tag_to_rq() can only ensure to return valid request in
following situation:
1) client send request message to server first
submit_bio
...
blk_mq_get_tag
...
blk_mq_get_driver_tag
...
nbd_queue_rq
nbd_handle_cmd
nbd_send_cmd
2) client receive respond message from server
recv_work
nbd_read_stat
blk_mq_tag_to_rq
If step 1) is missing, blk_mq_tag_to_rq() will return a stale
request, which might be freed. Thus convert to use
blk_mq_find_and_get_req() to make sure the returned request is not
freed.
But NBD_CMD_INFLIGHT has been added for checking if the reply is
expected, do we still need blk_mq_find_and_get_req() for covering
this issue? BTW, request and its payload is pre-allocated, so there
isn't real use-after-free.
Thanks,
Ming
On Thu, Sep 09, 2021 at 10:12:55PM +0800, Yu Kuai wrote:
quoted
blk_mq_tag_to_rq() can only ensure to return valid request in
following situation:
1) client send request message to server first
submit_bio
...
blk_mq_get_tag
...
blk_mq_get_driver_tag
...
nbd_queue_rq
nbd_handle_cmd
nbd_send_cmd
2) client receive respond message from server
recv_work
nbd_read_stat
blk_mq_tag_to_rq
If step 1) is missing, blk_mq_tag_to_rq() will return a stale
request, which might be freed. Thus convert to use
blk_mq_find_and_get_req() to make sure the returned request is not
freed.
But NBD_CMD_INFLIGHT has been added for checking if the reply is
expected, do we still need blk_mq_find_and_get_req() for covering
this issue? BTW, request and its payload is pre-allocated, so there
isn't real use-after-free.
Hi, Ming
Checking NBD_CMD_INFLIGHT relied on the request founded by tag is valid,
not the other way round.
nbd_read_stat
req = blk_mq_tag_to_rq()
cmd = blk_mq_rq_to_pdu(req)
mutex_lock(cmd->lock)
checking NBD_CMD_INFLIGHT
The checking doesn't have any effect on blk_mq_tag_to_rq().
Thanks,
Kuai
On Thu, Sep 09, 2021 at 10:12:52PM +0800, Yu Kuai wrote:
quoted
commit cddce0116058 ("nbd: Aovid double completion of a request")
try to fix that nbd_clear_que() and recv_work() can complete a
request concurrently. However, the problem still exists:
t1 t2 t3
nbd_disconnect_and_put
flush_workqueue
recv_work
blk_mq_complete_request
blk_mq_complete_request_remote -> this is true
WRITE_ONCE(rq->state, MQ_RQ_COMPLETE)
blk_mq_raise_softirq
blk_done_softirq
blk_complete_reqs
nbd_complete_rq
blk_mq_end_request
blk_mq_free_request
WRITE_ONCE(rq->state, MQ_RQ_IDLE)
nbd_clear_que
blk_mq_tagset_busy_iter
nbd_clear_req
__blk_mq_free_request
blk_mq_put_tag
blk_mq_complete_request -> complete again
There are three places where request can be completed in nbd:
recv_work(), nbd_clear_que() and nbd_xmit_timeout(). Since they
all hold cmd->lock before completing the request, it's easy to
avoid the problem by setting and checking a cmd flag.
Signed-off-by: Yu Kuai <redacted>
---
drivers/block/nbd.c | 11 +++++++++--
1 file changed, 9 insertions(+), 2 deletions(-)
On Tue, Sep 14, 2021 at 11:11:06AM +0800, yukuai (C) wrote:
On 2021/09/14 9:11, Ming Lei wrote:
quoted
On Thu, Sep 09, 2021 at 10:12:55PM +0800, Yu Kuai wrote:
quoted
blk_mq_tag_to_rq() can only ensure to return valid request in
following situation:
1) client send request message to server first
submit_bio
...
blk_mq_get_tag
...
blk_mq_get_driver_tag
...
nbd_queue_rq
nbd_handle_cmd
nbd_send_cmd
2) client receive respond message from server
recv_work
nbd_read_stat
blk_mq_tag_to_rq
If step 1) is missing, blk_mq_tag_to_rq() will return a stale
request, which might be freed. Thus convert to use
blk_mq_find_and_get_req() to make sure the returned request is not
freed.
But NBD_CMD_INFLIGHT has been added for checking if the reply is
expected, do we still need blk_mq_find_and_get_req() for covering
this issue? BTW, request and its payload is pre-allocated, so there
isn't real use-after-free.
Hi, Ming
Checking NBD_CMD_INFLIGHT relied on the request founded by tag is valid,
not the other way round.
nbd_read_stat
req = blk_mq_tag_to_rq()
cmd = blk_mq_rq_to_pdu(req)
mutex_lock(cmd->lock)
checking NBD_CMD_INFLIGHT
Request and its payload is pre-allocated, and either req->ref or cmd->lock can
serve the same purpose here. Once cmd->lock is held, you can check if the cmd is
inflight or not. If it isn't inflight, just return -ENOENT. Is there any
problem to handle in this way?
Thanks,
Ming
On Tue, Sep 14, 2021 at 11:11:06AM +0800, yukuai (C) wrote:
quoted
On 2021/09/14 9:11, Ming Lei wrote:
quoted
On Thu, Sep 09, 2021 at 10:12:55PM +0800, Yu Kuai wrote:
quoted
blk_mq_tag_to_rq() can only ensure to return valid request in
following situation:
1) client send request message to server first
submit_bio
...
blk_mq_get_tag
...
blk_mq_get_driver_tag
...
nbd_queue_rq
nbd_handle_cmd
nbd_send_cmd
2) client receive respond message from server
recv_work
nbd_read_stat
blk_mq_tag_to_rq
If step 1) is missing, blk_mq_tag_to_rq() will return a stale
request, which might be freed. Thus convert to use
blk_mq_find_and_get_req() to make sure the returned request is not
freed.
But NBD_CMD_INFLIGHT has been added for checking if the reply is
expected, do we still need blk_mq_find_and_get_req() for covering
this issue? BTW, request and its payload is pre-allocated, so there
isn't real use-after-free.
Hi, Ming
Checking NBD_CMD_INFLIGHT relied on the request founded by tag is valid,
not the other way round.
nbd_read_stat
req = blk_mq_tag_to_rq()
cmd = blk_mq_rq_to_pdu(req)
mutex_lock(cmd->lock)
checking NBD_CMD_INFLIGHT
Request and its payload is pre-allocated, and either req->ref or cmd->lock can
serve the same purpose here. Once cmd->lock is held, you can check if the cmd is
inflight or not. If it isn't inflight, just return -ENOENT. Is there any
problem to handle in this way?
Hi, Ming
in nbd_read_stat:
1) get a request by tag first
2) get nbd_cmd by the request
3) hold cmd->lock and check if cmd is inflight
If we want to check if the cmd is inflight in step 3), we have to do
setp 1) and 2) first. As I explained in patch 0, blk_mq_tag_to_rq()
can't make sure the returned request is not freed:
nbd_read_stat
blk_mq_sched_free_requests
blk_mq_free_rqs
blk_mq_tag_to_rq
-> get rq before clear mapping
blk_mq_clear_rq_mapping
__free_pages -> rq is freed
blk_mq_request_started -> UAF
Thanks,
Kuai
On Tue, Sep 14, 2021 at 03:13:38PM +0800, yukuai (C) wrote:
On 2021/09/14 14:44, Ming Lei wrote:
quoted
On Tue, Sep 14, 2021 at 11:11:06AM +0800, yukuai (C) wrote:
quoted
On 2021/09/14 9:11, Ming Lei wrote:
quoted
On Thu, Sep 09, 2021 at 10:12:55PM +0800, Yu Kuai wrote:
quoted
blk_mq_tag_to_rq() can only ensure to return valid request in
following situation:
1) client send request message to server first
submit_bio
...
blk_mq_get_tag
...
blk_mq_get_driver_tag
...
nbd_queue_rq
nbd_handle_cmd
nbd_send_cmd
2) client receive respond message from server
recv_work
nbd_read_stat
blk_mq_tag_to_rq
If step 1) is missing, blk_mq_tag_to_rq() will return a stale
request, which might be freed. Thus convert to use
blk_mq_find_and_get_req() to make sure the returned request is not
freed.
But NBD_CMD_INFLIGHT has been added for checking if the reply is
expected, do we still need blk_mq_find_and_get_req() for covering
this issue? BTW, request and its payload is pre-allocated, so there
isn't real use-after-free.
Hi, Ming
Checking NBD_CMD_INFLIGHT relied on the request founded by tag is valid,
not the other way round.
nbd_read_stat
req = blk_mq_tag_to_rq()
cmd = blk_mq_rq_to_pdu(req)
mutex_lock(cmd->lock)
checking NBD_CMD_INFLIGHT
Request and its payload is pre-allocated, and either req->ref or cmd->lock can
serve the same purpose here. Once cmd->lock is held, you can check if the cmd is
inflight or not. If it isn't inflight, just return -ENOENT. Is there any
problem to handle in this way?
Hi, Ming
in nbd_read_stat:
1) get a request by tag first
2) get nbd_cmd by the request
3) hold cmd->lock and check if cmd is inflight
If we want to check if the cmd is inflight in step 3), we have to do
setp 1) and 2) first. As I explained in patch 0, blk_mq_tag_to_rq()
can't make sure the returned request is not freed:
nbd_read_stat
blk_mq_sched_free_requests
blk_mq_free_rqs
blk_mq_tag_to_rq
-> get rq before clear mapping
blk_mq_clear_rq_mapping
__free_pages -> rq is freed
blk_mq_request_started -> UAF
If the above can happen, blk_mq_find_and_get_req() may not fix it too, just
wondering why not take the following simpler way for avoiding the UAF?
On Tue, Sep 14, 2021 at 03:13:38PM +0800, yukuai (C) wrote:
quoted
On 2021/09/14 14:44, Ming Lei wrote:
quoted
On Tue, Sep 14, 2021 at 11:11:06AM +0800, yukuai (C) wrote:
quoted
On 2021/09/14 9:11, Ming Lei wrote:
quoted
On Thu, Sep 09, 2021 at 10:12:55PM +0800, Yu Kuai wrote:
quoted
blk_mq_tag_to_rq() can only ensure to return valid request in
following situation:
1) client send request message to server first
submit_bio
...
blk_mq_get_tag
...
blk_mq_get_driver_tag
...
nbd_queue_rq
nbd_handle_cmd
nbd_send_cmd
2) client receive respond message from server
recv_work
nbd_read_stat
blk_mq_tag_to_rq
If step 1) is missing, blk_mq_tag_to_rq() will return a stale
request, which might be freed. Thus convert to use
blk_mq_find_and_get_req() to make sure the returned request is not
freed.
But NBD_CMD_INFLIGHT has been added for checking if the reply is
expected, do we still need blk_mq_find_and_get_req() for covering
this issue? BTW, request and its payload is pre-allocated, so there
isn't real use-after-free.
Hi, Ming
Checking NBD_CMD_INFLIGHT relied on the request founded by tag is valid,
not the other way round.
nbd_read_stat
req = blk_mq_tag_to_rq()
cmd = blk_mq_rq_to_pdu(req)
mutex_lock(cmd->lock)
checking NBD_CMD_INFLIGHT
Request and its payload is pre-allocated, and either req->ref or cmd->lock can
serve the same purpose here. Once cmd->lock is held, you can check if the cmd is
inflight or not. If it isn't inflight, just return -ENOENT. Is there any
problem to handle in this way?
Hi, Ming
in nbd_read_stat:
1) get a request by tag first
2) get nbd_cmd by the request
3) hold cmd->lock and check if cmd is inflight
If we want to check if the cmd is inflight in step 3), we have to do
setp 1) and 2) first. As I explained in patch 0, blk_mq_tag_to_rq()
can't make sure the returned request is not freed:
nbd_read_stat
blk_mq_sched_free_requests
blk_mq_free_rqs
blk_mq_tag_to_rq
-> get rq before clear mapping
blk_mq_clear_rq_mapping
__free_pages -> rq is freed
blk_mq_request_started -> UAF
If the above can happen, blk_mq_find_and_get_req() may not fix it too, just
Hi, Ming
Why can't blk_mq_find_and_get_req() fix it? I can't think of any
scenario that might have problem currently.
quoted hunk
wondering why not take the following simpler way for avoiding the UAF?
We can't make sure freeze_queue is called before this, thus this approch
can't fix the problem, right?
nbd_read_stat
blk_mq_tag_to_rq
elevator_switch
blk_mq_freeze_queue(q);
elevator_switch_mq
elevator_exit
blk_mq_sched_free_requests
blk_mq_request_started -> UAF
Thanks,
Kuai
quoted hunk
while (1) {
cmd = nbd_read_stat(nbd, args->index);
if (IS_ERR(cmd)) {
On Tue, Sep 14, 2021 at 03:13:38PM +0800, yukuai (C) wrote:
quoted
On 2021/09/14 14:44, Ming Lei wrote:
quoted
On Tue, Sep 14, 2021 at 11:11:06AM +0800, yukuai (C) wrote:
quoted
On 2021/09/14 9:11, Ming Lei wrote:
quoted
On Thu, Sep 09, 2021 at 10:12:55PM +0800, Yu Kuai wrote:
quoted
blk_mq_tag_to_rq() can only ensure to return valid request in
following situation:
1) client send request message to server first
submit_bio
...
blk_mq_get_tag
...
blk_mq_get_driver_tag
...
nbd_queue_rq
nbd_handle_cmd
nbd_send_cmd
2) client receive respond message from server
recv_work
nbd_read_stat
blk_mq_tag_to_rq
If step 1) is missing, blk_mq_tag_to_rq() will return a stale
request, which might be freed. Thus convert to use
blk_mq_find_and_get_req() to make sure the returned request is not
freed.
But NBD_CMD_INFLIGHT has been added for checking if the reply is
expected, do we still need blk_mq_find_and_get_req() for covering
this issue? BTW, request and its payload is pre-allocated, so there
isn't real use-after-free.
Hi, Ming
Checking NBD_CMD_INFLIGHT relied on the request founded by tag is
valid,
not the other way round.
nbd_read_stat
req = blk_mq_tag_to_rq()
cmd = blk_mq_rq_to_pdu(req)
mutex_lock(cmd->lock)
checking NBD_CMD_INFLIGHT
Request and its payload is pre-allocated, and either req->ref or
cmd->lock can
serve the same purpose here. Once cmd->lock is held, you can check
if the cmd is
inflight or not. If it isn't inflight, just return -ENOENT. Is there
any
problem to handle in this way?
Hi, Ming
in nbd_read_stat:
1) get a request by tag first
2) get nbd_cmd by the request
3) hold cmd->lock and check if cmd is inflight
If we want to check if the cmd is inflight in step 3), we have to do
setp 1) and 2) first. As I explained in patch 0, blk_mq_tag_to_rq()
can't make sure the returned request is not freed:
nbd_read_stat
blk_mq_sched_free_requests
blk_mq_free_rqs
blk_mq_tag_to_rq
-> get rq before clear mapping
blk_mq_clear_rq_mapping
__free_pages -> rq is freed
blk_mq_request_started -> UAF
If the above can happen, blk_mq_find_and_get_req() may not fix it too,
just
Hi, Ming
Why can't blk_mq_find_and_get_req() fix it? I can't think of any
scenario that might have problem currently.
quoted
wondering why not take the following simpler way for avoiding the UAF?
We can't make sure freeze_queue is called before this, thus this approch
can't fix the problem, right?
nbd_read_stat
blk_mq_tag_to_rq
elevator_switch
blk_mq_freeze_queue(q);
elevator_switch_mq
elevator_exit
blk_mq_sched_free_requests
blk_mq_request_started -> UAF
Hi, Ming
I forgot that if percpu_ref_tryget succeed here, blk_mq_free_queue()
will block untill blk_queue_exit() in nbd_read_stat().
Thanks,
Kuai
Thanks,
Kuai
quoted
while (1) {
cmd = nbd_read_stat(nbd, args->index);
if (IS_ERR(cmd)) {
Hi, Ming
This apporch is wrong.
If blk_mq_freeze_queue() is called, and nbd is waiting for all
request to complete. percpu_ref_tryget() will fail here, and deadlock
will occur because request can't complete in recv_work().
Thanks,
Kuai
On Tue, Sep 14, 2021 at 05:08:00PM +0800, yukuai (C) wrote:
On 2021/09/14 15:46, Ming Lei wrote:
quoted
On Tue, Sep 14, 2021 at 03:13:38PM +0800, yukuai (C) wrote:
quoted
On 2021/09/14 14:44, Ming Lei wrote:
quoted
On Tue, Sep 14, 2021 at 11:11:06AM +0800, yukuai (C) wrote:
quoted
On 2021/09/14 9:11, Ming Lei wrote:
quoted
On Thu, Sep 09, 2021 at 10:12:55PM +0800, Yu Kuai wrote:
quoted
blk_mq_tag_to_rq() can only ensure to return valid request in
following situation:
1) client send request message to server first
submit_bio
...
blk_mq_get_tag
...
blk_mq_get_driver_tag
...
nbd_queue_rq
nbd_handle_cmd
nbd_send_cmd
2) client receive respond message from server
recv_work
nbd_read_stat
blk_mq_tag_to_rq
If step 1) is missing, blk_mq_tag_to_rq() will return a stale
request, which might be freed. Thus convert to use
blk_mq_find_and_get_req() to make sure the returned request is not
freed.
But NBD_CMD_INFLIGHT has been added for checking if the reply is
expected, do we still need blk_mq_find_and_get_req() for covering
this issue? BTW, request and its payload is pre-allocated, so there
isn't real use-after-free.
Hi, Ming
Checking NBD_CMD_INFLIGHT relied on the request founded by tag is valid,
not the other way round.
nbd_read_stat
req = blk_mq_tag_to_rq()
cmd = blk_mq_rq_to_pdu(req)
mutex_lock(cmd->lock)
checking NBD_CMD_INFLIGHT
Request and its payload is pre-allocated, and either req->ref or cmd->lock can
serve the same purpose here. Once cmd->lock is held, you can check if the cmd is
inflight or not. If it isn't inflight, just return -ENOENT. Is there any
problem to handle in this way?
Hi, Ming
in nbd_read_stat:
1) get a request by tag first
2) get nbd_cmd by the request
3) hold cmd->lock and check if cmd is inflight
If we want to check if the cmd is inflight in step 3), we have to do
setp 1) and 2) first. As I explained in patch 0, blk_mq_tag_to_rq()
can't make sure the returned request is not freed:
nbd_read_stat
blk_mq_sched_free_requests
blk_mq_free_rqs
blk_mq_tag_to_rq
-> get rq before clear mapping
blk_mq_clear_rq_mapping
__free_pages -> rq is freed
blk_mq_request_started -> UAF
If the above can happen, blk_mq_find_and_get_req() may not fix it too, just
Hi, Ming
Why can't blk_mq_find_and_get_req() fix it? I can't think of any
scenario that might have problem currently.
The principle behind blk_mq_find_and_get_req() is that if one request's
ref is grabbed, the queue's usage counter is guaranteed to be grabbed,
and this way isn't straight-forward.
Yeah, it can fix the issue, but I don't think it is good to call it in
fast path cause tags->lock is required.
quoted
wondering why not take the following simpler way for avoiding the UAF?
We can't make sure freeze_queue is called before this, thus this approch
can't fix the problem, right?
nbd_read_stat
blk_mq_tag_to_rq
elevator_switch
blk_mq_freeze_queue(q);
elevator_switch_mq
elevator_exit
blk_mq_sched_free_requests
blk_mq_request_started -> UAF
No, blk_mq_freeze_queue() waits until .q_usage_counter becomes zero, so
there won't be any concurrent nbd_read_stat() during switching elevator
if ->q_usage_counter is grabbed in recv_work().
Thanks,
Ming
Hi, Ming
This apporch is wrong.
If blk_mq_freeze_queue() is called, and nbd is waiting for all
request to complete. percpu_ref_tryget() will fail here, and deadlock
will occur because request can't complete in recv_work().
No, percpu_ref_tryget() won't fail until ->q_usage_counter is zero, when
it is perfectly fine to do nothing in recv_work().
Thanks,
Ming
Hi, Ming
This apporch is wrong.
If blk_mq_freeze_queue() is called, and nbd is waiting for all
request to complete. percpu_ref_tryget() will fail here, and deadlock
will occur because request can't complete in recv_work().
No, percpu_ref_tryget() won't fail until ->q_usage_counter is zero, when
it is perfectly fine to do nothing in recv_work().
Hi Ming
This apporch is a good idea, however we should not get q_usage_counter
in reccv_work(), because It will block freeze queue.
How about get q_usage_counter in nbd_read_stat(), and put in error path
or after request completion?
Thanks
Kuai
Hi, Ming
This apporch is wrong.
If blk_mq_freeze_queue() is called, and nbd is waiting for all
request to complete. percpu_ref_tryget() will fail here, and deadlock
will occur because request can't complete in recv_work().
No, percpu_ref_tryget() won't fail until ->q_usage_counter is zero, when
it is perfectly fine to do nothing in recv_work().
Hi Ming
This apporch is a good idea, however we should not get q_usage_counter
in reccv_work(), because It will block freeze queue.
How about get q_usage_counter in nbd_read_stat(), and put in error path
or after request completion?
OK, looks I missed that nbd_read_stat() needs to wait for incoming reply
first, so how about the following change by partitioning nbd_read_stat()
into nbd_read_reply() and nbd_handle_reply()?
Hi, Ming
This apporch is wrong.
If blk_mq_freeze_queue() is called, and nbd is waiting for all
request to complete. percpu_ref_tryget() will fail here, and deadlock
will occur because request can't complete in recv_work().
No, percpu_ref_tryget() won't fail until ->q_usage_counter is zero, when
it is perfectly fine to do nothing in recv_work().
Hi Ming
This apporch is a good idea, however we should not get q_usage_counter
in reccv_work(), because It will block freeze queue.
How about get q_usage_counter in nbd_read_stat(), and put in error path
or after request completion?
OK, looks I missed that nbd_read_stat() needs to wait for incoming reply
first, so how about the following change by partitioning nbd_read_stat()
into nbd_read_reply() and nbd_handle_reply()?
Hi, Ming
The change looks good to me.
Do you want to send a patch to fix this?
Thanks,
Kuai
Hi, Ming
This apporch is wrong.
If blk_mq_freeze_queue() is called, and nbd is waiting for all
request to complete. percpu_ref_tryget() will fail here, and deadlock
will occur because request can't complete in recv_work().
No, percpu_ref_tryget() won't fail until ->q_usage_counter is zero, when
it is perfectly fine to do nothing in recv_work().
Hi Ming
This apporch is a good idea, however we should not get q_usage_counter
in reccv_work(), because It will block freeze queue.
How about get q_usage_counter in nbd_read_stat(), and put in error path
or after request completion?
OK, looks I missed that nbd_read_stat() needs to wait for incoming reply
first, so how about the following change by partitioning nbd_read_stat()
into nbd_read_reply() and nbd_handle_reply()?
Hi, Ming
The change looks good to me.
Do you want to send a patch to fix this?
I guess you may add inflight check or sort of change in nbd_read_stat(), so feel
free to fold it into your series.
Thanks,
Ming