Fwd: [powerpc/Baremetal]Kernel OOPS while executing memory hotplug on Power8 baremetal

8 messages, 4 authors, 2018-06-07 · open the first message on its own page

Fwd: [powerpc/Baremetal]Kernel OOPS while executing memory hotplug on Power8 baremetal

From: vrbagal1 <hidden>
Date: 2018-06-07 07:08:05

+scsi mailing list, and edited the subject line.

Pasting the traces here.

[ 2484.634761] Unable to handle kernel paging request for instruction 
fetch
[ 2484.634849] Faulting instruction address: 0x00000000
[ 2484.634862] Oops: Kernel access of bad area, sig: 11 [#1]
[ 2484.634905] LE SMP NR_CPUS=2048 NUMA PowerNV
[ 2484.634991] Dumping ftrace buffer:
[ 2484.635116]    (ftrace buffer empty)
[ 2484.635158] Modules linked in: binfmt_misc ipt_MASQUERADE 
nf_nat_masquerade_ipv4 tun bridge stp llc xt_tcpudp ipt_REJECT 
nf_reject_ipv4 xt_conntrack nfnetlink iptable_nat nf_conntrack_ipv4 
nf_defrag_ipv4 nf_nat_ipv4 nf_nat nf_conntrack iptable_mangle 
iptable_filter powernv_rng rng_core ipmi_powernv ipmi_devintf 
ipmi_msghandler vmx_crypto leds_powernv powernv_op_panel led_class 
kvm_hv nfsd kvm ip_tables x_tables autofs4
[ 2484.635528] CPU: 48 PID: 0 Comm: swapper/48 Not tainted 
4.17.0-autotest #1
[ 2484.635591] NIP:  0000000000000000 LR: c00000000014beb4 CTR: 
0000000000000000
[ 2484.635667] REGS: c000000ffff4f600 TRAP: 0400   Not tainted  
(4.17.0-autotest)
[ 2484.635741] MSR:  9000000040009033 <SF,HV,EE,ME,IR,DR,RI,LE>  CR: 
28028028  XER: 20000000
[ 2484.635835] CFAR: c000000000008934 SOFTE: 1
[ 2484.635835] GPR00: c00000000014c4ec c000000ffff4f880 c0000000010fb800 
c0000007eb6d1d10
[ 2484.635835] GPR04: 0000000000000003 0000000000000000 0000000000000000 
0000000000000000
[ 2484.635835] GPR08: c000000ffff4f910 0000000000000000 0000000000000000 
c000000f11557970
[ 2484.635835] GPR12: 0000000000000000 c000000ffffb1880 c000000f272dff90 
0000000000200042
[ 2484.635835] GPR16: 0000000100035566 c000000ffff4c000 0000000000000000 
0000000000000001
[ 2484.635835] GPR20: c000000000dd6c80 c000000001123b00 0000000000000005 
c000000ffff4f910
[ 2484.635835] GPR24: 0000000000000001 0000000000000000 0000000000000000 
0000000000000003
[ 2484.635835] GPR28: 0000000000000000 c0000007e03001b0 0000000000000000 
ffffffffffffffe8
[ 2484.636474] NIP [0000000000000000]           (null)
[ 2484.636530] LR [c00000000014beb4] __wake_up_common+0xe4/0x1e0
[ 2484.636593] Call Trace:
[ 2484.636620] [c000000ffff4f880] [c000000ffff4f8c0] 0xc000000ffff4f8c0 
(unreliable)
[ 2484.636698] [c000000ffff4f8f0] [c00000000014c4ec] 
__wake_up_common_lock+0xac/0x100
[ 2484.636776] [c000000ffff4f980] [c000000000243acc] 
mempool_free+0xcc/0xf0
[ 2484.636842] [c000000ffff4f9b0] [c00000000052c090] bio_free+0x50/0x90
[ 2484.636907] [c000000ffff4f9e0] [c000000000850dc0] 
dec_pending+0x130/0x310
[ 2484.636971] [c000000ffff4fa60] [c0000000008512fc] 
clone_endio+0xcc/0x180
[ 2484.637036] [c000000ffff4fae0] [c00000000052c2a4] 
bio_endio+0x164/0x280
[ 2484.637102] [c000000ffff4fb80] [c000000000538eb0] 
blk_update_request+0xf0/0x4a0
[ 2484.637179] [c000000ffff4fc20] [c0000000006ba600] 
scsi_end_request+0x50/0x270
[ 2484.637255] [c000000ffff4fc80] [c0000000006baa44] 
scsi_io_completion+0x224/0x6b0
[ 2484.637332] [c000000ffff4fd10] [c0000000006b09a8] 
scsi_finish_command+0x138/0x170
[ 2484.637408] [c000000ffff4fd50] [c0000000006b9b28] 
scsi_softirq_done+0x178/0x1d0
[ 2484.637485] [c000000ffff4fdd0] [c000000000543e88] 
blk_done_softirq+0xa8/0xd0
[ 2484.637562] [c000000ffff4fe10] [c000000000a1fdbc] 
__do_softirq+0x15c/0x3b4
[ 2484.637628] [c000000ffff4ff00] [c0000000000f7d98] irq_exit+0xf8/0x110
[ 2484.637693] [c000000ffff4ff20] [c000000000016e98] __do_irq+0x98/0x200
[ 2484.637759] [c000000ffff4ff90] [c000000000028bd4] 
call_do_irq+0x14/0x24
[ 2484.637823] [c000000f272dfa50] [c000000000017094] do_IRQ+0x94/0x110
[ 2484.637888] [c000000f272dfaa0] [c000000000008db8] 
hardware_interrupt_common+0x158/0x160
[ 2484.637967] --- interrupt: 501 at replay_interrupt_return+0x0/0x4
[ 2484.637967]     LR = arch_local_irq_restore+0x74/0x90
[ 2484.638068] [c000000f272dfd90] [c000000000871c88] 
menu_select+0xc8/0x7f0 (unreliable)
[ 2484.638145] [c000000f272dfdb0] [c00000000086fc88] 
cpuidle_enter_state+0x108/0x3c0
[ 2484.638222] [c000000f272dfe10] [c0000000001308e4] 
call_cpuidle+0x44/0x80
[ 2484.638286] [c000000f272dfe30] [c000000000130e78] do_idle+0x2f8/0x3a0
[ 2484.638350] [c000000f272dfec0] [c0000000001310f0] 
cpu_startup_entry+0x30/0x40
[ 2484.638427] [c000000f272dfef0] [c000000000043580] 
start_secondary+0x4d0/0x520
[ 2484.638504] [c000000f272dff90] [c00000000000b284] 
start_secondary_resume+0x10/0x14
[ 2484.638579] Instruction dump:
[ 2484.638618] XXXXXXXX XXXXXXXX XXXXXXXX XXXXXXXX XXXXXXXX XXXXXXXX 
XXXXXXXX XXXXXXXX
[ 2484.638697] XXXXXXXX XXXXXXXX XXXXXXXX XXXXXXXX XXXXXXXX XXXXXXXX 
XXXXXXXX XXXXXXXX
[ 2484.638778] ---[ end trace 9e63f5da6878e977 ]---
[ 2484.639008]
[ 2485.639101] Kernel panic - not syncing: Fatal exception in interrupt
[ 2485.639465] Dumping ftrace buffer:
[ 2485.639514]    (ftrace buffer empty)
[ 2485.639732] Rebooting in 10 seconds..


Cheers,
Venkat.
IBM Linux Technology Centre


-------- Original Message --------
Subject: [mainline] [powerpc/powervm]Kernel OOPS while executing memory 
hotplug on Power8 baremetal
Date: 2018-06-07 11:55
 From: vrbagal1 [off-list ref]
To: linuxppc-dev <redacted>
Cc: sachinp <redacted>, mpe@ellerman.id.au

Greetings!!!

Observing Kernel oops and machine reboots while executing memory hotplug 
test case, on Power8 Baremetal machine.

I see this is introduced some where between rc6 and 4.17.

Machine: Power 8 Baremetal
Kernel Version: Linux version 4.17.0-autotest
gcc Version: gcc version 4.8.5 20150623 (Red Hat 4.8.5-28) (GCC))


Attached is the .config file and traces found.

Cheers.

Re: Fwd: [powerpc/Baremetal]Kernel OOPS while executing memory hotplug on Power8 baremetal

From: Bart Van Assche <hidden>
Date: 2018-06-07 07:16:50

T24gVGh1LCAyMDE4LTA2LTA3IGF0IDEyOjM4ICswNTMwLCB2cmJhZ2FsMSB3cm90ZToNCj4gT2Jz
ZXJ2aW5nIEtlcm5lbCBvb3BzIGFuZCBtYWNoaW5lIHJlYm9vdHMgd2hpbGUgZXhlY3V0aW5nIG1l
bW9yeSBob3RwbHVnIA0KPiB0ZXN0IGNhc2UsIG9uIFBvd2VyOCBCYXJlbWV0YWwgbWFjaGluZS4N
Cj4gDQo+IEkgc2VlIHRoaXMgaXMgaW50cm9kdWNlZCBzb21lIHdoZXJlIGJldHdlZW4gcmM2IGFu
ZCA0LjE3Lg0KDQpQbGVhc2UgcHJvdmlkZSB0aGUgZXhhY3QgdmVyc2lvbnMgKGdpdCBjb21taXQg
SURzKSBvZiB0aGUga2VybmVsIHZlcnNpb25zDQp5b3UgaGF2ZSB0ZXN0ZWQuDQoNClRoYW5rcywN
Cg0KQmFydC4NCg0KDQoNCg0K

Re: Fwd: [powerpc/Baremetal]Kernel OOPS while executing memory hotplug on Power8 baremetal

From: Venkat Rao B <hidden>
Date: 2018-06-07 07:26:53


On Thursday 07 June 2018 12:46 PM, Bart Van Assche wrote:
On Thu, 2018-06-07 at 12:38 +0530, vrbagal1 wrote:
quoted
Observing Kernel oops and machine reboots while executing memory hotplug
test case, on Power8 Baremetal machine.

I see this is introduced some where between rc6 and 4.17.
Please provide the exact versions (git commit IDs) of the kernel versions
you have tested.
Commit Id ---> 5037be168f

Regards,
Venkat
IBM Linux Technology Centre

Thanks,

Bart.


Re: Fwd: [powerpc/Baremetal]Kernel OOPS while executing memory hotplug on Power8 baremetal

From: Bart Van Assche <hidden>
Date: 2018-06-07 07:42:51

T24gVGh1LCAyMDE4LTA2LTA3IGF0IDEyOjU2ICswNTMwLCBWZW5rYXQgUmFvIEIgd3JvdGU6DQo+
IE9uIFRodXJzZGF5IDA3IEp1bmUgMjAxOCAxMjo0NiBQTSwgQmFydCBWYW4gQXNzY2hlIHdyb3Rl
Og0KPiA+IE9uIFRodSwgMjAxOC0wNi0wNyBhdCAxMjozOCArMDUzMCwgdnJiYWdhbDEgd3JvdGU6
DQo+ID4gPiBPYnNlcnZpbmcgS2VybmVsIG9vcHMgYW5kIG1hY2hpbmUgcmVib290cyB3aGlsZSBl
eGVjdXRpbmcgbWVtb3J5IGhvdHBsdWcNCj4gPiA+IHRlc3QgY2FzZSwgb24gUG93ZXI4IEJhcmVt
ZXRhbCBtYWNoaW5lLg0KPiA+ID4gDQo+ID4gPiBJIHNlZSB0aGlzIGlzIGludHJvZHVjZWQgc29t
ZSB3aGVyZSBiZXR3ZWVuIHJjNiBhbmQgNC4xNy4NCj4gPiANCj4gPiBQbGVhc2UgcHJvdmlkZSB0
aGUgZXhhY3QgdmVyc2lvbnMgKGdpdCBjb21taXQgSURzKSBvZiB0aGUga2VybmVsIHZlcnNpb25z
DQo+ID4geW91IGhhdmUgdGVzdGVkLg0KPiANCj4gQ29tbWl0IElkIC0tLT4gNTAzN2JlMTY4Zg0K
DQpUaGUgcmVhc29uIEkgd2FzIGFza2luZyBmb3IgdGhlIGNvbW1pdCBJRCBpcyBiZWNhdXNlIEkg
c2F3IHRoYXQgY2xvbmVfZW5kaW8oKQ0Kb2NjdXJzIGluIHRoZSBvb3BzIHdoaWNoIG1lYW5zIHRo
YXQgdGhlIGRtIGRyaXZlciBpcyBpbnZvbHZlZC4gQW4gaW1wb3J0YW50IGZpeA0KZm9yIHRoZSBk
bSBkcml2ZXIgd2VudCB1cHN0cmVhbSByZWNlbnRseSwgbmFtZWx5IGQzNzc1MzU0MDU2OCAoImRt
OiBVc2Uga3phbGxvYw0KZm9yIGFsbCBzdHJ1Y3RzIHdpdGggZW1iZWRkZWQgYmlvc2V0cy9tZW1w
b29scyIpLiBDYW4geW91IGRvdWJsZSBjaGVjayB3aGV0aGVyDQp0aGF0IGNvbW1pdCBpdCBwcmVz
ZW50IGluIHlvdXIgdHJlZT8gSWYgaXQgaXMgbm90IHByZXNlbnQsIHBsZWFzZSB1cGRhdGUgdG8g
dGhlDQpsYXRlc3QgbWFzdGVyIGFuZCByZXRlc3QuIElmIGl0IGlzIHByZXNlbnQsIHBsZWFzZSBy
ZXBvcnQgaG93IHRvIHJlcHJvZHVjZQ0KdGhpcyBvb3BzIHRvIEtlbnQgT3ZlcnN0cmVldCwgSmVu
cyBBeGJvZSwgbGludXgtYmxvY2sgYW5kIE1pa2UgU25pdHplci4NCg0KVGhhbmtzLA0KDQpCYXJ0
Lg==

Re: Fwd: [powerpc/Baremetal]Kernel OOPS while executing memory hotplug on Power8 baremetal

From: vrbagal1 <hidden>
Date: 2018-06-07 10:37:23

On 2018-06-07 13:12, Bart Van Assche wrote:
On Thu, 2018-06-07 at 12:56 +0530, Venkat Rao B wrote:
quoted
On Thursday 07 June 2018 12:46 PM, Bart Van Assche wrote:
quoted
On Thu, 2018-06-07 at 12:38 +0530, vrbagal1 wrote:
quoted
Observing Kernel oops and machine reboots while executing memory hotplug
test case, on Power8 Baremetal machine.

I see this is introduced some where between rc6 and 4.17.
Please provide the exact versions (git commit IDs) of the kernel versions
you have tested.
Commit Id ---> 5037be168f
The reason I was asking for the commit ID is because I saw that 
clone_endio()
occurs in the oops which means that the dm driver is involved. An 
important fix
for the dm driver went upstream recently, namely d37753540568 ("dm: Use 
kzalloc
for all structs with embedded biosets/mempools"). Can you double check 
whether
that commit it present in your tree? If it is not present, please 
update to the
latest master and retest. If it is present, please report how to 
reproduce
this oops to Kent Overstreet, Jens Axboe, linux-block and Mike Snitzer.

Thanks,

Bart.

Yes, the fix is present in the tree, which I have tested.

Steps to reproduce:

Step1: Clone and Install avocado git clone 
https://github.com/avocado-framework/avocado.git
Step2: Clone 
https://github.com/avocado-framework-tests/avocado-misc-tests.git
        Test case is 
https://github.com/avocado-framework-tests/avocado-misc-tests/blob/master/memory/memhotplug.py
Step3: Command to run the test is avocado run 
avocado-misc-tests/memory/memhotplug.py

Regards,
Venkat.

Re: Fwd: [powerpc/Baremetal]Kernel OOPS while executing memory hotplug on Power8 baremetal

From: Michael Ellerman <mpe@ellerman.id.au>
Date: 2018-06-07 12:51:44

vrbagal1 [off-list ref] writes:
On 2018-06-07 13:12, Bart Van Assche wrote:
quoted
On Thu, 2018-06-07 at 12:56 +0530, Venkat Rao B wrote:
quoted
On Thursday 07 June 2018 12:46 PM, Bart Van Assche wrote:
quoted
On Thu, 2018-06-07 at 12:38 +0530, vrbagal1 wrote:
quoted
Observing Kernel oops and machine reboots while executing memory hotplug
test case, on Power8 Baremetal machine.

I see this is introduced some where between rc6 and 4.17.
Please provide the exact versions (git commit IDs) of the kernel versions
you have tested.
Commit Id ---> 5037be168f
The reason I was asking for the commit ID is because I saw that 
clone_endio()
occurs in the oops which means that the dm driver is involved. An 
important fix
for the dm driver went upstream recently, namely d37753540568 ("dm: Use 
kzalloc
for all structs with embedded biosets/mempools"). Can you double check 
whether
that commit it present in your tree? If it is not present, please 
update to the
latest master and retest. If it is present, please report how to 
reproduce
this oops to Kent Overstreet, Jens Axboe, linux-block and Mike Snitzer.

Thanks,

Bart.

Yes, the fix is present in the tree, which I have tested.

Steps to reproduce:

Step1: Clone and Install avocado git clone 
https://github.com/avocado-framework/avocado.git
Step2: Clone 
https://github.com/avocado-framework-tests/avocado-misc-tests.git
        Test case is 
https://github.com/avocado-framework-tests/avocado-misc-tests/blob/master/memory/memhotplug.py
Step3: Command to run the test is avocado run 
avocado-misc-tests/memory/memhotplug.py
That gave me:

  $ avocado run avocado-misc-tests/memory/memhotplug.py
  avocado: command not found

Was I meant to install it?

I tried this which worked (I think):

  $ ./scripts/avocado run avocado-misc-tests/memory/memhotplug.py
  Failed to load plugin from module "avocado_runner_vm": ImportError('No module named libvirt',)
  JOB ID     : 28deb5a455fb876a7e177deb2b46eab640f313c8
  JOB LOG    : /home/michael/avocado/job-results/job-2018-06-07T22.27-28deb5a/job.log
   (1/4) avocado-misc-tests/memory/memhotplug.py:memstress.test_hotplug_loop: PASS (10.62 s)
   (2/4) avocado-misc-tests/memory/memhotplug.py:memstress.test_hotplug_toggle: PASS (245.15 s)
   (3/4) avocado-misc-tests/memory/memhotplug.py:memstress.test_dlpar_mem_hotplug: PASS (0.37 s)
   (4/4) avocado-misc-tests/memory/memhotplug.py:memstress.test_hotplug_per_numa_node: PASS (41.09 s)
  RESULTS    : PASS 4 | ERROR 0 | FAIL 0 | SKIP 0 | WARN 0 | INTERRUPT 0 | CANCEL 0
  JOB TIME   : 323.45 s
  JOB HTML   : /home/michael/avocado/job-results/job-2018-06-07T22.27-28deb5a/results.html


So what's different about your system?

What does 'lsblk -O' say on your system?

cheers

Re: Fwd: [powerpc/Baremetal]Kernel OOPS while executing memory hotplug on Power8 baremetal

From: Jens Axboe <axboe@kernel.dk>
Date: 2018-06-07 14:45:41

On 6/7/18 4:37 AM, vrbagal1 wrote:
On 2018-06-07 13:12, Bart Van Assche wrote:
quoted
On Thu, 2018-06-07 at 12:56 +0530, Venkat Rao B wrote:
quoted
On Thursday 07 June 2018 12:46 PM, Bart Van Assche wrote:
quoted
On Thu, 2018-06-07 at 12:38 +0530, vrbagal1 wrote:
quoted
Observing Kernel oops and machine reboots while executing memory hotplug
test case, on Power8 Baremetal machine.

I see this is introduced some where between rc6 and 4.17.
Please provide the exact versions (git commit IDs) of the kernel versions
you have tested.
Commit Id ---> 5037be168f
The reason I was asking for the commit ID is because I saw that 
clone_endio()
occurs in the oops which means that the dm driver is involved. An 
important fix
for the dm driver went upstream recently, namely d37753540568 ("dm: Use 
kzalloc
for all structs with embedded biosets/mempools"). Can you double check 
whether
that commit it present in your tree? If it is not present, please 
update to the
latest master and retest. If it is present, please report how to 
reproduce
this oops to Kent Overstreet, Jens Axboe, linux-block and Mike Snitzer.

Thanks,

Bart.

Yes, the fix is present in the tree, which I have tested.

Steps to reproduce:

Step1: Clone and Install avocado git clone 
https://github.com/avocado-framework/avocado.git
Step2: Clone 
https://github.com/avocado-framework-tests/avocado-misc-tests.git
        Test case is 
https://github.com/avocado-framework-tests/avocado-misc-tests/blob/master/memory/memhotplug.py
Step3: Command to run the test is avocado run 
avocado-misc-tests/memory/memhotplug.py
Can you try with the below? Not a fully formed fix since I'd prefer
if the dm bioset copy stuff was changed instead, but worth a shot.

diff --git a/block/bio.c b/block/bio.c
index 595663e0281a..45bdee67d28b 100644
--- a/block/bio.c
+++ b/block/bio.c
@@ -1967,6 +1967,27 @@ int bioset_init(struct bio_set *bs,
 }
 EXPORT_SYMBOL(bioset_init);
 
+void bioset_move(struct bio_set *dst, struct bio_set *src)
+{
+	dst->bio_slab = src->bio_slab;
+	dst->front_pad = src->front_pad;
+	mempool_move(&dst->bio_pool, &src->bio_pool);
+	mempool_move(&dst->bvec_pool, &src->bvec_pool);
+#if defined(CONFIG_BLK_DEV_INTEGRITY)
+	mempool_move(&dst->bio_integrity_pool, &src->bio_integrity_pool);
+	mempool_move(&dst->bvec_integrity_pool, &src->bvec_integrity_pool);
+#endif
+	BUG_ON(!bio_list_empty(&src->rescue_list));
+	BUG_ON(work_pending(&src->rescue_work));
+	spin_lock_init(&dst->rescue_lock);
+	bio_list_init(&dst->rescue_list);
+	INIT_WORK(&dst->rescue_work, bio_alloc_rescue);
+	dst->rescue_workqueue = src->rescue_workqueue;
+
+	memset(src, 0, sizeof(*src));
+}
+EXPORT_SYMBOL(bioset_move);
+
 #ifdef CONFIG_BLK_CGROUP
 
 /**
diff --git a/drivers/md/dm.c b/drivers/md/dm.c
index 98dff36b89a3..87f636815baf 100644
--- a/drivers/md/dm.c
+++ b/drivers/md/dm.c
@@ -1982,10 +1982,8 @@ static void __bind_mempools(struct mapped_device *md, struct dm_table *t)
 	       bioset_initialized(&md->bs) ||
 	       bioset_initialized(&md->io_bs));
 
-	md->bs = p->bs;
-	memset(&p->bs, 0, sizeof(p->bs));
-	md->io_bs = p->io_bs;
-	memset(&p->io_bs, 0, sizeof(p->io_bs));
+	bioset_move(&md->bs, &p->bs);
+	bioset_move(&md->io_bs, &p->io_bs);
 out:
 	/* mempool bind completed, no longer need any mempools in the table */
 	dm_table_free_md_mempools(t);
diff --git a/include/linux/bio.h b/include/linux/bio.h
index 810a8bee8f85..7581231dd0a3 100644
--- a/include/linux/bio.h
+++ b/include/linux/bio.h
@@ -417,6 +417,7 @@ enum {
 extern int bioset_init(struct bio_set *, unsigned int, unsigned int, int flags);
 extern void bioset_exit(struct bio_set *);
 extern int biovec_init_pool(mempool_t *pool, int pool_entries);
+extern void bioset_move(struct bio_set *dst, struct bio_set *src);
 
 extern struct bio *bio_alloc_bioset(gfp_t, unsigned int, struct bio_set *);
 extern void bio_put(struct bio *);
diff --git a/include/linux/mempool.h b/include/linux/mempool.h
index 0c964ac107c2..20818919180c 100644
--- a/include/linux/mempool.h
+++ b/include/linux/mempool.h
@@ -47,6 +47,7 @@ extern int mempool_resize(mempool_t *pool, int new_min_nr);
 extern void mempool_destroy(mempool_t *pool);
 extern void *mempool_alloc(mempool_t *pool, gfp_t gfp_mask) __malloc;
 extern void mempool_free(void *element, mempool_t *pool);
+extern void mempool_move(mempool_t *dst, mempool_t *src);
 
 /*
  * A mempool_alloc_t and mempool_free_t that get the memory from
diff --git a/mm/mempool.c b/mm/mempool.c
index b54f2c20e5e0..dd402653367b 100644
--- a/mm/mempool.c
+++ b/mm/mempool.c
@@ -181,6 +181,8 @@ int mempool_init_node(mempool_t *pool, int min_nr, mempool_alloc_t *alloc_fn,
 		      mempool_free_t *free_fn, void *pool_data,
 		      gfp_t gfp_mask, int node_id)
 {
+	memset(pool, 0, sizeof(*pool));
+
 	spin_lock_init(&pool->lock);
 	pool->min_nr	= min_nr;
 	pool->pool_data = pool_data;
@@ -546,3 +548,19 @@ void mempool_free_pages(void *element, void *pool_data)
 	__free_pages(element, order);
 }
 EXPORT_SYMBOL(mempool_free_pages);
+
+void mempool_move(mempool_t *dst, mempool_t *src)
+{
+	BUG_ON(waitqueue_active(&src->wait));
+
+	spin_lock_init(&dst->lock);
+	dst->min_nr = src->min_nr;
+	dst->curr_nr = src->curr_nr;
+	memcpy(dst->elements, src->elements, sizeof(void *) * src->curr_nr);
+	dst->pool_data = src->pool_data;
+	dst->alloc = src->alloc;
+	dst->free = src->free;
+	init_waitqueue_head(&dst->wait);
+
+	memset(src, 0, sizeof(*src));
+}
-- 
Jens Axboe

Re: Fwd: [powerpc/Baremetal]Kernel OOPS while executing memory hotplug on Power8 baremetal

From: Jens Axboe <axboe@kernel.dk>
Date: 2018-06-07 15:40:47

On 6/7/18 8:45 AM, Jens Axboe wrote:
On 6/7/18 4:37 AM, vrbagal1 wrote:
quoted
On 2018-06-07 13:12, Bart Van Assche wrote:
quoted
On Thu, 2018-06-07 at 12:56 +0530, Venkat Rao B wrote:
quoted
On Thursday 07 June 2018 12:46 PM, Bart Van Assche wrote:
quoted
On Thu, 2018-06-07 at 12:38 +0530, vrbagal1 wrote:
quoted
Observing Kernel oops and machine reboots while executing memory hotplug
test case, on Power8 Baremetal machine.

I see this is introduced some where between rc6 and 4.17.
Please provide the exact versions (git commit IDs) of the kernel versions
you have tested.
Commit Id ---> 5037be168f
The reason I was asking for the commit ID is because I saw that 
clone_endio()
occurs in the oops which means that the dm driver is involved. An 
important fix
for the dm driver went upstream recently, namely d37753540568 ("dm: Use 
kzalloc
for all structs with embedded biosets/mempools"). Can you double check 
whether
that commit it present in your tree? If it is not present, please 
update to the
latest master and retest. If it is present, please report how to 
reproduce
this oops to Kent Overstreet, Jens Axboe, linux-block and Mike Snitzer.

Thanks,

Bart.

Yes, the fix is present in the tree, which I have tested.

Steps to reproduce:

Step1: Clone and Install avocado git clone 
https://github.com/avocado-framework/avocado.git
Step2: Clone 
https://github.com/avocado-framework-tests/avocado-misc-tests.git
        Test case is 
https://github.com/avocado-framework-tests/avocado-misc-tests/blob/master/memory/memhotplug.py
Step3: Command to run the test is avocado run 
avocado-misc-tests/memory/memhotplug.py
Can you try with the below? Not a fully formed fix since I'd prefer
if the dm bioset copy stuff was changed instead, but worth a shot.
This is closer to an actual fix, please try that instead.

diff --git a/block/bio.c b/block/bio.c
index 595663e0281a..0616d86b15c6 100644
--- a/block/bio.c
+++ b/block/bio.c
@@ -1967,6 +1967,21 @@ int bioset_init(struct bio_set *bs,
 }
 EXPORT_SYMBOL(bioset_init);
 
+int bioset_init_from_src(struct bio_set *new, struct bio_set *src)
+{
+	unsigned int pool_size = src->bio_pool.min_nr;
+	int flags;
+
+	flags = 0;
+	if (src->bvec_pool.min_nr)
+		flags |= BIOSET_NEED_BVECS;
+	if (src->rescue_workqueue)
+		flags |= BIOSET_NEED_RESCUER;
+
+	return bioset_init(new, pool_size, src->front_pad, flags);
+}
+EXPORT_SYMBOL(bioset_init_from_src);
+
 #ifdef CONFIG_BLK_CGROUP
 
 /**
diff --git a/drivers/md/dm.c b/drivers/md/dm.c
index 98dff36b89a3..20a8d63754bf 100644
--- a/drivers/md/dm.c
+++ b/drivers/md/dm.c
@@ -1953,9 +1953,10 @@ static void free_dev(struct mapped_device *md)
 	kvfree(md);
 }
 
-static void __bind_mempools(struct mapped_device *md, struct dm_table *t)
+static int __bind_mempools(struct mapped_device *md, struct dm_table *t)
 {
 	struct dm_md_mempools *p = dm_table_get_md_mempools(t);
+	int ret = 0;
 
 	if (dm_table_bio_based(t)) {
 		/*
@@ -1982,13 +1983,16 @@ static void __bind_mempools(struct mapped_device *md, struct dm_table *t)
 	       bioset_initialized(&md->bs) ||
 	       bioset_initialized(&md->io_bs));
 
-	md->bs = p->bs;
-	memset(&p->bs, 0, sizeof(p->bs));
-	md->io_bs = p->io_bs;
-	memset(&p->io_bs, 0, sizeof(p->io_bs));
+	ret = bioset_init_from_src(&md->bs, &p->bs);
+	if (ret)
+		goto out;
+	ret = bioset_init_from_src(&md->io_bs, &p->io_bs);
+	if (ret)
+		bioset_exit(&md->bs);
 out:
 	/* mempool bind completed, no longer need any mempools in the table */
 	dm_table_free_md_mempools(t);
+	return ret;
 }
 
 /*
@@ -2033,6 +2037,7 @@ static struct dm_table *__bind(struct mapped_device *md, struct dm_table *t,
 	struct request_queue *q = md->queue;
 	bool request_based = dm_table_request_based(t);
 	sector_t size;
+	int ret;
 
 	lockdep_assert_held(&md->suspend_lock);
 
@@ -2068,7 +2073,11 @@ static struct dm_table *__bind(struct mapped_device *md, struct dm_table *t,
 		md->immutable_target = dm_table_get_immutable_target(t);
 	}
 
-	__bind_mempools(md, t);
+	ret = __bind_mempools(md, t);
+	if (ret) {
+		old_map = ERR_PTR(ret);
+		goto out;
+	}
 
 	old_map = rcu_dereference_protected(md->map, lockdep_is_held(&md->suspend_lock));
 	rcu_assign_pointer(md->map, (void *)t);
@@ -2078,6 +2087,7 @@ static struct dm_table *__bind(struct mapped_device *md, struct dm_table *t,
 	if (old_map)
 		dm_sync_table(md);
 
+out:
 	return old_map;
 }
 
diff --git a/include/linux/bio.h b/include/linux/bio.h
index 810a8bee8f85..307682ac2f31 100644
--- a/include/linux/bio.h
+++ b/include/linux/bio.h
@@ -417,6 +417,7 @@ enum {
 extern int bioset_init(struct bio_set *, unsigned int, unsigned int, int flags);
 extern void bioset_exit(struct bio_set *);
 extern int biovec_init_pool(mempool_t *pool, int pool_entries);
+extern int bioset_init_from_src(struct bio_set *new, struct bio_set *src);
 
 extern struct bio *bio_alloc_bioset(gfp_t, unsigned int, struct bio_set *);
 extern void bio_put(struct bio *);
-- 
Jens Axboe
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help