RE: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

10 messages, 4 authors, 2021-08-18 · open the first message on its own page

RE: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

From: Finlayson, James M CIV (USA) <hidden>
Date: 2021-08-05 19:52:13

Sorry - again..I sent HTML instead of plain text

Resend - mailing list bounce  
All, 
Sorry for the delay - both work and life got into the way.   Here is some feedback:

BLUF upfront with 5.14rc3 kernel that our SA built - md0 a 10+1+1 RAID5 - 5.332 M IOPS 20.3GiB/s, md1 a 10+1+1 RAID5, 5.892M IOPS 22.5GiB/s  - best hero numbers I've ever seen on mdraid  RAID5 IOPS.   I think the kernel patch is good.  Prior was  socket0 1.263M IOPS 4934MiB/s, socket1 1.071M IOSP, 4183MiB/s....   I'm willing to help push this as hard as we can until we hit a bottleneck outside of our control.   

I need to verify the RAW IOPS - admittedly this is a different server and I didn't do any regression testing before the kernel, but my raw were  socket0: 13.2M IOPS and socket1  13.5M IOPS.   Prior was socket0 16.0M IOPS and socket1 13.5M IOPS.   - admittedly there appears to a regression in the socket0 "hero run" but what I don't know that since this is a different server, I don't know if I have a configuration management issue in my zealousness to test this patch or whether we have a regression.   I was so excited to have the attention of kernel developers that needed my help that I borrowed another system, because I didn't want to tear apart my "Frankenstein's monster" 32 partition mdraid LVM mess.   If I can switch kernels and reboot before work and life get back in the way, I'll follow  up..

I think I might have to give myself the action to run this to ground next week on the other server.   Without a doubt the mdraid lock improvement is worth taking forward.   I either have to find my error or point a finger as my raw hero numbers got worse.   I tend to see one socket outrun another -  the way HPE allocates the nvme drives to pcie root complexes  is not how I'd like to do it so the drives are unbalanced on the PCIe root complexes (drives are in 4 different root complexes on socket 0 and 3 on socket 1, so one would think socket0 will always be faster for hero runs  (an NPS4 numa mapping is the best way to show it:  
[root@gremlin04 hornet05]# cat *nps4
#filename=/dev/nvme0n1 0
#filename=/dev/nvme1n1 0
#filename=/dev/nvme2n1 1
#filename=/dev/nvme3n1 1
#filename=/dev/nvme4n1 2
#filename=/dev/nvme5n1 2
#filename=/dev/nvme6n1 2
#filename=/dev/nvme7n1 2
#filename=/dev/nvme8n1 3
#filename=/dev/nvme9n1 3
#filename=/dev/nvme10n1 3
#filename=/dev/nvme11n1 3
#filename=/dev/nvme12n1 4
#filename=/dev/nvme13n1 4
#filename=/dev/nvme14n1 4
#filename=/dev/nvme15n1 4
#filename=/dev/nvme17n1 5
#filename=/dev/nvme18n1 5
#filename=/dev/nvme19n1 5
#filename=/dev/nvme20n1 5
#filename=/dev/nvme21n1 6
#filename=/dev/nvme22n1 6
#filename=/dev/nvme23n1 6
#filename=/dev/nvme24n1 6


fio fiojim.hpdl385.nps1
socket0: (g=0): rw=randread, bs=(R) 4096B-4096B, (W) 4096B-4096B, (T) 4096B-4096B, ioengine=libaio, iodepth=128
...
socket1: (g=1): rw=randread, bs=(R) 4096B-4096B, (W) 4096B-4096B, (T) 4096B-4096B, ioengine=libaio, iodepth=128
...
socket0-md: (g=2): rw=randread, bs=(R) 4096B-4096B, (W) 4096B-4096B, (T) 4096B-4096B, ioengine=libaio, iodepth=128
...
socket1-md: (g=3): rw=randread, bs=(R) 4096B-4096B, (W) 4096B-4096B, (T) 4096B-4096B, ioengine=libaio, iodepth=128
...
fio-3.26
Starting 256 processes
Jobs: 128 (f=128): [_(128),r(128)][1.5%][r=42.8GiB/s][r=11.2M IOPS][eta 10h:40m:00s]        
socket0: (groupid=0, jobs=64): err= 0: pid=522428: Thu Aug  5 19:33:05 2021
  read: IOPS=13.2M, BW=50.2GiB/s (53.9GB/s)(14.7TiB/300005msec)
    slat (nsec): min=1312, max=8308.1k, avg=2206.72, stdev=1505.92
    clat (usec): min=14, max=42033, avg=619.56, stdev=671.45
     lat (usec): min=19, max=42045, avg=621.83, stdev=671.46
    clat percentiles (usec):
     |  1.00th=[  113],  5.00th=[  149], 10.00th=[  180], 20.00th=[  229],
     | 30.00th=[  273], 40.00th=[  310], 50.00th=[  351], 60.00th=[  408],
     | 70.00th=[  578], 80.00th=[  938], 90.00th=[ 1467], 95.00th=[ 1909],
     | 99.00th=[ 3163], 99.50th=[ 4178], 99.90th=[ 5800], 99.95th=[ 6390],
     | 99.99th=[ 8455]
   bw (  MiB/s): min=28741, max=61365, per=18.56%, avg=51489.80, stdev=82.09, samples=38016
   iops        : min=7357916, max=15709528, avg=13181362.22, stdev=21013.83, samples=38016
  lat (usec)   : 20=0.01%, 50=0.02%, 100=0.42%, 250=24.52%, 500=42.21%
  lat (usec)   : 750=7.94%, 1000=6.34%
  lat (msec)   : 2=14.26%, 4=3.74%, 10=0.54%, 20=0.01%, 50=0.01%
  cpu          : usr=14.58%, sys=47.48%, ctx=291912925, majf=0, minf=10492
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
     issued rwts: total=3949519687,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128
socket1: (groupid=1, jobs=64): err= 0: pid=522492: Thu Aug  5 19:33:05 2021
  read: IOPS=13.6M, BW=51.8GiB/s (55.7GB/s)(15.2TiB/300004msec)
    slat (nsec): min=1323, max=4335.7k, avg=2242.27, stdev=1608.25
    clat (usec): min=14, max=41341, avg=600.15, stdev=726.62
     lat (usec): min=20, max=41358, avg=602.46, stdev=726.64
    clat percentiles (usec):
     |  1.00th=[  115],  5.00th=[  151], 10.00th=[  184], 20.00th=[  231],
     | 30.00th=[  269], 40.00th=[  306], 50.00th=[  347], 60.00th=[  400],
     | 70.00th=[  506], 80.00th=[  799], 90.00th=[ 1303], 95.00th=[ 1909],
     | 99.00th=[ 3589], 99.50th=[ 4424], 99.90th=[ 7111], 99.95th=[ 7767],
     | 99.99th=[10290]
   bw (  MiB/s): min=28663, max=71847, per=21.11%, avg=53145.09, stdev=111.29, samples=38016
   iops        : min=7337860, max=18392866, avg=13605117.00, stdev=28491.19, samples=38016
  lat (usec)   : 20=0.01%, 50=0.02%, 100=0.36%, 250=24.52%, 500=44.77%
  lat (usec)   : 750=8.90%, 1000=6.37%
  lat (msec)   : 2=10.52%, 4=3.87%, 10=0.66%, 20=0.01%, 50=0.01%
  cpu          : usr=14.86%, sys=49.40%, ctx=282634154, majf=0, minf=10276
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
     issued rwts: total=4076360454,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128
socket0-md: (groupid=2, jobs=64): err= 0: pid=524061: Thu Aug  5 19:33:05 2021
  read: IOPS=5332k, BW=20.3GiB/s (21.8GB/s)(6102GiB/300002msec)
    slat (nsec): min=1633, max=17043k, avg=11123.38, stdev=8694.61
    clat (usec): min=186, max=18705, avg=1524.87, stdev=115.29
     lat (usec): min=200, max=18743, avg=1536.08, stdev=115.90
    clat percentiles (usec):
     |  1.00th=[ 1270],  5.00th=[ 1336], 10.00th=[ 1369], 20.00th=[ 1418],
     | 30.00th=[ 1467], 40.00th=[ 1500], 50.00th=[ 1532], 60.00th=[ 1549],
     | 70.00th=[ 1582], 80.00th=[ 1631], 90.00th=[ 1680], 95.00th=[ 1713],
     | 99.00th=[ 1795], 99.50th=[ 1811], 99.90th=[ 1893], 99.95th=[ 1926],
     | 99.99th=[ 2089]
   bw (  MiB/s): min=19030, max=21969, per=100.00%, avg=20843.43, stdev= 5.35, samples=38272
   iops        : min=4871687, max=5624289, avg=5335900.01, stdev=1370.43, samples=38272
  lat (usec)   : 250=0.01%, 500=0.01%, 750=0.01%, 1000=0.01%
  lat (msec)   : 2=99.97%, 4=0.02%, 10=0.01%, 20=0.01%
  cpu          : usr=5.56%, sys=77.91%, ctx=8118, majf=0, minf=9018
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
    issued rwts: total=1599503201,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128
socket1-md: (groupid=3, jobs=64): err= 0: pid=524125: Thu Aug  5 19:33:05 2021
  read: IOPS=5892k, BW=22.5GiB/s (24.1GB/s)(6743GiB/300002msec)
    slat (nsec): min=1663, max=1274.1k, avg=9896.09, stdev=7939.50
    clat (usec): min=236, max=11102, avg=1379.86, stdev=148.64
     lat (usec): min=239, max=11110, avg=1389.84, stdev=149.54
    clat percentiles (usec):
     |  1.00th=[ 1106],  5.00th=[ 1172], 10.00th=[ 1205], 20.00th=[ 1254],
     | 30.00th=[ 1287], 40.00th=[ 1336], 50.00th=[ 1369], 60.00th=[ 1401],
     | 70.00th=[ 1434], 80.00th=[ 1500], 90.00th=[ 1582], 95.00th=[ 1663],
     | 99.00th=[ 1811], 99.50th=[ 1860], 99.90th=[ 1942], 99.95th=[ 1958],
     | 99.99th=[ 2040]
   bw (  MiB/s): min=20982, max=24535, per=-82.15%, avg=23034.61, stdev=15.46, samples=38272
   iops        : min=5371404, max=6281119, avg=5896843.14, stdev=3958.21, samples=38272
  lat (usec)   : 250=0.01%, 500=0.01%, 750=0.01%, 1000=0.01%
  lat (msec)   : 2=99.97%, 4=0.02%, 10=0.01%, 20=0.01%
  cpu          : usr=6.55%, sys=74.98%, ctx=9833, majf=0, minf=8956
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
     issued rwts: total=1767618924,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128

Run status group 0 (all jobs):
   READ: bw=50.2GiB/s (53.9GB/s), 50.2GiB/s-50.2GiB/s (53.9GB/s-53.9GB/s), io=14.7TiB (16.2TB), run=300005-300005msec

Run status group 1 (all jobs):
   READ: bw=51.8GiB/s (55.7GB/s), 51.8GiB/s-51.8GiB/s (55.7GB/s-55.7GB/s), io=15.2TiB (16.7TB), run=300004-300004msec

Run status group 2 (all jobs):
   READ: bw=20.3GiB/s (21.8GB/s), 20.3GiB/s-20.3GiB/s (21.8GB/s-21.8GB/s), io=6102GiB (6552GB), run=300002-300002msec

Run status group 3 (all jobs):
   READ: bw=22.5GiB/s (24.1GB/s), 22.5GiB/s-22.5GiB/s (24.1GB/s-24.1GB/s), io=6743GiB (7240GB), run=300002-300002msec

Disk stats (read/write):
  nvme0n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme1n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme2n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme3n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme4n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme5n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme6n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme7n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme8n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme9n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme10n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme11n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme12n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme13n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme14n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme15n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme17n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme18n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme19n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme20n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme21n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme22n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme23n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme24n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  md0: ios=1599378656/0, merge=0/0, ticks=391992721/0, in_queue=391992721, util=100.00%
  md1: ios=1767484212/0, merge=0/0, ticks=427666887/0, in_queue=427666887, util=100.00%

From: Gal Ofri <redacted> 
Sent: Wednesday, July 28, 2021 5:43 AM
To: Finlayson, James M CIV (USA) <redacted>; 'linux-raid@vger.kernel.org' <redacted>
Subject: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

All active links contained in this email were disabled. Please verify the identity of the sender, and confirm the authenticity of all links contained within the message prior to copying and pasting the address to a Web browser. 
________________________________________

A recent commit raised the limit on raid5/6 read iops.
It's available in 5.14.
See Caution-https://github.com/torvalds/linux/commit/97ae27252f4962d0fcc38ee1d9f913d817a2024e < Caution-https://github.com/torvalds/linux/commit/97ae27252f4962d0fcc38ee1d9f913d817a2024e > 
commit 97ae27252f4962d0fcc38ee1d9f913d817a2024e
Author: Gal Ofri [off-list ref]
Date:   Mon Jun 7 14:07:03 2021 +0300
    md/raid5: avoid device_lock in read_one_chunk()

Please do share if you reach more iops in your env than described in the commit.

Cheers,
Gal, 
Volumez (formerly storing.io)

RE: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

From: Finlayson, James M CIV (USA) <hidden>
Date: 2021-08-05 20:50:58

As far as the slower hero numbers - false alarm on my part - rebooted with 4.18 RHEL 8.4 kernel
Socket0 hero - 13.2M IOPS, Socket1 hero 13.7M IOPS.   I have to figure out the differences either between my drives or my server.  Chances are, slot for slot I have PCIe cards that are in different slots between the two servers if I had to guess....

As a major flag though - with mdraid volumes I created under the 5.14rc3 kernel, I lock the system up solid when I try to access them under 4.18.....I'm not an expert on forcing NMI's and getting the stack traces, so I might have to leave that to others.....After two lockups, I returned to the 5.14 kernel.   If I need to run something - you have seen the config I have - I'm willing.   

I'm willing to push as hard as I can and to run anything that can help as long as it isn't urgent - I have a day job and have some constraints as a civil servant, however, I have the researcher, push push push mindset.   I want to really encourage the community to push as hard as possible on protected IOPS and I'm willing to help however I can....In my interactions with the processor and server OEMs - I'm encouraging them to get the Linux leaders in I/O development, the biggest baddest Server/SSD combinations they have early in the development.   I know they won't listen to me but I'm trying to help.

For those of you on Rome server, get with your server provider.   There are some things in the BIOS that can be tweaked for I/O.   


-----Original Message-----
From: Finlayson, James M CIV (USA) <redacted> 
Sent: Thursday, August 5, 2021 3:52 PM
To: 'linux-raid@vger.kernel.org' <redacted>
Cc: 'Gal Ofri' <redacted>; Finlayson, James M CIV (USA) <redacted>
Subject: RE: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

Sorry - again..I sent HTML instead of plain text

Resend - mailing list bounce
All,
Sorry for the delay - both work and life got into the way.   Here is some feedback:

BLUF upfront with 5.14rc3 kernel that our SA built - md0 a 10+1+1 RAID5 - 5.332 M IOPS 20.3GiB/s, md1 a 10+1+1 RAID5, 5.892M IOPS 22.5GiB/s  - best hero numbers I've ever seen on mdraid  RAID5 IOPS.   I think the kernel patch is good.  Prior was  socket0 1.263M IOPS 4934MiB/s, socket1 1.071M IOSP, 4183MiB/s....   I'm willing to help push this as hard as we can until we hit a bottleneck outside of our control.   

I need to verify the RAW IOPS - admittedly this is a different server and I didn't do any regression testing before the kernel, but my raw were  socket0: 13.2M IOPS and socket1  13.5M IOPS.   Prior was socket0 16.0M IOPS and socket1 13.5M IOPS.   - admittedly there appears to a regression in the socket0 "hero run" but what I don't know that since this is a different server, I don't know if I have a configuration management issue in my zealousness to test this patch or whether we have a regression.   I was so excited to have the attention of kernel developers that needed my help that I borrowed another system, because I didn't want to tear apart my "Frankenstein's monster" 32 partition mdraid LVM mess.   If I can switch kernels and reboot before work and life get back in the way, I'll follow  up..

I think I might have to give myself the action to run this to ground next week on the other server.   Without a doubt the mdraid lock improvement is worth taking forward.   I either have to find my error or point a finger as my raw hero numbers got worse.   I tend to see one socket outrun another -  the way HPE allocates the nvme drives to pcie root complexes  is not how I'd like to do it so the drives are unbalanced on the PCIe root complexes (drives are in 4 different root complexes on socket 0 and 3 on socket 1, so one would think socket0 will always be faster for hero runs  (an NPS4 numa mapping is the best way to show it:
[root@gremlin04 hornet05]# cat *nps4
#filename=/dev/nvme0n1 0
#filename=/dev/nvme1n1 0
#filename=/dev/nvme2n1 1
#filename=/dev/nvme3n1 1
#filename=/dev/nvme4n1 2
#filename=/dev/nvme5n1 2
#filename=/dev/nvme6n1 2
#filename=/dev/nvme7n1 2
#filename=/dev/nvme8n1 3
#filename=/dev/nvme9n1 3
#filename=/dev/nvme10n1 3
#filename=/dev/nvme11n1 3
#filename=/dev/nvme12n1 4
#filename=/dev/nvme13n1 4
#filename=/dev/nvme14n1 4
#filename=/dev/nvme15n1 4
#filename=/dev/nvme17n1 5
#filename=/dev/nvme18n1 5
#filename=/dev/nvme19n1 5
#filename=/dev/nvme20n1 5
#filename=/dev/nvme21n1 6
#filename=/dev/nvme22n1 6
#filename=/dev/nvme23n1 6
#filename=/dev/nvme24n1 6


fio fiojim.hpdl385.nps1
socket0: (g=0): rw=randread, bs=(R) 4096B-4096B, (W) 4096B-4096B, (T) 4096B-4096B, ioengine=libaio, iodepth=128 ...
socket1: (g=1): rw=randread, bs=(R) 4096B-4096B, (W) 4096B-4096B, (T) 4096B-4096B, ioengine=libaio, iodepth=128 ...
socket0-md: (g=2): rw=randread, bs=(R) 4096B-4096B, (W) 4096B-4096B, (T) 4096B-4096B, ioengine=libaio, iodepth=128 ...
socket1-md: (g=3): rw=randread, bs=(R) 4096B-4096B, (W) 4096B-4096B, (T) 4096B-4096B, ioengine=libaio, iodepth=128 ...
fio-3.26
Starting 256 processes
Jobs: 128 (f=128): [_(128),r(128)][1.5%][r=42.8GiB/s][r=11.2M IOPS][eta 10h:40m:00s]
socket0: (groupid=0, jobs=64): err= 0: pid=522428: Thu Aug  5 19:33:05 2021
  read: IOPS=13.2M, BW=50.2GiB/s (53.9GB/s)(14.7TiB/300005msec)
    slat (nsec): min=1312, max=8308.1k, avg=2206.72, stdev=1505.92
    clat (usec): min=14, max=42033, avg=619.56, stdev=671.45
     lat (usec): min=19, max=42045, avg=621.83, stdev=671.46
    clat percentiles (usec):
     |  1.00th=[  113],  5.00th=[  149], 10.00th=[  180], 20.00th=[  229],
     | 30.00th=[  273], 40.00th=[  310], 50.00th=[  351], 60.00th=[  408],
     | 70.00th=[  578], 80.00th=[  938], 90.00th=[ 1467], 95.00th=[ 1909],
     | 99.00th=[ 3163], 99.50th=[ 4178], 99.90th=[ 5800], 99.95th=[ 6390],
     | 99.99th=[ 8455]
   bw (  MiB/s): min=28741, max=61365, per=18.56%, avg=51489.80, stdev=82.09, samples=38016
   iops        : min=7357916, max=15709528, avg=13181362.22, stdev=21013.83, samples=38016
  lat (usec)   : 20=0.01%, 50=0.02%, 100=0.42%, 250=24.52%, 500=42.21%
  lat (usec)   : 750=7.94%, 1000=6.34%
  lat (msec)   : 2=14.26%, 4=3.74%, 10=0.54%, 20=0.01%, 50=0.01%
  cpu          : usr=14.58%, sys=47.48%, ctx=291912925, majf=0, minf=10492
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
     issued rwts: total=3949519687,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128
socket1: (groupid=1, jobs=64): err= 0: pid=522492: Thu Aug  5 19:33:05 2021
  read: IOPS=13.6M, BW=51.8GiB/s (55.7GB/s)(15.2TiB/300004msec)
    slat (nsec): min=1323, max=4335.7k, avg=2242.27, stdev=1608.25
    clat (usec): min=14, max=41341, avg=600.15, stdev=726.62
     lat (usec): min=20, max=41358, avg=602.46, stdev=726.64
    clat percentiles (usec):
     |  1.00th=[  115],  5.00th=[  151], 10.00th=[  184], 20.00th=[  231],
     | 30.00th=[  269], 40.00th=[  306], 50.00th=[  347], 60.00th=[  400],
     | 70.00th=[  506], 80.00th=[  799], 90.00th=[ 1303], 95.00th=[ 1909],
     | 99.00th=[ 3589], 99.50th=[ 4424], 99.90th=[ 7111], 99.95th=[ 7767],
     | 99.99th=[10290]
   bw (  MiB/s): min=28663, max=71847, per=21.11%, avg=53145.09, stdev=111.29, samples=38016
   iops        : min=7337860, max=18392866, avg=13605117.00, stdev=28491.19, samples=38016
  lat (usec)   : 20=0.01%, 50=0.02%, 100=0.36%, 250=24.52%, 500=44.77%
  lat (usec)   : 750=8.90%, 1000=6.37%
  lat (msec)   : 2=10.52%, 4=3.87%, 10=0.66%, 20=0.01%, 50=0.01%
  cpu          : usr=14.86%, sys=49.40%, ctx=282634154, majf=0, minf=10276
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
     issued rwts: total=4076360454,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128
socket0-md: (groupid=2, jobs=64): err= 0: pid=524061: Thu Aug  5 19:33:05 2021
  read: IOPS=5332k, BW=20.3GiB/s (21.8GB/s)(6102GiB/300002msec)
    slat (nsec): min=1633, max=17043k, avg=11123.38, stdev=8694.61
    clat (usec): min=186, max=18705, avg=1524.87, stdev=115.29
     lat (usec): min=200, max=18743, avg=1536.08, stdev=115.90
    clat percentiles (usec):
     |  1.00th=[ 1270],  5.00th=[ 1336], 10.00th=[ 1369], 20.00th=[ 1418],
     | 30.00th=[ 1467], 40.00th=[ 1500], 50.00th=[ 1532], 60.00th=[ 1549],
     | 70.00th=[ 1582], 80.00th=[ 1631], 90.00th=[ 1680], 95.00th=[ 1713],
     | 99.00th=[ 1795], 99.50th=[ 1811], 99.90th=[ 1893], 99.95th=[ 1926],
     | 99.99th=[ 2089]
   bw (  MiB/s): min=19030, max=21969, per=100.00%, avg=20843.43, stdev= 5.35, samples=38272
   iops        : min=4871687, max=5624289, avg=5335900.01, stdev=1370.43, samples=38272
  lat (usec)   : 250=0.01%, 500=0.01%, 750=0.01%, 1000=0.01%
  lat (msec)   : 2=99.97%, 4=0.02%, 10=0.01%, 20=0.01%
  cpu          : usr=5.56%, sys=77.91%, ctx=8118, majf=0, minf=9018
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
    issued rwts: total=1599503201,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128
socket1-md: (groupid=3, jobs=64): err= 0: pid=524125: Thu Aug  5 19:33:05 2021
  read: IOPS=5892k, BW=22.5GiB/s (24.1GB/s)(6743GiB/300002msec)
    slat (nsec): min=1663, max=1274.1k, avg=9896.09, stdev=7939.50
    clat (usec): min=236, max=11102, avg=1379.86, stdev=148.64
     lat (usec): min=239, max=11110, avg=1389.84, stdev=149.54
    clat percentiles (usec):
     |  1.00th=[ 1106],  5.00th=[ 1172], 10.00th=[ 1205], 20.00th=[ 1254],
     | 30.00th=[ 1287], 40.00th=[ 1336], 50.00th=[ 1369], 60.00th=[ 1401],
     | 70.00th=[ 1434], 80.00th=[ 1500], 90.00th=[ 1582], 95.00th=[ 1663],
     | 99.00th=[ 1811], 99.50th=[ 1860], 99.90th=[ 1942], 99.95th=[ 1958],
     | 99.99th=[ 2040]
   bw (  MiB/s): min=20982, max=24535, per=-82.15%, avg=23034.61, stdev=15.46, samples=38272
   iops        : min=5371404, max=6281119, avg=5896843.14, stdev=3958.21, samples=38272
  lat (usec)   : 250=0.01%, 500=0.01%, 750=0.01%, 1000=0.01%
  lat (msec)   : 2=99.97%, 4=0.02%, 10=0.01%, 20=0.01%
  cpu          : usr=6.55%, sys=74.98%, ctx=9833, majf=0, minf=8956
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
     issued rwts: total=1767618924,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128

Run status group 0 (all jobs):
   READ: bw=50.2GiB/s (53.9GB/s), 50.2GiB/s-50.2GiB/s (53.9GB/s-53.9GB/s), io=14.7TiB (16.2TB), run=300005-300005msec

Run status group 1 (all jobs):
   READ: bw=51.8GiB/s (55.7GB/s), 51.8GiB/s-51.8GiB/s (55.7GB/s-55.7GB/s), io=15.2TiB (16.7TB), run=300004-300004msec

Run status group 2 (all jobs):
   READ: bw=20.3GiB/s (21.8GB/s), 20.3GiB/s-20.3GiB/s (21.8GB/s-21.8GB/s), io=6102GiB (6552GB), run=300002-300002msec

Run status group 3 (all jobs):
   READ: bw=22.5GiB/s (24.1GB/s), 22.5GiB/s-22.5GiB/s (24.1GB/s-24.1GB/s), io=6743GiB (7240GB), run=300002-300002msec

Disk stats (read/write):
  nvme0n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme1n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme2n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme3n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme4n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme5n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme6n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme7n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme8n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme9n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme10n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme11n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme12n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme13n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme14n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme15n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme17n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme18n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme19n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme20n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme21n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme22n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme23n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme24n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  md0: ios=1599378656/0, merge=0/0, ticks=391992721/0, in_queue=391992721, util=100.00%
  md1: ios=1767484212/0, merge=0/0, ticks=427666887/0, in_queue=427666887, util=100.00%

From: Gal Ofri <redacted>
Sent: Wednesday, July 28, 2021 5:43 AM
To: Finlayson, James M CIV (USA) <redacted>; 'linux-raid@vger.kernel.org' <redacted>
Subject: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

All active links contained in this email were disabled. Please verify the identity of the sender, and confirm the authenticity of all links contained within the message prior to copying and pasting the address to a Web browser. 
________________________________________

A recent commit raised the limit on raid5/6 read iops.
It's available in 5.14.
See Caution-https://github.com/torvalds/linux/commit/97ae27252f4962d0fcc38ee1d9f913d817a2024e < Caution-https://github.com/torvalds/linux/commit/97ae27252f4962d0fcc38ee1d9f913d817a2024e > commit 97ae27252f4962d0fcc38ee1d9f913d817a2024e
Author: Gal Ofri [off-list ref]
Date:   Mon Jun 7 14:07:03 2021 +0300
    md/raid5: avoid device_lock in read_one_chunk()

Please do share if you reach more iops in your env than described in the commit.

Cheers,
Gal,
Volumez (formerly storing.io)

RE: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

From: Finlayson, James M CIV (USA) <hidden>
Date: 2021-08-05 21:11:21

Final spray from me for a few days.

In my strict numa adherence with mdraid, I see lots of variability between reboots/assembles.    Sometimes md0 wins, sometimes md1 wins, and in my earlier runs md0 and md1 are notionally balanced.   I change nothing but see this variance.   I just cranked up a week long extended run of these 10+1+1s under the 5.14rc3 kernel and right now   md0 is doing 5M IOPS and md1 6.3M - still totaling in the low 11's but quite the disparity.   Am I missing a tuning knob?   I shared everything I do and know in the earlier thread.   I just want to point this out while I have attention.   I've have seen this behavior over and over again.    The more I know about AMD, the more I think I can't depend on the HPC profile provided and I need to take full control of the BIOS.   I know there is a ton of power management going on under the covers, so maybe that is what I'm experiencing.    The more I type, the more I think I don't see it on Intel, but I don't have a modern Intel machine with modern SSDs to test.   I'll accept that there is nothing inherent in mdraid or the kernel to cause this and put my attention to the BIOS if the experts can confirm....

-----Original Message-----
From: Finlayson, James M CIV (USA) <redacted> 
Sent: Thursday, August 5, 2021 4:50 PM
To: 'linux-raid@vger.kernel.org' <redacted>
Cc: 'Gal Ofri' <redacted>; Finlayson, James M CIV (USA) <redacted>
Subject: RE: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

As far as the slower hero numbers - false alarm on my part - rebooted with 4.18 RHEL 8.4 kernel
Socket0 hero - 13.2M IOPS, Socket1 hero 13.7M IOPS.   I have to figure out the differences either between my drives or my server.  Chances are, slot for slot I have PCIe cards that are in different slots between the two servers if I had to guess....

As a major flag though - with mdraid volumes I created under the 5.14rc3 kernel, I lock the system up solid when I try to access them under 4.18.....I'm not an expert on forcing NMI's and getting the stack traces, so I might have to leave that to others.....After two lockups, I returned to the 5.14 kernel.   If I need to run something - you have seen the config I have - I'm willing.   

I'm willing to push as hard as I can and to run anything that can help as long as it isn't urgent - I have a day job and have some constraints as a civil servant, however, I have the researcher, push push push mindset.   I want to really encourage the community to push as hard as possible on protected IOPS and I'm willing to help however I can....In my interactions with the processor and server OEMs - I'm encouraging them to get the Linux leaders in I/O development, the biggest baddest Server/SSD combinations they have early in the development.   I know they won't listen to me but I'm trying to help.

For those of you on Rome server, get with your server provider.   There are some things in the BIOS that can be tweaked for I/O.   


-----Original Message-----
From: Finlayson, James M CIV (USA) <redacted> 
Sent: Thursday, August 5, 2021 3:52 PM
To: 'linux-raid@vger.kernel.org' <redacted>
Cc: 'Gal Ofri' <redacted>; Finlayson, James M CIV (USA) <redacted>
Subject: RE: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

Sorry - again..I sent HTML instead of plain text

Resend - mailing list bounce
All,
Sorry for the delay - both work and life got into the way.   Here is some feedback:

BLUF upfront with 5.14rc3 kernel that our SA built - md0 a 10+1+1 RAID5 - 5.332 M IOPS 20.3GiB/s, md1 a 10+1+1 RAID5, 5.892M IOPS 22.5GiB/s  - best hero numbers I've ever seen on mdraid  RAID5 IOPS.   I think the kernel patch is good.  Prior was  socket0 1.263M IOPS 4934MiB/s, socket1 1.071M IOSP, 4183MiB/s....   I'm willing to help push this as hard as we can until we hit a bottleneck outside of our control.   

I need to verify the RAW IOPS - admittedly this is a different server and I didn't do any regression testing before the kernel, but my raw were  socket0: 13.2M IOPS and socket1  13.5M IOPS.   Prior was socket0 16.0M IOPS and socket1 13.5M IOPS.   - admittedly there appears to a regression in the socket0 "hero run" but what I don't know that since this is a different server, I don't know if I have a configuration management issue in my zealousness to test this patch or whether we have a regression.   I was so excited to have the attention of kernel developers that needed my help that I borrowed another system, because I didn't want to tear apart my "Frankenstein's monster" 32 partition mdraid LVM mess.   If I can switch kernels and reboot before work and life get back in the way, I'll follow  up..

I think I might have to give myself the action to run this to ground next week on the other server.   Without a doubt the mdraid lock improvement is worth taking forward.   I either have to find my error or point a finger as my raw hero numbers got worse.   I tend to see one socket outrun another -  the way HPE allocates the nvme drives to pcie root complexes  is not how I'd like to do it so the drives are unbalanced on the PCIe root complexes (drives are in 4 different root complexes on socket 0 and 3 on socket 1, so one would think socket0 will always be faster for hero runs  (an NPS4 numa mapping is the best way to show it:
[root@gremlin04 hornet05]# cat *nps4
#filename=/dev/nvme0n1 0
#filename=/dev/nvme1n1 0
#filename=/dev/nvme2n1 1
#filename=/dev/nvme3n1 1
#filename=/dev/nvme4n1 2
#filename=/dev/nvme5n1 2
#filename=/dev/nvme6n1 2
#filename=/dev/nvme7n1 2
#filename=/dev/nvme8n1 3
#filename=/dev/nvme9n1 3
#filename=/dev/nvme10n1 3
#filename=/dev/nvme11n1 3
#filename=/dev/nvme12n1 4
#filename=/dev/nvme13n1 4
#filename=/dev/nvme14n1 4
#filename=/dev/nvme15n1 4
#filename=/dev/nvme17n1 5
#filename=/dev/nvme18n1 5
#filename=/dev/nvme19n1 5
#filename=/dev/nvme20n1 5
#filename=/dev/nvme21n1 6
#filename=/dev/nvme22n1 6
#filename=/dev/nvme23n1 6
#filename=/dev/nvme24n1 6


fio fiojim.hpdl385.nps1
socket0: (g=0): rw=randread, bs=(R) 4096B-4096B, (W) 4096B-4096B, (T) 4096B-4096B, ioengine=libaio, iodepth=128 ...
socket1: (g=1): rw=randread, bs=(R) 4096B-4096B, (W) 4096B-4096B, (T) 4096B-4096B, ioengine=libaio, iodepth=128 ...
socket0-md: (g=2): rw=randread, bs=(R) 4096B-4096B, (W) 4096B-4096B, (T) 4096B-4096B, ioengine=libaio, iodepth=128 ...
socket1-md: (g=3): rw=randread, bs=(R) 4096B-4096B, (W) 4096B-4096B, (T) 4096B-4096B, ioengine=libaio, iodepth=128 ...
fio-3.26
Starting 256 processes
Jobs: 128 (f=128): [_(128),r(128)][1.5%][r=42.8GiB/s][r=11.2M IOPS][eta 10h:40m:00s]
socket0: (groupid=0, jobs=64): err= 0: pid=522428: Thu Aug  5 19:33:05 2021
  read: IOPS=13.2M, BW=50.2GiB/s (53.9GB/s)(14.7TiB/300005msec)
    slat (nsec): min=1312, max=8308.1k, avg=2206.72, stdev=1505.92
    clat (usec): min=14, max=42033, avg=619.56, stdev=671.45
     lat (usec): min=19, max=42045, avg=621.83, stdev=671.46
    clat percentiles (usec):
     |  1.00th=[  113],  5.00th=[  149], 10.00th=[  180], 20.00th=[  229],
     | 30.00th=[  273], 40.00th=[  310], 50.00th=[  351], 60.00th=[  408],
     | 70.00th=[  578], 80.00th=[  938], 90.00th=[ 1467], 95.00th=[ 1909],
     | 99.00th=[ 3163], 99.50th=[ 4178], 99.90th=[ 5800], 99.95th=[ 6390],
     | 99.99th=[ 8455]
   bw (  MiB/s): min=28741, max=61365, per=18.56%, avg=51489.80, stdev=82.09, samples=38016
   iops        : min=7357916, max=15709528, avg=13181362.22, stdev=21013.83, samples=38016
  lat (usec)   : 20=0.01%, 50=0.02%, 100=0.42%, 250=24.52%, 500=42.21%
  lat (usec)   : 750=7.94%, 1000=6.34%
  lat (msec)   : 2=14.26%, 4=3.74%, 10=0.54%, 20=0.01%, 50=0.01%
  cpu          : usr=14.58%, sys=47.48%, ctx=291912925, majf=0, minf=10492
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
     issued rwts: total=3949519687,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128
socket1: (groupid=1, jobs=64): err= 0: pid=522492: Thu Aug  5 19:33:05 2021
  read: IOPS=13.6M, BW=51.8GiB/s (55.7GB/s)(15.2TiB/300004msec)
    slat (nsec): min=1323, max=4335.7k, avg=2242.27, stdev=1608.25
    clat (usec): min=14, max=41341, avg=600.15, stdev=726.62
     lat (usec): min=20, max=41358, avg=602.46, stdev=726.64
    clat percentiles (usec):
     |  1.00th=[  115],  5.00th=[  151], 10.00th=[  184], 20.00th=[  231],
     | 30.00th=[  269], 40.00th=[  306], 50.00th=[  347], 60.00th=[  400],
     | 70.00th=[  506], 80.00th=[  799], 90.00th=[ 1303], 95.00th=[ 1909],
     | 99.00th=[ 3589], 99.50th=[ 4424], 99.90th=[ 7111], 99.95th=[ 7767],
     | 99.99th=[10290]
   bw (  MiB/s): min=28663, max=71847, per=21.11%, avg=53145.09, stdev=111.29, samples=38016
   iops        : min=7337860, max=18392866, avg=13605117.00, stdev=28491.19, samples=38016
  lat (usec)   : 20=0.01%, 50=0.02%, 100=0.36%, 250=24.52%, 500=44.77%
  lat (usec)   : 750=8.90%, 1000=6.37%
  lat (msec)   : 2=10.52%, 4=3.87%, 10=0.66%, 20=0.01%, 50=0.01%
  cpu          : usr=14.86%, sys=49.40%, ctx=282634154, majf=0, minf=10276
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
     issued rwts: total=4076360454,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128
socket0-md: (groupid=2, jobs=64): err= 0: pid=524061: Thu Aug  5 19:33:05 2021
  read: IOPS=5332k, BW=20.3GiB/s (21.8GB/s)(6102GiB/300002msec)
    slat (nsec): min=1633, max=17043k, avg=11123.38, stdev=8694.61
    clat (usec): min=186, max=18705, avg=1524.87, stdev=115.29
     lat (usec): min=200, max=18743, avg=1536.08, stdev=115.90
    clat percentiles (usec):
     |  1.00th=[ 1270],  5.00th=[ 1336], 10.00th=[ 1369], 20.00th=[ 1418],
     | 30.00th=[ 1467], 40.00th=[ 1500], 50.00th=[ 1532], 60.00th=[ 1549],
     | 70.00th=[ 1582], 80.00th=[ 1631], 90.00th=[ 1680], 95.00th=[ 1713],
     | 99.00th=[ 1795], 99.50th=[ 1811], 99.90th=[ 1893], 99.95th=[ 1926],
     | 99.99th=[ 2089]
   bw (  MiB/s): min=19030, max=21969, per=100.00%, avg=20843.43, stdev= 5.35, samples=38272
   iops        : min=4871687, max=5624289, avg=5335900.01, stdev=1370.43, samples=38272
  lat (usec)   : 250=0.01%, 500=0.01%, 750=0.01%, 1000=0.01%
  lat (msec)   : 2=99.97%, 4=0.02%, 10=0.01%, 20=0.01%
  cpu          : usr=5.56%, sys=77.91%, ctx=8118, majf=0, minf=9018
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
    issued rwts: total=1599503201,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128
socket1-md: (groupid=3, jobs=64): err= 0: pid=524125: Thu Aug  5 19:33:05 2021
  read: IOPS=5892k, BW=22.5GiB/s (24.1GB/s)(6743GiB/300002msec)
    slat (nsec): min=1663, max=1274.1k, avg=9896.09, stdev=7939.50
    clat (usec): min=236, max=11102, avg=1379.86, stdev=148.64
     lat (usec): min=239, max=11110, avg=1389.84, stdev=149.54
    clat percentiles (usec):
     |  1.00th=[ 1106],  5.00th=[ 1172], 10.00th=[ 1205], 20.00th=[ 1254],
     | 30.00th=[ 1287], 40.00th=[ 1336], 50.00th=[ 1369], 60.00th=[ 1401],
     | 70.00th=[ 1434], 80.00th=[ 1500], 90.00th=[ 1582], 95.00th=[ 1663],
     | 99.00th=[ 1811], 99.50th=[ 1860], 99.90th=[ 1942], 99.95th=[ 1958],
     | 99.99th=[ 2040]
   bw (  MiB/s): min=20982, max=24535, per=-82.15%, avg=23034.61, stdev=15.46, samples=38272
   iops        : min=5371404, max=6281119, avg=5896843.14, stdev=3958.21, samples=38272
  lat (usec)   : 250=0.01%, 500=0.01%, 750=0.01%, 1000=0.01%
  lat (msec)   : 2=99.97%, 4=0.02%, 10=0.01%, 20=0.01%
  cpu          : usr=6.55%, sys=74.98%, ctx=9833, majf=0, minf=8956
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
     issued rwts: total=1767618924,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128

Run status group 0 (all jobs):
   READ: bw=50.2GiB/s (53.9GB/s), 50.2GiB/s-50.2GiB/s (53.9GB/s-53.9GB/s), io=14.7TiB (16.2TB), run=300005-300005msec

Run status group 1 (all jobs):
   READ: bw=51.8GiB/s (55.7GB/s), 51.8GiB/s-51.8GiB/s (55.7GB/s-55.7GB/s), io=15.2TiB (16.7TB), run=300004-300004msec

Run status group 2 (all jobs):
   READ: bw=20.3GiB/s (21.8GB/s), 20.3GiB/s-20.3GiB/s (21.8GB/s-21.8GB/s), io=6102GiB (6552GB), run=300002-300002msec

Run status group 3 (all jobs):
   READ: bw=22.5GiB/s (24.1GB/s), 22.5GiB/s-22.5GiB/s (24.1GB/s-24.1GB/s), io=6743GiB (7240GB), run=300002-300002msec

Disk stats (read/write):
  nvme0n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme1n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme2n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme3n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme4n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme5n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme6n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme7n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme8n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme9n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme10n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme11n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme12n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme13n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme14n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme15n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme17n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme18n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme19n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme20n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme21n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme22n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme23n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme24n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  md0: ios=1599378656/0, merge=0/0, ticks=391992721/0, in_queue=391992721, util=100.00%
  md1: ios=1767484212/0, merge=0/0, ticks=427666887/0, in_queue=427666887, util=100.00%

From: Gal Ofri <redacted>
Sent: Wednesday, July 28, 2021 5:43 AM
To: Finlayson, James M CIV (USA) <redacted>; 'linux-raid@vger.kernel.org' <redacted>
Subject: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

All active links contained in this email were disabled. Please verify the identity of the sender, and confirm the authenticity of all links contained within the message prior to copying and pasting the address to a Web browser. 
________________________________________

A recent commit raised the limit on raid5/6 read iops.
It's available in 5.14.
See Caution-https://github.com/torvalds/linux/commit/97ae27252f4962d0fcc38ee1d9f913d817a2024e < Caution-https://github.com/torvalds/linux/commit/97ae27252f4962d0fcc38ee1d9f913d817a2024e > commit 97ae27252f4962d0fcc38ee1d9f913d817a2024e
Author: Gal Ofri [off-list ref]
Date:   Mon Jun 7 14:07:03 2021 +0300
    md/raid5: avoid device_lock in read_one_chunk()

Please do share if you reach more iops in your env than described in the commit.

Cheers,
Gal,
Volumez (formerly storing.io)

Re: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

From: Gal Ofri <hidden>
Date: 2021-08-08 14:43:40

On Thu, 5 Aug 2021 21:10:40 +0000
"Finlayson, James M CIV (USA)" [off-list ref] wrote:
BLUF upfront with 5.14rc3 kernel that our SA built - md0 a 10+1+1 RAID5 - 5.332 M IOPS 20.3GiB/s, md1 a 10+1+1 RAID5, 5.892M IOPS 22.5GiB/s  - best hero numbers I've ever seen on mdraid  RAID5 IOPS.   I think the kernel patch is good.  Prior was  socket0 1.263M IOPS 4934MiB/s, socket1 1.071M IOSP, 4183MiB/s....   I'm willing to help push this as hard as we can until we hit a bottleneck outside of our control.   
That's great !
Thanks for sharing your results.
I'd appreciate if you could run a sequential-reads workload (128k/256k) so
that we get a better sense of the throughput potential here.
In my strict numa adherence with mdraid, I see lots of variability between reboots/assembles.    Sometimes md0 wins, sometimes md1 wins, and in my earlier runs md0 and md1 are notionally balanced.   I change nothing but see this variance.   I just cranked up a week long extended run of these 10+1+1s under the 5.14rc3 kernel and right now   md0 is doing 5M IOPS and md1 6.3M 
Given my humble experience with the code in question, I suspect that it is
not really optimized for numa awareness, so I find your findings quite
reasonable. I don't really have a good tip for that.

I'm focusing now on thin-provisioned logical volumes (lvm - it has a much
worse reads bottleneck actually), but we have plans for researching
md/raid5 again soon to improve write workloads.
I'll ping you when I have a patch that might be relevant.

Cheers,
Gal

RE: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

From: Finlayson, James M CIV (USA) <hidden>
Date: 2021-08-09 19:04:19

Sequential Performance:
BLUF, 1M sequential, direct I/O  reads, QD 128  - 85GiB/s across both 10+1+1 NUMA aware 128K striped LUNS.   Had the imbalance between NUMA 0 44.5GiB/s and NUMA 1 39.4GiB/s but still could be drifting power management on the AMD Rome cores.    I tried a 1280K blocksize to try to get a full stripe read, but Linux seems so unfriendly to non-power of 2 blocksizes.... performance decreased considerably (20GiB/s ?) with the 10x128KB blocksize....   I think I ran for about 40 minutes with the 1M reads...


socket0-md: (g=0): rw=read, bs=(R) 1024KiB-1024KiB, (W) 1024KiB-1024KiB, (T) 1024KiB-1024KiB, ioengine=libaio, iodepth=128
...
socket1-md: (g=1): rw=read, bs=(R) 1024KiB-1024KiB, (W) 1024KiB-1024KiB, (T) 1024KiB-1024KiB, ioengine=libaio, iodepth=128
...
fio-3.26
Starting 128 processes

fio: terminating on signal 2

socket0-md: (groupid=0, jobs=64): err= 0: pid=1645360: Mon Aug  9 18:53:36 2021
  read: IOPS=45.6k, BW=44.5GiB/s (47.8GB/s)(114TiB/2626961msec)
    slat (usec): min=12, max=4463, avg=24.86, stdev=15.58
    clat (usec): min=249, max=1904.8k, avg=179674.12, stdev=138190.51
     lat (usec): min=295, max=1904.8k, avg=179699.07, stdev=138191.00
    clat percentiles (msec):
     |  1.00th=[    3],  5.00th=[    5], 10.00th=[    7], 20.00th=[   17],
     | 30.00th=[  106], 40.00th=[  116], 50.00th=[  209], 60.00th=[  226],
     | 70.00th=[  236], 80.00th=[  321], 90.00th=[  351], 95.00th=[  372],
     | 99.00th=[  472], 99.50th=[  481], 99.90th=[ 1267], 99.95th=[ 1401],
     | 99.99th=[ 1586]
   bw (  MiB/s): min=  967, max=114322, per=8.68%, avg=45897.69, stdev=330.42, samples=333433
   iops        : min=  929, max=114304, avg=45879.39, stdev=330.41, samples=333433
  lat (usec)   : 250=0.01%, 500=0.01%, 750=0.05%, 1000=0.06%
  lat (msec)   : 2=0.49%, 4=4.36%, 10=9.43%, 20=7.52%, 50=3.48%
  lat (msec)   : 100=2.70%, 250=47.39%, 500=24.25%, 750=0.09%, 1000=0.01%
  lat (msec)   : 2000=0.15%
  cpu          : usr=0.07%, sys=1.83%, ctx=77483816, majf=0, minf=37747
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
     issued rwts: total=119750623,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128
socket1-md: (groupid=1, jobs=64): err= 0: pid=1645424: Mon Aug  9 18:53:36 2021
  read: IOPS=40.3k, BW=39.4GiB/s (42.3GB/s)(101TiB/2627054msec)
    slat (usec): min=12, max=57137, avg=23.77, stdev=27.80
    clat (usec): min=130, max=1746.1k, avg=203005.37, stdev=158045.10
     lat (usec): min=269, max=1746.1k, avg=203029.23, stdev=158045.27
    clat percentiles (usec):
     |  1.00th=[    570],  5.00th=[    693], 10.00th=[   2573],
     | 20.00th=[  21103], 30.00th=[ 102237], 40.00th=[ 143655],
     | 50.00th=[ 204473], 60.00th=[ 231736], 70.00th=[ 283116],
     | 80.00th=[ 320865], 90.00th=[ 421528], 95.00th=[ 455082],
     | 99.00th=[ 583009], 99.50th=[ 608175], 99.90th=[1061159],
     | 99.95th=[1166017], 99.99th=[1367344]
   bw (  MiB/s): min=  599, max=124821, per=-3.40%, avg=40571.79, stdev=319.36, samples=333904
   iops        : min=  568, max=124809, avg=40554.92, stdev=319.34, samples=333904
  lat (usec)   : 250=0.01%, 500=0.14%, 750=6.31%, 1000=2.60%
  lat (msec)   : 2=0.58%, 4=2.04%, 10=4.17%, 20=3.82%, 50=3.71%
  lat (msec)   : 100=5.91%, 250=32.86%, 500=33.81%, 750=3.81%, 1000=0.10%
  lat (msec)   : 2000=0.14%
  cpu          : usr=0.05%, sys=1.56%, ctx=71342745, majf=0, minf=37766
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
     issued rwts: total=105992570,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128

Run status group 0 (all jobs):
   READ: bw=44.5GiB/s (47.8GB/s), 44.5GiB/s-44.5GiB/s (47.8GB/s-47.8GB/s), io=114TiB (126TB), run=2626961-2626961msec

Run status group 1 (all jobs):
   READ: bw=39.4GiB/s (42.3GB/s), 39.4GiB/s-39.4GiB/s (42.3GB/s-42.3GB/s), io=101TiB (111TB), run=2627054-2627054msec

Disk stats (read/write):
    md0: ios=960804546/0, merge=0/0, ticks=18446744072288672424/0, in_queue=18446744072288672424, util=100.00%, aggrios=0/0, aggrmerge=0/0, aggrticks=0/0, aggrin_queue=0, aggrutil=0.00%
  nvme0n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme3n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme6n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme11n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme9n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme2n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme5n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme10n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme8n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme1n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme4n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme7n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
    md1: ios=850399203/0, merge=0/0, ticks=2118156441/0, in_queue=2118156441, util=100.00%, aggrios=0/0, aggrmerge=0/0, aggrticks=0/0, aggrin_queue=0, aggrutil=0.00%
  nvme15n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme18n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme20n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme23n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme14n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme17n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme22n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme13n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme19n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme21n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme12n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme24n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%

-----Original Message-----
From: Gal Ofri <redacted> 
Sent: Sunday, August 8, 2021 10:44 AM
To: Finlayson, James M CIV (USA) <redacted>
Cc: 'linux-raid@vger.kernel.org' <redacted>
Subject: Re: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

On Thu, 5 Aug 2021 21:10:40 +0000
"Finlayson, James M CIV (USA)" [off-list ref] wrote:
BLUF upfront with 5.14rc3 kernel that our SA built - md0 a 10+1+1 
RAID5 - 5.332 M IOPS 20.3GiB/s, md1 a 10+1+1 RAID5, 5.892M IOPS 22.5GiB/s  - best hero numbers I've ever seen on mdraid  RAID5 IOPS.   I think the kernel patch is good.  Prior was  socket0 1.263M IOPS 4934MiB/s, socket1 1.071M IOSP, 4183MiB/s....   I'm willing to help push this as hard as we can until we hit a bottleneck outside of our control.
That's great !
Thanks for sharing your results.
I'd appreciate if you could run a sequential-reads workload (128k/256k) so that we get a better sense of the throughput potential here.
In my strict numa adherence with mdraid, I see lots of variability between reboots/assembles.    Sometimes md0 wins, sometimes md1 wins, and in my earlier runs md0 and md1 are notionally balanced.   I change nothing but see this variance.   I just cranked up a week long extended run of these 10+1+1s under the 5.14rc3 kernel and right now   md0 is doing 5M IOPS and md1 6.3M 
Given my humble experience with the code in question, I suspect that it is not really optimized for numa awareness, so I find your findings quite reasonable. I don't really have a good tip for that.

I'm focusing now on thin-provisioned logical volumes (lvm - it has a much worse reads bottleneck actually), but we have plans for researching
md/raid5 again soon to improve write workloads.
I'll ping you when I have a patch that might be relevant.

Cheers,
Gal

RE: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

From: Finlayson, James M CIV (USA) <hidden>
Date: 2021-08-17 21:23:48

All,
A quick random performance update (this is the best I can do in "going for it" with all of the guidance from this list) - I'm thrilled.....

5.14rc4 kernel Gen 4 drives, all AMD Rome BIOS tuning to keep I/O from power throttling,  SMT turned on (off yielded higher performance but left no room for anything else),  15.36TB drives cut into 32 equal partitions,  32 NUMA aligned raid5 9+1s from the same partition on NUMA0 combined with an LVM concatenating all 32 RAID5's into one volume.    I then do the exact same thing on NUMA1.

4K random reads, SMT off, sustained bandwidth of > 90GB/s, sustained IOPS across both LVMs, ~23M - bad part, only 7% of the system left to do anything useful
4K random reads, SMT on, sustained bandwidth of > 84GB/s, sustained IOPS across both LVMs, ~21M - 46.7% idle (.73% users, 52.6% system time)
Takeaway - IMHO, no reason to turn off SMT, it helps way more than it hurts...

Without the partitioning and lvm shenanigans, with SMT on, 5.14rc4 kernel, most AMD BIOS tuning (not all), I'm at 46GB/s, 11.7M IOPS , 42.2% idle (3% user, 54.7% system time)

With stock RHEL 8.4, 4.18 kernel, SMT on, both partitioning and LVM shenanigans, most AMD BIOS tuning (not all), I'm at 81.5GB/s, 20.4M IOPS, 49% idle (5.5% user, 46.75% system time)

The question I have for the list, given my large drive sizes, it takes me a day to set up and build an mdraid/lvm configuration.    Has anybody found the "sweet spot" for how many partitions per drive?    I now have a script to generate the drive partitions, a script for building the mdraid volumes, and a procedure for unwinding from all of this and starting again.    

If anybody knows the point of diminishing return for the number of partitions per drive to max out at, it would save me a few days of letting 32 run for a day, reconfiguring for 16, 8, 4, 2, 1....I could just tear apart my LVMs and remake them with half as many RAID partitions, but depending upon how the nvme drive is "RAINed" across NAND chips, I might leave performance on the table.   The researcher in me says, start over, don't make ANY assumptions.

As an aside, on the server, I'm maintaining around 1.1M  NUMA aware IOPS per drive, when hitting all 24 drives individually without RAID, so I'm thrilled with the performance ceiling with the RAID, I just have to find a way to make it something somebody would be willing to maintain.   Somewhere is a sweet spot between sustainability and performance.   Once I find that I have to figure out if there is something useful to do with this new toy.....


Regards,
Jim




-----Original Message-----
From: Finlayson, James M CIV (USA) <redacted> 
Sent: Monday, August 9, 2021 3:02 PM
To: 'Gal Ofri' <redacted>; 'linux-raid@vger.kernel.org' <redacted>
Cc: Finlayson, James M CIV (USA) <redacted>
Subject: RE: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

Sequential Performance:
BLUF, 1M sequential, direct I/O  reads, QD 128  - 85GiB/s across both 10+1+1 NUMA aware 128K striped LUNS.   Had the imbalance between NUMA 0 44.5GiB/s and NUMA 1 39.4GiB/s but still could be drifting power management on the AMD Rome cores.    I tried a 1280K blocksize to try to get a full stripe read, but Linux seems so unfriendly to non-power of 2 blocksizes.... performance decreased considerably (20GiB/s ?) with the 10x128KB blocksize....   I think I ran for about 40 minutes with the 1M reads...


socket0-md: (g=0): rw=read, bs=(R) 1024KiB-1024KiB, (W) 1024KiB-1024KiB, (T) 1024KiB-1024KiB, ioengine=libaio, iodepth=128 ...
socket1-md: (g=1): rw=read, bs=(R) 1024KiB-1024KiB, (W) 1024KiB-1024KiB, (T) 1024KiB-1024KiB, ioengine=libaio, iodepth=128 ...
fio-3.26
Starting 128 processes

fio: terminating on signal 2

socket0-md: (groupid=0, jobs=64): err= 0: pid=1645360: Mon Aug  9 18:53:36 2021
  read: IOPS=45.6k, BW=44.5GiB/s (47.8GB/s)(114TiB/2626961msec)
    slat (usec): min=12, max=4463, avg=24.86, stdev=15.58
    clat (usec): min=249, max=1904.8k, avg=179674.12, stdev=138190.51
     lat (usec): min=295, max=1904.8k, avg=179699.07, stdev=138191.00
    clat percentiles (msec):
     |  1.00th=[    3],  5.00th=[    5], 10.00th=[    7], 20.00th=[   17],
     | 30.00th=[  106], 40.00th=[  116], 50.00th=[  209], 60.00th=[  226],
     | 70.00th=[  236], 80.00th=[  321], 90.00th=[  351], 95.00th=[  372],
     | 99.00th=[  472], 99.50th=[  481], 99.90th=[ 1267], 99.95th=[ 1401],
     | 99.99th=[ 1586]
   bw (  MiB/s): min=  967, max=114322, per=8.68%, avg=45897.69, stdev=330.42, samples=333433
   iops        : min=  929, max=114304, avg=45879.39, stdev=330.41, samples=333433
  lat (usec)   : 250=0.01%, 500=0.01%, 750=0.05%, 1000=0.06%
  lat (msec)   : 2=0.49%, 4=4.36%, 10=9.43%, 20=7.52%, 50=3.48%
  lat (msec)   : 100=2.70%, 250=47.39%, 500=24.25%, 750=0.09%, 1000=0.01%
  lat (msec)   : 2000=0.15%
  cpu          : usr=0.07%, sys=1.83%, ctx=77483816, majf=0, minf=37747
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
     issued rwts: total=119750623,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128
socket1-md: (groupid=1, jobs=64): err= 0: pid=1645424: Mon Aug  9 18:53:36 2021
  read: IOPS=40.3k, BW=39.4GiB/s (42.3GB/s)(101TiB/2627054msec)
    slat (usec): min=12, max=57137, avg=23.77, stdev=27.80
    clat (usec): min=130, max=1746.1k, avg=203005.37, stdev=158045.10
     lat (usec): min=269, max=1746.1k, avg=203029.23, stdev=158045.27
    clat percentiles (usec):
     |  1.00th=[    570],  5.00th=[    693], 10.00th=[   2573],
     | 20.00th=[  21103], 30.00th=[ 102237], 40.00th=[ 143655],
     | 50.00th=[ 204473], 60.00th=[ 231736], 70.00th=[ 283116],
     | 80.00th=[ 320865], 90.00th=[ 421528], 95.00th=[ 455082],
     | 99.00th=[ 583009], 99.50th=[ 608175], 99.90th=[1061159],
     | 99.95th=[1166017], 99.99th=[1367344]
   bw (  MiB/s): min=  599, max=124821, per=-3.40%, avg=40571.79, stdev=319.36, samples=333904
   iops        : min=  568, max=124809, avg=40554.92, stdev=319.34, samples=333904
  lat (usec)   : 250=0.01%, 500=0.14%, 750=6.31%, 1000=2.60%
  lat (msec)   : 2=0.58%, 4=2.04%, 10=4.17%, 20=3.82%, 50=3.71%
  lat (msec)   : 100=5.91%, 250=32.86%, 500=33.81%, 750=3.81%, 1000=0.10%
  lat (msec)   : 2000=0.14%
  cpu          : usr=0.05%, sys=1.56%, ctx=71342745, majf=0, minf=37766
  IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
     submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
     complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
     issued rwts: total=105992570,0,0,0 short=0,0,0,0 dropped=0,0,0,0
     latency   : target=0, window=0, percentile=100.00%, depth=128

Run status group 0 (all jobs):
   READ: bw=44.5GiB/s (47.8GB/s), 44.5GiB/s-44.5GiB/s (47.8GB/s-47.8GB/s), io=114TiB (126TB), run=2626961-2626961msec

Run status group 1 (all jobs):
   READ: bw=39.4GiB/s (42.3GB/s), 39.4GiB/s-39.4GiB/s (42.3GB/s-42.3GB/s), io=101TiB (111TB), run=2627054-2627054msec

Disk stats (read/write):
    md0: ios=960804546/0, merge=0/0, ticks=18446744072288672424/0, in_queue=18446744072288672424, util=100.00%, aggrios=0/0, aggrmerge=0/0, aggrticks=0/0, aggrin_queue=0, aggrutil=0.00%
  nvme0n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme3n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme6n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme11n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme9n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme2n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme5n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme10n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme8n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme1n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme4n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme7n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
    md1: ios=850399203/0, merge=0/0, ticks=2118156441/0, in_queue=2118156441, util=100.00%, aggrios=0/0, aggrmerge=0/0, aggrticks=0/0, aggrin_queue=0, aggrutil=0.00%
  nvme15n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme18n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme20n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme23n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme14n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme17n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme22n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme13n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme19n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme21n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme12n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
  nvme24n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%

-----Original Message-----
From: Gal Ofri <redacted>
Sent: Sunday, August 8, 2021 10:44 AM
To: Finlayson, James M CIV (USA) <redacted>
Cc: 'linux-raid@vger.kernel.org' <redacted>
Subject: Re: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

On Thu, 5 Aug 2021 21:10:40 +0000
"Finlayson, James M CIV (USA)" [off-list ref] wrote:
BLUF upfront with 5.14rc3 kernel that our SA built - md0 a 10+1+1
RAID5 - 5.332 M IOPS 20.3GiB/s, md1 a 10+1+1 RAID5, 5.892M IOPS 22.5GiB/s  - best hero numbers I've ever seen on mdraid  RAID5 IOPS.   I think the kernel patch is good.  Prior was  socket0 1.263M IOPS 4934MiB/s, socket1 1.071M IOSP, 4183MiB/s....   I'm willing to help push this as hard as we can until we hit a bottleneck outside of our control.
That's great !
Thanks for sharing your results.
I'd appreciate if you could run a sequential-reads workload (128k/256k) so that we get a better sense of the throughput potential here.
In my strict numa adherence with mdraid, I see lots of variability between reboots/assembles.    Sometimes md0 wins, sometimes md1 wins, and in my earlier runs md0 and md1 are notionally balanced.   I change nothing but see this variance.   I just cranked up a week long extended run of these 10+1+1s under the 5.14rc3 kernel and right now   md0 is doing 5M IOPS and md1 6.3M 
Given my humble experience with the code in question, I suspect that it is not really optimized for numa awareness, so I find your findings quite reasonable. I don't really have a good tip for that.

I'm focusing now on thin-provisioned logical volumes (lvm - it has a much worse reads bottleneck actually), but we have plans for researching
md/raid5 again soon to improve write workloads.
I'll ping you when I have a patch that might be relevant.

Cheers,
Gal

Re: [Non-DoD Source] Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

From: Matt Wallis <hidden>
Date: 2021-08-18 00:45:12

Hi Jim,

Awesome stuff. I’m looking to get access back to a server I was using before for my tests so I can play some more myself.
I did wonder about your use case, and if you were planning to present the storage over a network to another server, or intended to use it as local storage for an application.

The problem is basically that we’re limited no matter what we do. There’s no way with current PCIe+networking to get that bandwidth outside the box, and you don’t have much compute left inside the box.

You could simplify the configuration a little bit by using a parallel file system like BeeGFS. Parallel file systems like to stripe data over multiple targets anyway, so you could remove the LVM layer, and simply present 64 RAID volumes for BeeGFS to write to.  

Normal parallel file system operation is to export the volumes over a network, but BeeGFS does have an alternate mode called BeeOND, or BeeGFS on Demand, which builds up dynamic file systems using the local disks in multiple servers, you could potentially look at a single server BeeOND configuration and see if that worked, but I suspect you’d be exchanging bottlenecks.

There’s a new parallel FS on the market that might also be of interest, called MadFS. It’s based on another parallel file system but with certain parts re-written using the Rust language which significantly improved it’s ability to handle higher IOPs. 

Hmm, just realised the box I had access to before won’t help, it was built on an older Intel platform so bottlenecked by PCIe lanes. I’ll have to see if I can get something newer.

Matt.
On 18 Aug 2021, at 07:21, Finlayson, James M CIV (USA) [off-list ref] wrote:

All,
A quick random performance update (this is the best I can do in "going for it" with all of the guidance from this list) - I'm thrilled.....

5.14rc4 kernel Gen 4 drives, all AMD Rome BIOS tuning to keep I/O from power throttling,  SMT turned on (off yielded higher performance but left no room for anything else),  15.36TB drives cut into 32 equal partitions,  32 NUMA aligned raid5 9+1s from the same partition on NUMA0 combined with an LVM concatenating all 32 RAID5's into one volume.    I then do the exact same thing on NUMA1.

4K random reads, SMT off, sustained bandwidth of > 90GB/s, sustained IOPS across both LVMs, ~23M - bad part, only 7% of the system left to do anything useful
4K random reads, SMT on, sustained bandwidth of > 84GB/s, sustained IOPS across both LVMs, ~21M - 46.7% idle (.73% users, 52.6% system time)
Takeaway - IMHO, no reason to turn off SMT, it helps way more than it hurts...

Without the partitioning and lvm shenanigans, with SMT on, 5.14rc4 kernel, most AMD BIOS tuning (not all), I'm at 46GB/s, 11.7M IOPS , 42.2% idle (3% user, 54.7% system time)

With stock RHEL 8.4, 4.18 kernel, SMT on, both partitioning and LVM shenanigans, most AMD BIOS tuning (not all), I'm at 81.5GB/s, 20.4M IOPS, 49% idle (5.5% user, 46.75% system time)

The question I have for the list, given my large drive sizes, it takes me a day to set up and build an mdraid/lvm configuration.    Has anybody found the "sweet spot" for how many partitions per drive?    I now have a script to generate the drive partitions, a script for building the mdraid volumes, and a procedure for unwinding from all of this and starting again.    

If anybody knows the point of diminishing return for the number of partitions per drive to max out at, it would save me a few days of letting 32 run for a day, reconfiguring for 16, 8, 4, 2, 1....I could just tear apart my LVMs and remake them with half as many RAID partitions, but depending upon how the nvme drive is "RAINed" across NAND chips, I might leave performance on the table.   The researcher in me says, start over, don't make ANY assumptions.

As an aside, on the server, I'm maintaining around 1.1M  NUMA aware IOPS per drive, when hitting all 24 drives individually without RAID, so I'm thrilled with the performance ceiling with the RAID, I just have to find a way to make it something somebody would be willing to maintain.   Somewhere is a sweet spot between sustainability and performance.   Once I find that I have to figure out if there is something useful to do with this new toy.....


Regards,
Jim




-----Original Message-----
From: Finlayson, James M CIV (USA) <redacted> 
Sent: Monday, August 9, 2021 3:02 PM
To: 'Gal Ofri' <redacted>; 'linux-raid@vger.kernel.org' <redacted>
Cc: Finlayson, James M CIV (USA) <redacted>
Subject: RE: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

Sequential Performance:
BLUF, 1M sequential, direct I/O  reads, QD 128  - 85GiB/s across both 10+1+1 NUMA aware 128K striped LUNS.   Had the imbalance between NUMA 0 44.5GiB/s and NUMA 1 39.4GiB/s but still could be drifting power management on the AMD Rome cores.    I tried a 1280K blocksize to try to get a full stripe read, but Linux seems so unfriendly to non-power of 2 blocksizes.... performance decreased considerably (20GiB/s ?) with the 10x128KB blocksize....   I think I ran for about 40 minutes with the 1M reads...


socket0-md: (g=0): rw=read, bs=(R) 1024KiB-1024KiB, (W) 1024KiB-1024KiB, (T) 1024KiB-1024KiB, ioengine=libaio, iodepth=128 ...
socket1-md: (g=1): rw=read, bs=(R) 1024KiB-1024KiB, (W) 1024KiB-1024KiB, (T) 1024KiB-1024KiB, ioengine=libaio, iodepth=128 ...
fio-3.26
Starting 128 processes

fio: terminating on signal 2

socket0-md: (groupid=0, jobs=64): err= 0: pid=1645360: Mon Aug  9 18:53:36 2021
 read: IOPS=45.6k, BW=44.5GiB/s (47.8GB/s)(114TiB/2626961msec)
   slat (usec): min=12, max=4463, avg=24.86, stdev=15.58
   clat (usec): min=249, max=1904.8k, avg=179674.12, stdev=138190.51
    lat (usec): min=295, max=1904.8k, avg=179699.07, stdev=138191.00
   clat percentiles (msec):
    |  1.00th=[    3],  5.00th=[    5], 10.00th=[    7], 20.00th=[   17],
    | 30.00th=[  106], 40.00th=[  116], 50.00th=[  209], 60.00th=[  226],
    | 70.00th=[  236], 80.00th=[  321], 90.00th=[  351], 95.00th=[  372],
    | 99.00th=[  472], 99.50th=[  481], 99.90th=[ 1267], 99.95th=[ 1401],
    | 99.99th=[ 1586]
  bw (  MiB/s): min=  967, max=114322, per=8.68%, avg=45897.69, stdev=330.42, samples=333433
  iops        : min=  929, max=114304, avg=45879.39, stdev=330.41, samples=333433
 lat (usec)   : 250=0.01%, 500=0.01%, 750=0.05%, 1000=0.06%
 lat (msec)   : 2=0.49%, 4=4.36%, 10=9.43%, 20=7.52%, 50=3.48%
 lat (msec)   : 100=2.70%, 250=47.39%, 500=24.25%, 750=0.09%, 1000=0.01%
 lat (msec)   : 2000=0.15%
 cpu          : usr=0.07%, sys=1.83%, ctx=77483816, majf=0, minf=37747
 IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
    submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
    complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
    issued rwts: total=119750623,0,0,0 short=0,0,0,0 dropped=0,0,0,0
    latency   : target=0, window=0, percentile=100.00%, depth=128
socket1-md: (groupid=1, jobs=64): err= 0: pid=1645424: Mon Aug  9 18:53:36 2021
 read: IOPS=40.3k, BW=39.4GiB/s (42.3GB/s)(101TiB/2627054msec)
   slat (usec): min=12, max=57137, avg=23.77, stdev=27.80
   clat (usec): min=130, max=1746.1k, avg=203005.37, stdev=158045.10
    lat (usec): min=269, max=1746.1k, avg=203029.23, stdev=158045.27
   clat percentiles (usec):
    |  1.00th=[    570],  5.00th=[    693], 10.00th=[   2573],
    | 20.00th=[  21103], 30.00th=[ 102237], 40.00th=[ 143655],
    | 50.00th=[ 204473], 60.00th=[ 231736], 70.00th=[ 283116],
    | 80.00th=[ 320865], 90.00th=[ 421528], 95.00th=[ 455082],
    | 99.00th=[ 583009], 99.50th=[ 608175], 99.90th=[1061159],
    | 99.95th=[1166017], 99.99th=[1367344]
  bw (  MiB/s): min=  599, max=124821, per=-3.40%, avg=40571.79, stdev=319.36, samples=333904
  iops        : min=  568, max=124809, avg=40554.92, stdev=319.34, samples=333904
 lat (usec)   : 250=0.01%, 500=0.14%, 750=6.31%, 1000=2.60%
 lat (msec)   : 2=0.58%, 4=2.04%, 10=4.17%, 20=3.82%, 50=3.71%
 lat (msec)   : 100=5.91%, 250=32.86%, 500=33.81%, 750=3.81%, 1000=0.10%
 lat (msec)   : 2000=0.14%
 cpu          : usr=0.05%, sys=1.56%, ctx=71342745, majf=0, minf=37766
 IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
    submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
    complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
    issued rwts: total=105992570,0,0,0 short=0,0,0,0 dropped=0,0,0,0
    latency   : target=0, window=0, percentile=100.00%, depth=128

Run status group 0 (all jobs):
  READ: bw=44.5GiB/s (47.8GB/s), 44.5GiB/s-44.5GiB/s (47.8GB/s-47.8GB/s), io=114TiB (126TB), run=2626961-2626961msec

Run status group 1 (all jobs):
  READ: bw=39.4GiB/s (42.3GB/s), 39.4GiB/s-39.4GiB/s (42.3GB/s-42.3GB/s), io=101TiB (111TB), run=2627054-2627054msec

Disk stats (read/write):
   md0: ios=960804546/0, merge=0/0, ticks=18446744072288672424/0, in_queue=18446744072288672424, util=100.00%, aggrios=0/0, aggrmerge=0/0, aggrticks=0/0, aggrin_queue=0, aggrutil=0.00%
 nvme0n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme3n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme6n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme11n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme9n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme2n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme5n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme10n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme8n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme1n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme4n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme7n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
   md1: ios=850399203/0, merge=0/0, ticks=2118156441/0, in_queue=2118156441, util=100.00%, aggrios=0/0, aggrmerge=0/0, aggrticks=0/0, aggrin_queue=0, aggrutil=0.00%
 nvme15n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme18n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme20n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme23n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme14n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme17n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme22n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme13n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme19n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme21n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme12n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme24n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%

-----Original Message-----
From: Gal Ofri <redacted>
Sent: Sunday, August 8, 2021 10:44 AM
To: Finlayson, James M CIV (USA) <redacted>
Cc: 'linux-raid@vger.kernel.org' <redacted>
Subject: Re: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

On Thu, 5 Aug 2021 21:10:40 +0000
"Finlayson, James M CIV (USA)" [off-list ref] wrote:
quoted
BLUF upfront with 5.14rc3 kernel that our SA built - md0 a 10+1+1
RAID5 - 5.332 M IOPS 20.3GiB/s, md1 a 10+1+1 RAID5, 5.892M IOPS 22.5GiB/s  - best hero numbers I've ever seen on mdraid  RAID5 IOPS.   I think the kernel patch is good.  Prior was  socket0 1.263M IOPS 4934MiB/s, socket1 1.071M IOSP, 4183MiB/s....   I'm willing to help push this as hard as we can until we hit a bottleneck outside of our control.
That's great !
Thanks for sharing your results.
I'd appreciate if you could run a sequential-reads workload (128k/256k) so that we get a better sense of the throughput potential here.
quoted
In my strict numa adherence with mdraid, I see lots of variability between reboots/assembles.    Sometimes md0 wins, sometimes md1 wins, and in my earlier runs md0 and md1 are notionally balanced.   I change nothing but see this variance.   I just cranked up a week long extended run of these 10+1+1s under the 5.14rc3 kernel and right now   md0 is doing 5M IOPS and md1 6.3M 
Given my humble experience with the code in question, I suspect that it is not really optimized for numa awareness, so I find your findings quite reasonable. I don't really have a good tip for that.

I'm focusing now on thin-provisioned logical volumes (lvm - it has a much worse reads bottleneck actually), but we have plans for researching
md/raid5 again soon to improve write workloads.
I'll ping you when I have a patch that might be relevant.

Cheers,
Gal

RE: [Non-DoD Source] Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

From: Finlayson, James M CIV (USA) <hidden>
Date: 2021-08-18 10:21:38

All,
I'm happy to be in the position to pioneer some of this "tuning" if nobody has done this prior.   After updating this thread and then providing a status report to my leadership, it hit me on what we're really balancing is "how many mdraid kernel worker threads" it takes to hit max IOPS.  I'll go find that out.     If real world testing becomes my contribution , so be it.     I was an O/S developer originally working in the I/O subsystem, but early in my career that effort was deprecated, so I've only been an integrator of COTS and open source for the last 30 years and my programming skills have minimized to just perl and bash.   I don't have the skills necessary to make coding contributions.

Where I'd like mdraid to get is such that we don't need to do this, but this is a marathon, not a sprint.

As far as the PCIe lanes, AMD has situations where 160 gen 4 lanes are available (deleting 1 XGMI2 socket to socket interconnect).   If you have NUMA awareness, the box seems highly capable.    

Regards,
Jim



-----Original Message-----
From: Matt Wallis <redacted> 
Sent: Tuesday, August 17, 2021 8:45 PM
To: Finlayson, James M CIV (USA) <redacted>
Cc: linux-raid@vger.kernel.org
Subject: Re: [Non-DoD Source] Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

Hi Jim,

Awesome stuff. I’m looking to get access back to a server I was using before for my tests so I can play some more myself.
I did wonder about your use case, and if you were planning to present the storage over a network to another server, or intended to use it as local storage for an application.

The problem is basically that we’re limited no matter what we do. There’s no way with current PCIe+networking to get that bandwidth outside the box, and you don’t have much compute left inside the box.

You could simplify the configuration a little bit by using a parallel file system like BeeGFS. Parallel file systems like to stripe data over multiple targets anyway, so you could remove the LVM layer, and simply present 64 RAID volumes for BeeGFS to write to.  

Normal parallel file system operation is to export the volumes over a network, but BeeGFS does have an alternate mode called BeeOND, or BeeGFS on Demand, which builds up dynamic file systems using the local disks in multiple servers, you could potentially look at a single server BeeOND configuration and see if that worked, but I suspect you’d be exchanging bottlenecks.

There’s a new parallel FS on the market that might also be of interest, called MadFS. It’s based on another parallel file system but with certain parts re-written using the Rust language which significantly improved it’s ability to handle higher IOPs. 

Hmm, just realised the box I had access to before won’t help, it was built on an older Intel platform so bottlenecked by PCIe lanes. I’ll have to see if I can get something newer.

Matt.
On 18 Aug 2021, at 07:21, Finlayson, James M CIV (USA) [off-list ref] wrote:

All,
A quick random performance update (this is the best I can do in "going for it" with all of the guidance from this list) - I'm thrilled.....

5.14rc4 kernel Gen 4 drives, all AMD Rome BIOS tuning to keep I/O from power throttling,  SMT turned on (off yielded higher performance but left no room for anything else),  15.36TB drives cut into 32 equal partitions,  32 NUMA aligned raid5 9+1s from the same partition on NUMA0 combined with an LVM concatenating all 32 RAID5's into one volume.    I then do the exact same thing on NUMA1.

4K random reads, SMT off, sustained bandwidth of > 90GB/s, sustained 
IOPS across both LVMs, ~23M - bad part, only 7% of the system left to 
do anything useful 4K random reads, SMT on, sustained bandwidth of > 84GB/s, sustained IOPS across both LVMs, ~21M - 46.7% idle (.73% users, 52.6% system time) Takeaway - IMHO, no reason to turn off SMT, it helps way more than it hurts...

Without the partitioning and lvm shenanigans, with SMT on, 5.14rc4 
kernel, most AMD BIOS tuning (not all), I'm at 46GB/s, 11.7M IOPS , 
42.2% idle (3% user, 54.7% system time)

With stock RHEL 8.4, 4.18 kernel, SMT on, both partitioning and LVM 
shenanigans, most AMD BIOS tuning (not all), I'm at 81.5GB/s, 20.4M 
IOPS, 49% idle (5.5% user, 46.75% system time)

The question I have for the list, given my large drive sizes, it takes me a day to set up and build an mdraid/lvm configuration.    Has anybody found the "sweet spot" for how many partitions per drive?    I now have a script to generate the drive partitions, a script for building the mdraid volumes, and a procedure for unwinding from all of this and starting again.    

If anybody knows the point of diminishing return for the number of partitions per drive to max out at, it would save me a few days of letting 32 run for a day, reconfiguring for 16, 8, 4, 2, 1....I could just tear apart my LVMs and remake them with half as many RAID partitions, but depending upon how the nvme drive is "RAINed" across NAND chips, I might leave performance on the table.   The researcher in me says, start over, don't make ANY assumptions.

As an aside, on the server, I'm maintaining around 1.1M  NUMA aware IOPS per drive, when hitting all 24 drives individually without RAID, so I'm thrilled with the performance ceiling with the RAID, I just have to find a way to make it something somebody would be willing to maintain.   Somewhere is a sweet spot between sustainability and performance.   Once I find that I have to figure out if there is something useful to do with this new toy.....


Regards,
Jim




-----Original Message-----
From: Finlayson, James M CIV (USA) <redacted>
Sent: Monday, August 9, 2021 3:02 PM
To: 'Gal Ofri' <redacted>; 'linux-raid@vger.kernel.org' 
[off-list ref]
Cc: Finlayson, James M CIV (USA) <redacted>
Subject: RE: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

Sequential Performance:
BLUF, 1M sequential, direct I/O  reads, QD 128  - 85GiB/s across both 10+1+1 NUMA aware 128K striped LUNS.   Had the imbalance between NUMA 0 44.5GiB/s and NUMA 1 39.4GiB/s but still could be drifting power management on the AMD Rome cores.    I tried a 1280K blocksize to try to get a full stripe read, but Linux seems so unfriendly to non-power of 2 blocksizes.... performance decreased considerably (20GiB/s ?) with the 10x128KB blocksize....   I think I ran for about 40 minutes with the 1M reads...


socket0-md: (g=0): rw=read, bs=(R) 1024KiB-1024KiB, (W) 1024KiB-1024KiB, (T) 1024KiB-1024KiB, ioengine=libaio, iodepth=128 ...
socket1-md: (g=1): rw=read, bs=(R) 1024KiB-1024KiB, (W) 1024KiB-1024KiB, (T) 1024KiB-1024KiB, ioengine=libaio, iodepth=128 ...
fio-3.26
Starting 128 processes

fio: terminating on signal 2

socket0-md: (groupid=0, jobs=64): err= 0: pid=1645360: Mon Aug  9 
18:53:36 2021
 read: IOPS=45.6k, BW=44.5GiB/s (47.8GB/s)(114TiB/2626961msec)
   slat (usec): min=12, max=4463, avg=24.86, stdev=15.58
   clat (usec): min=249, max=1904.8k, avg=179674.12, stdev=138190.51
    lat (usec): min=295, max=1904.8k, avg=179699.07, stdev=138191.00
   clat percentiles (msec):
    |  1.00th=[    3],  5.00th=[    5], 10.00th=[    7], 20.00th=[   17],
    | 30.00th=[  106], 40.00th=[  116], 50.00th=[  209], 60.00th=[  226],
    | 70.00th=[  236], 80.00th=[  321], 90.00th=[  351], 95.00th=[  372],
    | 99.00th=[  472], 99.50th=[  481], 99.90th=[ 1267], 99.95th=[ 1401],
    | 99.99th=[ 1586]
  bw (  MiB/s): min=  967, max=114322, per=8.68%, avg=45897.69, stdev=330.42, samples=333433
  iops        : min=  929, max=114304, avg=45879.39, stdev=330.41, samples=333433
 lat (usec)   : 250=0.01%, 500=0.01%, 750=0.05%, 1000=0.06%
 lat (msec)   : 2=0.49%, 4=4.36%, 10=9.43%, 20=7.52%, 50=3.48%
 lat (msec)   : 100=2.70%, 250=47.39%, 500=24.25%, 750=0.09%, 1000=0.01%
 lat (msec)   : 2000=0.15%
 cpu          : usr=0.07%, sys=1.83%, ctx=77483816, majf=0, minf=37747
 IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
    submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
    complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
    issued rwts: total=119750623,0,0,0 short=0,0,0,0 dropped=0,0,0,0
    latency   : target=0, window=0, percentile=100.00%, depth=128
socket1-md: (groupid=1, jobs=64): err= 0: pid=1645424: Mon Aug  9 
18:53:36 2021
 read: IOPS=40.3k, BW=39.4GiB/s (42.3GB/s)(101TiB/2627054msec)
   slat (usec): min=12, max=57137, avg=23.77, stdev=27.80
   clat (usec): min=130, max=1746.1k, avg=203005.37, stdev=158045.10
    lat (usec): min=269, max=1746.1k, avg=203029.23, stdev=158045.27
   clat percentiles (usec):
    |  1.00th=[    570],  5.00th=[    693], 10.00th=[   2573],
    | 20.00th=[  21103], 30.00th=[ 102237], 40.00th=[ 143655],
    | 50.00th=[ 204473], 60.00th=[ 231736], 70.00th=[ 283116],
    | 80.00th=[ 320865], 90.00th=[ 421528], 95.00th=[ 455082],
    | 99.00th=[ 583009], 99.50th=[ 608175], 99.90th=[1061159],
    | 99.95th=[1166017], 99.99th=[1367344]
  bw (  MiB/s): min=  599, max=124821, per=-3.40%, avg=40571.79, stdev=319.36, samples=333904
  iops        : min=  568, max=124809, avg=40554.92, stdev=319.34, samples=333904
 lat (usec)   : 250=0.01%, 500=0.14%, 750=6.31%, 1000=2.60%
 lat (msec)   : 2=0.58%, 4=2.04%, 10=4.17%, 20=3.82%, 50=3.71%
 lat (msec)   : 100=5.91%, 250=32.86%, 500=33.81%, 750=3.81%, 1000=0.10%
 lat (msec)   : 2000=0.14%
 cpu          : usr=0.05%, sys=1.56%, ctx=71342745, majf=0, minf=37766
 IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%, >=64=100.0%
    submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.0%
    complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%, >=64=0.1%
    issued rwts: total=105992570,0,0,0 short=0,0,0,0 dropped=0,0,0,0
    latency   : target=0, window=0, percentile=100.00%, depth=128

Run status group 0 (all jobs):
  READ: bw=44.5GiB/s (47.8GB/s), 44.5GiB/s-44.5GiB/s 
(47.8GB/s-47.8GB/s), io=114TiB (126TB), run=2626961-2626961msec

Run status group 1 (all jobs):
  READ: bw=39.4GiB/s (42.3GB/s), 39.4GiB/s-39.4GiB/s 
(42.3GB/s-42.3GB/s), io=101TiB (111TB), run=2627054-2627054msec

Disk stats (read/write):
   md0: ios=960804546/0, merge=0/0, ticks=18446744072288672424/0, 
in_queue=18446744072288672424, util=100.00%, aggrios=0/0, 
aggrmerge=0/0, aggrticks=0/0, aggrin_queue=0, aggrutil=0.00%
 nvme0n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme3n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme6n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme11n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme9n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme2n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme5n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme10n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme8n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme1n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme4n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme7n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
   md1: ios=850399203/0, merge=0/0, ticks=2118156441/0, 
in_queue=2118156441, util=100.00%, aggrios=0/0, aggrmerge=0/0, 
aggrticks=0/0, aggrin_queue=0, aggrutil=0.00%
 nvme15n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme18n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme20n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme23n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme14n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme17n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme22n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme13n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme19n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme21n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme12n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme24n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%

-----Original Message-----
From: Gal Ofri <redacted>
Sent: Sunday, August 8, 2021 10:44 AM
To: Finlayson, James M CIV (USA) <redacted>
Cc: 'linux-raid@vger.kernel.org' <redacted>
Subject: Re: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

On Thu, 5 Aug 2021 21:10:40 +0000
"Finlayson, James M CIV (USA)" [off-list ref] wrote:
quoted
BLUF upfront with 5.14rc3 kernel that our SA built - md0 a 10+1+1
RAID5 - 5.332 M IOPS 20.3GiB/s, md1 a 10+1+1 RAID5, 5.892M IOPS 22.5GiB/s  - best hero numbers I've ever seen on mdraid  RAID5 IOPS.   I think the kernel patch is good.  Prior was  socket0 1.263M IOPS 4934MiB/s, socket1 1.071M IOSP, 4183MiB/s....   I'm willing to help push this as hard as we can until we hit a bottleneck outside of our control.
That's great !
Thanks for sharing your results.
I'd appreciate if you could run a sequential-reads workload (128k/256k) so that we get a better sense of the throughput potential here.
quoted
In my strict numa adherence with mdraid, I see lots of variability between reboots/assembles.    Sometimes md0 wins, sometimes md1 wins, and in my earlier runs md0 and md1 are notionally balanced.   I change nothing but see this variance.   I just cranked up a week long extended run of these 10+1+1s under the 5.14rc3 kernel and right now   md0 is doing 5M IOPS and md1 6.3M 
Given my humble experience with the code in question, I suspect that it is not really optimized for numa awareness, so I find your findings quite reasonable. I don't really have a good tip for that.

I'm focusing now on thin-provisioned logical volumes (lvm - it has a 
much worse reads bottleneck actually), but we have plans for 
researching
md/raid5 again soon to improve write workloads.
I'll ping you when I have a patch that might be relevant.

Cheers,
Gal

Re: [Non-DoD Source] Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

From: Doug Ledford <hidden>
Date: 2021-08-18 19:49:38

On Wed, 2021-08-18 at 10:20 +0000, Finlayson, James M CIV (USA) wrote:
All,
I'm happy to be in the position to pioneer some of this "tuning" if
nobody has done this prior.
Here's what I would be interested to know: how does btrfs do using these
drives bare in raid5 mode?  You have to do the metadata in raid1 mode,
but you can tear down and retest btrfs filesystems on this in a matter
of minutes because it doesn't have to initialize the array.  So you
could try a btrfs per NUMA node, one big btrfs, or other configurations.

Now, let me explain why I think this would be interesting.  I'm a long
time user and developer on the MD raid stack, going all the way back to
the first SSE implementation of the raid5 xor operations.  I've always
used mdraid, and later lvm + mdraid, to build my boxes.  But I've come
to believe that there is an inherint weakness to the mdraid + lvm +
filesystem stack that btrfs (and zfs) overcome by building their raid
code into the filesystem itself.  The inherint weakness is that the
filesystem is the source of truth for what blocks on the device have or
have not been allocated, and what their contents should be.  The mdraid
stack therefore has to do things like initialize the array because it
doesn't know what's written and what isn't.  This also impacts
reconstruction and error recovery similarly.  But, more importantly, it
means that in an attempt to avoid always having huge latency penalties
caused by read-modify-write cycles, the mdraid subsystem maintains its
own cache layer (the stripe cache) separate from the official page cache
of the kernel.  Although I haven't instrumented things to see for sure
if I'm right, my suspicion is that the stripe cache sometimes gets
blocked up under memory pressure and stalls writes to the array.  The
symptom I see is that when I'm copying a large file to the server via
10Gig Ethernet, it will start at 900MB/s and may stay that fast for the
entire operation, but other times the copy will stall, sometimes going
all the way to 0MB/s, for a random period of time.  My suspicion is that
when this happens, there is memory pressure and the raid5 code is having
trouble reading in blocks for read-modify-write operations when the
write it needs to perform is not a full stripe wide write.  This is
avoided when the filesystem is aware of the multi drive layout and
issues the reads itself.  So I strongly suspect that when I build my
next iteration of my home server, it's going to be btrfs (in fact, I
have a test install on it already, but I haven't had the time to do all
the testing needed to confirm it actually solves the problem of the
previous generation of my server).
   After updating this thread and then providing a status report to my
leadership, it hit me on what we're really balancing is "how many
mdraid kernel worker threads" it takes to hit max IOPS.  I'll go find
that out.     If real world testing becomes my contribution , so be
it.     I was an O/S developer originally working in the I/O
subsystem, but early in my career that effort was deprecated, so I've
only been an integrator of COTS and open source for the last 30 years
and my programming skills have minimized to just perl and bash.   I
don't have the skills necessary to make coding contributions.

Where I'd like mdraid to get is such that we don't need to do this,
but this is a marathon, not a sprint.

As far as the PCIe lanes, AMD has situations where 160 gen 4 lanes are
available (deleting 1 XGMI2 socket to socket interconnect).   If you
have NUMA awareness, the box seems highly capable.    

Regards,
Jim



-----Original Message-----
From: Matt Wallis <redacted> 
Sent: Tuesday, August 17, 2021 8:45 PM
To: Finlayson, James M CIV (USA) <redacted>
Cc: linux-raid@vger.kernel.org
Subject: Re: [Non-DoD Source] Can't get RAID5/RAID6 NVMe randomread
IOPS - AMD ROME what am I missing?????

Hi Jim,

Awesome stuff. I’m looking to get access back to a server I was using
before for my tests so I can play some more myself.
I did wonder about your use case, and if you were planning to present
the storage over a network to another server, or intended to use it as
local storage for an application.

The problem is basically that we’re limited no matter what we do.
There’s no way with current PCIe+networking to get that bandwidth
outside the box, and you don’t have much compute left inside the box.

You could simplify the configuration a little bit by using a parallel
file system like BeeGFS. Parallel file systems like to stripe data
over multiple targets anyway, so you could remove the LVM layer, and
simply present 64 RAID volumes for BeeGFS to write to.  

Normal parallel file system operation is to export the volumes over a
network, but BeeGFS does have an alternate mode called BeeOND, or
BeeGFS on Demand, which builds up dynamic file systems using the local
disks in multiple servers, you could potentially look at a single
server BeeOND configuration and see if that worked, but I suspect
you’d be exchanging bottlenecks.

There’s a new parallel FS on the market that might also be of
interest, called MadFS. It’s based on another parallel file system but
with certain parts re-written using the Rust language which
significantly improved it’s ability to handle higher IOPs. 

Hmm, just realised the box I had access to before won’t help, it was
built on an older Intel platform so bottlenecked by PCIe lanes. I’ll
have to see if I can get something newer.

Matt.
quoted
On 18 Aug 2021, at 07:21, Finlayson, James M CIV (USA) <
james.m.finlayson4.civ@mail.mil> wrote:

All,
A quick random performance update (this is the best I can do in
"going for it" with all of the guidance from this list) - I'm
thrilled.....

5.14rc4 kernel Gen 4 drives, all AMD Rome BIOS tuning to keep I/O
from power throttling,  SMT turned on (off yielded higher
performance but left no room for anything else),  15.36TB drives cut
into 32 equal partitions,  32 NUMA aligned raid5 9+1s from the same
partition on NUMA0 combined with an LVM concatenating all 32 RAID5's
into one volume.    I then do the exact same thing on NUMA1.

4K random reads, SMT off, sustained bandwidth of > 90GB/s, sustained
IOPS across both LVMs, ~23M - bad part, only 7% of the system left
to 
do anything useful 4K random reads, SMT on, sustained bandwidth of >
84GB/s, sustained IOPS across both LVMs, ~21M - 46.7% idle (.73%
users, 52.6% system time) Takeaway - IMHO, no reason to turn off
SMT, it helps way more than it hurts...

Without the partitioning and lvm shenanigans, with SMT on, 5.14rc4 
kernel, most AMD BIOS tuning (not all), I'm at 46GB/s, 11.7M IOPS , 
42.2% idle (3% user, 54.7% system time)

With stock RHEL 8.4, 4.18 kernel, SMT on, both partitioning and LVM 
shenanigans, most AMD BIOS tuning (not all), I'm at 81.5GB/s, 20.4M 
IOPS, 49% idle (5.5% user, 46.75% system time)

The question I have for the list, given my large drive sizes, it
takes me a day to set up and build an mdraid/lvm configuration.   
Has anybody found the "sweet spot" for how many partitions per
drive?    I now have a script to generate the drive partitions, a
script for building the mdraid volumes, and a procedure for
unwinding from all of this and starting again.    

If anybody knows the point of diminishing return for the number of
partitions per drive to max out at, it would save me a few days of
letting 32 run for a day, reconfiguring for 16, 8, 4, 2, 1....I
could just tear apart my LVMs and remake them with half as many RAID
partitions, but depending upon how the nvme drive is "RAINed" across
NAND chips, I might leave performance on the table.   The researcher
in me says, start over, don't make ANY assumptions.

As an aside, on the server, I'm maintaining around 1.1M  NUMA aware
IOPS per drive, when hitting all 24 drives individually without
RAID, so I'm thrilled with the performance ceiling with the RAID, I
just have to find a way to make it something somebody would be
willing to maintain.   Somewhere is a sweet spot between
sustainability and performance.   Once I find that I have to figure
out if there is something useful to do with this new toy.....


Regards,
Jim




-----Original Message-----
From: Finlayson, James M CIV (USA) <redacted>
Sent: Monday, August 9, 2021 3:02 PM
To: 'Gal Ofri' <redacted>; 'linux-raid@vger.kernel.org' 
[off-list ref]
Cc: Finlayson, James M CIV (USA) <redacted>
Subject: RE: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe
randomread IOPS - AMD ROME what am I missing?????

Sequential Performance:
BLUF, 1M sequential, direct I/O  reads, QD 128  - 85GiB/s across
both 10+1+1 NUMA aware 128K striped LUNS.   Had the imbalance
between NUMA 0 44.5GiB/s and NUMA 1 39.4GiB/s but still could be
drifting power management on the AMD Rome cores.    I tried a 1280K
blocksize to try to get a full stripe read, but Linux seems so
unfriendly to non-power of 2 blocksizes.... performance decreased
considerably (20GiB/s ?) with the 10x128KB blocksize....   I think I
ran for about 40 minutes with the 1M reads...


socket0-md: (g=0): rw=read, bs=(R) 1024KiB-1024KiB, (W) 1024KiB-
1024KiB, (T) 1024KiB-1024KiB, ioengine=libaio, iodepth=128 ...
socket1-md: (g=1): rw=read, bs=(R) 1024KiB-1024KiB, (W) 1024KiB-
1024KiB, (T) 1024KiB-1024KiB, ioengine=libaio, iodepth=128 ...
fio-3.26
Starting 128 processes

fio: terminating on signal 2

socket0-md: (groupid=0, jobs=64): err= 0: pid=1645360: Mon Aug  9 
18:53:36 2021
 read: IOPS=45.6k, BW=44.5GiB/s (47.8GB/s)(114TiB/2626961msec)
   slat (usec): min=12, max=4463, avg=24.86, stdev=15.58
   clat (usec): min=249, max=1904.8k, avg=179674.12, stdev=138190.51
    lat (usec): min=295, max=1904.8k, avg=179699.07, stdev=138191.00
   clat percentiles (msec):
    |  1.00th=[    3],  5.00th=[    5], 10.00th=[    7], 20.00th=[  
17],
    | 30.00th=[  106], 40.00th=[  116], 50.00th=[  209], 60.00th=[ 
226],
    | 70.00th=[  236], 80.00th=[  321], 90.00th=[  351], 95.00th=[ 
372],
    | 99.00th=[  472], 99.50th=[  481], 99.90th=[ 1267], 99.95th=[
1401],
    | 99.99th=[ 1586]
  bw (  MiB/s): min=  967, max=114322, per=8.68%, avg=45897.69,
stdev=330.42, samples=333433
  iops        : min=  929, max=114304, avg=45879.39, stdev=330.41,
samples=333433
 lat (usec)   : 250=0.01%, 500=0.01%, 750=0.05%, 1000=0.06%
 lat (msec)   : 2=0.49%, 4=4.36%, 10=9.43%, 20=7.52%, 50=3.48%
 lat (msec)   : 100=2.70%, 250=47.39%, 500=24.25%, 750=0.09%,
1000=0.01%
 lat (msec)   : 2000=0.15%
 cpu          : usr=0.07%, sys=1.83%, ctx=77483816, majf=0,
minf=37747
 IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%,
quoted
=64=100.0%
    submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%,
quoted
=64=0.0%
    complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%,
quoted
=64=0.1%
    issued rwts: total=119750623,0,0,0 short=0,0,0,0 dropped=0,0,0,0
    latency   : target=0, window=0, percentile=100.00%, depth=128
socket1-md: (groupid=1, jobs=64): err= 0: pid=1645424: Mon Aug  9 
18:53:36 2021
 read: IOPS=40.3k, BW=39.4GiB/s (42.3GB/s)(101TiB/2627054msec)
   slat (usec): min=12, max=57137, avg=23.77, stdev=27.80
   clat (usec): min=130, max=1746.1k, avg=203005.37, stdev=158045.10
    lat (usec): min=269, max=1746.1k, avg=203029.23, stdev=158045.27
   clat percentiles (usec):
    |  1.00th=[    570],  5.00th=[    693], 10.00th=[   2573],
    | 20.00th=[  21103], 30.00th=[ 102237], 40.00th=[ 143655],
    | 50.00th=[ 204473], 60.00th=[ 231736], 70.00th=[ 283116],
    | 80.00th=[ 320865], 90.00th=[ 421528], 95.00th=[ 455082],
    | 99.00th=[ 583009], 99.50th=[ 608175], 99.90th=[1061159],
    | 99.95th=[1166017], 99.99th=[1367344]
  bw (  MiB/s): min=  599, max=124821, per=-3.40%, avg=40571.79,
stdev=319.36, samples=333904
  iops        : min=  568, max=124809, avg=40554.92, stdev=319.34,
samples=333904
 lat (usec)   : 250=0.01%, 500=0.14%, 750=6.31%, 1000=2.60%
 lat (msec)   : 2=0.58%, 4=2.04%, 10=4.17%, 20=3.82%, 50=3.71%
 lat (msec)   : 100=5.91%, 250=32.86%, 500=33.81%, 750=3.81%,
1000=0.10%
 lat (msec)   : 2000=0.14%
 cpu          : usr=0.05%, sys=1.56%, ctx=71342745, majf=0,
minf=37766
 IO depths    : 1=0.1%, 2=0.1%, 4=0.1%, 8=0.1%, 16=0.1%, 32=0.1%,
quoted
=64=100.0%
    submit    : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%,
quoted
=64=0.0%
    complete  : 0=0.0%, 4=100.0%, 8=0.0%, 16=0.0%, 32=0.0%, 64=0.0%,
quoted
=64=0.1%
    issued rwts: total=105992570,0,0,0 short=0,0,0,0 dropped=0,0,0,0
    latency   : target=0, window=0, percentile=100.00%, depth=128

Run status group 0 (all jobs):
  READ: bw=44.5GiB/s (47.8GB/s), 44.5GiB/s-44.5GiB/s 
(47.8GB/s-47.8GB/s), io=114TiB (126TB), run=2626961-2626961msec

Run status group 1 (all jobs):
  READ: bw=39.4GiB/s (42.3GB/s), 39.4GiB/s-39.4GiB/s 
(42.3GB/s-42.3GB/s), io=101TiB (111TB), run=2627054-2627054msec

Disk stats (read/write):
   md0: ios=960804546/0, merge=0/0, ticks=18446744072288672424/0, 
in_queue=18446744072288672424, util=100.00%, aggrios=0/0, 
aggrmerge=0/0, aggrticks=0/0, aggrin_queue=0, aggrutil=0.00%
 nvme0n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme3n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme6n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme11n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme9n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme2n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme5n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme10n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme8n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme1n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme4n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme7n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
   md1: ios=850399203/0, merge=0/0, ticks=2118156441/0, 
in_queue=2118156441, util=100.00%, aggrios=0/0, aggrmerge=0/0, 
aggrticks=0/0, aggrin_queue=0, aggrutil=0.00%
 nvme15n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme18n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme20n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme23n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme14n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme17n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme22n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme13n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme19n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme21n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme12n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%
 nvme24n1: ios=0/0, merge=0/0, ticks=0/0, in_queue=0, util=0.00%

-----Original Message-----
From: Gal Ofri <redacted>
Sent: Sunday, August 8, 2021 10:44 AM
To: Finlayson, James M CIV (USA) <redacted>
Cc: 'linux-raid@vger.kernel.org' <redacted>
Subject: Re: [Non-DoD Source] Re: Can't get RAID5/RAID6 NVMe
randomread IOPS - AMD ROME what am I missing?????

On Thu, 5 Aug 2021 21:10:40 +0000
"Finlayson, James M CIV (USA)" [off-list ref]
wrote:
quoted
BLUF upfront with 5.14rc3 kernel that our SA built - md0 a 10+1+1
RAID5 - 5.332 M IOPS 20.3GiB/s, md1 a 10+1+1 RAID5, 5.892M IOPS
22.5GiB/s  - best hero numbers I've ever seen on mdraid  RAID5
IOPS.   I think the kernel patch is good.  Prior was  socket0
1.263M IOPS 4934MiB/s, socket1 1.071M IOSP, 4183MiB/s....   I'm
willing to help push this as hard as we can until we hit a
bottleneck outside of our control.
That's great !
Thanks for sharing your results.
I'd appreciate if you could run a sequential-reads workload
(128k/256k) so that we get a better sense of the throughput
potential here.
quoted
In my strict numa adherence with mdraid, I see lots of variability
between reboots/assembles.    Sometimes md0 wins, sometimes md1
wins, and in my earlier runs md0 and md1 are notionally
balanced.   I change nothing but see this variance.   I just
cranked up a week long extended run of these 10+1+1s under the
5.14rc3 kernel and right now   md0 is doing 5M IOPS and md1 6.3M 
Given my humble experience with the code in question, I suspect that
it is not really optimized for numa awareness, so I find your
findings quite reasonable. I don't really have a good tip for that.

I'm focusing now on thin-provisioned logical volumes (lvm - it has a
much worse reads bottleneck actually), but we have plans for 
researching
md/raid5 again soon to improve write workloads.
I'll ping you when I have a patch that might be relevant.

Cheers,
Gal
-- 
Doug Ledford [off-list ref]
    GPG KeyID: B826A3330E572FDD
    Fingerprint = AE6B 1BDA 122B 23B4 265B  1274 B826 A333 0E57 2FDD

Re: [Non-DoD Source] Can't get RAID5/RAID6 NVMe randomread IOPS - AMD ROME what am I missing?????

From: Doug Ledford <hidden>
Date: 2021-08-18 19:59:13

quoted
The question I have for the list, given my large drive sizes, it
takes me a day to set up and build an mdraid/lvm configuration.   
Has anybody found the "sweet spot" for how many partitions per
drive?    I now have a script to generate the drive partitions, a
script for building the mdraid volumes, and a procedure for
unwinding from all of this and starting again.    
I don't have a feeling for the sweet spot on the number of partitions,
but if you put too many devices in a raid5/6 array, you virtually
guarantee all writes will have to be read-modify-write writes instead of
full stripe writes.

So, when dealing with keeping the parity on the array in sync, a full
stripe write allows you to simply write all blocks in the stripe,
calculate the parity as you do so, and then write the parity out.  For a
partial stripe write, you either have to read in the blocks you aren't
writing and then treat it as a full stripe write and calculate the
parity, or you have to read in the blocks being written and the current
parity block, xor the blocks being over written out of the existing
parity block and then xor the blocks you are writing over the old ones
into the parity block, then write the new blocks and new parity out.

For that reason, I usually try to keep my arrays to no more than 7 or 8
members.  A lot of times, for streaming testing, really high numbers of
drives in a parity raid array will seem to perform fine, but when under
real world conditions might not do so well.  There are also several
filesystems that will optimize their metadata layout when put on an
mdraid device (xfs and ext4), but I'm pretty sure that gets blocked when
you put lvm between the filesystem and the mdraid device.

-- 
Doug Ledford [off-list ref]
    GPG KeyID: B826A3330E572FDD
    Fingerprint = AE6B 1BDA 122B 23B4 265B  1274 B826 A333 0E57 2FDD
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help