From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-19 06:11:19
On power9, Virtual Accelerator Switchboard (VAS) allows user space or
kernel to communicate with Nest Accelerator (NX) directly using COPY/PASTE
instructions. NX provides various functionalities such as compression,
encryption and etc. But only compression (842 and GZIP formats) is
supported in Linux kernel on power9.
842 compression driver (drivers/crypto/nx/nx-842-powernv.c)
is already included in Linux. Only GZIP support will be available from
user space.
Applications can issue GZIP compression / decompression requests to NX with
COPY/PASTE instructions. When NX is processing these requests, can hit
fault on the request buffer (not in memory). It issues an interrupt and
pastes fault CRB in fault FIFO. Expects kernel to handle this fault and
return credits for both send and fault windows after processing.
This patch series adds IRQ and fault window setup, and NX fault handling:
- Alloc IRQ and trigger port address, and configure IRQ per VAS instance.
- Set port# for each window to generate an interrupt when noticed fault.
- Set fault window and FIFO on which NX paste fault CRB.
- Setup IRQ thread fault handler per VAS instance.
- When receiving an interrupt, Read CRBs from fault FIFO and update
coprocessor_status_block (CSB) in the corresponding CRB with translation
failure (CSB_CC_TRANSLATION). After issuing NX requests, process polls
on CSB address. When it sees translation error, can touch the request
buffer to bring the page in to memory and reissue NX request.
- If copy_to_user fails on user space CSB address, OS sends SEGV signal.
Tested these patches with NX-GZIP support and will be posting this series
soon.
Patches 1 & 2: Define alloc IRQ and get port address per chip which are needed
to alloc IRQ per VAS instance.
Patch 3: Define nx_fault_stamp on which NX writes fault status for the fault
CRB
Patch 4: Alloc and setup IRQ and trigger port address for each VAS instance
Patch 5: Setup fault window per each VAS instance. This window is used for
NX to paste fault CRB in FIFO.
Patches 6 & 7: Setup threaded IRQ per VAS and register NX with fault window
ID and port number for each send window so that NX paste fault CRB
in this window.
Patch 8: Reference to pid and mm so that pid is not used until window closed.
Needed for multi thread application where child can open a window
and can be used by parent later.
Patches 9 and 10: Process CRBs from fault FIFO and notify tasks by
updating CSB or through signals.
Patches 11 and 12: Return credits for send and fault windows after handling
faults.
Patch 14:Fix closing send window after all credits are returned. This issue
happens only for user space requests. No page faults on kernel
request buffer.
Changelog:
V2:
- Use threaded IRQ instead of own kernel thread handler
- Use pswid instead of user space CSB address to find valid CRB
- Removed unused macros and other changes as suggested by Christoph Hellwig
V3:
- Rebased to 5.5-rc2
- Use struct pid * instead of pid_t for vas_window tgid
- Code cleanup as suggested by Christoph Hellwig
V4:
- Define xive alloc and get IRQ info based on chip ID and use these
functions for IRQ setup per VAS instance. It eliminates skiboot
dependency as suggested by Oliver.
V5:
- Do not update CSB if the process is exiting (patch9)
V6:
- Add interrupt handler instead of default one and return IRQ_HANDLED
if the fault handling thread is already in progress. (Patch6)
- Use platform send window ID and CCW[0] bit to find valid CRB in
fault FIFO (Patch6).
- Return fault address to user space in BE and other changes as
suggested by Michael Neuling. (patch9)
- Rebased to 5.6-rc4
V7:
- Fix sparse warnings (patches 6,9 and 10)
V8:
- Move mm_context_remove_copro() before mmdrop() (patch8)
- Move barrier before csb.flags store and add WARN_ON_ONCE() checks (patch9)
Haren Myneni (14):
powerpc/xive: Define xive_native_alloc_irq_on_chip()
powerpc/xive: Define xive_native_alloc_get_irq_info()
powerpc/vas: Define nx_fault_stamp in coprocessor_request_block
powerpc/vas: Alloc and setup IRQ and trigger port address
powerpc/vas: Setup fault window per VAS instance
powerpc/vas: Setup thread IRQ handler per VAS instance
powerpc/vas: Register NX with fault window ID and IRQ port value
powerpc/vas: Take reference to PID and mm for user space windows
powerpc/vas: Update CSB and notify process for fault CRBs
powerpc/vas: Print CRB and FIFO values
powerpc/vas: Do not use default credits for receive window
powerpc/vas: Return credits after handling fault
powerpc/vas: Display process stuck message
powerpc/vas: Free send window in VAS instance after credits returned
arch/powerpc/include/asm/icswx.h | 18 +-
arch/powerpc/include/asm/xive.h | 11 +-
arch/powerpc/platforms/powernv/Makefile | 2 +-
arch/powerpc/platforms/powernv/ocxl.c | 20 +-
arch/powerpc/platforms/powernv/vas-debug.c | 2 +-
arch/powerpc/platforms/powernv/vas-fault.c | 332 ++++++++++++++++++++++++++++
arch/powerpc/platforms/powernv/vas-window.c | 185 ++++++++++++++--
arch/powerpc/platforms/powernv/vas.c | 101 ++++++++-
arch/powerpc/platforms/powernv/vas.h | 51 ++++-
arch/powerpc/sysdev/xive/native.c | 29 ++-
10 files changed, 704 insertions(+), 47 deletions(-)
create mode 100644 arch/powerpc/platforms/powernv/vas-fault.c
--
1.8.3.1
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-19 06:15:41
This function allocates IRQ on a specific chip. VAS needs per chip
IRQ allocation and will have IRQ handler per VAS instance.
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/include/asm/xive.h | 9 ++++++++-
arch/powerpc/sysdev/xive/native.c | 6 +++---
2 files changed, 11 insertions(+), 4 deletions(-)
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-19 06:17:47
pnv_ocxl_alloc_xive_irq() in ocxl.c allocates IRQ and gets trigger port
address. VAS also needs this function, but based on chip ID. So moved
this common function to xive/native.c.
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/include/asm/xive.h | 2 ++
arch/powerpc/platforms/powernv/ocxl.c | 20 ++------------------
arch/powerpc/sysdev/xive/native.c | 23 +++++++++++++++++++++++
3 files changed, 27 insertions(+), 18 deletions(-)
@@ -487,24 +487,8 @@ int pnv_ocxl_spa_remove_pe_from_cache(void *platform_data, int pe_handle)intpnv_ocxl_alloc_xive_irq(u32*irq,u64*trigger_addr){-__be64flags,trigger_page;-s64rc;-u32hwirq;--hwirq=xive_native_alloc_irq();-if(!hwirq)-return-ENOENT;--rc=opal_xive_get_irq_info(hwirq,&flags,NULL,&trigger_page,NULL,-NULL);-if(rc||!trigger_page){-xive_native_free_irq(hwirq);-return-ENOENT;-}-*irq=hwirq;-*trigger_addr=be64_to_cpu(trigger_page);-return0;-+returnxive_native_alloc_get_irq_info(OPAL_XIVE_ANY_CHIP,irq,+trigger_addr);}EXPORT_SYMBOL_GPL(pnv_ocxl_alloc_xive_irq);
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-19 06:19:40
Kernel sets fault address and status in CRB for NX page fault on user
space address after processing page fault. User space gets the signal
and handles the fault mentioned in CRB by bringing the page in to
memory and send NX request again.
Signed-off-by: Sukadev Bhattiprolu <redacted>
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/include/asm/icswx.h | 18 +++++++++++++++++-
1 file changed, 17 insertions(+), 1 deletion(-)
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-19 06:21:26
Alloc IRQ and get trigger port address for each VAS instance. Kernel
register this IRQ per VAS instance and sets this port for each send
window. NX interrupts the kernel when it sees page fault.
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/platforms/powernv/vas.c | 34 ++++++++++++++++++++++++++++------
arch/powerpc/platforms/powernv/vas.h | 2 ++
2 files changed, 30 insertions(+), 6 deletions(-)
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-19 06:24:55
Setup thread IRQ handler per each VAS instance. When NX sees a fault
on CRB, kernel gets an interrupt and vas_fault_handler will be
executed to process fault CRBs. Read all valid CRBs from fault FIFO,
determine the corresponding send window from CRB and process fault
requests.
Signed-off-by: Sukadev Bhattiprolu <redacted>
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/platforms/powernv/vas-fault.c | 90 +++++++++++++++++++++++++++++
arch/powerpc/platforms/powernv/vas-window.c | 60 +++++++++++++++++++
arch/powerpc/platforms/powernv/vas.c | 49 +++++++++++++++-
arch/powerpc/platforms/powernv/vas.h | 6 ++
4 files changed, 204 insertions(+), 1 deletion(-)
@@ -1254,3 +1263,54 @@ int vas_win_close(struct vas_window *window)return0;}EXPORT_SYMBOL_GPL(vas_win_close);++structvas_window*vas_pswid_to_window(structvas_instance*vinst,+uint32_tpswid)+{+structvas_window*window;+intwinid;++if(!pswid){+pr_devel("%s: called for pswid 0!\n",__func__);+returnERR_PTR(-ESRCH);+}++decode_pswid(pswid,NULL,&winid);++if(winid>=VAS_WINDOWS_PER_CHIP)+returnERR_PTR(-ESRCH);++/*+*Ifapplicationclosesthewindowbeforethehardware+*returnsthefaultCRB,weshouldwaitinvas_win_close()+*forthependingrequests.sothewindowmustbeactive+*andtheprocessalive.+*+*Ifitsakernelprocess,weshouldnotgetanyfaultsand+*shouldnotgethere.+*/+window=vinst->windows[winid];++if(!window){+pr_err("PSWID decode: Could not find window for winid %d pswid %d vinst 0x%p\n",+winid,pswid,vinst);+returnNULL;+}++/*+*Dosomesanitychecksonthedecodedwindow.Windowshouldbe+*NXGZIPusersendwindow.FTWwindowsshouldnotincurfaults+*sincetheirCRBsareignored(notqueuedonFIFOorprocessed+*byNX).+*/+if(!window->tx_win||!window->user_win||!window->nx_win||+window->cop==VAS_COP_TYPE_FAULT||+window->cop==VAS_COP_TYPE_FTW){+pr_err("PSWID decode: id %d, tx %d, user %d, nx %d, cop %d\n",+winid,window->tx_win,window->user_win,+window->nx_win,window->cop);+WARN_ON(1);+}++returnwindow;+}
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-19 06:27:01
For each user space send window, register NX with fault window ID
and port value so that NX paste CRBs in this fault FIFO when it
sees fault on the request buffer.
Signed-off-by: Sukadev Bhattiprolu <redacted>
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/platforms/powernv/vas-window.c | 15 +++++++++++++--
arch/powerpc/platforms/powernv/vas.h | 15 +++++++++++++++
2 files changed, 28 insertions(+), 2 deletions(-)
@@ -373,7 +373,7 @@ int init_winctx_regs(struct vas_window *window, struct vas_winctx *winctx)init_xlate_regs(window,winctx->user_win);val=0ULL;-val=SET_FIELD(VAS_FAULT_TX_WIN,val,0);+val=SET_FIELD(VAS_FAULT_TX_WIN,val,winctx->fault_win_id);write_hvwc_reg(window,VREG(FAULT_TX_WIN),val);/* In PowerNV, interrupts go to HV. */
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-19 06:28:40
When process opens a window, its pid and tgid will be saved in vas_window
struct. This window will be closed when the process exits. Kernel handles
NX faults by updating CSB or send SEGV signal to pid if user space csb_addr
is invalid.
In multi-thread applications, a window can be opened by child thread, but
it will not be closed when this thread exits. Expects parent to clean up
all resources including NX windows. Child thread can send requests using
this window and can be killed before they are completed. But the pid
assigned to this thread can be reused for other task while requests are
pending. If the csb_addr passed in these requests is invalid, kernel will
end up sending signal to the wrong task.
To prevent reusing the pid, take references to pid and mm when the window
is opened and release them during window close.
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/platforms/powernv/vas-debug.c | 2 +-
arch/powerpc/platforms/powernv/vas-window.c | 53 ++++++++++++++++++++++++++---
arch/powerpc/platforms/powernv/vas.h | 9 ++++-
3 files changed, 57 insertions(+), 7 deletions(-)
@@ -1068,8 +1067,43 @@ struct vas_window *vas_tx_win_open(int vasid, enum vas_cop_type cop,gotofree_window;}-set_vinst_win(vinst,txwin);+if(txwin->user_win){+/*+*Windowopenedbychildthreadmaynotbeclosedwhen+*itexits.Sotakereferencetoitspidandreleaseit+*whenthewindowisfreebyparentthread.+*Acquireareferencetothetask'spidtomakesure+*pidwillnotbere-used-neededonlyformultithread+*applications.+*/+txwin->pid=get_task_pid(current,PIDTYPE_PID);+/*+*Acquireareferencetothetask'smm.+*/+txwin->mm=get_task_mm(current);+if(!txwin->mm){+put_pid(txwin->pid);+pr_err("VAS: pid(%d): mm_struct is not found\n",+current->pid);+rc=-EPERM;+gotofree_window;+}++mmgrab(txwin->mm);+mmput(txwin->mm);+mm_context_add_copro(txwin->mm);+/*+*Processcloseswindowduringexit.Inthecaseof+*multithreadapplication,childcanopenwindowand+*canexitwithoutclosingit.Expectsparentthread+*touseandclosethewindow.Sodonotneedtotake+*pidreferenceforparentthread.+*/+txwin->tgid=find_get_pid(task_tgid_vnr(current));+}++set_vinst_win(vinst,txwin);returntxwin;free_window:
@@ -1266,8 +1300,17 @@ int vas_win_close(struct vas_window *window)poll_window_castout(window);/* if send window, drop reference to matching receive window */-if(window->tx_win)+if(window->tx_win){+if(window->user_win){+/* Drop references to pid and mm */+put_pid(window->pid);+if(window->mm){+mm_context_remove_copro(window->mm);+mmdrop(window->mm);+}+}put_rx_win(window->rxwin);+}vas_window_free(window);
@@ -353,7 +353,9 @@ struct vas_window {booluser_win;/* True if user space window */void*hvwc_map;/* HV window context */void*uwc_map;/* OS/User window context */-pid_tpid;/* Linux process id of owner */+structpid*pid;/* Linux process id of owner */+structpid*tgid;/* Thread group ID of owner */+structmm_struct*mm;/* Linux process mm_struct */intwcreds_max;/* Window credits */char*dbgname;
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-19 06:30:04
For each fault CRB, update fault address in CRB (fault_storage_addr)
and translation error status in CSB so that user space can touch the
fault address and resend the request. If the user space passed invalid
CSB address send signal to process with SIGSEGV.
Signed-off-by: Sukadev Bhattiprolu <redacted>
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/platforms/powernv/vas-fault.c | 115 +++++++++++++++++++++++++++++
1 file changed, 115 insertions(+)
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-19 06:33:39
System checkstops if RxFIFO overruns with more requests than the
maximum possible number of CRBs allowed in FIFO at any time. So
max credits value (rxattr.wcreds_max) is set and is passed to
vas_rx_win_open() by the the driver.
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/platforms/powernv/vas-window.c | 4 ++--
arch/powerpc/platforms/powernv/vas.h | 2 --
2 files changed, 2 insertions(+), 4 deletions(-)
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-19 06:35:12
NX expects OS to return credit for send window after processing each
fault. Also credit has to be returned even for fault window.
Signed-off-by: Sukadev Bhattiprolu <redacted>
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/platforms/powernv/vas-fault.c | 9 +++++++++
arch/powerpc/platforms/powernv/vas-window.c | 17 +++++++++++++++++
arch/powerpc/platforms/powernv/vas.h | 1 +
3 files changed, 27 insertions(+)
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-19 06:37:15
Process can not close send window until all requests are processed.
Means wait until window state is not busy and send credits are
returned. Display debug messages in case taking longer to close the
window.
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/platforms/powernv/vas-window.c | 28 ++++++++++++++++++++++++++++
1 file changed, 28 insertions(+)
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-19 06:38:36
NX may be processing requests while trying to close window. Wait until
all credits are returned and then free send window from VAS instance.
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/platforms/powernv/vas-window.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
@@ -1317,14 +1317,14 @@ int vas_win_close(struct vas_window *window)unmap_paste_region(window);-clear_vinst_win(window);-poll_window_busy_state(window);unpin_close_window(window);poll_window_credits(window);+clear_vinst_win(window);+poll_window_castout(window);/* if send window, drop reference to matching receive window */
From: Nicholas Piggin <npiggin@gmail.com> Date: 2020-03-23 00:36:32
Haren Myneni's on March 19, 2020 4:13 pm:
quoted hunk
Kernel sets fault address and status in CRB for NX page fault on user
space address after processing page fault. User space gets the signal
and handles the fault mentioned in CRB by bringing the page in to
memory and send NX request again.
Signed-off-by: Sukadev Bhattiprolu <redacted>
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/include/asm/icswx.h | 18 +++++++++++++++++-
1 file changed, 17 insertions(+), 1 deletion(-)
"icswx" is not a thing anymore, after 6ff4d3e96652 ("powerpc: Remove old
unused icswx based coprocessor support"). I guess NX is reusing some
things from it, but it would be good to get rid of the cruft and re-name
this file and and relevant names.
NX already uses this file, so I guesss that can happen after this series.
Thanks,
Nick
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-23 01:00:51
On Mon, 2020-03-23 at 10:30 +1000, Nicholas Piggin wrote:
Haren Myneni's on March 19, 2020 4:13 pm:
quoted
Kernel sets fault address and status in CRB for NX page fault on user
space address after processing page fault. User space gets the signal
and handles the fault mentioned in CRB by bringing the page in to
memory and send NX request again.
Signed-off-by: Sukadev Bhattiprolu <redacted>
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/include/asm/icswx.h | 18 +++++++++++++++++-
1 file changed, 17 insertions(+), 1 deletion(-)
"icswx" is not a thing anymore, after 6ff4d3e96652 ("powerpc: Remove old
unused icswx based coprocessor support"). I guess NX is reusing some
things from it, but it would be good to get rid of the cruft and re-name
this file and and relevant names.
NX already uses this file, so I guesss that can happen after this series.
But NX uses icswx on P8 and icswx.h has only NX specific macros right
now.
From: Nicholas Piggin <npiggin@gmail.com> Date: 2020-03-23 01:12:27
Haren Myneni's on March 19, 2020 4:14 pm:
Alloc IRQ and get trigger port address for each VAS instance. Kernel
register this IRQ per VAS instance and sets this port for each send
window. NX interrupts the kernel when it sees page fault.
Again, should cc Cedric and Greg for XIVE / interrupt stuff. And
for patch 2/14.
The changelogs could use a bit of work. They're hard to read, and it can
be a bit hard to decipher "why".
Allocate a xive irq on each chip with a vas instance. The NX
coprocessor raises a host CPU interrupt via vas if it encounters a
page fault on an effective address. Subsequent patches register the
trigger port with the NX coprocessor, and create a vas fault handler
for this interrupt mapping.
Don't know if the technical details are correct, but something like that
in structure.
Thanks,
Nick
From: Nicholas Piggin <npiggin@gmail.com> Date: 2020-03-23 01:36:32
Haren Myneni's on March 23, 2020 10:57 am:
On Mon, 2020-03-23 at 10:30 +1000, Nicholas Piggin wrote:
quoted
Haren Myneni's on March 19, 2020 4:13 pm:
quoted
Kernel sets fault address and status in CRB for NX page fault on user
space address after processing page fault. User space gets the signal
and handles the fault mentioned in CRB by bringing the page in to
memory and send NX request again.
Signed-off-by: Sukadev Bhattiprolu <redacted>
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/include/asm/icswx.h | 18 +++++++++++++++++-
1 file changed, 17 insertions(+), 1 deletion(-)
"icswx" is not a thing anymore, after 6ff4d3e96652 ("powerpc: Remove old
unused icswx based coprocessor support"). I guess NX is reusing some
things from it, but it would be good to get rid of the cruft and re-name
this file and and relevant names.
NX already uses this file, so I guesss that can happen after this series.
But NX uses icswx on P8 and icswx.h has only NX specific macros right
now.
Oh P8 in kernel is still using it? Ignore me then.
Thanks,
Nick
From: Nicholas Piggin <npiggin@gmail.com> Date: 2020-03-23 02:29:12
Haren Myneni's on March 19, 2020 4:15 pm:
Setup thread IRQ handler per each VAS instance. When NX sees a fault
on CRB, kernel gets an interrupt and vas_fault_handler will be
executed to process fault CRBs. Read all valid CRBs from fault FIFO,
determine the corresponding send window from CRB and process fault
requests.
Perhaps some more overview/why.
"If NX encounters a translation error when accessing the CRB or one
of addresses in the request, it raises an interrupt on the CPU to
handle the fault.
The below comment could just be moved to replace the one at the top of the
function. Can you explain slightly more about how the faults work, and
be more clear about what the coprocessor does versus what the host does? The
use of VAS and NX is a bit confusing too. VAS doesn't interrupt with
page faults, does it? NX has the page fault(s), and it requests VAS to
interrupt the host?
+
+ /*
+ * VAS can interrupt with multiple page faults. So process all
+ * valid CRBs within fault FIFO until reaches invalid CRB.
When NX encounters a fault accessing a memory address for a particular
CRB, it updates the nx_fault_stamp field in the CRB (to what?), and
copies the CRB to the fault FIFO memory, then raises an interrupt on the
CPU (memory ordering on the store and load sides are provided how?). NX
can store multiple faults into the FIFO per interrupt (does it proceed
asynchronously after the interrupt? what's the stopping condition?).
When the CPU takes this interrupt, it reads the faulting CRBs from the
FIFO and processes them in order until it reaches an invalid entry, FIFO
empty (memory ordering how?). After each FIFO entry is processed, store
to mark them as invalid. (How does NX resume after this?)
How is the fault actually even "handled" here? Nothing seems to be
actually done for them.
+ * NX updates nx_fault_stamp in CRB and pastes in fault FIFO.
+ * kernel retrives send window from parition send window ID
+ * (pswid) in nx_fault_stamp. So pswid should be valid and
+ * ccw[0] (in be) should be zero since this bit is reserved.
+ * If user space touches this bit, NX returns with "CRB format
+ * error".
+ *
+ * After reading CRB entry, invalidate it with pswid (set
+ * 0xffffffff) and ccw[0] (set to 1).
Al this is very busy and hard to decipher unambiguously. It should read
more like a spec, a precise sequence of things happening.
+ *
+ * In case kernel receives another interrupt with different page
+ * fault, CRBs are already processed by the previous handling. So
+ * will be returned from this function when it sees invalid CRB.
+ */
Ambiguous at best. Assuming the NX continues running asynchronously and
it's a usual kind of FIFO, I assume this means if the kernel gets
another interrupt for a page fault corresponding to a FIFO entry that
has already been processed by this fault.
+ do {
Can you make this 'while (true)' or 'for (;;)' so you don't need to go
to the bottom to see it's an infinite loop.
+ mutex_lock(&vinst->mutex);
What does this protect? Threaded handlers don't run concurrently for the
same request_threaded_irq?
+
+ spin_lock_irqsave(&vinst->fault_lock, flags);
+ /*
+ * Advance the fault fifo pointer to next CRB.
The code below the comment isn't advancing the fault fifo pointer, it's
grabbing the current one. The pointer (fault_crbs) is advanced later.
You presumabl don't want to advance over an invalid entry.
+ * Use CRB_SIZE rather than sizeof(*crb) since the latter is
+ * aligned to CRB_ALIGN (256) but the CRB written to by VAS is
+ * only CRB_SIZE in len.
+ */
+ fifo = vinst->fault_fifo + (vinst->fault_crbs * CRB_SIZE);
+ entry = fifo;
Don't think you should really do this. It may be harmless in this case,
but the compiler expects the type to be aligned. Make it another type,
like coprocessor_fault_block or something?
So what does the fault_lock protect? The only data it protects is
faults_in_progress (vs the hard interrupt handler), which doesn't
achieve anything by itself, so I guess it also prevents the hard irq
handler from returning until the handler here has checked that the
fault FIFO is empty then returns IRQ_HANDLED? That seems fine (so long
as memory ordering details are okay), but it should be documented
that way.
Also why is the hard handler in a different file? Makes it harder to
see how this works at a glance.
faults_in_progress does not have to be atomic because it's always
accessed under the lock. And IMO it should have a better name. If the
NX can be causing more faults as we go, it really doesn't indicate
anything about faults. It's whether or not the threaded handler is
currently woken and processing faults.
+ mutex_unlock(&vinst->mutex);
+ return IRQ_HANDLED;
+ }
+
+ spin_unlock_irqrestore(&vinst->fault_lock, flags);
+ vinst->fault_crbs++;
+ if (vinst->fault_crbs == (vinst->fault_fifo_size / CRB_SIZE))
+ vinst->fault_crbs = 0;
+
+ memcpy(crb, fifo, CRB_SIZE);
+ entry->stamp.nx.pswid = cpu_to_be32(FIFO_INVALID_ENTRY);
+ entry->ccw |= cpu_to_be32(CCW0_INVALID);
+ mutex_unlock(&vinst->mutex);
+
+ pr_devel("VAS[%d] fault_fifo %p, fifo %p, fault_crbs %d\n",
+ vinst->vas_id, vinst->fault_fifo, fifo,
+ vinst->fault_crbs);
+
+ window = vas_pswid_to_window(vinst,
+ be32_to_cpu(crb->stamp.nx.pswid));
+
+ if (IS_ERR(window)) {
+ /*
+ * We got an interrupt about a specific send
+ * window but we can't find that window and we can't
+ * even clean it up (return credit).
+ * But we should not get here.
+ */
+ pr_err("VAS[%d] fault_fifo %p, fifo %p, pswid 0x%x, fault_crbs %d bad CRB?\n",
+ vinst->vas_id, vinst->fault_fifo, fifo,
+ be32_to_cpu(crb->stamp.nx.pswid),
+ vinst->fault_crbs);
+
+ WARN_ON_ONCE(1);
+ atomic_set(&vinst->faults_in_progress, 0);
+ return IRQ_HANDLED;
Shouldn't get here but you have a handler for it, so it should try to
be graceful. Keep processing the rest of the FIFO until it's empty
otherwise you have a missed wakeup here? Probably less code too, just
delete the last 2 lines.
Thanks,
Nick
quoted hunk
+ }
+
+ } while (true);
+}
+
+/*
* Fault window is opened per VAS instance. NX pastes fault CRB in fault
* FIFO upon page faults.
*/
@@ -1254,3 +1263,54 @@ int vas_win_close(struct vas_window *window)return0;}EXPORT_SYMBOL_GPL(vas_win_close);++structvas_window*vas_pswid_to_window(structvas_instance*vinst,+uint32_tpswid)+{+structvas_window*window;+intwinid;++if(!pswid){+pr_devel("%s: called for pswid 0!\n",__func__);+returnERR_PTR(-ESRCH);+}++decode_pswid(pswid,NULL,&winid);++if(winid>=VAS_WINDOWS_PER_CHIP)+returnERR_PTR(-ESRCH);++/*+*Ifapplicationclosesthewindowbeforethehardware+*returnsthefaultCRB,weshouldwaitinvas_win_close()+*forthependingrequests.sothewindowmustbeactive+*andtheprocessalive.+*+*Ifitsakernelprocess,weshouldnotgetanyfaultsand+*shouldnotgethere.+*/+window=vinst->windows[winid];++if(!window){+pr_err("PSWID decode: Could not find window for winid %d pswid %d vinst 0x%p\n",+winid,pswid,vinst);+returnNULL;+}++/*+*Dosomesanitychecksonthedecodedwindow.Windowshouldbe+*NXGZIPusersendwindow.FTWwindowsshouldnotincurfaults+*sincetheirCRBsareignored(notqueuedonFIFOorprocessed+*byNX).+*/+if(!window->tx_win||!window->user_win||!window->nx_win||+window->cop==VAS_COP_TYPE_FAULT||+window->cop==VAS_COP_TYPE_FTW){+pr_err("PSWID decode: id %d, tx %d, user %d, nx %d, cop %d\n",+winid,window->tx_win,window->user_win,+window->nx_win,window->cop);+WARN_ON(1);+}++returnwindow;+}
From: Nicholas Piggin <npiggin@gmail.com> Date: 2020-03-23 02:39:57
Haren Myneni's on March 19, 2020 4:16 pm:
When process opens a window, its pid and tgid will be saved in vas_window
struct. This window will be closed when the process exits. Kernel handles
NX faults by updating CSB or send SEGV signal to pid if user space csb_addr
is invalid.
Bit of a nitpick, but can you use articles consistently ("the", "a")? I
won't keep nitpicking changelogs but I think they could be made easier
to read. I'm happy to help proof read and suggest things offline when
you're happy with the technical content of them, let me know.
In multi-thread applications, a window can be opened by child thread, but
it will not be closed when this thread exits. Expects parent to clean up
all resources including NX windows. Child thread can send requests using
this window and can be killed before they are completed. But the pid
assigned to this thread can be reused for other task while requests are
pending. If the csb_addr passed in these requests is invalid, kernel will
end up sending signal to the wrong task.
To prevent reusing the pid, take references to pid and mm when the window
is opened and release them during window close.
We went over this together a while back, but task management isn't
something I look at every day and it's complicated and easy to introduce
bugs. I suggest if we can get the changelog and comments written well
and understandable for someone who does not know or care about vas,
then cc linux-kernel and the maintainers, and hopefully someone will
take a look. It's not a large patch so if assumptions and concurrency
etc is documented, then it shouldn't be too much work.
Thanks,
Nick
@@ -1068,8 +1067,43 @@ struct vas_window *vas_tx_win_open(int vasid, enum vas_cop_type cop,gotofree_window;}-set_vinst_win(vinst,txwin);+if(txwin->user_win){+/*+*Windowopenedbychildthreadmaynotbeclosedwhen+*itexits.Sotakereferencetoitspidandreleaseit+*whenthewindowisfreebyparentthread.+*Acquireareferencetothetask'spidtomakesure+*pidwillnotbere-used-neededonlyformultithread+*applications.+*/+txwin->pid=get_task_pid(current,PIDTYPE_PID);+/*+*Acquireareferencetothetask'smm.+*/+txwin->mm=get_task_mm(current);+if(!txwin->mm){+put_pid(txwin->pid);+pr_err("VAS: pid(%d): mm_struct is not found\n",+current->pid);+rc=-EPERM;+gotofree_window;+}++mmgrab(txwin->mm);+mmput(txwin->mm);+mm_context_add_copro(txwin->mm);+/*+*Processcloseswindowduringexit.Inthecaseof+*multithreadapplication,childcanopenwindowand+*canexitwithoutclosingit.Expectsparentthread+*touseandclosethewindow.Sodonotneedtotake+*pidreferenceforparentthread.+*/+txwin->tgid=find_get_pid(task_tgid_vnr(current));+}++set_vinst_win(vinst,txwin);returntxwin;free_window:
@@ -1266,8 +1300,17 @@ int vas_win_close(struct vas_window *window)poll_window_castout(window);/* if send window, drop reference to matching receive window */-if(window->tx_win)+if(window->tx_win){+if(window->user_win){+/* Drop references to pid and mm */+put_pid(window->pid);+if(window->mm){+mm_context_remove_copro(window->mm);+mmdrop(window->mm);+}+}put_rx_win(window->rxwin);+}vas_window_free(window);
@@ -353,7 +353,9 @@ struct vas_window {booluser_win;/* True if user space window */void*hvwc_map;/* HV window context */void*uwc_map;/* OS/User window context */-pid_tpid;/* Linux process id of owner */+structpid*pid;/* Linux process id of owner */+structpid*tgid;/* Thread group ID of owner */+structmm_struct*mm;/* Linux process mm_struct */intwcreds_max;/* Window credits */char*dbgname;
From: Nicholas Piggin <npiggin@gmail.com> Date: 2020-03-23 02:43:05
Haren Myneni's on March 19, 2020 4:17 pm:
For each fault CRB, update fault address in CRB (fault_storage_addr)
and translation error status in CSB so that user space can touch the
fault address and resend the request. If the user space passed invalid
CSB address send signal to process with SIGSEGV.
This is where the actual fault handling is done? Does this need to be
split from the other patch? Why not merge them and put it after the
reference counting one?
I'll wait until comments and questions on the first fault handling patch
are resolved before I look at this one.
Thanks,
Nick
From: Nicholas Piggin <npiggin@gmail.com> Date: 2020-03-23 02:46:40
Haren Myneni's on March 19, 2020 4:18 pm:
System checkstops if RxFIFO overruns with more requests than the
maximum possible number of CRBs allowed in FIFO at any time. So
max credits value (rxattr.wcreds_max) is set and is passed to
vas_rx_win_open() by the the driver.
This seems like it should be a bug fix or merged in the NX fault
window register patch or something.
Thanks,
Nick
Any chance of a little bit of explanation how the credit system works?
Or is it in the code somewhere already?
I don't suppose there is a chance to batch credit updates with multiple
faults? (maybe the MMIO is insignificant)
From: Cédric Le Goater <clg@kaod.org> Date: 2020-03-23 08:51:35
On 3/19/20 7:12 AM, Haren Myneni wrote:
This function allocates IRQ on a specific chip. VAS needs per chip
IRQ allocation and will have IRQ handler per VAS instance.
The pool of generic interrupt source (IPI) numbers is generally used
by user space application which generally do not care on which chip
the interrupt is allocated. It's used by the CXL driver and KVM for
the guest interrupts. The CPU IPI are the exceptions.
The underlying FW call will try to allocate on the chip of the CPU
first and then on the others. If you specify a chip id, there is no
fallback. Is it what you want ?
Why do you need to allocate a generic interrupt source (IPI) from
a specific chip ? Is it a VAS requirement ?
Could you explain a bit more how it is used because there might be
similar request.
The code is fine.
Thanks,
C.
From: Cédric Le Goater <clg@kaod.org> Date: 2020-03-23 08:55:32
On 3/19/20 7:13 AM, Haren Myneni wrote:
pnv_ocxl_alloc_xive_irq() in ocxl.c allocates IRQ and gets trigger port
address. VAS also needs this function, but based on chip ID. So moved
this common function to xive/native.c.
We now have two drivers using the lowlevel routines of the machine
irqchip driver. I am not sure OCXL is doing the right thing by calling
opal_xive_get_irq_info() and not xive_native_populate_irq_data().
C.
@@ -487,24 +487,8 @@ int pnv_ocxl_spa_remove_pe_from_cache(void *platform_data, int pe_handle)intpnv_ocxl_alloc_xive_irq(u32*irq,u64*trigger_addr){-__be64flags,trigger_page;-s64rc;-u32hwirq;--hwirq=xive_native_alloc_irq();-if(!hwirq)-return-ENOENT;--rc=opal_xive_get_irq_info(hwirq,&flags,NULL,&trigger_page,NULL,-NULL);-if(rc||!trigger_page){-xive_native_free_irq(hwirq);-return-ENOENT;-}-*irq=hwirq;-*trigger_addr=be64_to_cpu(trigger_page);-return0;-+returnxive_native_alloc_get_irq_info(OPAL_XIVE_ANY_CHIP,irq,+trigger_addr);}EXPORT_SYMBOL_GPL(pnv_ocxl_alloc_xive_irq);
From: Cédric Le Goater <clg@kaod.org> Date: 2020-03-23 09:03:35
On 3/19/20 7:08 AM, Haren Myneni wrote:
On power9, Virtual Accelerator Switchboard (VAS) allows user space or
kernel to communicate with Nest Accelerator (NX) directly using COPY/PASTE
instructions. NX provides various functionalities such as compression,
encryption and etc. But only compression (842 and GZIP formats) is
supported in Linux kernel on power9.
842 compression driver (drivers/crypto/nx/nx-842-powernv.c)
is already included in Linux. Only GZIP support will be available from
user space.
Applications can issue GZIP compression / decompression requests to NX with
COPY/PASTE instructions. When NX is processing these requests, can hit
fault on the request buffer (not in memory). It issues an interrupt and
pastes fault CRB in fault FIFO. Expects kernel to handle this fault and
return credits for both send and fault windows after processing.
This patch series adds IRQ and fault window setup, and NX fault handling:
- Alloc IRQ and trigger port address, and configure IRQ per VAS instance.
Is the model similar to OCXL ?
If so, I suppose that the IRQ for fault handling is allocated by skiboot,
exposed in the DT and automatically mapped in Linux when the driver is
loaded.
Are there other interrupts ? Such as for job completion ?
Thanks,
C.
- Set port# for each window to generate an interrupt when noticed fault.
- Set fault window and FIFO on which NX paste fault CRB.
- Setup IRQ thread fault handler per VAS instance.
- When receiving an interrupt, Read CRBs from fault FIFO and update
coprocessor_status_block (CSB) in the corresponding CRB with translation
failure (CSB_CC_TRANSLATION). After issuing NX requests, process polls
on CSB address. When it sees translation error, can touch the request
buffer to bring the page in to memory and reissue NX request.
- If copy_to_user fails on user space CSB address, OS sends SEGV signal.
Tested these patches with NX-GZIP support and will be posting this series
soon.
Patches 1 & 2: Define alloc IRQ and get port address per chip which are needed
to alloc IRQ per VAS instance.
Patch 3: Define nx_fault_stamp on which NX writes fault status for the fault
CRB
Patch 4: Alloc and setup IRQ and trigger port address for each VAS instance
Patch 5: Setup fault window per each VAS instance. This window is used for
NX to paste fault CRB in FIFO.
Patches 6 & 7: Setup threaded IRQ per VAS and register NX with fault window
ID and port number for each send window so that NX paste fault CRB
in this window.
Patch 8: Reference to pid and mm so that pid is not used until window closed.
Needed for multi thread application where child can open a window
and can be used by parent later.
Patches 9 and 10: Process CRBs from fault FIFO and notify tasks by
updating CSB or through signals.
Patches 11 and 12: Return credits for send and fault windows after handling
faults.
Patch 14:Fix closing send window after all credits are returned. This issue
happens only for user space requests. No page faults on kernel
request buffer.
Changelog:
V2:
- Use threaded IRQ instead of own kernel thread handler
- Use pswid instead of user space CSB address to find valid CRB
- Removed unused macros and other changes as suggested by Christoph Hellwig
V3:
- Rebased to 5.5-rc2
- Use struct pid * instead of pid_t for vas_window tgid
- Code cleanup as suggested by Christoph Hellwig
V4:
- Define xive alloc and get IRQ info based on chip ID and use these
functions for IRQ setup per VAS instance. It eliminates skiboot
dependency as suggested by Oliver.
V5:
- Do not update CSB if the process is exiting (patch9)
V6:
- Add interrupt handler instead of default one and return IRQ_HANDLED
if the fault handling thread is already in progress. (Patch6)
- Use platform send window ID and CCW[0] bit to find valid CRB in
fault FIFO (Patch6).
- Return fault address to user space in BE and other changes as
suggested by Michael Neuling. (patch9)
- Rebased to 5.6-rc4
V7:
- Fix sparse warnings (patches 6,9 and 10)
V8:
- Move mm_context_remove_copro() before mmdrop() (patch8)
- Move barrier before csb.flags store and add WARN_ON_ONCE() checks (patch9)
Haren Myneni (14):
powerpc/xive: Define xive_native_alloc_irq_on_chip()
powerpc/xive: Define xive_native_alloc_get_irq_info()
powerpc/vas: Define nx_fault_stamp in coprocessor_request_block
powerpc/vas: Alloc and setup IRQ and trigger port address
powerpc/vas: Setup fault window per VAS instance
powerpc/vas: Setup thread IRQ handler per VAS instance
powerpc/vas: Register NX with fault window ID and IRQ port value
powerpc/vas: Take reference to PID and mm for user space windows
powerpc/vas: Update CSB and notify process for fault CRBs
powerpc/vas: Print CRB and FIFO values
powerpc/vas: Do not use default credits for receive window
powerpc/vas: Return credits after handling fault
powerpc/vas: Display process stuck message
powerpc/vas: Free send window in VAS instance after credits returned
arch/powerpc/include/asm/icswx.h | 18 +-
arch/powerpc/include/asm/xive.h | 11 +-
arch/powerpc/platforms/powernv/Makefile | 2 +-
arch/powerpc/platforms/powernv/ocxl.c | 20 +-
arch/powerpc/platforms/powernv/vas-debug.c | 2 +-
arch/powerpc/platforms/powernv/vas-fault.c | 332 ++++++++++++++++++++++++++++
arch/powerpc/platforms/powernv/vas-window.c | 185 ++++++++++++++--
arch/powerpc/platforms/powernv/vas.c | 101 ++++++++-
arch/powerpc/platforms/powernv/vas.h | 51 ++++-
arch/powerpc/sysdev/xive/native.c | 29 ++-
10 files changed, 704 insertions(+), 47 deletions(-)
create mode 100644 arch/powerpc/platforms/powernv/vas-fault.c
From: Cédric Le Goater <clg@kaod.org> Date: 2020-03-23 09:26:35
On 3/19/20 7:14 AM, Haren Myneni wrote:
Alloc IRQ and get trigger port address for each VAS instance. Kernel
register this IRQ per VAS instance and sets this port for each send
window. NX interrupts the kernel when it sees page fault.
I don't understand why this is not done by the OPAL driver for each VAS
of the system. Is the VAS unit very different from OpenCAPI regarding
the fault ?
C.
From: Michael Ellerman <mpe@ellerman.id.au> Date: 2020-03-23 11:34:51
Nicholas Piggin [off-list ref] writes:
Haren Myneni's on March 19, 2020 4:13 pm:
quoted
Kernel sets fault address and status in CRB for NX page fault on user
space address after processing page fault. User space gets the signal
and handles the fault mentioned in CRB by bringing the page in to
memory and send NX request again.
Signed-off-by: Sukadev Bhattiprolu <redacted>
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/include/asm/icswx.h | 18 +++++++++++++++++-
1 file changed, 17 insertions(+), 1 deletion(-)
"icswx" is not a thing anymore, after 6ff4d3e96652 ("powerpc: Remove old
unused icswx based coprocessor support").
Yeah that commit ripped out some parts of the previous attempt at a user
visible API for this sort of "coprocessor" stuff. VAS is yet another
attempt to do something useful with most of the same pieces but some
slightly different details.
I guess NX is reusing some
things from it, but it would be good to get rid of the cruft and re-name
this file and and relevant names.
NX already uses this file, so I guesss that can happen after this series.
A lot of the CRB/CSB stuff is still the same, and P8 still uses icswx.
But I'd be happy if the header was renamed eventually, as icswx is now a
legacy name.
cheers
From: Cédric Le Goater <clg@kaod.org> Date: 2020-03-23 11:59:02
On 3/23/20 10:06 AM, Cédric Le Goater wrote:
On 3/19/20 7:14 AM, Haren Myneni wrote:
quoted
Alloc IRQ and get trigger port address for each VAS instance. Kernel
register this IRQ per VAS instance and sets this port for each send
window. NX interrupts the kernel when it sees page fault.
I don't understand why this is not done by the OPAL driver for each VAS
of the system. Is the VAS unit very different from OpenCAPI regarding
the fault ?
I checked the previous patchsets and I see that v3 was more like I expected
it: one interrupt for faults allocated by the skiboot driver and exposed
in the DT.
What made you change your mind ?
This version is hijacking the lowlevel routines of the XIVE irqchip which
is not the best approach. OCXL is doing that because it needs to allocate
interrupts for the user space processes using the AFU and we should rework
that part.
However, the translation fault interrupt is allocated by skiboot.
Sorry for the noise, I would like to understand more how this works. I also
have passthrough in mind.
C.
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-23 18:19:13
On Mon, 2020-03-23 at 22:32 +1100, Michael Ellerman wrote:
Nicholas Piggin [off-list ref] writes:
quoted
Haren Myneni's on March 19, 2020 4:13 pm:
quoted
Kernel sets fault address and status in CRB for NX page fault on user
space address after processing page fault. User space gets the signal
and handles the fault mentioned in CRB by bringing the page in to
memory and send NX request again.
Signed-off-by: Sukadev Bhattiprolu <redacted>
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/include/asm/icswx.h | 18 +++++++++++++++++-
1 file changed, 17 insertions(+), 1 deletion(-)
"icswx" is not a thing anymore, after 6ff4d3e96652 ("powerpc: Remove old
unused icswx based coprocessor support").
Yeah that commit ripped out some parts of the previous attempt at a user
visible API for this sort of "coprocessor" stuff. VAS is yet another
attempt to do something useful with most of the same pieces but some
slightly different details.
quoted
I guess NX is reusing some
things from it, but it would be good to get rid of the cruft and re-name
this file and and relevant names.
quoted
NX already uses this file, so I guesss that can happen after this series.
A lot of the CRB/CSB stuff is still the same, and P8 still uses icswx.
But I'd be happy if the header was renamed eventually, as icswx is now a
legacy name.
We can move all macros and struct definitions to vas.h and remove
icswx.h. Can I do this after this series?
Thanks
Haren
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-23 19:06:24
On Mon, 2020-03-23 at 10:27 +0100, Cédric Le Goater wrote:
On 3/23/20 10:06 AM, Cédric Le Goater wrote:
quoted
On 3/19/20 7:14 AM, Haren Myneni wrote:
quoted
Alloc IRQ and get trigger port address for each VAS instance. Kernel
register this IRQ per VAS instance and sets this port for each send
window. NX interrupts the kernel when it sees page fault.
I don't understand why this is not done by the OPAL driver for each VAS
of the system. Is the VAS unit very different from OpenCAPI regarding
the fault ?
I checked the previous patchsets and I see that v3 was more like I expected
it: one interrupt for faults allocated by the skiboot driver and exposed
in the DT.
What made you change your mind ?
This version is hijacking the lowlevel routines of the XIVE irqchip which
is not the best approach. OCXL is doing that because it needs to allocate
interrupts for the user space processes using the AFU and we should rework
that part.
However, the translation fault interrupt is allocated by skiboot.
Sorry my mistake. I should have CC you earlier.
Each VAS instance will generate fault interrupt which is per chip. There
won't be other job completion interrupts.
Correct, V3 used allocating interrupts per chip in skiboot and exposed
in DT. Since XIVE code has similar feature, exploited this approach so
that we do not need skiboot changes.
Thanks
Haren
Sorry for the noise, I would like to understand more how this works. I also
have passthrough in mind.
C.
On Mon, Mar 23, 2020 at 8:28 PM Cédric Le Goater [off-list ref] wrote:
On 3/23/20 10:06 AM, Cédric Le Goater wrote:
quoted
On 3/19/20 7:14 AM, Haren Myneni wrote:
quoted
Alloc IRQ and get trigger port address for each VAS instance. Kernel
register this IRQ per VAS instance and sets this port for each send
window. NX interrupts the kernel when it sees page fault.
I don't understand why this is not done by the OPAL driver for each VAS
of the system. Is the VAS unit very different from OpenCAPI regarding
the fault ?
I checked the previous patchsets and I see that v3 was more like I expected
it: one interrupt for faults allocated by the skiboot driver and exposed
in the DT.
What made you change your mind ?
From init_vas_inst() in arch/powerpc/platforms/powernv/vas.c:
if (pdev->num_resources != 4) {
pr_err("Unexpected DT configuration for [%s, %d]\n",
pdev->name, vasid);
return -ENODEV;
}
This code should never have been written, but here we are. Due to the
above adding an interrupt in the DT makes the driver unable to bind on
older kernels. In an older version of the patches (don't think it was
posted) Haren was using a non-standard interrupt property and we could
work around the problem by going back to that.
However, we already have the OPAL calls for allocating / freeing
hardware interrupt numbers so why not do that? If we ever want to take
advantage of the job completion interrupts we'd want to have the
ability to allocate them since the completion interrupts are
per-window rather than per-VAS.
This version is hijacking the lowlevel routines of the XIVE irqchip which
is not the best approach. OCXL is doing that because it needs to allocate
interrupts for the user space processes using the AFU and we should rework
that part.
What'd you have in mind for the reworking the oxcl interrupt
allocation? I didn't find it that objectionable since it's more or
less the same as what happens when allocating IPIs.
However, the translation fault interrupt is allocated by skiboot.
Sorry for the noise, I would like to understand more how this works. I also
have passthrough in mind.
C.
From: Cédric Le Goater <clg@kaod.org> Date: 2020-03-24 12:29:24
On 3/23/20 8:02 PM, Haren Myneni wrote:
On Mon, 2020-03-23 at 10:27 +0100, Cédric Le Goater wrote:
quoted
On 3/23/20 10:06 AM, Cédric Le Goater wrote:
quoted
On 3/19/20 7:14 AM, Haren Myneni wrote:
quoted
Alloc IRQ and get trigger port address for each VAS instance. Kernel
register this IRQ per VAS instance and sets this port for each send
window. NX interrupts the kernel when it sees page fault.
I don't understand why this is not done by the OPAL driver for each VAS
of the system. Is the VAS unit very different from OpenCAPI regarding
the fault ?
I checked the previous patchsets and I see that v3 was more like I expected
it: one interrupt for faults allocated by the skiboot driver and exposed
in the DT.
What made you change your mind ?
This version is hijacking the lowlevel routines of the XIVE irqchip which
is not the best approach. OCXL is doing that because it needs to allocate
interrupts for the user space processes using the AFU and we should rework
that part.
However, the translation fault interrupt is allocated by skiboot.
Sorry my mistake. I should have CC you earlier.
Each VAS instance will generate fault interrupt which is per chip. There
won't be other job completion interrupts.
That's a very good reason to set everything in the skiboot driver and
advertise the interrupt number in the DT. The interrupt will be mapped
automatically by OF routines and the driver will only have to install
an interrupt handler.
Correct, V3 used allocating interrupts per chip in skiboot and exposed
in DT. Since XIVE code has similar feature, exploited this approach so
that we do not need skiboot changes.
It's not the same. These are the low level (OPAL) interface used by the
XIVE driver. The exception is the KVM XIVE device which needs a finer
grain.
C.
From: Cédric Le Goater <clg@kaod.org> Date: 2020-03-24 13:55:28
On 3/24/20 3:26 AM, Oliver O'Halloran wrote:
On Mon, Mar 23, 2020 at 8:28 PM Cédric Le Goater [off-list ref] wrote:
quoted
On 3/23/20 10:06 AM, Cédric Le Goater wrote:
quoted
On 3/19/20 7:14 AM, Haren Myneni wrote:
quoted
Alloc IRQ and get trigger port address for each VAS instance. Kernel
register this IRQ per VAS instance and sets this port for each send
window. NX interrupts the kernel when it sees page fault.
I don't understand why this is not done by the OPAL driver for each VAS
of the system. Is the VAS unit very different from OpenCAPI regarding
the fault ?
I checked the previous patchsets and I see that v3 was more like I expected
it: one interrupt for faults allocated by the skiboot driver and exposed
in the DT.
What made you change your mind ?
From init_vas_inst() in arch/powerpc/platforms/powernv/vas.c:
if (pdev->num_resources != 4) {
pr_err("Unexpected DT configuration for [%s, %d]\n",
pdev->name, vasid);
return -ENODEV;
}
This code should never have been written, but here we are. Due to the
above adding an interrupt in the DT makes the driver unable to bind on
older kernels. In an older version of the patches (don't think it was
posted) Haren was using a non-standard interrupt property and we could
work around the problem by going back to that.
ok ... :/ I didn't know. Don't we have a rule on LinuxPPC for such
things ? Such as, the culprit should send a croissant to everyone
involved.
However, we already have the OPAL calls for allocating / freeing
hardware interrupt numbers so why not do that?
It's a good way to work around the problem but we are bypassing the
irqchip which does other things for the driver.
If we ever want to take
advantage of the job completion interrupts we'd want to have the
ability to allocate them since the completion interrupts are
per-window rather than per-VAS.
Yes. That's what I thought it was about to begin with. OCXL has a
first implementation of such interrupts.
quoted
This version is hijacking the lowlevel routines of the XIVE irqchip which
is not the best approach. OCXL is doing that because it needs to allocate
interrupts for the user space processes using the AFU and we should rework
that part.
What'd you have in mind for the reworking the oxcl interrupt allocation?
I didn't find it that objectionable since it's more or less the same as
what happens when allocating IPIs.
I think we need to work a bit more on the concepts, on the interfaces,
internal at the platform kernel level and at the user space level, and
on the configuration, with chip affinity in mind. There are bunch of
information on the sources that are retrieved from the firmware or
hypervisor that we care about. An irqchip might be the best option
for the moment.
At the same time, it would be good to keep in mind user interrupts.
C.
quoted
However, the translation fault interrupt is allocated by skiboot.
Sorry for the noise, I would like to understand more how this works. I also
have passthrough in mind.
C.
From: Cédric Le Goater <clg@kaod.org> Date: 2020-03-24 14:08:41
On 3/19/20 7:13 AM, Haren Myneni wrote:
pnv_ocxl_alloc_xive_irq() in ocxl.c allocates IRQ and gets trigger port
address. VAS also needs this function, but based on chip ID. So moved
this common function to xive/native.c.
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
I think we should work on a new interface for generic IPI use.
This is a beginning.
Reviewed-by: Cédric Le Goater <clg@kaod.org>
Thanks,
C.
@@ -487,24 +487,8 @@ int pnv_ocxl_spa_remove_pe_from_cache(void *platform_data, int pe_handle)intpnv_ocxl_alloc_xive_irq(u32*irq,u64*trigger_addr){-__be64flags,trigger_page;-s64rc;-u32hwirq;--hwirq=xive_native_alloc_irq();-if(!hwirq)-return-ENOENT;--rc=opal_xive_get_irq_info(hwirq,&flags,NULL,&trigger_page,NULL,-NULL);-if(rc||!trigger_page){-xive_native_free_irq(hwirq);-return-ENOENT;-}-*irq=hwirq;-*trigger_addr=be64_to_cpu(trigger_page);-return0;-+returnxive_native_alloc_get_irq_info(OPAL_XIVE_ANY_CHIP,irq,+trigger_addr);
From: Cédric Le Goater <clg@kaod.org> Date: 2020-03-24 14:32:14
On 3/19/20 7:12 AM, Haren Myneni wrote:
This function allocates IRQ on a specific chip. VAS needs per chip
IRQ allocation and will have IRQ handler per VAS instance.
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
Reviewed-by: Cédric Le Goater <clg@kaod.org>
Thanks,
C.
From: Cédric Le Goater <clg@kaod.org> Date: 2020-03-24 15:10:00
On 3/19/20 7:14 AM, Haren Myneni wrote:
quoted hunk
Alloc IRQ and get trigger port address for each VAS instance. Kernel
register this IRQ per VAS instance and sets this port for each send
window. NX interrupts the kernel when it sees page fault.
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/platforms/powernv/vas.c | 34 ++++++++++++++++++++++++++++------
arch/powerpc/platforms/powernv/vas.h | 2 ++
2 files changed, 30 insertions(+), 6 deletions(-)
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-24 21:09:28
On Tue, 2020-03-24 at 15:48 +0100, Cédric Le Goater wrote:
On 3/19/20 7:14 AM, Haren Myneni wrote:
quoted
Alloc IRQ and get trigger port address for each VAS instance. Kernel
register this IRQ per VAS instance and sets this port for each send
window. NX interrupts the kernel when it sees page fault.
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/platforms/powernv/vas.c | 34 ++++++++++++++++++++++++++++------
arch/powerpc/platforms/powernv/vas.h | 2 ++
2 files changed, 30 insertions(+), 6 deletions(-)
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-25 03:02:41
On Mon, 2020-03-23 at 12:23 +1000, Nicholas Piggin wrote:
Haren Myneni's on March 19, 2020 4:15 pm:
quoted
Setup thread IRQ handler per each VAS instance. When NX sees a fault
on CRB, kernel gets an interrupt and vas_fault_handler will be
executed to process fault CRBs. Read all valid CRBs from fault FIFO,
determine the corresponding send window from CRB and process fault
requests.
Perhaps some more overview/why.
"If NX encounters a translation error when accessing the CRB or one
of addresses in the request, it raises an interrupt on the CPU to
handle the fault.
Are page faults the only reason why VAS would raise this interrupt? Is
NX really the only possible user of this, so you can have NX specifics
in here?
Yes, When NX sees page faults, it generates interrupts on specific VAS
instance. Right now NX is the only user. So trying to make as
generalized as possible.
The below comment could just be moved to replace the one at the top of the
function. Can you explain slightly more about how the faults work, and
be more clear about what the coprocessor does versus what the host does? The
use of VAS and NX is a bit confusing too. VAS doesn't interrupt with
page faults, does it? NX has the page fault(s), and it requests VAS to
interrupt the host?
NX is the one who raises interrupt on specific VAS port that is
registered with fault window. I will make it clear in the comment.
quoted
+
+ /*
+ * VAS can interrupt with multiple page faults. So process all
+ * valid CRBs within fault FIFO until reaches invalid CRB.
When NX encounters a fault accessing a memory address for a particular
CRB, it updates the nx_fault_stamp field in the CRB (to what?), and
copies the CRB to the fault FIFO memory, then raises an interrupt on the
CPU (memory ordering on the store and load sides are provided how?). NX
can store multiple faults into the FIFO per interrupt (does it proceed
asynchronously after the interrupt? what's the stopping condition?).
User space fills CRB and sends request (CRB). NX processes the request
and update CSB (in CRB struct). If NX sees any page fault either on
request buffers or csb address, updates nx_fault_stamp struct (in user
space CRB) and pastes CRB in fault_fifo. Then raises interrupt on port
defined in fault_window (which is per VAS instance).
NX can raise single interrupt for multiple faults and can paste in fault
FIFO. But OS and NX use credits to control fault FIFO.
Initially FIFO_SIZE/CRB_SIZE credits are available for fault window. For
example, When NX pastes CRB in fault fifo, credits will be reduced by 1.
It can continue paste CRBs in FIFO until credits reached to 0.
When OS handles the fault CRB, increments index in FIFO and returns the
credit so that NX knows one more CRB slot is available.
struct coprocessor_request_block {
__be32 ccw;
__be32 flags;
__be64 csb_addr;
struct data_descriptor_entry source;
struct data_descriptor_entry target;
struct coprocessor_completion_block ccb;
union {
struct nx_fault_stamp nx;
u8 reserved[16];
} stamp;
u8 reserved[32];
struct coprocessor_status_block csb;
} __packed __aligned(CRB_ALIGN);
When the CPU takes this interrupt, it reads the faulting CRBs from the
FIFO and processes them in order until it reaches an invalid entry, FIFO
empty (memory ordering how?). After each FIFO entry is processed, store
to mark them as invalid. (How does NX resume after this?)
NX should do atomic copy of CRB in fault FIFO. But we had barrier (as
part of spin_unlock()) for the safe side which is suggested by HW team.
NX stops pasting CRBs if credits are not available and start when credit
is returned by OS after handling fault.
How is the fault actually even "handled" here? Nothing seems to be
actually done for them.
quoted
+ * NX updates nx_fault_stamp in CRB and pastes in fault FIFO.
+ * kernel retrives send window from parition send window ID
+ * (pswid) in nx_fault_stamp. So pswid should be valid and
+ * ccw[0] (in be) should be zero since this bit is reserved.
+ * If user space touches this bit, NX returns with "CRB format
+ * error".
+ *
+ * After reading CRB entry, invalidate it with pswid (set
+ * 0xffffffff) and ccw[0] (set to 1).
Al this is very busy and hard to decipher unambiguously. It should read
more like a spec, a precise sequence of things happening.
Sure, will make it clear
quoted
+ *
+ * In case kernel receives another interrupt with different page
+ * fault, CRBs are already processed by the previous handling. So
+ * will be returned from this function when it sees invalid CRB.
+ */
Ambiguous at best. Assuming the NX continues running asynchronously and
it's a usual kind of FIFO, I assume this means if the kernel gets
another interrupt for a page fault corresponding to a FIFO entry that
has already been processed by this fault.
quoted
+ do {
Can you make this 'while (true)' or 'for (;;)' so you don't need to go
to the bottom to see it's an infinite loop.
quoted
+ mutex_lock(&vinst->mutex);
What does this protect? Threaded handlers don't run concurrently for the
same request_threaded_irq?
We can remove this mutex_lock.
quoted
+
+ spin_lock_irqsave(&vinst->fault_lock, flags);
+ /*
+ * Advance the fault fifo pointer to next CRB.
The code below the comment isn't advancing the fault fifo pointer, it's
grabbing the current one. The pointer (fault_crbs) is advanced later.
You presumabl don't want to advance over an invalid entry.
Advancing to next entry which is invalid means start processing from
this entry in the next fault.
quoted
+ * Use CRB_SIZE rather than sizeof(*crb) since the latter is
+ * aligned to CRB_ALIGN (256) but the CRB written to by VAS is
+ * only CRB_SIZE in len.
+ */
+ fifo = vinst->fault_fifo + (vinst->fault_crbs * CRB_SIZE);
+ entry = fifo;
Don't think you should really do this. It may be harmless in this case,
but the compiler expects the type to be aligned. Make it another type,
like coprocessor_fault_block or something?
Can add like "entry = (struct coprocessor_request_block *)fifo;
So what does the fault_lock protect? The only data it protects is
faults_in_progress (vs the hard interrupt handler), which doesn't
achieve anything by itself, so I guess it also prevents the hard irq
handler from returning until the handler here has checked that the
fault FIFO is empty then returns IRQ_HANDLED? That seems fine (so long
as memory ordering details are okay), but it should be documented
that way.
In the case of using default_handler, wakes up thread if it is not in
progress and checks whether interrupts are handled with in some
duration. If un_handled interrupts reached 99000, display bad_irq trace
and disables IRQ. In our case we do not have one fault per interrupt.
We ran some test case which continuously generates NX faults, the
handler thread is busy processing fault CRBs in FIFO and the later
interrupts are not handled.
So added own handler which checks whether fault_thread is in progress.
If so returned IRQ_HANDLED. fault_lock is used to check valid entry
section and check faults_in_progress in handler.
Added comment in vas_fault_handler().
Also why is the hard handler in a different file? Makes it harder to
see how this works at a glance.
OK, Will change. I thought IRQ handler per VAS is added in vas.c since
it has VAS initialization and vas_fault.c is only for fault handling.
faults_in_progress does not have to be atomic because it's always
accessed under the lock. And IMO it should have a better name. If the
NX can be causing more faults as we go, it really doesn't indicate
anything about faults. It's whether or not the threaded handler is
currently woken and processing faults.
Used atomic since not using spin_lock/unlock for failing case when
window (from CRB) is not valid.
Sure, How about fault_thread_in_progress? As it is long name, used
faults_in_progress.
quoted
+ mutex_unlock(&vinst->mutex);
+ return IRQ_HANDLED;
+ }
+
+ spin_unlock_irqrestore(&vinst->fault_lock, flags);
+ vinst->fault_crbs++;
+ if (vinst->fault_crbs == (vinst->fault_fifo_size / CRB_SIZE))
+ vinst->fault_crbs = 0;
+
+ memcpy(crb, fifo, CRB_SIZE);
+ entry->stamp.nx.pswid = cpu_to_be32(FIFO_INVALID_ENTRY);
+ entry->ccw |= cpu_to_be32(CCW0_INVALID);
+ mutex_unlock(&vinst->mutex);
+
+ pr_devel("VAS[%d] fault_fifo %p, fifo %p, fault_crbs %d\n",
+ vinst->vas_id, vinst->fault_fifo, fifo,
+ vinst->fault_crbs);
+
+ window = vas_pswid_to_window(vinst,
+ be32_to_cpu(crb->stamp.nx.pswid));
+
+ if (IS_ERR(window)) {
+ /*
+ * We got an interrupt about a specific send
+ * window but we can't find that window and we can't
+ * even clean it up (return credit).
+ * But we should not get here.
+ */
+ pr_err("VAS[%d] fault_fifo %p, fifo %p, pswid 0x%x, fault_crbs %d bad CRB?\n",
+ vinst->vas_id, vinst->fault_fifo, fifo,
+ be32_to_cpu(crb->stamp.nx.pswid),
+ vinst->fault_crbs);
+
+ WARN_ON_ONCE(1);
+ atomic_set(&vinst->faults_in_progress, 0);
+ return IRQ_HANDLED;
Shouldn't get here but you have a handler for it, so it should try to
be graceful. Keep processing the rest of the FIFO until it's empty
otherwise you have a missed wakeup here? Probably less code too, just
delete the last 2 lines.
Derive window address from pswid which is pasted by NX. So if window
address is not valid means a bug, we should not be reached this. If
getting this failure, printed data in FIFO for around 10 CRBs. Not sure
whether proceeding with this failure. We may end up receiving lots of
messages on the console if we see similar failures in later CRBs.
Thinking may be disable IRQ so that kernel will not receive any
interrupts. I have not tested this case. I can change it to continue now
and Can I add disable IRQ as TODO?
Thanks for your detailed review.
Thanks,
Nick
quoted
+ }
+
+ } while (true);
+}
+
+/*
* Fault window is opened per VAS instance. NX pastes fault CRB in fault
* FIFO upon page faults.
*/
@@ -1254,3 +1263,54 @@ int vas_win_close(struct vas_window *window)return0;}EXPORT_SYMBOL_GPL(vas_win_close);++structvas_window*vas_pswid_to_window(structvas_instance*vinst,+uint32_tpswid)+{+structvas_window*window;+intwinid;++if(!pswid){+pr_devel("%s: called for pswid 0!\n",__func__);+returnERR_PTR(-ESRCH);+}++decode_pswid(pswid,NULL,&winid);++if(winid>=VAS_WINDOWS_PER_CHIP)+returnERR_PTR(-ESRCH);++/*+*Ifapplicationclosesthewindowbeforethehardware+*returnsthefaultCRB,weshouldwaitinvas_win_close()+*forthependingrequests.sothewindowmustbeactive+*andtheprocessalive.+*+*Ifitsakernelprocess,weshouldnotgetanyfaultsand+*shouldnotgethere.+*/+window=vinst->windows[winid];++if(!window){+pr_err("PSWID decode: Could not find window for winid %d pswid %d vinst 0x%p\n",+winid,pswid,vinst);+returnNULL;+}++/*+*Dosomesanitychecksonthedecodedwindow.Windowshouldbe+*NXGZIPusersendwindow.FTWwindowsshouldnotincurfaults+*sincetheirCRBsareignored(notqueuedonFIFOorprocessed+*byNX).+*/+if(!window->tx_win||!window->user_win||!window->nx_win||+window->cop==VAS_COP_TYPE_FAULT||+window->cop==VAS_COP_TYPE_FTW){+pr_err("PSWID decode: id %d, tx %d, user %d, nx %d, cop %d\n",+winid,window->tx_win,window->user_win,+window->nx_win,window->cop);+WARN_ON(1);+}++returnwindow;+}
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-25 03:07:22
On Mon, 2020-03-23 at 12:40 +1000, Nicholas Piggin wrote:
Haren Myneni's on March 19, 2020 4:18 pm:
quoted
System checkstops if RxFIFO overruns with more requests than the
maximum possible number of CRBs allowed in FIFO at any time. So
max credits value (rxattr.wcreds_max) is set and is passed to
vas_rx_win_open() by the the driver.
This seems like it should be a bug fix or merged in the NX fault
window register patch or something.
Yes, it is a bug fix and can affect with any VAS windows, Not related to
NX fault window. Hence added as separate patch.
From: Haren Myneni <haren@linux.ibm.com> Date: 2020-03-25 03:38:12
On Mon, 2020-03-23 at 12:44 +1000, Nicholas Piggin wrote:
Haren Myneni's on March 19, 2020 4:19 pm:
quoted
NX expects OS to return credit for send window after processing each
fault. Also credit has to be returned even for fault window.
And this should be merged in the fault handler function.
credits are assigned and used per VAS window - default value is 1024 for
user space windows, and fault_fifo_size/CRb_SIZE for fault window.
When user space submits request, credit is taken on specific window (by
VAS). After successful processing of this request, NX return credit. In
case if NX sees fault, expects OS return credit for the corresponding
user space window after handling fault CRB.
Similarly NX takes credit on fault window after pasting fault CRB and
expects return credit after handling fault CRB. NX workbook has on
credits usage and How this credit system works.
Thought vas_return_credit() is unique function and added as separate
patch so that easy to review.
Any chance of a little bit of explanation how the credit system works?
Or is it in the code somewhere already?
Sure will add few comments on credit usage.
I don't suppose there is a chance to batch credit updates with multiple
faults? (maybe the MMIO is insignificant)
Yes, we return credit after processing each CRB. In the case of fault
window, NX can continue pasting fault CRB whenever the credit is
available.
Thanks
Haren
From: Michael Ellerman <mpe@ellerman.id.au> Date: 2020-03-25 11:15:24
Haren Myneni [off-list ref] writes:
On Mon, 2020-03-23 at 22:32 +1100, Michael Ellerman wrote:
quoted
Nicholas Piggin [off-list ref] writes:
quoted
Haren Myneni's on March 19, 2020 4:13 pm:
quoted
Kernel sets fault address and status in CRB for NX page fault on user
space address after processing page fault. User space gets the signal
and handles the fault mentioned in CRB by bringing the page in to
memory and send NX request again.
Signed-off-by: Sukadev Bhattiprolu <redacted>
Signed-off-by: Haren Myneni <haren@linux.ibm.com>
---
arch/powerpc/include/asm/icswx.h | 18 +++++++++++++++++-
1 file changed, 17 insertions(+), 1 deletion(-)
"icswx" is not a thing anymore, after 6ff4d3e96652 ("powerpc: Remove old
unused icswx based coprocessor support").
Yeah that commit ripped out some parts of the previous attempt at a user
visible API for this sort of "coprocessor" stuff. VAS is yet another
attempt to do something useful with most of the same pieces but some
slightly different details.
quoted
I guess NX is reusing some
things from it, but it would be good to get rid of the cruft and re-name
this file and and relevant names.
quoted
NX already uses this file, so I guesss that can happen after this series.
A lot of the CRB/CSB stuff is still the same, and P8 still uses icswx.
But I'd be happy if the header was renamed eventually, as icswx is now a
legacy name.
We can move all macros and struct definitions to vas.h and remove
icswx.h. Can I do this after this series?
Well they're still needed by the non-vas Power8 code, so that wouldn't
be quite right either :)
But yeah we can do whatever movement later as a cleanup.
cheers