Thread (18 messages) flat view 18 messages, 3 authors, 2013-07-11

Re: [PATCH 2/2] KVM: PPC: Book3E: Get vcpu's last instruction for emulation

From: Scott Wood <hidden>
Date: 2013-07-09 18:47:27
Also in: kvm

On 07/09/2013 12:44:32 PM, Alexander Graf wrote:
On 07/09/2013 07:13 PM, Scott Wood wrote:
quoted
On 07/08/2013 08:39:05 AM, Alexander Graf wrote:
quoted
=20
On 28.06.2013, at 11:20, Mihai Caraman wrote:
=20
quoted
lwepx faults needs to be handled by KVM and this implies =20
additional code
quoted
in DO_KVM macro to identify the source of the exception =20
originated from
quoted
host context. This requires to check the Exception Syndrome =20
Register
quoted
(ESR[EPID]) and External PID Load Context Register (EPLC[EGS]) =20
for DTB_MISS,
quoted
DSI and LRAT exceptions which is too intrusive for the host.

Get rid of lwepx and acquire last instuction in =20
kvmppc_handle_exit() by
quoted
searching for the physical address and kmap it. This fixes an =20
infinite loop
=20
What's the difference in speed for this?
=20
Also, could we call lwepx later in host code, when =20
kvmppc_get_last_inst() gets invoked?
=20
Any use of lwepx is problematic unless we want to add overhead to =20
the main Linux TLB miss handler.
=20
What exactly would be missing?
If lwepx faults, it goes to the normal host TLB miss handler.  Without =20
adding code to it to recognize that it's an external-PID fault, it will =20
try to search the normal Linux page tables and insert a normal host =20
entry.  If it thinks it has succeeded, it will retry the instruction =20
rather than search for an exception handler.  The instruction will =20
fault again, and you get a hang.
I'd also still like to see some performance benchmarks on this to =20
make sure we're not walking into a bad direction.
I doubt it'll be significantly different.  There's overhead involved in =20
setting up for lwepx as well.  It doesn't hurt to test, though this is =20
a functional correctness issue, so I'm not sure what better =20
alternatives we have.  I don't want to slow down non-KVM TLB misses for =20
this.
quoted
quoted
quoted
+    addr =3D (mas7_mas3 & (~0ULL << psize_shift)) |
+           (geaddr & ((1ULL << psize_shift) - 1ULL));
+
+    /* Map a page and get guest's instruction */
+    page =3D pfn_to_page(addr >> PAGE_SHIFT);
=20
So it seems to me like you're jumping through a lot of hoops to =20
make sure this works for LRAT and non-LRAT at the same time. Can't =20
we just treat them as the different things they are?
=20
What if we have different MMU backends for LRAT and non-LRAT? The =20
non-LRAT case could then try lwepx, if that fails, fall back to =20
read the shadow TLB. For the LRAT case, we'd do lwepx, if that =20
fails fall back to this logic.
=20
This isn't about LRAT; it's about hardware threads.  It also fixes =20
the handling of execute-only pages on current chips.
=20
On non-LRAT systems we could always check our shadow copy of the =20
guest's TLB, no? I'd really like to know what the performance =20
difference would be for the 2 approaches.
I suspect that tlbsx is faster, or at worst similar.  And unlike =20
comparing tlbsx to lwepx (not counting a fix for the threading =20
problem), we don't already have code to search the guest TLB, so =20
testing would be more work.

-Scott=
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help