From: Michael Neuling <hidden> Date: 2018-09-12 05:21:40
This stops us from doing code patching in init sections after they've
been freed.
In this chain:
kvm_guest_init() ->
kvm_use_magic_page() ->
fault_in_pages_readable() ->
__get_user() ->
__get_user_nocheck() ->
barrier_nospec();
We have a code patching location at barrier_nospec() and
kvm_guest_init() is an init function. This whole chain gets inlined,
so when we free the init section (hence kvm_guest_init()), this code
goes away and hence should no longer be patched.
We seen this as userspace memory corruption when using a memory
checker while doing partition migration testing on powervm (this
starts the code patching post migration via
/sys/kernel/mobility/migration). In theory, it could also happen when
using /sys/kernel/debug/powerpc/barrier_nospec.
With this patch there is a small change of a race if we code patch
between the init section being freed and setting SYSTEM_RUNNING (in
kernel_init()) but that seems like an impractical time and small
window for any code patching to occur.
cc: stable@vger.kernel.org # 4.13+
Signed-off-by: Michael Neuling <redacted>
---
For stable I've marked this as v4.13+ since that's when we refactored
code-patching.c but it could go back even further than that. In
reality though, I think we can only hit this since the first
spectre/meltdown changes.
v2:
Print when we skip an address
---
arch/powerpc/lib/code-patching.c | 22 ++++++++++++++++++++++
1 file changed, 22 insertions(+)
This stops us from doing code patching in init sections after they've
been freed.
In this chain:
kvm_guest_init() ->
kvm_use_magic_page() ->
fault_in_pages_readable() ->
__get_user() ->
__get_user_nocheck() ->
barrier_nospec();
We have a code patching location at barrier_nospec() and
kvm_guest_init() is an init function. This whole chain gets inlined,
so when we free the init section (hence kvm_guest_init()), this code
goes away and hence should no longer be patched.
We seen this as userspace memory corruption when using a memory
checker while doing partition migration testing on powervm (this
starts the code patching post migration via
/sys/kernel/mobility/migration). In theory, it could also happen when
using /sys/kernel/debug/powerpc/barrier_nospec.
With this patch there is a small change of a race if we code patch
between the init section being freed and setting SYSTEM_RUNNING (in
kernel_init()) but that seems like an impractical time and small
window for any code patching to occur.
cc: stable@vger.kernel.org # 4.13+
Signed-off-by: Michael Neuling <redacted>
---
For stable I've marked this as v4.13+ since that's when we refactored
code-patching.c but it could go back even further than that. In
reality though, I think we can only hit this since the first
spectre/meltdown changes.
v2:
Print when we skip an address
---
arch/powerpc/lib/code-patching.c | 22 ++++++++++++++++++++++
1 file changed, 22 insertions(+)
I would call this function differently, for instance init_is_finished(),
because as you mentionned it doesn't exactly mean that init memory is freed.
static int __patch_instruction(unsigned int *exec_addr, unsigned int instr,
unsigned int *patch_addr)
{
int err;
+ /* Make sure we aren't patching a freed init section */
+ if (in_init_section(patch_addr) && init_freed()) {
The test must be done on exec_addr, not on patch_addr, as patch_addr is
the address where the instruction as been remapped RW for allowing its
modification.
Also I think it should be tested the other way round, because the
init_freed() is a simpler test which will be false most of the time once
the system is running so it should be checked first.
I think it would be better to put this verification in
patch_instruction() instead, to avoid RW mapping/unmapping the
instruction to patch when we are not going to do the patching.
Christophe
=20
I would call this function differently, for instance init_is_finished(),=
=20
because as you mentionned it doesn't exactly mean that init memory is fre=
ed.
Talking to Nick and mpe offline I think we are going to have to add a flag =
when
we free init mem rather than doing what we have now since what we have now =
has a
potential race. That change will eliminate the function entirely.
quoted
static int __patch_instruction(unsigned int *exec_addr, unsigned int
instr,
unsigned int *patch_addr)
{
int err;
=20
+ /* Make sure we aren't patching a freed init section */
+ if (in_init_section(patch_addr) && init_freed()) {
=20
The test must be done on exec_addr, not on patch_addr, as patch_addr is=
=20
the address where the instruction as been remapped RW for allowing its=
=20
modification.
Thanks for the catch
Also I think it should be tested the other way round, because the=20
init_freed() is a simpler test which will be false most of the time once=
=20
the system is running so it should be checked first.
=20
I think it would be better to put this verification in=20
patch_instruction() instead, to avoid RW mapping/unmapping the=20
instruction to patch when we are not going to do the patching.
If we do it there then we miss the raw_patch_intruction case.
IMHO I don't think we need to optimise this rare and non-critical path.=20
Mikey
I would call this function differently, for instance init_is_finished(),
because as you mentionned it doesn't exactly mean that init memory is freed.
Talking to Nick and mpe offline I think we are going to have to add a flag when
we free init mem rather than doing what we have now since what we have now has a
potential race. That change will eliminate the function entirely.
quoted
quoted
static int __patch_instruction(unsigned int *exec_addr, unsigned int
instr,
unsigned int *patch_addr)
{
int err;
+ /* Make sure we aren't patching a freed init section */
+ if (in_init_section(patch_addr) && init_freed()) {
The test must be done on exec_addr, not on patch_addr, as patch_addr is
the address where the instruction as been remapped RW for allowing its
modification.
Thanks for the catch
quoted
Also I think it should be tested the other way round, because the
init_freed() is a simpler test which will be false most of the time once
the system is running so it should be checked first.
Sorry I can't see what's wrong. You're (or Cody :-P) going to have to spell it
this out for me...
I suspect that the suggestion is the opening parenthesis of "(unsigned long)" should sit directly under the "K" of "KERN_DEBUG". I'm pretty sure Documentation/process/coding-style.rst is very adamant that all identation is always 8 characters and spaces should never be used, but there still seems to be a lot of places/suggestions that argument lists that spill over multiple lines should be space indented to align with the very first argument at the top level. So, I guess I'm not sure what the desire is here. Although moving to pr_debug might fit it to a single line anyways. ;)
-Tyrel
I think it would be better to put this verification in
patch_instruction() instead, to avoid RW mapping/unmapping the
instruction to patch when we are not going to do the patching.
If we do it there then we miss the raw_patch_intruction case.
IMHO I don't think we need to optimise this rare and non-critical path.
Mikey
Sorry I can't see what's wrong. You're (or Cody :-P) going to have to spell it
this out for me...
I suspect that the suggestion is the opening parenthesis of "(unsigned long)" should sit directly under the "K" of "KERN_DEBUG". I'm pretty sure Documentation/process/coding-style.rst is very adamant that all identation is always 8 characters and spaces should never be used, but there still seems to be a lot of places/suggestions that argument lists that spill over multiple lines should be space indented to align with the very first argument at the top level. So, I guess I'm not sure what the desire is here. Although moving to pr_debug might fit it to a single line anyways. ;)
I would call this function differently, for instance init_is_finished(),
because as you mentionned it doesn't exactly mean that init memory is freed.
Talking to Nick and mpe offline I think we are going to have to add a flag when
we free init mem rather than doing what we have now since what we have now has a
potential race. That change will eliminate the function entirely.
quoted
quoted
static int __patch_instruction(unsigned int *exec_addr, unsigned int
instr,
unsigned int *patch_addr)
{
int err;
+ /* Make sure we aren't patching a freed init section */
+ if (in_init_section(patch_addr) && init_freed()) {
The test must be done on exec_addr, not on patch_addr, as patch_addr is
the address where the instruction as been remapped RW for allowing its
modification.
Thanks for the catch
quoted
Also I think it should be tested the other way round, because the
init_freed() is a simpler test which will be false most of the time once
the system is running so it should be checked first.
I think it would be better to put this verification in
patch_instruction() instead, to avoid RW mapping/unmapping the
instruction to patch when we are not going to do the patching.
If we do it there then we miss the raw_patch_intruction case.
raw_patch_instruction() can only be used during init. Once kernel memory
has been marked readonly, raw_patch_instruction() cannot be used
anymore. And mark_readonly() is called immediately after free_initmem()
Christophe
IMHO I don't think we need to optimise this rare and non-critical path.
Mikey