Thread (3 messages) flat view 3 messages, 2 authors, 2015-07-17

Re: [RFC PATCH] getcpu_cache system call: caching current CPU number (x86)

From: Andy Lutomirski <hidden>
Date: 2015-07-16 19:27:10

On Thu, Jul 16, 2015 at 11:08 AM, Mathieu Desnoyers
[off-list ref] wrote:
----- On Jul 14, 2015, at 5:34 AM, Ben Maurer bmaurer-b10kYP2dOMg@public.gmane.org wrote:
quoted
Mathieu Desnoyers wrote:
quoted
If we invoke this per-thread registration directly in the glibc NPTL
implementation,
in start_thread, do you think it would fit your requirements ?
I guess this would basically be transparent to the user -- we'd just need to
make sure that the registration happens very early, before any chance of
calling malloc.
Yes, this is my thinking too.
quoted
That said, having the ability for the kernel to understand that TLS
implementation are laid out using the same offset on each thread seems like
something that could be valuable long term. Doing so makes it possible to build
other TLS-based features without forcing each thread to be registered.
AFAIU, using a fixed hardcoded ABI between kernel and user-space might make
transition from the pre-existing ABI (where this memory area is not
reserved) a bit tricky without registering the area, or getting a "feature"
flag, through a system call.

The related question then becomes: should we issue this system call once
per process, or once per thread at thread creation ? Issuing it once per
thread is marginally more costly for thread creation, but seems to be
easier to deal with internally within the kernel.

We could however ensure that only a single system call is needed per new-coming
thread, rather than one system call per feature. One way to do this would be
to register an area that may contain more than just the CPU id. It could
consist of an expandable structure with fixed offsets. When registered, we
could pass the size of that structure as an argument to the system call, so
the kernel knows which features are expected by user-space.
If we actually bit the bullet and implemented per-cpu mappings, we
could have this be completely flexible because there would be no
format at all.  Similarly, if we implemented per-cpu segments,
userspace would need to agree with *itself* how to arbitrate it, but
the kernel wouldn't need to be involved.

With this kind of memory poking, it's definitely messier, which is unfortunate.

--Andy
Keyboard shortcuts
hback out one level
jnext message in thread
kprevious message in thread
ldrill in
Escclose help / fold thread tree
?toggle this help