Current vDSO64 implementation does not have support for coarse clocks
(CLOCK_MONOTONIC_COARSE, CLOCK_REALTIME_COARSE), for which it falls back
to system call, increasing the response time, vDSO implementation reduces
the cycle time. Below is a benchmark of the difference in execution time
with and without vDSO support.
(Non-coarse clocks are also included just for completion)
Without vDSO support:
--------------------
clock-gettime-realtime: syscall: 172 nsec/call
clock-gettime-realtime: libc: 26 nsec/call
clock-gettime-realtime: vdso: 21 nsec/call
clock-gettime-monotonic: syscall: 170 nsec/call
clock-gettime-monotonic: libc: 30 nsec/call
clock-gettime-monotonic: vdso: 24 nsec/call
clock-gettime-realtime-coarse: syscall: 153 nsec/call
clock-gettime-realtime-coarse: libc: 15 nsec/call
clock-gettime-realtime-coarse: vdso: 9 nsec/call
clock-gettime-monotonic-coarse: syscall: 167 nsec/call
clock-gettime-monotonic-coarse: libc: 15 nsec/call
clock-gettime-monotonic-coarse: vdso: 11 nsec/call
CC: Benjamin Herrenschmidt <benh@kernel.crashing.org>
Signed-off-by: Santosh Sivaraj <redacted>
---
arch/powerpc/kernel/asm-offsets.c | 2 ++
arch/powerpc/kernel/vdso64/gettimeofday.S | 56 +++++++++++++++++++++++++++++++
2 files changed, 58 insertions(+)
@@ -396,6 +396,8 @@ int main(void)/* Other bits used by the vdso */DEFINE(CLOCK_REALTIME,CLOCK_REALTIME);DEFINE(CLOCK_MONOTONIC,CLOCK_MONOTONIC);+DEFINE(CLOCK_REALTIME_COARSE,CLOCK_REALTIME_COARSE);+DEFINE(CLOCK_MONOTONIC_COARSE,CLOCK_MONOTONIC_COARSE);DEFINE(NSEC_PER_SEC,NSEC_PER_SEC);DEFINE(CLOCK_REALTIME_RES,MONOTONIC_RES_NSEC);
@@ -112,6 +117,57 @@ V_FUNCTION_BEGIN(__kernel_clock_gettime)1:bgecr1,80faddir4,r4,-1addr5,r5,r7+b80f++/*+*Forcoarseclockswegetdatadirectlyfromthevdsodatapage,so+*wedon't need to call __do_get_tspec, but we still need to do the+*countertrick.+*/+65:blV_LOCAL_FUNC(__get_datapage)/*getdatapage*/+70:ldr8,CFG_TB_UPDATE_COUNT(r3)+andi.r0,r8,1/*pendingupdate?loop*/+bne-70b+xorr0,r8,r8/*createdependency*/+addr3,r3,r0++/*+*CLOCK_REALTIME_COARSE,belowvaluesareneededforMONOTONIC_COARSE+*too+*/+ldr4,STAMP_XTIME+TSPC64_TV_SEC(r3)+ldr5,STAMP_XTIME+TSPC64_TV_NSEC(r3)+bnecr1,78f++/*CLOCK_MONOTONIC_COARSE*/+lwar6,WTOM_CLOCK_SEC(r3)+lwar9,WTOM_CLOCK_NSEC(r3)++/*checkifcounterhasupdated*/+78:orr0,r6,r9+xorr0,r0,r0+addr3,r3,r0+ldr0,CFG_TB_UPDATE_COUNT(r3)+cmpldcr0,r0,r8/*checkifupdated*/+bne-70b++/*Counterhasnotupdated,socontinuecalculatingpropervaluesfor+*secandnsecifmonotoniccoarse,orjustreturnwiththeproper+*valuesforrealtime.+*/+bnecr1,80f++/*Addwall->monotonicoffsetandcheckforoverfloworunderflow*/+addr4,r4,r6+addr5,r5,r9+cmpdcr0,r5,r7+cmpdicr1,r5,0+blt79f+subfr5,r7,r5+addir4,r4,1+79:bgecr1,80f+addir4,r4,-1+addr5,r5,r780:stdr4,TSPC64_TV_SEC(r11)stdr5,TSPC64_TV_NSEC(r11)
'beq', followed by a 'b' looks weird without considering the next patch.
I think this can be organized better to not have to update r7/r11/r12 if
using the system call. See next patch for my comments.
.cfi_register lr,r12
If you move the mflr, you should move the above line along with it.
- Naveen
- mr r11,r4 /* r11 saves tp */
- bl V_LOCAL_FUNC(__get_datapage) /* get data page */
- lis r7,NSEC_PER_SEC@h /* want nanoseconds */
- ori r7,r7,NSEC_PER_SEC@l
+49: bl V_LOCAL_FUNC(__get_datapage) /* get data page */
50: bl V_LOCAL_FUNC(__do_get_tspec) /* get time from tb & kernel */
bne cr1,80f /* if not monotonic, all done */
--
2.13.5
From: Naveen N. Rao <hidden> Date: 2017-10-06 09:28:45
On 2017/09/18 09:23AM, Santosh Sivaraj wrote:
quoted hunk
Current vDSO64 implementation does not have support for coarse clocks
(CLOCK_MONOTONIC_COARSE, CLOCK_REALTIME_COARSE), for which it falls back
to system call, increasing the response time, vDSO implementation reduces
the cycle time. Below is a benchmark of the difference in execution time
with and without vDSO support.
(Non-coarse clocks are also included just for completion)
Without vDSO support:
--------------------
clock-gettime-realtime: syscall: 172 nsec/call
clock-gettime-realtime: libc: 26 nsec/call
clock-gettime-realtime: vdso: 21 nsec/call
clock-gettime-monotonic: syscall: 170 nsec/call
clock-gettime-monotonic: libc: 30 nsec/call
clock-gettime-monotonic: vdso: 24 nsec/call
clock-gettime-realtime-coarse: syscall: 153 nsec/call
clock-gettime-realtime-coarse: libc: 15 nsec/call
clock-gettime-realtime-coarse: vdso: 9 nsec/call
clock-gettime-monotonic-coarse: syscall: 167 nsec/call
clock-gettime-monotonic-coarse: libc: 15 nsec/call
clock-gettime-monotonic-coarse: vdso: 11 nsec/call
CC: Benjamin Herrenschmidt <benh@kernel.crashing.org>
Signed-off-by: Santosh Sivaraj <redacted>
---
arch/powerpc/kernel/asm-offsets.c | 2 ++
arch/powerpc/kernel/vdso64/gettimeofday.S | 56 +++++++++++++++++++++++++++++++
2 files changed, 58 insertions(+)
@@ -396,6 +396,8 @@ int main(void)/* Other bits used by the vdso */DEFINE(CLOCK_REALTIME,CLOCK_REALTIME);DEFINE(CLOCK_MONOTONIC,CLOCK_MONOTONIC);+DEFINE(CLOCK_REALTIME_COARSE,CLOCK_REALTIME_COARSE);+DEFINE(CLOCK_MONOTONIC_COARSE,CLOCK_MONOTONIC_COARSE);DEFINE(NSEC_PER_SEC,NSEC_PER_SEC);DEFINE(CLOCK_REALTIME_RES,MONOTONIC_RES_NSEC);
If you use cr5-7 here, you should be able to re-organize this to not
have to update r4/r11/r12 if we're taking the syscall path. Not
necessarily a huge win by itself, but can also help reuse some of the
other code between the _COARSE and the regular variants.
- Naveen
quoted hunk
+
b 99f /* Fallback to syscall */
.cfi_register lr,r12
49: bl V_LOCAL_FUNC(__get_datapage) /* get data page */
@@ -112,6 +117,57 @@ V_FUNCTION_BEGIN(__kernel_clock_gettime) 1: bge cr1,80f addi r4,r4,-1 add r5,r5,r7+ b 80f++ /*+ * For coarse clocks we get data directly from the vdso data page, so+ * we don't need to call __do_get_tspec, but we still need to do the+ * counter trick.+ */+65: bl V_LOCAL_FUNC(__get_datapage) /* get data page */+70: ld r8,CFG_TB_UPDATE_COUNT(r3)+ andi. r0,r8,1 /* pending update ? loop */+ bne- 70b+ xor r0,r8,r8 /* create dependency */+ add r3,r3,r0++ /*+ * CLOCK_REALTIME_COARSE, below values are needed for MONOTONIC_COARSE+ * too+ */+ ld r4,STAMP_XTIME+TSPC64_TV_SEC(r3)+ ld r5,STAMP_XTIME+TSPC64_TV_NSEC(r3)+ bne cr1,78f++ /* CLOCK_MONOTONIC_COARSE */+ lwa r6,WTOM_CLOCK_SEC(r3)+ lwa r9,WTOM_CLOCK_NSEC(r3)++ /* check if counter has updated */+78: or r0,r6,r9+ xor r0,r0,r0+ add r3,r3,r0+ ld r0,CFG_TB_UPDATE_COUNT(r3)+ cmpld cr0,r0,r8 /* check if updated */+ bne- 70b++ /* Counter has not updated, so continue calculating proper values for+ * sec and nsec if monotonic coarse, or just return with the proper+ * values for realtime.+ */+ bne cr1,80f++ /* Add wall->monotonic offset and check for overflow or underflow */+ add r4,r4,r6+ add r5,r5,r9+ cmpd cr0,r5,r7+ cmpdi cr1,r5,0+ blt 79f+ subf r5,r7,r5+ addi r4,r4,1+79: bge cr1,80f+ addi r4,r4,-1+ add r5,r5,r7 80: std r4,TSPC64_TV_SEC(r11) std r5,TSPC64_TV_NSEC(r11)
'beq', followed by a 'b' looks weird without considering the next patch.
I think this can be organized better to not have to update r7/r11/r12 if
using the system call. See next patch for my comments.
quoted
.cfi_register lr,r12
If you move the mflr, you should move the above line along with it.
s/should/must/.
It literally says "lr is saved in r12".
cheers
From: Naveen N. Rao <hidden> Date: 2017-10-06 11:25:41
On 2017/09/18 09:23AM, Santosh Sivaraj wrote:
quoted hunk
Current vDSO64 implementation does not have support for coarse clocks
(CLOCK_MONOTONIC_COARSE, CLOCK_REALTIME_COARSE), for which it falls back
to system call, increasing the response time, vDSO implementation reduces
the cycle time. Below is a benchmark of the difference in execution time
with and without vDSO support.
(Non-coarse clocks are also included just for completion)
Without vDSO support:
--------------------
clock-gettime-realtime: syscall: 172 nsec/call
clock-gettime-realtime: libc: 26 nsec/call
clock-gettime-realtime: vdso: 21 nsec/call
clock-gettime-monotonic: syscall: 170 nsec/call
clock-gettime-monotonic: libc: 30 nsec/call
clock-gettime-monotonic: vdso: 24 nsec/call
clock-gettime-realtime-coarse: syscall: 153 nsec/call
clock-gettime-realtime-coarse: libc: 15 nsec/call
clock-gettime-realtime-coarse: vdso: 9 nsec/call
clock-gettime-monotonic-coarse: syscall: 167 nsec/call
clock-gettime-monotonic-coarse: libc: 15 nsec/call
clock-gettime-monotonic-coarse: vdso: 11 nsec/call
CC: Benjamin Herrenschmidt <benh@kernel.crashing.org>
Signed-off-by: Santosh Sivaraj <redacted>
---
arch/powerpc/kernel/asm-offsets.c | 2 ++
arch/powerpc/kernel/vdso64/gettimeofday.S | 56 +++++++++++++++++++++++++++++++
2 files changed, 58 insertions(+)
@@ -396,6 +396,8 @@ int main(void)/* Other bits used by the vdso */DEFINE(CLOCK_REALTIME,CLOCK_REALTIME);DEFINE(CLOCK_MONOTONIC,CLOCK_MONOTONIC);+DEFINE(CLOCK_REALTIME_COARSE,CLOCK_REALTIME_COARSE);+DEFINE(CLOCK_MONOTONIC_COARSE,CLOCK_MONOTONIC_COARSE);DEFINE(NSEC_PER_SEC,NSEC_PER_SEC);DEFINE(CLOCK_REALTIME_RES,MONOTONIC_RES_NSEC);
@@ -112,6 +117,57 @@ V_FUNCTION_BEGIN(__kernel_clock_gettime)1:bgecr1,80faddir4,r4,-1addr5,r5,r7+b80f++/*+*Forcoarseclockswegetdatadirectlyfromthevdsodatapage,so+*wedon't need to call __do_get_tspec, but we still need to do the+*countertrick.+*/+65:blV_LOCAL_FUNC(__get_datapage)/*getdatapage*/+70:ldr8,CFG_TB_UPDATE_COUNT(r3)+andi.r0,r8,1/*pendingupdate?loop*/+bne-70b+xorr0,r8,r8/*createdependency*/+addr3,r3,r0++/*+*CLOCK_REALTIME_COARSE,belowvaluesareneededforMONOTONIC_COARSE+*too+*/+ldr4,STAMP_XTIME+TSPC64_TV_SEC(r3)+ldr5,STAMP_XTIME+TSPC64_TV_NSEC(r3)+bnecr1,78f++/*CLOCK_MONOTONIC_COARSE*/+lwar6,WTOM_CLOCK_SEC(r3)+lwar9,WTOM_CLOCK_NSEC(r3)++/*checkifcounterhasupdated*/+78:orr0,r6,r9+xorr0,r0,r0+addr3,r3,r0+ldr0,CFG_TB_UPDATE_COUNT(r3)+cmpldcr0,r0,r8/*checkifupdated*/+bne-70b
Don't you need a dependency on r4/r5 here for REALTIME_COARSE?
Something like:
/* check if counter has updated */
or r0,r6,r9
78: or r0,r4,r5
xor r0,r0,r0
+
+ /* Counter has not updated, so continue calculating proper values for
+ * sec and nsec if monotonic coarse, or just return with the proper
+ * values for realtime.
+ */
+ bne cr1,80f
+
I think the below hunk can surely be shared across the _COARSE and
regular clocks, if not more.
- Naveen
* Naveen N. Rao [off-list ref] wrote (on 2017-10-06 11:25:28 +0000):
On 2017/09/18 09:23AM, Santosh Sivaraj wrote:
quoted
Current vDSO64 implementation does not have support for coarse clocks
(CLOCK_MONOTONIC_COARSE, CLOCK_REALTIME_COARSE), for which it falls back
to system call, increasing the response time, vDSO implementation reduces
the cycle time. Below is a benchmark of the difference in execution time
with and without vDSO support.
(Non-coarse clocks are also included just for completion)
Without vDSO support:
--------------------
clock-gettime-realtime: syscall: 172 nsec/call
clock-gettime-realtime: libc: 26 nsec/call
clock-gettime-realtime: vdso: 21 nsec/call
clock-gettime-monotonic: syscall: 170 nsec/call
clock-gettime-monotonic: libc: 30 nsec/call
clock-gettime-monotonic: vdso: 24 nsec/call
clock-gettime-realtime-coarse: syscall: 153 nsec/call
clock-gettime-realtime-coarse: libc: 15 nsec/call
clock-gettime-realtime-coarse: vdso: 9 nsec/call
clock-gettime-monotonic-coarse: syscall: 167 nsec/call
clock-gettime-monotonic-coarse: libc: 15 nsec/call
clock-gettime-monotonic-coarse: vdso: 11 nsec/call
CC: Benjamin Herrenschmidt <benh@kernel.crashing.org>
Signed-off-by: Santosh Sivaraj <redacted>
---
arch/powerpc/kernel/asm-offsets.c | 2 ++
arch/powerpc/kernel/vdso64/gettimeofday.S | 56 +++++++++++++++++++++++++++++++
2 files changed, 58 insertions(+)
@@ -396,6 +396,8 @@ int main(void)/* Other bits used by the vdso */DEFINE(CLOCK_REALTIME,CLOCK_REALTIME);DEFINE(CLOCK_MONOTONIC,CLOCK_MONOTONIC);+DEFINE(CLOCK_REALTIME_COARSE,CLOCK_REALTIME_COARSE);+DEFINE(CLOCK_MONOTONIC_COARSE,CLOCK_MONOTONIC_COARSE);DEFINE(NSEC_PER_SEC,NSEC_PER_SEC);DEFINE(CLOCK_REALTIME_RES,MONOTONIC_RES_NSEC);
@@ -112,6 +117,57 @@ V_FUNCTION_BEGIN(__kernel_clock_gettime)1:bgecr1,80faddir4,r4,-1addr5,r5,r7+b80f++/*+*Forcoarseclockswegetdatadirectlyfromthevdsodatapage,so+*wedon't need to call __do_get_tspec, but we still need to do the+*countertrick.+*/+65:blV_LOCAL_FUNC(__get_datapage)/*getdatapage*/+70:ldr8,CFG_TB_UPDATE_COUNT(r3)+andi.r0,r8,1/*pendingupdate?loop*/+bne-70b+xorr0,r8,r8/*createdependency*/+addr3,r3,r0++/*+*CLOCK_REALTIME_COARSE,belowvaluesareneededforMONOTONIC_COARSE+*too+*/+ldr4,STAMP_XTIME+TSPC64_TV_SEC(r3)+ldr5,STAMP_XTIME+TSPC64_TV_NSEC(r3)+bnecr1,78f++/*CLOCK_MONOTONIC_COARSE*/+lwar6,WTOM_CLOCK_SEC(r3)+lwar9,WTOM_CLOCK_NSEC(r3)++/*checkifcounterhasupdated*/+78:orr0,r6,r9+xorr0,r0,r0+addr3,r3,r0+ldr0,CFG_TB_UPDATE_COUNT(r3)+cmpldcr0,r0,r8/*checkifupdated*/+bne-70b
Don't you need a dependency on r4/r5 here for REALTIME_COARSE?
Something like:
/* check if counter has updated */
or r0,r6,r9
78: or r0,r4,r5
xor r0,r0,r0
Yes, we would need it. Will update in v2.
quoted
+
+ /* Counter has not updated, so continue calculating proper values for
+ * sec and nsec if monotonic coarse, or just return with the proper
+ * values for realtime.
+ */
+ bne cr1,80f
+
I think the below hunk can surely be shared across the _COARSE and
regular clocks, if not more.
Yes, except for the label its the same for both monotonic and
monotonic_coarse, will update in the next set.
Thanks,
Santosh
* Naveen N. Rao [off-list ref] wrote (on 2017-10-06 09:28:30 +0000):
On 2017/09/18 09:23AM, Santosh Sivaraj wrote:
quoted
Current vDSO64 implementation does not have support for coarse clocks
(CLOCK_MONOTONIC_COARSE, CLOCK_REALTIME_COARSE), for which it falls back
to system call, increasing the response time, vDSO implementation reduces
the cycle time. Below is a benchmark of the difference in execution time
with and without vDSO support.
(Non-coarse clocks are also included just for completion)
Without vDSO support:
--------------------
clock-gettime-realtime: syscall: 172 nsec/call
clock-gettime-realtime: libc: 26 nsec/call
clock-gettime-realtime: vdso: 21 nsec/call
clock-gettime-monotonic: syscall: 170 nsec/call
clock-gettime-monotonic: libc: 30 nsec/call
clock-gettime-monotonic: vdso: 24 nsec/call
clock-gettime-realtime-coarse: syscall: 153 nsec/call
clock-gettime-realtime-coarse: libc: 15 nsec/call
clock-gettime-realtime-coarse: vdso: 9 nsec/call
clock-gettime-monotonic-coarse: syscall: 167 nsec/call
clock-gettime-monotonic-coarse: libc: 15 nsec/call
clock-gettime-monotonic-coarse: vdso: 11 nsec/call
CC: Benjamin Herrenschmidt <benh@kernel.crashing.org>
Signed-off-by: Santosh Sivaraj <redacted>
---
arch/powerpc/kernel/asm-offsets.c | 2 ++
arch/powerpc/kernel/vdso64/gettimeofday.S | 56 +++++++++++++++++++++++++++++++
2 files changed, 58 insertions(+)
@@ -396,6 +396,8 @@ int main(void)/* Other bits used by the vdso */DEFINE(CLOCK_REALTIME,CLOCK_REALTIME);DEFINE(CLOCK_MONOTONIC,CLOCK_MONOTONIC);+DEFINE(CLOCK_REALTIME_COARSE,CLOCK_REALTIME_COARSE);+DEFINE(CLOCK_MONOTONIC_COARSE,CLOCK_MONOTONIC_COARSE);DEFINE(NSEC_PER_SEC,NSEC_PER_SEC);DEFINE(CLOCK_REALTIME_RES,MONOTONIC_RES_NSEC);
If you use cr5-7 here, you should be able to re-organize this to not
have to update r4/r11/r12 if we're taking the syscall path. Not
necessarily a huge win by itself, but can also help reuse some of the
other code between the _COARSE and the regular variants.
If we are going to use cr5-7, then the first patch is no longer required, we
don't have to do a re-org of the intial clock_id checks. I will send the
updated patch.
Thanks,
Santosh
- Naveen
quoted
+
b 99f /* Fallback to syscall */
.cfi_register lr,r12
49: bl V_LOCAL_FUNC(__get_datapage) /* get data page */
@@ -112,6 +117,57 @@ V_FUNCTION_BEGIN(__kernel_clock_gettime) 1: bge cr1,80f addi r4,r4,-1 add r5,r5,r7+ b 80f++ /*+ * For coarse clocks we get data directly from the vdso data page, so+ * we don't need to call __do_get_tspec, but we still need to do the+ * counter trick.+ */+65: bl V_LOCAL_FUNC(__get_datapage) /* get data page */+70: ld r8,CFG_TB_UPDATE_COUNT(r3)+ andi. r0,r8,1 /* pending update ? loop */+ bne- 70b+ xor r0,r8,r8 /* create dependency */+ add r3,r3,r0++ /*+ * CLOCK_REALTIME_COARSE, below values are needed for MONOTONIC_COARSE+ * too+ */+ ld r4,STAMP_XTIME+TSPC64_TV_SEC(r3)+ ld r5,STAMP_XTIME+TSPC64_TV_NSEC(r3)+ bne cr1,78f++ /* CLOCK_MONOTONIC_COARSE */+ lwa r6,WTOM_CLOCK_SEC(r3)+ lwa r9,WTOM_CLOCK_NSEC(r3)++ /* check if counter has updated */+78: or r0,r6,r9+ xor r0,r0,r0+ add r3,r3,r0+ ld r0,CFG_TB_UPDATE_COUNT(r3)+ cmpld cr0,r0,r8 /* check if updated */+ bne- 70b++ /* Counter has not updated, so continue calculating proper values for+ * sec and nsec if monotonic coarse, or just return with the proper+ * values for realtime.+ */+ bne cr1,80f++ /* Add wall->monotonic offset and check for overflow or underflow */+ add r4,r4,r6+ add r5,r5,r9+ cmpd cr0,r5,r7+ cmpdi cr1,r5,0+ blt 79f+ subf r5,r7,r5+ addi r4,r4,1+79: bge cr1,80f+ addi r4,r4,-1+ add r5,r5,r7 80: std r4,TSPC64_TV_SEC(r11) std r5,TSPC64_TV_NSEC(r11)
Current vDSO64 implementation does not have support for coarse clocks
(CLOCK_MONOTONIC_COARSE, CLOCK_REALTIME_COARSE), for which it falls back
to system call, increasing the response time, vDSO implementation reduces
the cycle time. Below is a benchmark of the difference in execution times.
(Non-coarse clocks are also included just for completion)
clock-gettime-realtime: syscall: 172 nsec/call
clock-gettime-realtime: libc: 28 nsec/call
clock-gettime-realtime: vdso: 22 nsec/call
clock-gettime-monotonic: syscall: 171 nsec/call
clock-gettime-monotonic: libc: 30 nsec/call
clock-gettime-monotonic: vdso: 25 nsec/call
clock-gettime-realtime-coarse: syscall: 153 nsec/call
clock-gettime-realtime-coarse: libc: 16 nsec/call
clock-gettime-realtime-coarse: vdso: 10 nsec/call
clock-gettime-monotonic-coarse: syscall: 167 nsec/call
clock-gettime-monotonic-coarse: libc: 17 nsec/call
clock-gettime-monotonic-coarse: vdso: 11 nsec/call
CC: Benjamin Herrenschmidt <benh@kernel.crashing.org>
Signed-off-by: Santosh Sivaraj <redacted>
---
arch/powerpc/kernel/asm-offsets.c | 2 +
arch/powerpc/kernel/vdso64/gettimeofday.S | 67 ++++++++++++++++++++++++++-----
2 files changed, 58 insertions(+), 11 deletions(-)
@@ -396,6 +396,8 @@ int main(void)/* Other bits used by the vdso */DEFINE(CLOCK_REALTIME,CLOCK_REALTIME);DEFINE(CLOCK_MONOTONIC,CLOCK_MONOTONIC);+DEFINE(CLOCK_REALTIME_COARSE,CLOCK_REALTIME_COARSE);+DEFINE(CLOCK_MONOTONIC_COARSE,CLOCK_MONOTONIC_COARSE);DEFINE(NSEC_PER_SEC,NSEC_PER_SEC);DEFINE(CLOCK_REALTIME_RES,MONOTONIC_RES_NSEC);
@@ -97,19 +104,57 @@ V_FUNCTION_BEGIN(__kernel_clock_gettime)ldr0,CFG_TB_UPDATE_COUNT(r3)cmpldcr0,r0,r8/*checkifupdated*/bne-50b+b78f-/*Addwall->monotonicoffsetandcheckforoverfloworunderflow.+/*+*Forcoarseclockswegetdatadirectlyfromthevdsodatapage,so+*wedon't need to call __do_get_tspec, but we still need to do the+*countertrick.*/-addr4,r4,r6-addr5,r5,r9-cmpdcr0,r5,r7-cmpdicr1,r5,0-blt1f-subfr5,r7,r5-addir4,r4,1-1:bgecr1,80f-addir4,r4,-1-addr5,r5,r7+70:ldr8,CFG_TB_UPDATE_COUNT(r3)+andi.r0,r8,1/*pendingupdate?loop*/+bne-70b+xorr0,r8,r8/*createdependency*/+addr3,r3,r0++/*+*CLOCK_REALTIME_COARSE,belowvaluesareneededforMONOTONIC_COARSE+*too+*/+ldr4,STAMP_XTIME+TSPC64_TV_SEC(r3)+ldr5,STAMP_XTIME+TSPC64_TV_NSEC(r3)+bnecr6,75f++/*CLOCK_MONOTONIC_COARSE*/+lwar6,WTOM_CLOCK_SEC(r3)+lwar9,WTOM_CLOCK_NSEC(r3)++/*checkifcounterhasupdated*/+75:orr0,r6,r9+orr0,r4,r5+xorr0,r0,r0+addr3,r3,r0+ldr0,CFG_TB_UPDATE_COUNT(r3)+cmpldcr0,r0,r8/*checkifupdated*/+bne-70b++/*Counterhasnotupdated,socontinuecalculatingpropervaluesfor+*secandnsecifmonotoniccoarse,orjustreturnwiththeproper+*valuesforrealtime.+*/+bnecr6,80f++/*Addwall->monotonicoffsetandcheckforoverfloworunderflow*/+78:addr4,r4,r6+addr5,r5,r9+cmpdcr0,r5,r7+cmpdicr1,r5,0+blt79f+subfr5,r7,r5+addir4,r4,1+79:bgecr1,80f+addir4,r4,-1+addr5,r5,r780:stdr4,TSPC64_TV_SEC(r11)stdr5,TSPC64_TV_NSEC(r11)
From: Naveen N. Rao <hidden> Date: 2017-10-09 10:39:31
On 2017/10/09 08:09AM, Santosh Sivaraj wrote:
quoted hunk
Current vDSO64 implementation does not have support for coarse clocks
(CLOCK_MONOTONIC_COARSE, CLOCK_REALTIME_COARSE), for which it falls back
to system call, increasing the response time, vDSO implementation reduces
the cycle time. Below is a benchmark of the difference in execution times.
(Non-coarse clocks are also included just for completion)
clock-gettime-realtime: syscall: 172 nsec/call
clock-gettime-realtime: libc: 28 nsec/call
clock-gettime-realtime: vdso: 22 nsec/call
clock-gettime-monotonic: syscall: 171 nsec/call
clock-gettime-monotonic: libc: 30 nsec/call
clock-gettime-monotonic: vdso: 25 nsec/call
clock-gettime-realtime-coarse: syscall: 153 nsec/call
clock-gettime-realtime-coarse: libc: 16 nsec/call
clock-gettime-realtime-coarse: vdso: 10 nsec/call
clock-gettime-monotonic-coarse: syscall: 167 nsec/call
clock-gettime-monotonic-coarse: libc: 17 nsec/call
clock-gettime-monotonic-coarse: vdso: 11 nsec/call
CC: Benjamin Herrenschmidt <benh@kernel.crashing.org>
Signed-off-by: Santosh Sivaraj <redacted>
---
arch/powerpc/kernel/asm-offsets.c | 2 +
arch/powerpc/kernel/vdso64/gettimeofday.S | 67 ++++++++++++++++++++++++++-----
2 files changed, 58 insertions(+), 11 deletions(-)
@@ -396,6 +396,8 @@ int main(void)/* Other bits used by the vdso */DEFINE(CLOCK_REALTIME,CLOCK_REALTIME);DEFINE(CLOCK_MONOTONIC,CLOCK_MONOTONIC);+DEFINE(CLOCK_REALTIME_COARSE,CLOCK_REALTIME_COARSE);+DEFINE(CLOCK_MONOTONIC_COARSE,CLOCK_MONOTONIC_COARSE);DEFINE(NSEC_PER_SEC,NSEC_PER_SEC);DEFINE(CLOCK_REALTIME_RES,MONOTONIC_RES_NSEC);
@@ -97,19 +104,57 @@ V_FUNCTION_BEGIN(__kernel_clock_gettime)ldr0,CFG_TB_UPDATE_COUNT(r3)cmpldcr0,r0,r8/*checkifupdated*/bne-50b+b78f-/*Addwall->monotonicoffsetandcheckforoverfloworunderflow.+/*+*Forcoarseclockswegetdatadirectlyfromthevdsodatapage,so+*wedon't need to call __do_get_tspec, but we still need to do the+*countertrick.*/-addr4,r4,r6-addr5,r5,r9-cmpdcr0,r5,r7-cmpdicr1,r5,0-blt1f-subfr5,r7,r5-addir4,r4,1-1:bgecr1,80f-addir4,r4,-1-addr5,r5,r7+70:ldr8,CFG_TB_UPDATE_COUNT(r3)+andi.r0,r8,1/*pendingupdate?loop*/+bne-70b+xorr0,r8,r8/*createdependency*/+addr3,r3,r0++/*+*CLOCK_REALTIME_COARSE,belowvaluesareneededforMONOTONIC_COARSE+*too+*/+ldr4,STAMP_XTIME+TSPC64_TV_SEC(r3)+ldr5,STAMP_XTIME+TSPC64_TV_NSEC(r3)+bnecr6,75f++/*CLOCK_MONOTONIC_COARSE*/+lwar6,WTOM_CLOCK_SEC(r3)+lwar9,WTOM_CLOCK_NSEC(r3)++/*checkifcounterhasupdated*/+75:orr0,r6,r9+orr0,r4,r5+xorr0,r0,r0
The label '75:' should be on the second instruction since we don't need
to worry about r6/r9 for REALTIME_COARSE.
Also, the above hunk should actually be:
or r0,r6,r9
or r0,r0,r4
or r0,r0,r5
xor r0,r0,r0
Otherwise, the first 'or' will be skipped. I realized this after I
replied to your previous version, but missed letting you know...
I also notice that the code for dealing with CLOCK_MONOTONIC is similar
for _COARSE and regular clocks. If possible, we should reuse that as
well.
- Naveen
* Naveen N. Rao [off-list ref] wrote (on 2017-10-09 10:39:18 +0000):
On 2017/10/09 08:09AM, Santosh Sivaraj wrote:
quoted
Current vDSO64 implementation does not have support for coarse clocks
(CLOCK_MONOTONIC_COARSE, CLOCK_REALTIME_COARSE), for which it falls back
to system call, increasing the response time, vDSO implementation reduces
the cycle time. Below is a benchmark of the difference in execution times.
(Non-coarse clocks are also included just for completion)
clock-gettime-realtime: syscall: 172 nsec/call
clock-gettime-realtime: libc: 28 nsec/call
clock-gettime-realtime: vdso: 22 nsec/call
clock-gettime-monotonic: syscall: 171 nsec/call
clock-gettime-monotonic: libc: 30 nsec/call
clock-gettime-monotonic: vdso: 25 nsec/call
clock-gettime-realtime-coarse: syscall: 153 nsec/call
clock-gettime-realtime-coarse: libc: 16 nsec/call
clock-gettime-realtime-coarse: vdso: 10 nsec/call
clock-gettime-monotonic-coarse: syscall: 167 nsec/call
clock-gettime-monotonic-coarse: libc: 17 nsec/call
clock-gettime-monotonic-coarse: vdso: 11 nsec/call
CC: Benjamin Herrenschmidt <benh@kernel.crashing.org>
Signed-off-by: Santosh Sivaraj <redacted>
---
arch/powerpc/kernel/asm-offsets.c | 2 +
arch/powerpc/kernel/vdso64/gettimeofday.S | 67 ++++++++++++++++++++++++++-----
2 files changed, 58 insertions(+), 11 deletions(-)
@@ -396,6 +396,8 @@ int main(void)/* Other bits used by the vdso */DEFINE(CLOCK_REALTIME,CLOCK_REALTIME);DEFINE(CLOCK_MONOTONIC,CLOCK_MONOTONIC);+DEFINE(CLOCK_REALTIME_COARSE,CLOCK_REALTIME_COARSE);+DEFINE(CLOCK_MONOTONIC_COARSE,CLOCK_MONOTONIC_COARSE);DEFINE(NSEC_PER_SEC,NSEC_PER_SEC);DEFINE(CLOCK_REALTIME_RES,MONOTONIC_RES_NSEC);
@@ -97,19 +104,57 @@ V_FUNCTION_BEGIN(__kernel_clock_gettime)ldr0,CFG_TB_UPDATE_COUNT(r3)cmpldcr0,r0,r8/*checkifupdated*/bne-50b+b78f-/*Addwall->monotonicoffsetandcheckforoverfloworunderflow.+/*+*Forcoarseclockswegetdatadirectlyfromthevdsodatapage,so+*wedon't need to call __do_get_tspec, but we still need to do the+*countertrick.*/-addr4,r4,r6-addr5,r5,r9-cmpdcr0,r5,r7-cmpdicr1,r5,0-blt1f-subfr5,r7,r5-addir4,r4,1-1:bgecr1,80f-addir4,r4,-1-addr5,r5,r7+70:ldr8,CFG_TB_UPDATE_COUNT(r3)+andi.r0,r8,1/*pendingupdate?loop*/+bne-70b+xorr0,r8,r8/*createdependency*/+addr3,r3,r0++/*+*CLOCK_REALTIME_COARSE,belowvaluesareneededforMONOTONIC_COARSE+*too+*/+ldr4,STAMP_XTIME+TSPC64_TV_SEC(r3)+ldr5,STAMP_XTIME+TSPC64_TV_NSEC(r3)+bnecr6,75f++/*CLOCK_MONOTONIC_COARSE*/+lwar6,WTOM_CLOCK_SEC(r3)+lwar9,WTOM_CLOCK_NSEC(r3)++/*checkifcounterhasupdated*/+75:orr0,r6,r9+orr0,r4,r5+xorr0,r0,r0
The label '75:' should be on the second instruction since we don't need
to worry about r6/r9 for REALTIME_COARSE.
Also, the above hunk should actually be:
or r0,r6,r9
or r0,r0,r4
or r0,r0,r5
xor r0,r0,r0
Otherwise, the first 'or' will be skipped. I realized this after I
replied to your previous version, but missed letting you know...
I also notice that the code for dealing with CLOCK_MONOTONIC is similar
for _COARSE and regular clocks. If possible, we should reuse that as
well.
In this case we will be adding more checks and branches in order to reuse
the code. If we want to keep the code common we will have to do a lot of
jumping around, code will contain a bunch of branches, which I feel will make
the code/flow hard to understand. (Q: Does lot of branches have bad effect on
branch prediction?)
Will wait for your thoughts, before respinning.
Thanks,
Santosh
I also notice that the code for dealing with CLOCK_MONOTONIC is similar
for _COARSE and regular clocks. If possible, we should reuse that as
well.
In this case we will be adding more checks and branches in order to reuse
the code. If we want to keep the code common we will have to do a lot of
jumping around, code will contain a bunch of branches, which I feel will make
the code/flow hard to understand. (Q: Does lot of branches have bad effect on
branch prediction?)
Right - like we discussed offline, if it hurts readability, that's a
good enough reason not to do this. We are only talking about a few
instructions here anyway, so no need to worry too much.
- Naveen
Current vDSO64 implementation does not have support for coarse clocks
(CLOCK_MONOTONIC_COARSE, CLOCK_REALTIME_COARSE), for which it falls back
to system call, increasing the response time, vDSO implementation reduces
the cycle time. Below is a benchmark of the difference in execution times.
(Non-coarse clocks are also included just for completion)
clock-gettime-realtime: syscall: 172 nsec/call
clock-gettime-realtime: libc: 28 nsec/call
clock-gettime-realtime: vdso: 22 nsec/call
clock-gettime-monotonic: syscall: 171 nsec/call
clock-gettime-monotonic: libc: 30 nsec/call
clock-gettime-monotonic: vdso: 25 nsec/call
clock-gettime-realtime-coarse: syscall: 153 nsec/call
clock-gettime-realtime-coarse: libc: 16 nsec/call
clock-gettime-realtime-coarse: vdso: 10 nsec/call
clock-gettime-monotonic-coarse: syscall: 167 nsec/call
clock-gettime-monotonic-coarse: libc: 17 nsec/call
clock-gettime-monotonic-coarse: vdso: 11 nsec/call
CC: Benjamin Herrenschmidt <benh@kernel.crashing.org>
Signed-off-by: Santosh Sivaraj <redacted>
---
arch/powerpc/kernel/asm-offsets.c | 2 +
arch/powerpc/kernel/vdso64/gettimeofday.S | 67 ++++++++++++++++++++++++++-----
2 files changed, 58 insertions(+), 11 deletions(-)
@@ -396,6 +396,8 @@ int main(void)/* Other bits used by the vdso */DEFINE(CLOCK_REALTIME,CLOCK_REALTIME);DEFINE(CLOCK_MONOTONIC,CLOCK_MONOTONIC);+DEFINE(CLOCK_REALTIME_COARSE,CLOCK_REALTIME_COARSE);+DEFINE(CLOCK_MONOTONIC_COARSE,CLOCK_MONOTONIC_COARSE);DEFINE(NSEC_PER_SEC,NSEC_PER_SEC);DEFINE(CLOCK_REALTIME_RES,MONOTONIC_RES_NSEC);
@@ -97,19 +104,57 @@ V_FUNCTION_BEGIN(__kernel_clock_gettime)ldr0,CFG_TB_UPDATE_COUNT(r3)cmpldcr0,r0,r8/*checkifupdated*/bne-50b+b78f-/*Addwall->monotonicoffsetandcheckforoverfloworunderflow.+/*+*Forcoarseclockswegetdatadirectlyfromthevdsodatapage,so+*wedon't need to call __do_get_tspec, but we still need to do the+*countertrick.*/-addr4,r4,r6-addr5,r5,r9-cmpdcr0,r5,r7-cmpdicr1,r5,0-blt1f-subfr5,r7,r5-addir4,r4,1-1:bgecr1,80f-addir4,r4,-1-addr5,r5,r7+70:ldr8,CFG_TB_UPDATE_COUNT(r3)+andi.r0,r8,1/*pendingupdate?loop*/+bne-70b+xorr0,r8,r8/*createdependency*/+addr3,r3,r0++/*+*CLOCK_REALTIME_COARSE,belowvaluesareneededforMONOTONIC_COARSE+*too+*/+ldr4,STAMP_XTIME+TSPC64_TV_SEC(r3)+ldr5,STAMP_XTIME+TSPC64_TV_NSEC(r3)+bnecr6,75f++/*CLOCK_MONOTONIC_COARSE*/+lwar6,WTOM_CLOCK_SEC(r3)+lwar9,WTOM_CLOCK_NSEC(r3)++/*checkifcounterhasupdated*/+orr0,r6,r9+75:orr0,r4,r5+xorr0,r0,r0+addr3,r3,r0+ldr0,CFG_TB_UPDATE_COUNT(r3)+cmpldcr0,r0,r8/*checkifupdated*/+bne-70b++/*Counterhasnotupdated,socontinuecalculatingpropervaluesfor+*secandnsecifmonotoniccoarse,orjustreturnwiththeproper+*valuesforrealtime.+*/+bnecr6,80f++/*Addwall->monotonicoffsetandcheckforoverfloworunderflow*/+78:addr4,r4,r6+addr5,r5,r9+cmpdcr0,r5,r7+cmpdicr1,r5,0+blt79f+subfr5,r7,r5+addir4,r4,1+79:bgecr1,80f+addir4,r4,-1+addr5,r5,r780:stdr4,TSPC64_TV_SEC(r11)stdr5,TSPC64_TV_NSEC(r11)
From: Naveen N. Rao <hidden> Date: 2017-10-11 07:04:56
Hi Santosh,
This seems to have gone from v4 to v6 -- did I miss v5?
On 2017/10/10 11:10PM, Santosh Sivaraj wrote:
Current vDSO64 implementation does not have support for coarse clocks
(CLOCK_MONOTONIC_COARSE, CLOCK_REALTIME_COARSE), for which it falls back
to system call, increasing the response time, vDSO implementation reduces
the cycle time. Below is a benchmark of the difference in execution times.
(Non-coarse clocks are also included just for completion)
clock-gettime-realtime: syscall: 172 nsec/call
clock-gettime-realtime: libc: 28 nsec/call
clock-gettime-realtime: vdso: 22 nsec/call
clock-gettime-monotonic: syscall: 171 nsec/call
clock-gettime-monotonic: libc: 30 nsec/call
clock-gettime-monotonic: vdso: 25 nsec/call
clock-gettime-realtime-coarse: syscall: 153 nsec/call
clock-gettime-realtime-coarse: libc: 16 nsec/call
clock-gettime-realtime-coarse: vdso: 10 nsec/call
clock-gettime-monotonic-coarse: syscall: 167 nsec/call
clock-gettime-monotonic-coarse: libc: 17 nsec/call
clock-gettime-monotonic-coarse: vdso: 11 nsec/call
CC: Benjamin Herrenschmidt <benh@kernel.crashing.org>
Signed-off-by: Santosh Sivaraj <redacted>
---
arch/powerpc/kernel/asm-offsets.c | 2 +
arch/powerpc/kernel/vdso64/gettimeofday.S | 67 ++++++++++++++++++++++++++-----
2 files changed, 58 insertions(+), 11 deletions(-)
... and no changes since the last rev?
It is better to post new versions in a separate thread and to include
the changelog for easier review.
- Naveen
* Naveen N. Rao [off-list ref] wrote (on 2017-10-11 07:04:43 +0000):
Hi Santosh,
This seems to have gone from v4 to v6 -- did I miss v5?
Nope, this is indeed v5, a typo :-(
On 2017/10/10 11:10PM, Santosh Sivaraj wrote:
quoted
Current vDSO64 implementation does not have support for coarse clocks
(CLOCK_MONOTONIC_COARSE, CLOCK_REALTIME_COARSE), for which it falls back
to system call, increasing the response time, vDSO implementation reduces
the cycle time. Below is a benchmark of the difference in execution times.
(Non-coarse clocks are also included just for completion)
clock-gettime-realtime: syscall: 172 nsec/call
clock-gettime-realtime: libc: 28 nsec/call
clock-gettime-realtime: vdso: 22 nsec/call
clock-gettime-monotonic: syscall: 171 nsec/call
clock-gettime-monotonic: libc: 30 nsec/call
clock-gettime-monotonic: vdso: 25 nsec/call
clock-gettime-realtime-coarse: syscall: 153 nsec/call
clock-gettime-realtime-coarse: libc: 16 nsec/call
clock-gettime-realtime-coarse: vdso: 10 nsec/call
clock-gettime-monotonic-coarse: syscall: 167 nsec/call
clock-gettime-monotonic-coarse: libc: 17 nsec/call
clock-gettime-monotonic-coarse: vdso: 11 nsec/call
CC: Benjamin Herrenschmidt <benh@kernel.crashing.org>
Signed-off-by: Santosh Sivaraj <redacted>
---
arch/powerpc/kernel/asm-offsets.c | 2 +
arch/powerpc/kernel/vdso64/gettimeofday.S | 67 ++++++++++++++++++++++++++-----
2 files changed, 58 insertions(+), 11 deletions(-)
... and no changes since the last rev?
There is that one line change of label. But otherwise its the same.
It is better to post new versions in a separate thread and to include
the changelog for easier review.