git: f12b76524d7b - main - callout: do not retry a try-lock callout sooner than a tick

From: R. Christian McDonald <rcm_at_FreeBSD.org>
Date: Sat, 03 Oct 2026 23:58:27 UTC
The branch main has been updated by rcm:

URL: https://cgit.FreeBSD.org/src/commit/?id=f12b76524d7bbf3d6c1a890834b6f8d018704ef0

commit f12b76524d7bbf3d6c1a890834b6f8d018704ef0
Author:     R. Christian McDonald <rcm@FreeBSD.org>
AuthorDate: 2026-10-03 23:55:32 +0000
Commit:     R. Christian McDonald <rcm@FreeBSD.org>
CommitDate: 2026-10-03 23:58:05 +0000

    callout: do not retry a try-lock callout sooner than a tick
    
    When softclock_call_cc() fails to acquire the lock of a CALLOUT_TRYLOCK
    callout, it reschedules the callout half its precision after
    cc_lastscan and halves the precision. Repeated failures shrink the
    delay toward zero, and once the precision reaches 1 the callout is due
    immediately: the timer fires again at once and softclock retries the
    lock in a tight loop for as long as the lock is held.
    
    If the lock owner runs on the callout's CPU and no other CPU is idle,
    the softclock thread preempts it on every attempt, starving the thread
    it is waiting on. On an 8-CPU arm64 VM, a test module holding the lock
    saw 760,000 attempts per second, each with its own timer interrupt, and
    progressed at 38% of its normal rate. On a 4-core amd64 system under
    loopback TCP load, a netisr thread holding an inpcb lock made no
    progress for 12 minutes while the TCP timer callout was retried 830,000
    times per second.
    
    Keep the half-precision retry, but never schedule it less than one
    tick from now. Base the deadline on sbinuptime() rather than
    cc_lastscan: cc_lastscan is the time of the last callout_process()
    scan, which can be a tick or more in the past by the time softclock
    runs the callout, leaving cc_lastscan + tick_sbt already expired.
    
    Floor the precision at a tick as well. With a precision of 1 the
    retries of many contended callouts cannot share a timer interrupt: the
    timer is armed for the earliest retry, each interrupt collects only
    those already due, and every callout_process() call walks all the
    contended callouts in the callwheel bucket. On the same VM, with
    10,000 callouts on one lock, the owner took 5.0 to 5.8 s to finish 1 s
    of work with a precision of 1 and 1.3 s with the floor, and
    callout_process() ran about 58,000 times against about 900. A single
    contended callout is now retried about every two ticks.
    
    Reviewed by:            kib
    Fixes:                  efcb2ec8cb81 ("callout: provide CALLOUT_TRYLOCK flag")
    MFC after:              2 weeks
    Sponsored by:           Rubicon Communications, LLC ("Netgate")
    Differential Revision:  https://reviews.freebsd.org/D60246
---
 sys/kern/kern_timeout.c | 23 ++++++++++++++++++++---
 1 file changed, 20 insertions(+), 3 deletions(-)

diff --git a/sys/kern/kern_timeout.c b/sys/kern/kern_timeout.c
index 1f97e4f813a1..05856921df1d 100644
--- a/sys/kern/kern_timeout.c
+++ b/sys/kern/kern_timeout.c
@@ -634,6 +634,7 @@ softclock_call_cc(struct callout *c, struct callout_cpu *cc,
 	struct lock_class *class;
 	struct lock_object *c_lock;
 	uintptr_t lock_status;
+	sbintime_t prec;
 	int c_iflags;
 #ifdef SMP
 	struct callout_cpu *new_cc;
@@ -676,10 +677,26 @@ softclock_call_cc(struct callout *c, struct callout_cpu *cc,
 		if (c_iflags & CALLOUT_TRYLOCK) {
 			if (__predict_false(class->lc_trylock(c_lock,
 			    lock_status) == 0)) {
+				/*
+				 * Retry after half of the precision, but not
+				 * sooner than a tick from the current time.
+				 * Every failure halves the precision, so the
+				 * retry delay would otherwise shrink to zero.
+				 * Immediate retries can starve a
+				 * lower-priority lock owner on the same CPU,
+				 * preventing it from running to release the
+				 * lock.
+				 *
+				 * Use the current time because cc_lastscan may
+				 * be stale by the time the lock is attempted.
+				 * Keep the precision at least a tick as well,
+				 * so that the retries of many contended
+				 * callouts can share a timer interrupt.
+				 */
 				cc_exec_curr(cc, direct) = NULL;
-				callout_cc_add(c, cc,
-				    cc->cc_lastscan + c->c_precision / 2,
-				    qmax(c->c_precision / 2, 1), c_func, c_arg,
+				prec = qmax(c->c_precision / 2, tick_sbt);
+				callout_cc_add(c, cc, sbinuptime() + prec,
+				    prec, c_func, c_arg,
 				    (direct) ? C_DIRECT_EXEC : 0);
 				return;
 			}