You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: .github/workflows/build_pull_request.yml
+8-1Lines changed: 8 additions & 1 deletion
Original file line number
Diff line number
Diff line change
@@ -47,6 +47,7 @@ jobs:
47
47
# The measurement: ProcessAsync must scale from 1 to 4 workers. A process-wide serialization point holds the ratio
48
48
# near 1.3 on any core count. Diagnostic on the shared runner (its core count and isolation are not promised, so
49
49
# this job is not required and continues on error); the release gate is `--scale 3 16 4.0` on the reference machine.
50
+
# The per-serializer diagnostics after the gate (one run per cell, never part of the exit code) go to their own summary section.
50
51
scaling:
51
52
runs-on: ubuntu-latest
52
53
continue-on-error: true
@@ -67,8 +68,14 @@ jobs:
67
68
{
68
69
echo "## ProcessAsync scaling (diagnostic, not required)"
69
70
echo
70
-
grep -E '^\|' scale.txt
71
+
sed '/^Diagnostics/,$d' scale.txt | grep -E '^\|'
71
72
echo
72
73
[ "$status" -eq 0 ] && echo "pass: 4/1 at least 2.0" || echo "**flag: 4/1 below 2.0 on this runner; reproduce with --scale 3 16 4.0 on the reference machine before reading it as a regression**"
74
+
echo
75
+
echo "### Per-serializer diagnostics (one run per cell, not gated)"
76
+
echo
77
+
sed -n '/^Diagnostics/,$p' scale.txt | grep -E '^\|'
Copy file name to clipboardExpand all lines: CHANGELOG.md
+1Lines changed: 1 addition & 0 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -27,6 +27,7 @@ behaviour: a breaking change to either means a new major version.
27
27
-`protected JsonRpcService(bool autoBind)`: a subclass constructed with `base(false)` binds itself nowhere, for services that a host or an explicit `BindService` call binds.
28
28
-`SECURITY.md` (private vulnerability reporting) and this changelog.
29
29
-`TestServer_Console --scale` is the release gate for the `ProcessAsync` path. It measures the inline rows at 1, 2 and N workers in three paired runs, takes the medians and fails when N/1 is below the threshold. `--kestrel [seconds] async` runs the host with `EnableAsyncMethods = true`. The README adds `--async` rows for `ProcessAsync` at 1 and 16 workers. The 1.x string overloads' thread-pool benchmark is now the `t` menu entry and no longer the default.
30
+
-`--scale` prints per-serializer and yielding-row diagnostics after the gate; the Kestrel `EnableAsyncMethods = true` row with methods that suspend once is a release-required regression row: re-measured before each release against the previous release's figure, with no absolute floor.
Copy file name to clipboardExpand all lines: README.md
+3-3Lines changed: 3 additions & 3 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -619,7 +619,7 @@ The Sync, Async, Legacy and Kestrel results were measured on 2026-09-25 on an id
619
619
```
620
620
dotnet run -c Release --project TestServer_Console -- --sync 3 # library only, 1..N threads (add a thread count, e.g. --sync 3 1, for one row)
621
621
dotnet run -c Release --project TestServer_Console -- --async 3 16 # ProcessAsync from 16 awaited workers, one row per registration (--async 3 1 for the 1-worker column)
622
-
dotnet run -c Release --project TestServer_Console -- --scale 3 16 4.0 # release gate: ProcessAsync must scale at least 4x from 1 to 16 workers
622
+
dotnet run -c Release --project TestServer_Console -- --scale 3 16 4.0 # release gate: ProcessAsync must scale at least 4x from 1 to 16 workers, then per-serializer diagnostics (--no-diagnostics skips them)
623
623
dotnet run -c Release --project TestServer_Console -- --kestrel 3 # through the AspNetCore package, HTTP and TCP (add `async` for EnableAsyncMethods = true)
624
624
dotnet run -c Release --project TestServer_Console -- --compare 3 # the same calls through StreamJsonRpc and gRPC for .NET, side by side
625
625
dotnet run -c Release --project TestServer_Console -- --sweep 2 benchmarks/charts/sweep.json # every library and transport at 1, 2, 4, 8, 16 connections; one file per run
@@ -675,7 +675,7 @@ The 16-worker rows are an equal-weight mix of the five requests, except `yieldsO
675
675
676
676
With `RpcContextFlow.None`, the dispatcher adds no allocation to a method that completes inline. Allocations in the `Task<T>` None row come from the service's own `Task.FromResult` (`Task<int>` for 8 comes from the runtime's cache). The 32 bytes of `StringMe` are its result string. Flow allocates the `InvocationState` and the execution-context bridge on every call, inline or not. A real suspension allocates the method's own async state plus completion state in the result writer, the request handler and the document processor. The `yieldsOnce` None row measured 559 B per request at one worker, including the method's own allocations. The 7 to 10 M target for a hosted server applies to methods that complete inline. A method that suspends also costs a continuation per request.
677
677
678
-
Before 2.0.0, the `ProcessAsync` path was capped near 4 M RPC/s at every worker count because each document took one lock on the shared scratch pool. No single-threaded benchmark could see this limit. Each thread now caches one scratch in front of that pool. `--scale [seconds] [workers] [threshold]` is the gate that catches the next such limit. It measures the inline None rows at 1, 2 and 16 workers in three paired runs and takes the medians. It exits non-zero when any 16/1 ratio is below 4.0. The lock gave 1.3; the cache gives 7.1 to 7.3 on the idle reference machine (the gate's medians, 2026-09-25). Run it on the reference machine before a release and paste its table into the release notes. The pull-request build runs a diagnostic `--scale 3 4 2.0` on the shared runner. It also checks that every `lock`, `Interlocked`, `Volatile.Write`, thread-static and writable static field on the request-path files of the core and both companion serializers is listed in `.github/request-path-sync.allowlist` with a reason (per-thread, miss-path, registration-only, read-only-after-init).
678
+
Before 2.0.0, the `ProcessAsync` path was capped near 4 M RPC/s at every worker count because each document took one lock on the shared scratch pool. No single-threaded benchmark could see this limit. Each thread now caches one scratch in front of that pool. `--scale [seconds] [workers] [threshold]` is the gate that catches the next such limit. It measures the inline None rows at 1, 2 and 16 workers in three paired runs and takes the medians. It exits non-zero when any 16/1 ratio is below 4.0. After the gate it prints a diagnostics table that is never gated: the three inline None rows and the `yieldsOnce` None row at 1 and 16 workers under each serializer (`jsmn`, `stj`, `newtonsoft`), one run per cell, with the 16/1 ratio and the bytes per request at one worker, so that a regression in one serializer's path or in the suspending path shows up on its own. A final `--no-diagnostics` argument skips the table. The lock gave 1.3; the cache gives 7.1 to 7.3 on the idle reference machine (the gate's medians, 2026-09-25). Run it on the reference machine before a release and paste its gate table into the release notes. The same release run measures the `--kestrel 3 async` TCP row with methods that suspend once and compares it with the previous release's published figure: a drop larger than the paired-run spread is a release blocker, and the row has no absolute floor (see [Kestrel](#kestrel-through-the-aspnetcore-package)). The pull-request build runs a diagnostic `--scale 3 4 2.0` on the shared runner. It also checks that every `lock`, `Interlocked`, `Volatile.Write`, thread-static and writable static field on the request-path files of the core and both companion serializers is listed in `.github/request-path-sync.allowlist` with a reason (per-thread, miss-path, registration-only, read-only-after-init).
@@ -710,7 +710,7 @@ This mode is slower than the byte modes because it measures the cost of the .NET
710
710
| TCP, 256 pipelined, `EnableAsyncMethods = true`, methods that complete inline | 15.0 M to 15.4 M |`--kestrel 3 async`, two runs; inside the spread of the `false` row |
711
711
| TCP, 256 pipelined, `EnableAsyncMethods = true`, methods that suspend once | 1.25 M to 1.29 M | five `async Task<T>` methods awaiting `Task.Yield()`|
712
712
713
-
The TCP client keeps 256 requests in flight per connection and refills from a precomputed ring of request bytes with one `Send` per refill. On the server, `JsonFramer` feeds the same `Process` call the HTTP endpoint makes. With `EnableAsyncMethods = true`, the connection handler processes each connection's documents one at a time, in order. So 256 pipelined requests are 256 sequential invocations, and a method that suspends pays that cost per request. Concurrency comes from the 16 connections.
713
+
The TCP client keeps 256 requests in flight per connection and refills from a precomputed ring of request bytes with one `Send` per refill. On the server, `JsonFramer` feeds the same `Process` call the HTTP endpoint makes. With `EnableAsyncMethods = true`, the connection handler processes each connection's documents one at a time, in order. So 256 pipelined requests are 256 sequential invocations, and a method that suspends pays that cost per request. Concurrency comes from the 16 connections. The last row is a release-required regression row: it is measured with `--kestrel 3 async` on the reference machine before every release and compared with the previous release's published figure, a drop larger than the paired-run spread blocks the release, and it has no absolute floor.
print($"| Serializer / registration | 1 worker | {workers} workers | {workers}/1 | B per request |");
176
+
print("| --- | ---: | ---: | ---: | ---: |");
177
+
foreach(varrowinrows)print(row);
178
+
print("");
179
+
if(failure!=null)
180
+
print($"Diagnostics stopped at {where}: {failure.GetType().Name}: {failure.Message}. The gate result above stands; the diagnostics are not part of it.");
print($"{shape} / {flow} / shape {i+1}: {totalBytes} B total over 2000 requests ({totalBytes/2000.0:F1} B/RPC), including method allocations.");
218
+
if(report)print($"{shape} / {flow} / shape {i+1}: {totalBytes} B total over 2000 requests ({totalBytes/2000.0:F1} B/RPC), including method allocations.");
0 commit comments