You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Commit 7c83fd3
Browse filesBrowse the repository at this point in the historyBrowse files
From an audit of every performance figure quoted outside a table: the
Performance paragraph says 4.3 and 10.3 times like the table; the conditions
paragraph states each set's date and run policy (the comparison, in-process and
WebAssembly sets stay on 2026-09-23); the yielding row's allocation is the
table's 559 B at one worker, including the service's own; the async ranges use
the table's precision; the sweep chart's whiskers span five runs, not two; the
WasmHost conditions say which column is one run and which the better of two.
Copy file name to clipboardExpand all lines: CHANGELOG.md
+1-1Lines changed: 1 addition & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -30,7 +30,7 @@ behaviour: a breaking change to either means a new major version.
30
30
### Changed
31
31
32
32
- The core no longer depends on Json.NET.
33
-
-`ProcessAsync` no longer serializes the process on one lock per document. Each thread now caches one async scratch (input copy, reader, staged output) in front of the shared pool, which handles only misses and overflow. On 16 threads, the inline rows went from about 4 M to 22 M to 32 M RPC/s, and the row with a real suspension went from 3.9 M to 9 M. The library retains one scratch per thread that has run `ProcessAsync` plus 64 shared, with buffers at most 64 KiB each.
33
+
-`ProcessAsync` no longer serializes the process on one lock per document. Each thread now caches one async scratch (input copy, reader, staged output) in front of the shared pool, which handles only misses and overflow. At 16 workers, the inline rows went from about 4 M to 22.2 M to 32.1 M RPC/s across the registrations, and the row with a real suspension from 3.9 M to 8.96 M (one run per row, 2026-09-25). The library retains one scratch per thread that has run `ProcessAsync` plus 64 shared, with buffers at most 64 KiB each.
34
34
- The request path looks sessions up without creating them. A request for a session id that was never registered answers `-32601` for every call and leaves the registry untouched; sessions are created by binding and by the per-session `Config` setters. Registration adds the session before publishing the registry version, so a thread that misses its snapshot consults the master registry and cannot answer `-32601` for a session that exists.
35
35
- The AspNetCore host binds every registered service, `JsonRpcService` subclasses included, to its effective session (the registration's session, then `JsonRpcOptions.SessionId`, then the default). It no longer skips a subclass on the default session.
36
36
- The core package's description says "no JSON library dependency" instead of "no dependencies". The session registry uses the framework's `ConcurrentDictionary`; the `NonBlocking` package reference is gone, so the core has no dependencies on `net8.0` and `net10.0` (measured with `SessionRegistryBenchmarks`: unknown-id lookups and register/destroy cycles got faster, stable lookups and dispatch are unchanged).
Copy file name to clipboardExpand all lines: README.md
+9-9Lines changed: 9 additions & 9 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -8,11 +8,11 @@ Version 2.0 rebuilt the pipeline around UTF-8 bytes and made the JSON serializer
8
8
9
9
## Performance
10
10
11
-
2.0 answers the same five requests 4 times faster than 1.2.3 through the 1.x string API and 10 times faster through the byte entry points its hosts use. The byte path handles 31.7 M requests per second on one 8-core desktop, with no allocation for a numeric request. All results come from the same machine, the same requests and one session.
11
+
2.0 answers the same five requests 4.3 times faster than 1.2.3 through the 1.x string API and 10.3 times faster through the byte entry points its hosts use. The synchronous byte path handles 31.7 M requests per second on 16 threads of one 8-core desktop, with no allocation for a numeric request. All results come from the same machine, the same requests and one session, one run per row.
<imgalt="JSON-RPC.Net 1.2.3 and 2.0 on one machine: 3.08 M requests per second through the 1.2.3 string API, 13.3 M through the same API on 2.0, and 31.7 M and 32.1 M through the 2.0 byte entry points"src="benchmarks/charts/headline-1x-vs-2.svg">
15
+
<imgalt="JSON-RPC.Net 1.2.3 and 2.0 on one machine, one run per row on 2026-09-25: 3.08 M requests per second through the 1.2.3 string API, 13.3 M through the same API on 2.0, and 31.7 M and 32.1 M through the 2.0 byte entry points"src="benchmarks/charts/headline-1x-vs-2.svg">
16
16
</picture>
17
17
18
18
| Path | RPC/s | Against 1.2.3 |
@@ -22,7 +22,7 @@ Version 2.0 rebuilt the pipeline around UTF-8 bytes and made the JSON serializer
The request path uses a span tokenizer over the UTF-8 bytes and invokers compiled against the concrete reader and writer. It writes the response straight into a pooled buffer, with no `Task`, result string or continuation per request. A numeric request never touches the GC. Over Kestrel TCP, the AspNetCore package holds 14.3 M to 16.5 M with 256 requests in flight per connection. [Benchmarks](#benchmarks) has the full tables, conditions and day-to-day ranges. The [explorer](https://astn.github.io/JSON-RPC.NET/benchmarks/charts/explorer.html) has the exact values. `benchmarks/Baseline` re-runs the 1.2.3 row from NuGet with the same loop as the 2.0 harness.
25
+
The request path uses a span tokenizer over the UTF-8 bytes and invokers compiled against the concrete reader and writer. It writes the response straight into a pooled buffer, with no `Task`, result string or continuation per request. A numeric request never touches the GC. Over Kestrel TCP, the AspNetCore package holds 14.3 M to 16.5 M with 256 requests in flight per connection. [Benchmarks](#benchmarks) has the full tables, conditions and the ranges over each day's runs. The [explorer](https://astn.github.io/JSON-RPC.NET/benchmarks/charts/explorer.html) has the exact values. `benchmarks/Baseline` re-runs the 1.2.3 row from NuGet with the same loop as the 2.0 harness.
26
26
27
27
28
28
It is a server only. There are no client proxies and no server-to-client calls. If you need a bidirectional RPC framework, look at [StreamJsonRpc](https://www.nuget.org/packages/StreamJsonRpc); the benchmarks compare the two.
@@ -308,7 +308,7 @@ In the second `-32602` row, "could not convert" means the serializer refused the
308
308
309
309
A method may return `Task`, `Task<T>`, `ValueTask` or `ValueTask<T>`. Call it through `JsonRpcProcessor.ProcessAsync`. The processor awaits the method and writes its result; `Task` and `ValueTask` answer `null`. A method that completes synchronously runs inline and the byte overloads then return `Task.CompletedTask`.
310
310
311
-
**Cost.** With `RpcContextFlow.None`, the built-in numeric fast path adds no dispatcher allocation when the invocation completes inline. The service's own allocations (a `Task.FromResult`, a result string) are separate and included in the harness figures. Flow allocates an `InvocationState` even when the call completes inline. A method that suspends may allocate its own async state plus completion state in the result writer, the request handler and the document processor. The yielding benchmark measured about 560 B per request under None, plus a continuation per request. The figures are in [Async](#async-processasync-awaited-workers).
311
+
**Cost.** With `RpcContextFlow.None`, the built-in numeric fast path adds no dispatcher allocation when the invocation completes inline. The service's own allocations (a `Task.FromResult`, a result string) are separate and included in the harness figures. Flow allocates an `InvocationState` even when the call completes inline. A method that suspends may allocate its own async state plus completion state in the result writer, the request handler and the document processor. The yielding benchmark measured 559 B per request under None in a single run at one worker on 2026-09-25, including the service's allocations. A suspension also costs a continuation per request. The figures are in [Async](#async-processasync-awaited-workers).
312
312
313
313
```csharp
314
314
[JsonRpcMethod("lookup")]
@@ -510,7 +510,7 @@ The `jsonrpc` member policy (`Lenient` by default) is a compatibility setting, n
510
510
| Kestrel HTTP, one request per POST |`EnableAsyncMethods = false`| 123 k to 182 k | HTTP/1.1 round trips dominate |
511
511
| Legacy string API, thread pool |`Task<string> Process(string)`, batches of 6,000 | 12.4 M to 13.3 M |[Legacy](#legacy-string-api-scheduled-synchronous-execution); the 1.x overloads, not the byte path |
512
512
513
-
All numbers below are from an AMD Ryzen 7 7800X3D (8 cores / 16 threads, 4.2 GHz), 64 GB, Windows 11, .NET 10, Release, Server GC, with the built-in serializer, measured 2026-09-25 on an idle machine. The exceptions are the StreamJsonRpc and gRPC comparison, the connection sweep and the WebAssembly rows, which carry their own dates. Where a row gives two figures, they are the low and high over that day's runs: three runs for the Sync, Legacy and default-mode Kestrel rows, two for the `EnableAsyncMethods = true` rows. A single figure is one 3 s run. The WSL virtual machine takes 15 to 25 % of the box when idle and was shut down for the comparison runs. A single benchmark thread on this machine varies with whatever else lands on its core's SMT sibling, so the 1-thread rows are from runs on an idle core.
513
+
The Sync, Async, Legacy and Kestrel results were measured on 2026-09-25 on an idle AMD Ryzen 7 7800X3D (8 cores / 16 threads, 4.2 GHz), 64 GB, Windows 11, .NET 10, Release, Server GC, with the built-in serializer. Throughput ranges are the low and high over that day's runs: three runs for Sync, Legacy and default-mode Kestrel, two for the `EnableAsyncMethods = true` rows. Each Async throughput cell is one 3 s run; the summary's inline Async spread is across registrations, not runs. Each Sync ns figure is the reported cost from one of that row's runs. The StreamJsonRpc and gRPC comparison and its in-process rows are from 2026-09-23, with the WSL virtual machine shut down. Their ranges cover two runs; the comparison's single throughput figure is the value both runs rounded to. The WebAssembly results are also from 2026-09-23: one interpreter run and the better of two AOT runs per row. The connection sweep is a separate session with WSL and other work active, so its absolute figures are not comparable with the tables. A single benchmark thread on this machine varies with whatever else lands on its core's SMT sibling, so the 1-thread rows are from runs on an idle core.
514
514
515
515
`TestServer_Console` is the benchmark harness. It binds one service with five small methods (`add`, `addInt`, `NullableFloatToNullableFloat`, `Test2`, `StringMe`), drives the same five requests through the server, checks every response is a `result` rather than an error, and ends each mode with a bar chart of RPC/s. For one-request timings with an allocation column, the numbers to check before merging a change to the dispatch path, see [benchmarks/Micro](benchmarks/Micro/README.md).
516
516
@@ -571,9 +571,9 @@ The 16-worker rows are an equal-weight mix of the five requests, except `yieldsO
With `RpcContextFlow.None`, the dispatcher adds no allocation to a method that completes inline. Allocations in the `Task<T>` None row come from the service's own `Task.FromResult` (`Task<int>` for 8 comes from the runtime's cache). The 32 bytes of `StringMe` are its result string. Flow allocates the `InvocationState` and the execution-context bridge on every call, inline or not. A real suspension allocates the method's own async state plus completion state in the result writer, the request handler and the document processor. This costs about 560 B per request in the `yieldsOnce` None row. The 7 to 10 M target for a hosted server applies to methods that complete inline. A method that suspends also costs a continuation per request.
574
+
With `RpcContextFlow.None`, the dispatcher adds no allocation to a method that completes inline. Allocations in the `Task<T>` None row come from the service's own `Task.FromResult` (`Task<int>` for 8 comes from the runtime's cache). The 32 bytes of `StringMe` are its result string. Flow allocates the `InvocationState` and the execution-context bridge on every call, inline or not. A real suspension allocates the method's own async state plus completion state in the result writer, the request handler and the document processor. The `yieldsOnce` None row measured 559 B per request at one worker, including the method's own allocations. The 7 to 10 M target for a hosted server applies to methods that complete inline. A method that suspends also costs a continuation per request.
575
575
576
-
Before 2.0.0, the `ProcessAsync` path was capped near 4 M RPC/s at every worker count because each document took one lock on the shared scratch pool. No single-threaded benchmark could see this limit. Each thread now caches one scratch in front of that pool. `--scale [seconds] [workers] [threshold]` is the gate that catches the next such limit. It measures the inline None rows at 1, 2 and 16 workers in three paired runs and takes the medians. It exits non-zero when any 16/1 ratio is below 4.0. The lock gave 1.3; the cache gives 7.1 to 7.3 on the idle reference machine. Run it on the reference machine before a release and paste its table into the release notes. The pull-request build runs a diagnostic `--scale 3 4 2.0` on the shared runner. It also checks that every `lock`, `Interlocked`, `Volatile.Write`, thread-static and writable static field on the request-path files of the core and both companion serializers is listed in `.github/request-path-sync.allowlist` with a reason (per-thread, miss-path, registration-only, read-only-after-init).
576
+
Before 2.0.0, the `ProcessAsync` path was capped near 4 M RPC/s at every worker count because each document took one lock on the shared scratch pool. No single-threaded benchmark could see this limit. Each thread now caches one scratch in front of that pool. `--scale [seconds] [workers] [threshold]` is the gate that catches the next such limit. It measures the inline None rows at 1, 2 and 16 workers in three paired runs and takes the medians. It exits non-zero when any 16/1 ratio is below 4.0. The lock gave 1.3; the cache gives 7.1 to 7.3 on the idle reference machine (the gate's medians, 2026-09-25). Run it on the reference machine before a release and paste its table into the release notes. The pull-request build runs a diagnostic `--scale 3 4 2.0` on the shared runner. It also checks that every `lock`, `Interlocked`, `Volatile.Write`, thread-static and writable static field on the request-path files of the core and both companion serializers is listed in `.github/request-path-sync.allowlist` with a reason (per-thread, miss-path, registration-only, read-only-after-init).
<imgalt="Every library and transport by client connections, 1 to 16: three panels on a shared log axis, one per library, with a marker shape and dash per setting and whiskers spanning two runs"src="benchmarks/charts/compare-connections.svg">
655
+
<imgalt="Every library and transport by client connections, 1 to 16: three panels on a shared log axis, one per library, with a marker shape and dash per setting and whiskers spanning five runs"src="benchmarks/charts/compare-connections.svg">
656
656
</picture>
657
657
658
658
`--sweep` runs every one of those paths at 1, 2, 4, 8 and 16 client connections (gRPC: channels) and writes one JSON file per run; the chart above is five 2 s runs per point, the marker at the median and the whisker from the lowest to the highest run. It is a separate session from the tables: the WSL virtual machine was running and other work was active, so its absolute figures sit below the table rows (JSON-RPC.Net over TCP 10.9 M to 13.0 M at 16 connections against 13.7 M to 14.6 M in the table), and its gRPC unary figure runs higher (360 k to 404 k against 192 k to 198 k; the cause is not pinned down, and the table keeps the `--compare` figure). What the sweep adds is the shape: JSON-RPC.Net over TCP and batched HTTP climb almost linearly with connections, StreamJsonRpc gains 8 to 10× from one connection to sixteen, and gRPC's .NET client is flat from two channels on because it competes with the server for the same eight cores.
@@ -671,7 +671,7 @@ simdjson was evaluated as a fourth parser and not adopted: through the only main
671
671
672
672
The charts, the explorer page and the figures in this file come from one data file, [benchmarks/charts/benchmarks.json](benchmarks/charts/benchmarks.json); how they are rendered and checked is under [Building](#charts).
673
673
674
-
On 2026-09-25 the `ProcessAsync` path was found capped near 4 M RPC/s at every worker count. Every document took one lock on the shared scratch pool, which the single-threaded micro-benchmarks could not detect. A one-slot per-thread cache in front of the pool took the inline rows to 22 M to 32 M at 16 workers, against 31.7 M for the synchronous entry point in the same session. The `--scale` gate and the request-path allowlist exist to catch the next such limit before a release.
674
+
On 2026-09-25 the `ProcessAsync` path was found capped near 4 M RPC/s at every worker count. Every document took one lock on the shared scratch pool, which the single-threaded micro-benchmarks could not detect. A one-slot per-thread cache in front of the pool took the inline rows to 22.2 M to 32.1 M at 16 workers (one run per registration), against 31.7 M for the synchronous entry point in the same session. The `--scale` gate and the request-path allowlist exist to catch the next such limit before a release.
675
675
676
676
The Performance section at the top reports one session on 2026-09-25. The last 1.x release on NuGet, 1.2.3, was driven by the same loop as the `t` entry ([benchmarks/Baseline](benchmarks/Baseline/Program.cs)). It reached 3.08 M at its best batch size of 1,200 and fell to 1.5 M at two million. In the same session, 2.0 peaked at 13.3 M through the same string API, and the byte entry points ran at 31.7 M and 32.1 M.
0 commit comments