Skip to content

Runtime & performance

Targets: net10.0, csharp-14 · Last reviewed: 2026-09-14 · Sources: stephen-toub, dotnet-blog, ms-learn, steve-gordon, jetbrains-dotnet, jon-skeet, andrew-lock, ardalis

The .NET 10 JIT rewards idiomatic code. Optimize by measuring, not by folklore.

Opinions

Immutable collections

Pick the immutable collection by how the collection is built, not by the word "immutable". ImmutableList<T> and ImmutableDictionary<K,V> pay for incremental change: their tree structure exists so Add can return a new collection sharing most of the old one, and every read walks that tree. Most application state isn't built that way: it's assembled once at startup or per refresh, then read constantly and replaced wholesale. For that shape use ImmutableArray<T> and FrozenDictionary<K,V>, which trade construction cost for flat, fast reads. Skeet measured a validation pass over election data drop from 5.5ms to 0.826ms on the switch, "due to it performing lots of read accesses". (Skeet: Changing Immutable Collections)

Keep the Immutable* builders only where the incremental-change semantics are the point (an accumulating snapshot handed to concurrent readers between edits). One migration cost to expect: ImmutableArray<T> is a struct, so default is a valid-but-unusable value where ImmutableList<T> would simply have been null. A nullable ImmutableArray<T>? field needs unwrapping via .Value or an is pattern before use.

Async

  • Return Task, not void; await, don't block. async void is for event handlers only, and .Result / .Wait() on async work is a deadlock and thread-starvation hazard, because the compiler-generated state machine expects to resume via continuations, not blocked threads. (Toub: How Async/Await Really Works in C#)
  • ConfigureAwait(false) in library code; omit it in application code. Libraries can't know their caller's context, so they should not capture it. Not capturing avoids deadlocks with context-blocking callers and skips a needless context hop. Application code on ASP.NET Core has no SynchronizationContext to capture, so ConfigureAwait(false) there is noise; UI application code usually wants the context. Apply one rule per layer rather than deciding case by case. (Toub: ConfigureAwait FAQ)
  • Default to Task<T>; reserve ValueTask<T> for hot APIs that usually complete synchronously. ValueTask is worth it only when profiling shows Task allocations matter and the synchronous path dominates (e.g. a buffered read). Its contract is stricter: await it exactly once, never concurrently, never twice. When in doubt, Task is the safe, composable choice. (Toub: Understanding the Whys, Whats, and Whens of ValueTask)

Span<T> and ArrayPool<T>

Before, with one string allocation per field:

static int SumCsvAllocating(string line)
{
    var sum = 0;
    foreach (var field in line.Split(','))
    {
        sum += int.Parse(field);
    }
    return sum;
}

After, with zero allocations, because Split over a span yields Ranges rather than strings:

static int SumCsv(ReadOnlySpan<char> line)
{
    var sum = 0;
    foreach (var range in line.Split(','))
    {
        sum += int.Parse(line[range]);
    }
    return sum;
}

For transient large buffers, rent from the shared pool instead of allocating per call, and always return in a finally:

using System.Buffers;

static async Task CopyAsync(Stream source, Stream destination, CancellationToken cancellationToken = default)
{
    var buffer = ArrayPool<byte>.Shared.Rent(81920);
    try
    {
        int read;
        while ((read = await source.ReadAsync(buffer, cancellationToken)) > 0)
        {
            await destination.WriteAsync(buffer.AsMemory(0, read), cancellationToken);
        }
    }
    finally
    {
        ArrayPool<byte>.Shared.Return(buffer);
    }
}

Never use a rented buffer after returning it, and never assume Rent gives exactly the requested size: slice to what you used.

GC modes

Keep the defaults. ASP.NET Core defaults to server GC with DATAS (dynamic heap-count adaptation, on by default since .NET 9), which gives throughput under load and shrinks the heap when idle; everything else defaults to workstation GC. Override in exactly two situations, and verify with memory/latency measurements:

  • Many .NET services on one small node (dense containers, sidecars): force workstation GC (<ServerGarbageCollection>false</ServerGarbageCollection>). Per-core server heaps multiply across processes and DATAS only softens, not removes, that footprint.
  • A CPU-bound batch/worker app that isn't ASP.NET Core: opt in to server GC for throughput.

(Microsoft Learn: Workstation and server GC, Microsoft Learn: Dynamic adaptation to application sizes (DATAS), Toub: Performance Improvements in .NET 10)

Set GC knobs in runtimeconfig.json or MSBuild, because the environment-variable form is hexadecimal and gets this wrong silently. DOTNET_GCHeapHardLimitPercent=60 does not cap the heap at 60%. The value is parsed as hex, so it reads as 0x60, which is 96%, and the cap you thought you set is barely a cap at all. Written as System.GC.HeapHardLimitPercent in runtimeconfig.json the same number is decimal and means what it says. The rule covers the numeric GC settings generally, heap count and LOH threshold and high-memory percent among them. Usually there is nothing to switch on in the first place: under a container memory limit the GC already treats that limit as total physical memory and defaults the hard limit to 75% of it, so a value of your own tightens a default rather than enabling one. (Microsoft Learn: Garbage collector config settings, reviewed 2026-02-09; Smith: Top 10 ways to reduce .NET memory usage in Kubernetes)

Environment.ProcessorCount tells you what this process may use, not what the machine has. Since .NET 6 it honours process affinity and container CPU limits, so under a cgroup quota it reports the quota. That is the number you want for sizing a thread pool or a Parallel loop, and the wrong one for reporting host capacity or counting cores for a licence. The BCL exposes no host total, so the platform call is the only route: GetActiveProcessorCount on Windows, sysctlbyname("hw.logicalcpu") on macOS, and parsing /sys/devices/system/cpu/online on Linux. Cache it behind a singleton, since it cannot change while the process lives, and declare the import with [LibraryImport] so the marshalling is source-generated and survives trimming. (Lock: Finding the total number of processors on a machine with .NET)

Native AOT

Use Native AOT for short-lived and size-sensitive workloads such as CLI tools, serverless functions and sidecars; keep the JIT for long-running services. AOT wins startup (milliseconds, no JIT warmup) and disk/memory footprint; the JIT wins steady-state throughput via tiered compilation and dynamic PGO, and tolerates reflection-heavy libraries that AOT's trimming breaks. Going AOT means the whole dependency graph must be trim/AOT-safe (source-generated JSON, no runtime codegen). Audit IsAotCompatible warnings before committing, and run the test suite as an AOT build too, because the reflection failures this causes appear nowhere else (see testing.md). .NET 10 file-based apps make the CLI-tool case trivial: dotnet publish app.cs produces a Native AOT binary by default. (Microsoft Learn: Native AOT deployment, Microsoft Learn: File-based apps)

Decide invariant globalization separately from AOT. dotnet new webapiaot sets <InvariantGlobalization>true</InvariantGlobalization> next to <PublishAot>true</PublishAot>, so the property tends to enter a codebase attached to a decision that has nothing to do with it, and it changes what every ToString and Compare in the app does. Keep it only if you meant it (see globalization.md).