Learn learning medium confidence

Microsoft Research Presents LLM-42 for Selective Deterministic Inference

The SOSP 2026 paper uses a decode-verify-rollback scheduler to reproduce outputs for selected requests without forcing an entire serving batch onto deterministic kernels.

Edited by Tyronne Panaino

Microsoft Research has listed LLM-42 as a SOSP 2026 paper, highlighting a systems approach for making selected large-language-model requests deterministic without imposing the same constraint on every request in a serving batch. Microsoft's research index dated the listing September 29, 2026.

The work matters to teams that need repeatable evaluation, auditing or regression tests from shared inference infrastructure. It is not a new model or a generally available cloud feature. The underlying preprint was submitted in January and last revised on January 30, so the current development is its conference-publication listing rather than the first disclosure of the method.

Why identical requests can still diverge

A fixed prompt and fixed sampling settings do not by themselves guarantee an identical output. The paper traces one source of serving-level variation to floating-point arithmetic, dynamic batching and GPU kernels that change their reduction order with the shape of a batch. Small numerical differences can alter a token decision; because generation is autoregressive, one changed token can send the rest of a sequence down a different path.

One established response is to use batch-invariant kernels. The authors argue that this approach couples reproducibility to specialized kernel implementations and can make all requests in a shared batch pay a performance cost, even when only one workload needs deterministic output.

How the verify-rollback design works

LLM-42 keeps a normal high-throughput decode path for candidate tokens and adds a verifier that replays a fixed-size window under a consistent reduction schedule. Tokens that match are committed. When verification finds a mismatch, the system rolls back to the last consistent token, discards the divergent continuation and resumes generation from the repaired state.

The design also replaces the relevant key-value cache entries with values from the verification pass. That step matters because matching visible tokens would not be sufficient if the hidden cache state still differed and caused a later divergence. In the paper's design, determinism is selectable per request rather than a property imposed on every request sharing a batch.

What the authors measured

The researchers implemented LLM-42 on SGLang. In one experiment they ran 350 Llama-3.1-8B-Instruct requests drawn from the ShareGPT dataset at six queries per second, fixing each output at 512 tokens. They report that many sequences initially matched for hundreds of tokens, but that a first mismatch usually propagated quickly through the remaining output.

For a serving comparison in which one of 11 requests required deterministic output, the authors report 911 decoded tokens per second. They describe that as 2.2 times the deterministic-mode throughput of the SGLang comparison and within 3% of their best non-deterministic mode. Those figures are results from the paper's own setup, not independent performance guarantees for other models, accelerators, batch patterns or production services.

Where selective determinism could help

The proposal is most relevant where reproducing an output is part of the job: comparing model versions, investigating a regression, checking a safety evaluation or preserving an auditable run. Creative and exploratory workloads may not need bit-level repeatability, so a request-level control could avoid forcing them onto the same verification path.

The next useful checkpoint is replication outside the authors' environment. The fetched evidence does not establish cross-hardware behavior, operational reliability at production scale, availability in a commercial service or the maintenance cost of integrating the scheduler with rapidly changing serving engines.

Status

Learning. Internal confidence is medium. The method and evaluation are documented by the authors and listed by Microsoft Research for SOSP 2026, but no independent replication or production deployment evidence was fetched.

Sources

Update note: Last reviewed 2026-09-30. We will revise this explainer if independent evaluations or a production implementation become available.

Sources

Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.

More Learn coverage