Speculative Decoding Shipped Because It Doesn't Change the Output.

ai inference speculative-decoding mechanism performance

Liquid AI shipped LFM2.5-DSpark on August 20, 2026: three draft models targeting LFM2.5 1.2B-Instruct, LFM2.5 2.6B, and LFM2.5 8B-A1B (an MoE), under the lfm1.0 license. The headline numbers are 2.67× mean speedup on H100 and 2.27× mean on M4 Max, against the 2.6B target, with a 3.06× peak on MATH500. The headline is the least interesting thing about the release. The interesting thing is why this kind of release can ship at all, in production, on day one, without an eval-suite expansion or a rollback story.

What speculative decoding is

The pattern is straightforward. A small draft model proposes a short sequence of tokens; the target model verifies them in parallel, accepts what matches, rejects the rest, and continues from the first rejected position. Three numbers characterise a cycle: draft length (how many tokens the drafter proposes per step), acceptance rate (how many the verifier accepts), and effective tokens-per-step (the realised throughput gain).

The crucial property is exactness under greedy decoding. When the verifier uses argmax and accepts a draft token, that token is byte-identical to what the target model would have produced alone. The output stream is byte-identical to non-speculative greedy decoding. The whole family of EAGLE, Medusa, DFlash, DSpark trades on this property. Without it, speculative decoding would be a quantisation problem with worse tooling, not a free lunch.

What DSpark actually does

The interesting design choice is what is not there. The paper (arXiv:2607.05147, Cheng et al., DeepSeek-AI + PKU) describes a confidence-scheduled Markov-head draft, layered on the DFlash / EAGLE-3 lineage. A confidence scheduler was built to dynamically trim low-confidence drafts at inference time. It was left dormant. Dynamic trimming hurt more than helped. DSpark ships the simpler, static version because the simpler version wins on real benchmarks.

This is a posture statement. The speculative-decoding family has been chasing dynamic adaptation for three generations. EAGLE-2 introduced the tree-of-tokens idea, where the drafter proposes multiple candidates at branching positions and the verifier picks the longest accepted prefix. EAGLE-3 refined the head; DFlash (arXiv:2602.06036) generalised the Markov head to a confidence-scheduled mechanism. DSpark ships static. The dynamic machinery is built and dormant. The static version won. Posture matters when the family has been climbing for three generations and the release ships the floor.

The numbers

The table is the spine.

Draft targetAcceptance lengthH100 meanM4 Max meanMATH500 peak
LFM2.5-1.2B-Instruct————
LFM2.5-2.6B4.812.67×2.27×3.06×
LFM2.5-8B-A1B (MoE)——1.18×—

Two extras worth naming. BFCL function-calling latency: 57% reduction on M4 Max against the 2.6B target (BFCL is in PMLR v267). Confidence-scheduled verification: built but not used; the released checkpoint uses static drafting.

The honest weakness is the 8B-A1B on M4 Max row: 1.18×. MoE routing plus draft-model memory bandwidth saturates the unified-memory bus before the draft path helps. The release shipped the 8B-A1B draft anyway, which is the right call: an honest number on a real workload is more useful than a cherry-picked one. The 8B-A1B draft is not unusable; it is just not the headline.

Why exactness under greedy is the whole game

The portable rule. Any inference optimisation that does not change outputs is a free lunch: A/B parity, no eval-suite re-runs, no golden regression traces, no rollback plan. Optimisations that do change outputs (post-training quantisation, distillation, pruning, kernel fusion with numerical drift) need eval suites, golden traces, and a rollback story. Speculative decoding under greedy needs none of that.

The cost calculus is asymmetric. Exactness ships. Non-exactness budgets. That is why every production serving stack I have shipped or audited has speculative decoding on the hot path, and almost none of them have post-training quantisation on the hot path without an eval gate. The decision was not made on speedup numbers; it was made on whether the change needed a quality gate. Speculative decoding does not. The 2.67× is the bonus.

This is the inverse of the byte-identical-still-broken property. There, a transformation was supposed to be byte-identical to a reference and was not; the breakage was discovered under load. Here, byte-identity is the load-bearing property, and it is exact under the specific decoder configuration the family targets (greedy). Both posts are about the difference between a property the system claims to have and the property the system has. DSpark lives on the side where the claim matches the property. That is the reason it can ship under lfm1.0 with SGLang integration in flight (PR #31041) and Metal kernels in llama.cpp.

Where this generalises

The reader should leave with one transferable test. When you evaluate an inference optimisation, ask one question: does it change outputs under greedy decoding?

If no, ship it without an eval suite. If yes, budget for one.

The question applies to every serving decision for the next two years: KV-cache compression, paged attention, continuous batching, prefix caching, quantisation kernels, draft models, MoE expert pruning. Each one either changes outputs or does not, and that single property determines whether the rollout needs a quality gate. Speedup is irrelevant until the question is answered. Most of the time, the answer is “no” (the optimisation is exact under greedy), and the answer is the green light. Some of the time, the answer is “yes,” and the question is whether the speedup is worth the eval-suite tax. Both questions live downstream of the first one.

The trap

The trap is to read the speedup numbers as the result. They are the side effect. The result is that Liquid AI shipped three draft models under a permissive license, targeting three different target sizes, on day one, with SGLang integration in flight and Metal kernels in llama.cpp. The result is that the bottleneck for serving LFM2.5 on Apple Silicon is no longer compute; it is memory bandwidth on the draft path. The result is that a confidence scheduler was built, evaluated, and left dormant because the static version was the right call.

Post-Omarchy continuity. Visible properties (exactness under greedy) shipped. Hidden properties (dynamic adaptation) were shelved. The whole release is a single argument: the technique is shippable because the property it claims is the property it has, and the property it claims is exactness under a specific decoder. Speculative decoding has always claimed that property. DSpark is the release where the claim matches the benchmark, the benchmark matches the workload, and the workload matches the production stack. The 2.67× is the headline. The byte-identity is the result.

$ cat GIT .md
· 7 min read

84 Repositories Vanished. The Fix Was mkdir.

S3 has no empty directories, and in a fully packed git repository the refs directories are exactly that: empty. A file-by-file sync carried every object intact and dropped the two directories git requires to call something a repository, so 84 of 517 came back unreadable while the database still said they had commits. The migration's health checks stayed green the whole time, because none of them ever open a repository.

git aws s3 gitlab mechanism
$ cat KUBERNETES .md
· 7 min read

The Constraint Was Satisfied. The Zone Was Empty Anyway.

A topology spread constraint is only ever a statement about the population its selector matches, and an operator-set label can quietly pool sibling Deployments into one shared count. Three collectors each satisfied their hard zone spread while the union of them left a zone empty, and the standard repair converged on the same wrong answer every time.

kubernetes scheduling opentelemetry finops mechanism