- by x32x01 ||
Running large language models can be expensive, especially when the same model has to handle millions of requests. But a new Apple research paper suggests that reducing AI inference costs does not always require a smaller model, quantization, or retraining.
The researchers introduced LoopCD, a training-free decoding technique designed for Looped Transformers. In their experiments, LoopCD allowed models to match or outperform full-depth baselines while using fewer recurrent iterations, reducing forward compute by 22.5% to 48.2%.
The idea is surprisingly simple: instead of using only the final state produced by a Looped Transformer, LoopCD also uses an earlier intermediate state as a weaker reference.
A Looped Transformer repeatedly executes a shared block. Each iteration refines the model's internal representation, so earlier iterations contain less computation while later iterations contain more.
Normally, the model uses only the final state to predict the next token.
LoopCD takes advantage of the information that would otherwise be discarded.
It compares the stronger final prediction with an earlier prediction and uses that difference to guide the final token selection.
💡 The important part: LoopCD does not require additional model training or an auxiliary model.
Instead of changing the model's parameters, it changes how the model's existing recurrent states are used during decoding.
The researchers found that LoopCD could improve prediction quality at the same recurrent depth. More importantly, those gains allowed them to reduce the number of recurrent iterations while still matching or exceeding the original full-depth model.
That is where the potential cost savings come from.
The final state
LoopCD also uses an earlier state, such as
The method then pushes the final prediction away from the weaker prediction and toward the information learned during later recurrent refinement.
The paper describes two main versions:
It combines the early and final hidden states before they pass through the model's final output layers.
One major advantage is that the output layers only need to run once, resulting in zero additional output overhead compared with the normal final prediction.
It calculates predictions from both the earlier and final states and then combines them to guide token selection.
This approach requires one additional pass through the model's output layers, but it can provide strong accuracy improvements.
For example, with Ouro-2.6B-Thinking, adaptive LoopCD-Logits increased AIME 2024 pass@1 from 61.88% to 73.33% at the same recurrent depth.
That's an improvement of 11.45 percentage points.
The technique also improved code-generation performance.
For Huginn at 32 recurrent iterations, LoopCD-Hidden increased HumanEval pass@1 from 22.56% to 31.71%.
These results were reported across four different Looped Transformer families: Ouro, Huginn, Parcae, and Looped-Qwen3. The evaluation covered mathematical reasoning, code generation, and multiple-choice benchmarks.
The researchers tested LoopCD with approximately half the recurrent iterations.
Despite doing substantially less computation, the guided models could still match or outperform the original models running at full recurrent depth.
Across the evaluated configurations, this reduced forward FLOPs by 22.5% to 48.2%.
In other words, the goal is not simply:
“Make the same model run faster.”
It is closer to:
“Use better decoding so the model does not need as many recurrent computations.”
That distinction matters.
The paper reports a reduction of 22.5%-48.2% in forward FLOPs under its evaluated configurations. That should not automatically be interpreted as a 48% reduction in the real-world price of every LLM API or production workload.
Actual cost depends on many factors, including:
LoopCD demonstrates that substantial inference-compute savings may be possible by improving decoding efficiency rather than changing the trained model.
Potential applications include:
For years, much of the focus has been on making models smaller, compressing their weights, or training more efficient architectures.
But another possibility is to keep the model largely intact and make better use of the computation it already performs.
The difference can be summarized like this:
LoopCD is particularly interesting because it is training-free. The researchers use information already produced inside the recurrent computation instead of requiring an additional training stage.
It is the idea that LLM efficiency is not only about building smaller models.
There may be significant room to improve how existing models use their computation.
A model can perform multiple internal refinement steps, but not every step necessarily needs to be executed if better decoding can extract more useful information from the states that are already available.
That opens an interesting question:
For now, LoopCD provides an interesting example of how much can potentially be gained without retraining the model itself. The Apple-affiliated research reports improved benchmark performance at full depth and substantial compute reductions when recurrent depth is reduced.
It uses an earlier recurrent state as a weaker reference and contrasts it with the final state to improve token selection.
The researchers report:
Making LLMs cheaper may not always require making the models smaller. Sometimes, the bigger opportunity is learning how to use their existing computation more intelligently.
The researchers introduced LoopCD, a training-free decoding technique designed for Looped Transformers. In their experiments, LoopCD allowed models to match or outperform full-depth baselines while using fewer recurrent iterations, reducing forward compute by 22.5% to 48.2%.
What Is LoopCD?
LoopCD stands for Looped Contrastive Decoding.The idea is surprisingly simple: instead of using only the final state produced by a Looped Transformer, LoopCD also uses an earlier intermediate state as a weaker reference.
A Looped Transformer repeatedly executes a shared block. Each iteration refines the model's internal representation, so earlier iterations contain less computation while later iterations contain more.
Normally, the model uses only the final state to predict the next token.
LoopCD takes advantage of the information that would otherwise be discarded.
It compares the stronger final prediction with an earlier prediction and uses that difference to guide the final token selection.
💡 The important part: LoopCD does not require additional model training or an auxiliary model.
Why This Can Reduce LLM Inference Costs
There are several common ways to make LLM inference cheaper:- Use a smaller model.
- Apply quantization.
- Reduce precision.
- Optimize the hardware or inference engine.
- Reduce the amount of computation performed by the model.
- Train a model specifically for a more efficient deployment.
Instead of changing the model's parameters, it changes how the model's existing recurrent states are used during decoding.
The researchers found that LoopCD could improve prediction quality at the same recurrent depth. More importantly, those gains allowed them to reduce the number of recurrent iterations while still matching or exceeding the original full-depth model.
That is where the potential cost savings come from.
How Does LoopCD Work?
Consider a Looped Transformer that performs several recurrent iterations: Code:
h0 → h1 → h2 → ... → hR hR normally produces the next-token prediction.LoopCD also uses an earlier state, such as
h1, as a weaker reference.The method then pushes the final prediction away from the weaker prediction and toward the information learned during later recurrent refinement.
The paper describes two main versions:
LoopCD-Hidden
LoopCD-Hidden performs the contrast directly in the hidden-state space.It combines the early and final hidden states before they pass through the model's final output layers.
One major advantage is that the output layers only need to run once, resulting in zero additional output overhead compared with the normal final prediction.
LoopCD-Logits
LoopCD-Logits performs the contrast after the model produces its logits.It calculates predictions from both the earlier and final states and then combines them to guide token selection.
This approach requires one additional pass through the model's output layers, but it can provide strong accuracy improvements.
The Results Are More Interesting Than the 48% Number
The headline result is the reduction in forward FLOPs, but the accuracy improvements are also significant.For example, with Ouro-2.6B-Thinking, adaptive LoopCD-Logits increased AIME 2024 pass@1 from 61.88% to 73.33% at the same recurrent depth.
That's an improvement of 11.45 percentage points.
The technique also improved code-generation performance.
For Huginn at 32 recurrent iterations, LoopCD-Hidden increased HumanEval pass@1 from 22.56% to 31.71%.
| Model / Task | Baseline | LoopCD | Improvement |
|---|---|---|---|
| Ouro-2.6B-Thinking - AIME 2024 | 61.88% | 73.33% | +11.45 |
| Huginn - HumanEval | 22.56% | 31.71% | +9.15 |
The Real Trick: Use Fewer Iterations
This is where the research becomes particularly interesting for inference efficiency.The researchers tested LoopCD with approximately half the recurrent iterations.
Despite doing substantially less computation, the guided models could still match or outperform the original models running at full recurrent depth.
Across the evaluated configurations, this reduced forward FLOPs by 22.5% to 48.2%.
In other words, the goal is not simply:
“Make the same model run faster.”
It is closer to:
“Use better decoding so the model does not need as many recurrent computations.”
That distinction matters.
Does This Mean LLMs Can Become 48% Cheaper?
Not necessarily.The paper reports a reduction of 22.5%-48.2% in forward FLOPs under its evaluated configurations. That should not automatically be interpreted as a 48% reduction in the real-world price of every LLM API or production workload.
Actual cost depends on many factors, including:
- Hardware utilization
- Memory bandwidth
- Batch size
- KV-cache behavior
- Token generation patterns
- Model architecture
- Serving infrastructure
- Power consumption
- Software and kernel efficiency
LoopCD demonstrates that substantial inference-compute savings may be possible by improving decoding efficiency rather than changing the trained model.
Why This Matters for AI Systems
If techniques like this become practical across more architectures, the impact could extend beyond individual research benchmarks.Potential applications include:
- 🤖 AI agents that generate many model calls
- 🔄 Multi-step AI pipelines
- 🏢 Enterprise AI systems
- 🌐 LLM APIs serving large numbers of requests
- 💻 Local AI inference
- 🧠 Reasoning models that use repeated computation
- ⚙️ Large-scale inference infrastructure
Smaller Models vs. More Efficient Inference
This research also highlights an important direction in LLM optimization.For years, much of the focus has been on making models smaller, compressing their weights, or training more efficient architectures.
But another possibility is to keep the model largely intact and make better use of the computation it already performs.
The difference can be summarized like this:
| Approach | Main Idea |
|---|---|
| Smaller model | Reduce parameter count |
| Quantization | Reduce numerical precision |
| Distillation | Train a smaller model from a larger one |
| Pruning | Remove parts of the model |
| LoopCD | Improve decoding and reduce recurrent computation |
What This Could Mean for the Future of LLMs
The bigger lesson may not be the exact 48.2% figure.It is the idea that LLM efficiency is not only about building smaller models.
There may be significant room to improve how existing models use their computation.
A model can perform multiple internal refinement steps, but not every step necessarily needs to be executed if better decoding can extract more useful information from the states that are already available.
That opens an interesting question:
The answer may ultimately be both.Will the next generation of LLM optimization come mainly from building smaller models, or from making existing models compute more efficiently?
For now, LoopCD provides an interesting example of how much can potentially be gained without retraining the model itself. The Apple-affiliated research reports improved benchmark performance at full depth and substantial compute reductions when recurrent depth is reduced.
The Bottom Line
🚀 LoopCD is a training-free contrastive decoding technique for Looped Transformers.It uses an earlier recurrent state as a weaker reference and contrasts it with the final state to improve token selection.
The researchers report:
- 22.5%-48.2% lower forward FLOPs in reduced-depth experiments.
- Improved reasoning performance on benchmarks such as AIME.
- Improved code-generation performance on HumanEval.
- No additional model training.
- No auxiliary model.
- A hidden-state variant with no extra output-layer pass.
Making LLMs cheaper may not always require making the models smaller. Sometimes, the bigger opportunity is learning how to use their existing computation more intelligently.