Less from More.
Reinforcing Sparse Video Reasoning
from Dense References
University of Amsterdam
Learn from dense video. Reason with sparse frames.
benchmarks
model sizes
frame-rate settings
0.1 · 0.2 · 0.5 · 1.0 · 2.0 fpsMore evidence in training.
Less input at inference.
Dense and sparse views share one policy. The stronger dense prediction guides its sparse counterpart.

Two temporal views of the same video, trained together.
Align with the dense prediction only when its grounding reward is higher.
At inference, only sparse frames are needed.
Abstract
Video-language models commonly assume that more temporal observations lead to more reliable reasoning. We question this assumption and argue that the key challenge is not merely processing more video frames efficiently, but learning to reason reliably under limited temporal evidence. We propose SAVER, a dense-to-sparse post-training framework that uses dense video views as training-time references for sparse-frame inference. During reinforcement post-training, paired dense and sparse views are optimized with grounding rewards and a reliability-gated reference reward, encouraging sparse view predictions to preserve task-relevant temporal evidence. Across three temporal grounding benchmarks and six video question-answering benchmarks, SAVER consistently improves performance across frame budgets. In particular, SAVER can match or surpass dense-frame Qwen3.5 baselines while using substantially fewer frames. These results show that temporal grounding can serve as an effective evidence-localization proxy for learning sparse video reasoning that transfers to broader video understanding tasks.
Small frame budgets.
Stronger video reasoning.
SAVER and Qwen3.5, compared at the same model size and frame rate.
Qwen3.5 SAVER
Loading benchmark results…
Explore individual benchmarks
Macro-averages over 3 grounding and 6 video-QA benchmarks. Individual results vary by task.
Nearly dense performance, with sparse input.
At 0.1 fps, SAVER-2B reaches 50.8% average video-QA accuracy, close to Qwen3.5-2B at 2.0 fps (51.0%), while using 38 rather than 139 frames on average.
TABLE 8 / INFERENCE COST
Less input. Lower cost.
Qwen3.5-2B · 2.0 → 0.1 fps
visual tokens
33,113 → 4,386peak GPU memory
11.30 → 5.73 GiBend-to-end latency
6.58 → 2.40 s| Benchmark | Visual tokens | Peak memory (GiB) | E2E latency (s) |
|---|---|---|---|
| Video-MME | 69,750 → 3,060 | 18.72 → 4.80 | 10.03 → 1.41 |
| LongVideoBench | 92,160 → 13,950 | 20.65 → 7.89 | 16.59 → 4.31 |
| VideoMMMU | 89,160 → 19,137 | 21.54 → 12.99 | 20.62 → 8.02 |
| Average · 9 benchmarks | 33,113 → 4,386 | 11.30 → 5.73 | 6.58 → 2.40 |
Base-model profiling on one H100 (96 GB), BF16, batch size 1; not a SAVER-specific speed measurement.
Preserving temporal evidence.
In this qualitative example, sparse Qwen3.5 predicts a broad interval. SAVER produces a tighter interval closer to the ground truth.

What remains challenging? Sparse frames can miss a key event entirely or confuse repeated actions. SAVER improves aggregate performance but cannot recover visual evidence that was never observed.
Three sizes.
One sparse-input recipe.
Qwen3.5-based checkpoints, available on Hugging Face.
Citation
@misc{sun2026less,
title = {Less from More: Reinforcing Sparse Video Reasoning from Dense References},
author = {Wenfang Sun and Yingjun Du and Cees G. M. Snoek},
year = {2026},
eprint = {2610.10893},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2610.10893}
}