Less from More.

Reinforcing Sparse Video Reasoning
from Dense References

Wenfang Sun · Yingjun Du · Cees G. M. Snoek

University of Amsterdam

Learn from dense video. Reason with sparse frames.

9

benchmarks

3

model sizes

5

frame-rate settings

0.1 · 0.2 · 0.5 · 1.0 · 2.0 fps

More evidence in training.
Less input at inference.

Dense and sparse views share one policy. The stronger dense prediction guides its sparse counterpart.

SAVER framework with paired sparse and dense video views, a shared Qwen policy, and format, grounding, and reliability-gated reference rewards.
Dense-to-sparse reinforcement post-training
Dense view
Sparse view
πShared policy

Two temporal views of the same video, trained together.

Abstract

Video-language models commonly assume that more temporal observations lead to more reliable reasoning. We question this assumption and argue that the key challenge is not merely processing more video frames efficiently, but learning to reason reliably under limited temporal evidence. We propose SAVER, a dense-to-sparse post-training framework that uses dense video views as training-time references for sparse-frame inference. During reinforcement post-training, paired dense and sparse views are optimized with grounding rewards and a reliability-gated reference reward, encouraging sparse view predictions to preserve task-relevant temporal evidence. Across three temporal grounding benchmarks and six video question-answering benchmarks, SAVER consistently improves performance across frame budgets. In particular, SAVER can match or surpass dense-frame Qwen3.5 baselines while using substantially fewer frames. These results show that temporal grounding can serve as an effective evidence-localization proxy for learning sparse video reasoning that transfers to broader video understanding tasks.

Small frame budgets.
Stronger video reasoning.

SAVER and Qwen3.5, compared at the same model size and frame rate.

Inference frame rate

Qwen3.5 SAVER

Loading benchmark results…

Explore individual benchmarks

Macro-averages over 3 grounding and 6 video-QA benchmarks. Individual results vary by task.

Nearly dense performance, with sparse input.

At 0.1 fps, SAVER-2B reaches 50.8% average video-QA accuracy, close to Qwen3.5-2B at 2.0 fps (51.0%), while using 38 rather than 139 frames on average.

TABLE 8 / INFERENCE COST

Less input. Lower cost.

Qwen3.5-2B · 2.0 → 0.1 fps

−86.8%

visual tokens

33,113 → 4,386
−49.3%

peak GPU memory

11.30 → 5.73 GiB
−63.6%

end-to-end latency

6.58 → 2.40 s
Selected benchmarks · dense → sparse
BenchmarkVisual tokensPeak memory (GiB)E2E latency (s)
Video-MME69,750 → 3,06018.72 → 4.8010.03 → 1.41
LongVideoBench92,160 → 13,95020.65 → 7.8916.59 → 4.31
VideoMMMU89,160 → 19,13721.54 → 12.9920.62 → 8.02
Average · 9 benchmarks33,113 → 4,38611.30 → 5.736.58 → 2.40

Base-model profiling on one H100 (96 GB), BF16, batch size 1; not a SAVER-specific speed measurement.

Preserving temporal evidence.

In this qualitative example, sparse Qwen3.5 predicts a broad interval. SAVER produces a tighter interval closer to the ground truth.

Pouring orange juice example: ground truth 7.5 to 13 seconds, sparse Qwen3.5 prediction 5.7 to 13.9 seconds, SAVER prediction 6.3 to 12.3 seconds.
Qualitative temporal grounding example from the manuscript

What remains challenging? Sparse frames can miss a key event entirely or confuse repeated actions. SAVER improves aggregate performance but cannot recover visual evidence that was never observed.

Three sizes.
One sparse-input recipe.

Qwen3.5-based checkpoints, available on Hugging Face.

Citation

@misc{sun2026less,
  title         = {Less from More: Reinforcing Sparse Video Reasoning from Dense References},
  author        = {Wenfang Sun and Yingjun Du and Cees G. M. Snoek},
  year          = {2026},
  eprint        = {2610.10893},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2610.10893}
}