Towards Human-Like Agentic
Streaming Video Understanding

Boyuan Sun1 Deshui Miao2Wenzhao Gao2Shaoyong Jia2Shaohui Jiao2 Chi Liu3Yan Shu3Derek Li3Bryan Dai3 Shengsheng Qian4Qibin Hou1
1 VCIP, Nankai University 2 ByteDance Inc.3 IQuest Research 4 Institute of Automation, Chinese Academy of Sciences
Watch demo GitHub Application

See it in action

From observation to an informed answer

Watch visual changes, structured memory, and query-driven reasoning unfold across three video examples.

Download video

Music: “Meanwhile” by Scott Buckley · CC BY 4.0

Quantitative results

OVO-Bench performance

Evaluation across real-time visual perception, backward tracing, and forward active responding.

OVO-Bench group averages and overall scores, as reported in the paper.
ModelReal-Time
Avg.
Backward
Avg.
RT+BT
Avg.
Forward
Avg.
Overall
Human
Human Agents93.2092.3392.7792.9092.81
Proprietary MLLMs
Gemini 1.5 Pro69.3262.5465.9357.1563.00
Gemini 3 Flash72.4462.4767.4657.1564.02
GPT-4o64.4660.7562.6153.4059.54
Open-Source Offline MLLMs
Qwen2-VL-72B61.9256.9559.4449.3056.27
LLaVA-Video-7B63.5240.4051.9654.8252.91
LLaVA-OneVision-7B64.0243.7153.8750.5052.74
Qwen2-VL-7B55.9846.4651.2248.7450.39
InternVL-V2-8B60.3943.4451.9246.6050.15
LongVU-7B57.6135.0146.3147.5046.71
Qwen3-VL-8B65.5246.1855.8555.7355.81
Limited Open-Source Online MLLMs
Stream-IT71.30————
HERMES69.0049.4059.20——
VideoLLM-online-8B20.7917.7319.26——
OASIS78.1457.2167.68——
SimpleBaseline79.9054.9067.40——
LatentStream68.5060.0064.20——
JoyAI-VL-Interaction68.4048.6058.50——
Mage-VL79.8448.1564.00——
FOLIO82.0069.1075.55——
Full Open-Source Online MLLMs & Video Agents
Flash-VStream-7B28.3727.3827.8845.0933.61
Dispider54.5536.0645.3134.7241.78
ReKV62.5045.3053.9048.0052.00
StreamForest-7B61.2052.0256.6153.4955.57
Streamo67.4049.2058.3057.0057.90
VST67.2056.7061.9554.0059.30
ViSpeak66.2857.5261.9054.2561.08
StreamAgent61.3041.7051.5045.4049.40
EventMemAgent68.2958.0363.1655.9260.75
DeltaStreamer72.2367.7269.9858.2566.07

RT+BT averages the Real-Time and Backward groups. — denotes an unreported result.

The idea

Video keeps changing. Memory should keep up.

In streaming video, questions can arrive after the relevant moment has passed. A retained description may miss the detail needed to answer. DeltaStreamer records visual changes alongside scene descriptions, organizes them into structured textual memory, and uses a query-driven planner to gather the evidence an answer needs.

  1. 01

    Observe changes

    Compare consecutive frames and record explicit visual deltas, preserving transitions that similar scene descriptions can overlook.

  2. 02

    Organize memory

    Connect working observations, episode summaries, and pinned facts with pointers to their original visual evidence.

  3. 03

    Plan and look back

    Identify what the question requires, search memory, and revisit historical frames when a missing visual detail matters.

DeltaStreamer framework connecting observation of consecutive frames, explicit visual deltas, structured textual memory, query-driven planning, and visual evidence retrieval.
The DeltaStreamer framework. Observation builds memory; a question directs the search for supporting evidence.

Qualitative examples

Remember the change. Recover the missing detail.

Recorded changes help recover an action sequence. When textual memory lacks the letters visible in an earlier scene, planning and visual replay recover the evidence needed for an answer.

Two real paper examples. The first traces an action through visual changes stored in memory. The second follows a plan to revisit an earlier video scene and read visual details missing from memory.
Two examples from the paper illustrate complementary roles of change-aware memory and targeted visual replay.

Reference

Citation

@misc{sun2026deltastreamer,
  title  = {DeltaStreamer: Towards Human-Like Agentic Streaming Video Understanding},
  author = {Boyuan Sun and Deshui Miao and Wenzhao Gao and Shaoyong Jia and
            Shaohui Jiao and Chi Liu and Yan Shu and Derek Li and Bryan Dai and
            Shengsheng Qian and Qibin Hou},
  year   = {2026}
}