Thursday, October 1, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Earth Science

Frozen AI Models Learn to Watch Hours of Video Without Any Training

October 1, 2026
in Earth Science
Violet Maxwell
By Violet Maxwell Scienmag Editorial Profile - Natural Hazards
Reading Time: 5 mins read
0
Frozen AI Models Learn to Watch Hours of Video Without Any Training

Frozen AI Models Learn to Watch Hours of Video Without Any Training

Frozen AI Models Learn to Watch Hours of Video Without Any Training

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Every minute, more than 500 hours of video land on YouTube alone, and the flood of long-form footage—from lectures and live streams to surveillance archives and professional analytics—has exposed a hard truth about today’s artificial intelligence: even the most powerful vision-language models buckle when asked to make sense of hours of continuous visual content. A comprehensive new survey published in the open-access journal Vicinagearth maps out a rapidly growing counter-movement that promises to change this. Instead of retraining giant models for every new video task, researchers are building training-free pipelines that squeeze remarkable long-video understanding out of frozen, pretrained systems, orchestrating what the models already know rather than teaching them something new.

The survey, led by Jingren Liu of the Institute of Artificial Intelligence (TeleAI) at China Telecom together with colleagues from Tianjin University, City University of Hong Kong, and several other institutions, organizes the field around three stubborn obstacles. First is visual-token redundancy: when a video is chopped into thousands of frames, each frame is converted into hundreds or thousands of visual tokens, and the sheer volume overwhelms hardware long before any useful reasoning can happen. Second is the context-window problem: most architectures can only attend to a fixed span of tokens at once, which fragments a continuous video into disconnected snippets and destroys long-range temporal structure. Third is the reasoning gap: questions about causality, event prediction, and narrative disambiguation demand multi-step, abstract inference that goes far beyond simple perception.

The authors argue that the answer lies not in ever-larger models but in smarter inference-time engineering, and they sort existing solutions into three methodological paradigms. Selection-based methods attack redundancy directly. Systems such as VidCom² compute a distinctiveness score for each frame based on inter-frame similarity and allocate a token budget through a softmax distribution before pruning low-value tokens. METok goes further with an event-aware, multi-stage pipeline that segments continuous events by cosine similarity during encoding, prunes tokens hierarchically by attention strength during prefill, and discards low-impact entries from the key-value cache during decoding to cut both floating-point operations and memory use. DYTO achieves zero-shot compression by clustering frames into dynamic key-token groups and merging them from coarse to fine through a binary process.

Retrieval and frame-selection techniques form a second layer of the selection paradigm. APVR performs pivot frame retrieval, scores candidate frames, and then applies attention-based token filtering within the survivors, fusing spatio-temporal semantic confidence with query-aware attention. The T* framework reframes temporal retrieval as a spatial search problem, adaptively adjusting granularity across space and time to locate keyframes under extreme frame budgets. Meanwhile, adaptive keyframe sampling methods such as AKS balance relevance and coverage with a recursive judge-and-split strategy, and VSLS constructs logical triples—spatial co-occurrence, temporal proximity, attribute dependency, and causal order—to iteratively optimize which frames a model actually sees. Nar-KFC even casts keyframe selection as an integer quadratic programming problem, jointly optimizing relevance and diversity with a greedy solver while inserting coherent textual narratives to smooth temporal discontinuities.

The second paradigm replaces flat inputs with memory. Adaptive hierarchical methods such as VideoTree build a multilevel tree of key frames in real time, expanding breadth and depth according to relevance, while HEM-LLM partitions long videos into logical events and maintains multigranularity memories within and between events. Memory-augmented designs push this further: GlobalCom2 uses a global thumbnail to judge the importance of each frame’s representation and adaptively allocates compression budgets to local regions, and ∞-VIDEO borrows the idea of human sticky memory, continuously consolidating attention across a single pass through a video. For live streams, systems like LiveVLM generate and compress key-value tensors on the fly, and QuickVideo combines parallel CPU decoding, intra-group prefilling, cache pruning, and overlapping CPU-GPU execution to slash end-to-end latency.

The third and most recent paradigm treats video understanding as an active, agentic process. LVAgent deploys multiple multimodal agents in select-perceive-act-reflect cycles, collaborating across rounds without fixed sampling. VideoAgent uses a large language model as a central coordinator that iteratively identifies and compiles key information, while VCAgent adds curiosity-driven self-exploration, autonomously navigating segments to build understanding. VideoAgent2 introduces an uncertainty-aware chain-of-thought that decides when to plan, adjust, and acquire more evidence. Complementary chain-of-thought frameworks impose structure on the reasoning itself: VideoChat-A1’s chain-of-shot paradigm, which divides videos into shots and reasons across multi-round dialogues, reaches state-of-the-art accuracy of 77.0 percent on Video-MME and 70.1 percent on EgoSchema, while Self-ReS uses self-reflection and sparse attention to cut inference time by 46 percent. Retrieval-augmented pipelines such as Video-RAG and AdaVideoRAG round out the toolkit by fetching task-relevant context, with the latter adapting retrieval granularity to query complexity.

Whether these tricks actually work is the question the survey’s benchmark analysis answers, and the picture is sobering as well as encouraging. On the widely used Video-MME benchmark, the best configuration—Qwen2.5-VL-72B paired with the FlexSelect token-selection strategy—scores 74.4 overall, a modest gain over the base model’s 73.4. Larger language models clearly help: VILA* with Frame-Voyager improves from 50.5 to 60.0 when scaled from 8 billion to 34 billion parameters, and LLaVA-OneVision’s 72-billion-parameter version scores 66.3 versus 56.5 for the 7-billion variant. More input frames help too, with METok jumping from 36.4 to 46.6 on medium-length tasks when given 128 frames instead of 32. But the most striking pattern is universal: every model collapses on long videos. GPT-4o scores 71.4 on short clips yet only 55.2 on long ones, a 16.2-point gap that the authors identify as the field’s core unsolved problem.

The survey’s taxonomy of evaluation suites sharpens that diagnosis. General benchmarks such as MVBench, Perception Test, VideoVista, and Video-MME probe foundational comprehension and reveal that models score 85 to 88 percent on low-level perception but plunge to 30 to 40 percent on high-level reasoning. Hour-long stress tests like MovieChat, LVBench, MLVU, and LongVideoBench show steep degradation in contextual coherence as duration grows. Reasoning-centric benchmarks dig deeper: V-STaR forces models to produce explicit what-when-where reasoning chains, TimeLogic tests formal temporal logic, CARVE and MECD+ target causal discovery, and PhysBench spans 10,002 entries covering everything from mechanics to electromagnetism. Knowledge-grounded suites expose a particularly worrying flaw—systematic overconfidence, with models reporting confidence scores of 0.7 to 0.9 even when accuracy falls to 30 or 40 percent. Egocentric benchmarks are the harshest of all. EgoSchema shows that reliable comprehension of first-person video demands a median of 100 seconds of evidence, nearly six times the 18 seconds typical of third-person footage, and training-free systems drop from 65 to 70 percent accuracy on 30-second segments to a mere 25 to 33 percent on extended egocentric sequences.

Why pursue a paradigm with such visible weaknesses? The authors point to practical advantages that retraining cannot match. Training-free pipelines decouple capability from training budgets, adapt to new tasks without parameter updates, avoid catastrophic forgetting, and—crucially—produce auditable intermediate artifacts: the selected frames, memory traces, and reasoning chains can be inspected, so failures can be localized to specific stages such as sampling, retrieval, or reasoning. That transparency matters for the deployment scenarios the survey envisions, from professional analytics to embodied human-AI interaction and safety-critical decision making. Real-world constraints remain formidable, however: edge devices impose strict limits on compute, memory, and power; interactive applications like augmented reality demand low latency; and large-scale deployment incurs bandwidth and maintenance costs that task-oriented communication and utility-aware load shedding can only partially offset.

The survey closes with a research agenda that reads as a bridge between paradigms. Hybrid approaches could layer lightweight, parameter-efficient adaptations such as LoRA modules or mixture-of-experts on top of frozen backbones, refining only the components that need it. Adaptive memory and compression architectures should become more interpretable and resource-aware, and future benchmarks ought to measure latency, token efficiency, and long-horizon robustness alongside accuracy. Cross-modal orchestration—coordinating text, audio, and vision dynamically rather than treating video as a purely visual stream—stands out as essential for real-world generality. The overall trajectory, the authors conclude, marks a pivotal inflection: ad hoc engineering heuristics are giving way to principled, end-to-end frameworks that pair scalable computation with cognitively grounded reasoning, offering a sustainable path toward machines that can genuinely watch, remember, and reason over the endless video the modern world produces.

Subject of Research: Training-free methods, benchmarks, and open challenges for long video understanding with large multimodal models

Article Title: Towards training-free long video understanding: methods, benchmarks, and open challenges

Article References: Liu, J., Wang, Y., Zhang, L., Wang, Y., Xu, S., Wang, L., Yan, J., Zhang, D., & Chen, X. (2025). Towards training-free long video understanding: methods, benchmarks, and open challenges. Vicinagearth, 2(1), Article 6. https://doi.org/10.1007/s44336-025-00017-w

Image Credits: AI Generated

DOI: 10.1007/s44336-025-00017-w

Keywords: long video understanding, training-free methods, vision-language models, token compression, keyframe selection, agent-based reasoning, chain-of-thought, memory architectures, benchmarks, egocentric video, multimodal AI, retrieval-augmented generation

Cite Scienmag News

Violet Maxwell. (October 1, 2026). Frozen AI Models Learn to Watch Hours of Video Without Any Training. Scienmag. https://scienmag.com/frozen-ai-models-learn-to-watch-hours-of-video-without-any-training/

Violet Maxwell. "Frozen AI Models Learn to Watch Hours of Video Without Any Training." Scienmag, 1 October 2026, https://scienmag.com/frozen-ai-models-learn-to-watch-hours-of-video-without-any-training/. Accessed 1 October 2026.

Violet Maxwell. "Frozen AI Models Learn to Watch Hours of Video Without Any Training." Scienmag. October 1, 2026. https://scienmag.com/frozen-ai-models-learn-to-watch-hours-of-video-without-any-training/

Tags: agent-based reasoningAI model scalabilityBenchmarkschain-of-thoughtcontext-window limitationsegocentric videoFrozen AI modelskeyframe selectionlarge-scale video processinglong video understandinglong-form video analysismemory architecturesmultimodal AIpretrained AI systemsretrieval-augmented generationtoken compressiontraining-free AI pipelinestraining-free methodsvideo analytics in surveillance and streamingvideo content comprehensionvideo token redundancyvideo understandingvision-language models
Share26Tweet16
Previous Post

Waste Sawdust Becomes Powerful Water Catalyst When Nitrogen and Boron Are Perfectly Balanced

Next Post

New Questionnaire Aims to Measure How Staff Judge Mechanical Restraint in Psychiatric Care

Related Posts

Fish Livers Reveal Hidden Metal Loads in Pristine Andaman Waters
Earth Science

Fish Livers Reveal Hidden Metal Loads in Pristine Andaman Waters

October 1, 2026
Hidden Drought Memory: Fractal Analysis Reveals Brazil’s Cerrado Is Far More Unstable Than Rainfall Trends Suggest
Earth Science

Hidden Drought Memory: Fractal Analysis Reveals Brazil’s Cerrado Is Far More Unstable Than Rainfall Trends Suggest

October 1, 2026
Drying Rice Paddies on Purpose Slashes Methane and Reshapes Soil Chemistry in Vietnam
Earth Science

Drying Rice Paddies on Purpose Slashes Methane and Reshapes Soil Chemistry in Vietnam

October 1, 2026
Machine Learning Predicts How Earthquake-Damaged Concrete Columns Will Behave Next
Earth Science

Machine Learning Predicts How Earthquake-Damaged Concrete Columns Will Behave Next

October 1, 2026
Machine Learning Pinpoints Prime Sites for Check Dams in Iran’s Semi-Arid Mountains
Earth Science

Machine Learning Pinpoints Prime Sites for Check Dams in Iran’s Semi-Arid Mountains

October 1, 2026
Plastic Chemicals in Nigerian Streams and Fish Exceed Safe Health Limits, Study Warns
Earth Science

Plastic Chemicals in Nigerian Streams and Fish Exceed Safe Health Limits, Study Warns

October 1, 2026
Next Post
New Questionnaire Aims to Measure How Staff Judge Mechanical Restraint in Psychiatric Care

New Questionnaire Aims to Measure How Staff Judge Mechanical Restraint in Psychiatric Care

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Fish Livers Reveal Hidden Metal Loads in Pristine Andaman Waters
  • Hidden Drought Memory: Fractal Analysis Reveals Brazil’s Cerrado Is Far More Unstable Than Rainfall Trends Suggest
  • AI Diffusion Network Learns to Read Faint Light Pulses From Neutrino Detectors
  • Autophagy’s Double-Edged Role in Womb Scarring and Age-Related Fertility Decline

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading