---
title: "VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary"
canonical_url: "https://www.modelscope.cn/papers/126245"
md_url: "https://www.modelscope.cn/papers/126245.md"
arxiv_id: 2503.09402
published: 2025-03-12
last_updated: 2025-03-12
authors:
  - "Kevin Qinghong Lin"
  - "Mike Zheng Shou"
model_name: VLog
model_developer: "新加坡国立大学显示实验室"
domain:
  - "计算机视觉"
  - "自然语言处理"
  - "深度学习"
type:
  - "计算机视觉"
  - "自然语言处理"
  - "深度学习"
  - "Computer Vision and Pattern Recognition (cs.CV)"
arxiv_url: "https://arxiv.org/abs/2503.09402"
pdf_url: "https://arxiv.org/pdf/2503.09402.pdf"
code_link: "https://github.com/showlab/VLog"
---

# VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary

> Human daily activities can be concisely narrated as sequences of routine events (e.g., turning off an alarm) in video streams, forming an event vocabulary. Motivated by this, we introduce VLog, a novel video understanding framework that define video…

「VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary」是 ModelScope 魔搭社区收录的论文，arXiv 2503.09402，作者为 Kevin Qinghong Lin, Mike Zheng Shou，发表于 2025-03-12，属于 计算机视觉、自然语言处理、深度学习 领域。

- **ArXiv**: 2503.09402
- **Published**: 2025-03-12
- **Authors**: Kevin Qinghong Lin, Mike Zheng Shou
- **Model**: VLog
- **Developer**: 新加坡国立大学显示实验室
- **Domain**: 计算机视觉, 自然语言处理, 深度学习
- **ArXiv URL**: https://arxiv.org/abs/2503.09402
- **PDF**: https://arxiv.org/pdf/2503.09402.pdf
- **Code**: https://github.com/showlab/VLog

Source: https://www.modelscope.cn/papers/126245

---

> VLog：用生成式检索重新定义视频叙述，实现高效与精准兼备的视频理解

## 摘要

本文的研究背景在于当前视频理解模型主要依赖于大型语言模型的子词词汇，这些模型在推理阶段逐词生成描述，效率低下且缺乏对复杂关系的建模能力。此外，传统的子词词汇虽然涵盖广泛的语言信息，但在视频理解任务中缺乏视觉可解释性。为解决这些问题，本文提出了一种名为VLog的新框架。VLog通过引入叙述词汇（Narration Vocabulary）重新定义了视频描述方式，采用轻量级GPT-2语言模型并结合对比检索模型的优势，实现了高效且复杂的推理能力。具体而言，VLog提出了三项创新：1）一种生成式检索模型，结合了语言模型的推理能力和检索模型的高效相似性搜索；2）基于大规模视频叙述数据构建的分层叙述词汇，能够快速索引特定事件；3）一种词汇更新策略，用于扩展推理过程中遇到的新事件。为了验证VLog的有效性，作者构建了一个新的开发集VidCap-Eval，并在多个公开数据集（如EgoSchema、COIN和HiREST）上进行了实验。实验结果表明，VLog能够在保证准确性和上下文相关性的前提下，显著提高长视频处理速度，为视频理解提供了全新的视角。

## Abstract

Human daily activities can be concisely narrated as sequences of routine events (e.g., turning off an alarm) in video streams, forming an event vocabulary. Motivated by this, we introduce VLog, a novel video understanding framework that define video narrations as vocabulary, going beyond the typical subword vocabularies in existing generative video-language models. Built on the lightweight language model GPT-2, VLog feature three key innovations: (i) A generative retrieval model, marrying language model's complex reasoning capabilities with contrastive retrieval's efficient similarity search. (ii) A hierarchical vocabulary derived from large-scale video narrations using our narration pair encoding algorithm, enabling efficient indexing of specific events (e.g., cutting a tomato) by identifying broader scenarios (e.g., kitchen) with expressive postfixes (e.g., by the left hand). (iii) A vocabulary update strategy leveraging generative models to extend the vocabulary for novel events encountered during inference. To validate our approach, we introduce VidCap-Eval, a development set requiring concise narrations with reasoning relationships (e.g., before and after). Experiments on EgoSchema, COIN, and HiREST further demonstrate the effectiveness of VLog, highlighting its ability to generate concise, contextually accurate, and efficient narrations, offering a novel perspective on video understanding. Codes are released at this https URL.
