---
title: "Lighting-grounded Video Generation with Renderer-based Agent Reasoning"
canonical_url: "https://www.modelscope.cn/papers/265642"
md_url: "https://www.modelscope.cn/papers/265642.md"
arxiv_id: 2604.07966
published: 2026-04-09
last_updated: 2026-04-09
authors:
  - "Ziqi Cai"
  - "Taoyu Yang"
  - "Zheng Chang"
  - "Si Li"
  - "Han Jiang"
  - "Shuchen Weng"
  - "Boxin Shi"
model_name: LiVER
model_developer: "北京大学、北京邮电大学、OpenBayes Information Technology Co.、Ltd.、北京智源人工智能研究院"
domain:
  - "计算机视觉"
  - "视频生成"
  - "可控生成"
  - "光照编辑"
  - "扩散模型"
type:
  - "计算机视觉"
  - "视频生成"
  - "可控生成"
  - "光照编辑"
  - "扩散模型"
  - "Computer Vision and Pattern Recognition"
arxiv_url: "https://arxiv.org/abs/2604.07966"
pdf_url: "https://arxiv.org/pdf/2604.07966"
---

# Lighting-grounded Video Generation with Renderer-based Agent Reasoning

> Diffusion models have achieved remarkable progress in video generation, but their controllability remains a major limitation. Key scene factors such as layout, lighting, and camera trajectory are often entangled or only weakly modeled, restricting their…

「Lighting-grounded Video Generation with Renderer-based Agent Reasoning」是 ModelScope 魔搭社区收录的论文，arXiv 2604.07966，作者为 Ziqi Cai, Taoyu Yang, Zheng Chang et al.，发表于 2026-04-09，属于 计算机视觉、视频生成、可控生成 领域。

- **ArXiv**: 2604.07966
- **Published**: 2026-04-09
- **Authors**: Ziqi Cai, Taoyu Yang, Zheng Chang, Si Li, Han Jiang, Shuchen Weng, Boxin Shi
- **Model**: LiVER
- **Developer**: 北京大学、北京邮电大学、OpenBayes Information Technology Co.、Ltd.、北京智源人工智能研究院
- **Domain**: 计算机视觉, 视频生成, 可控生成, 光照编辑, 扩散模型
- **ArXiv URL**: https://arxiv.org/abs/2604.07966
- **PDF**: https://arxiv.org/pdf/2604.07966

Source: https://www.modelscope.cn/papers/265642

---

> LiVER：基于渲染器代理推理与光照接地场景代理的物理真实视频生成新范式

## 摘要

本文针对当前文本到视频（T2V）扩散模型在物理可控性上的不足，尤其是对光照、场景布局和相机轨迹等关键因素建模弱、纠缠严重的问题，提出了LiVER（Lighting-grounded Video genERation）框架。研究背景在于影视制作与虚拟制片等高要求场景亟需显式、解耦、物理真实的三维控制能力。作者提出以渲染器为驱动的智能代理（renderer-based agent），将文本指令解析为粗粒度3D场景图，并联合HDR环境光、相机轨迹与材质属性，通过物理渲染引擎（Blender）生成包含Diffuse、Rough GGX和Glossy GGX三通道的‘光照感知场景代理’（lighting-grounded scene proxy）。该代理作为显式物理条件注入预训练视频扩散模型（Wan2.2-5B），辅以轻量级条件编码器与适配器及三阶段渐进训练策略，实现高质量、高保真、高一致性的视频合成。为支撑训练与评估，作者构建了首个兼顾真实与合成的光照标注视频数据集LiVERSet（11K视频，含几何、HDR、相机、文本等密集标注）。实验表明，LiVER在光照真实性（如阴影、反射、材质响应）、布局忠实度与相机轨迹对齐方面达到SOTA水平，显著提升了可控视频生成的物理可信度。

## Abstract

Diffusion models have achieved remarkable progress in video generation, but their controllability remains a major limitation. Key scene factors such as layout, lighting, and camera trajectory are often entangled or only weakly modeled, restricting their applicability in domains like filmmaking and virtual production where explicit scene control is essential. We present LiVER, a diffusion-based framework for scene-controllable video generation. To achieve this, we introduce a novel framework that conditions video synthesis on explicit 3D scene properties, supported by a new large-scale dataset with dense annotations of object layout, lighting, and camera parameters. Our method disentangles these properties by rendering control signals from a unified 3D representation. We propose a lightweight conditioning module and a progressive training strategy to integrate these signals into a foundational video diffusion model, ensuring stable convergence and high fidelity. Our framework enables a wide range of applications, including image-to-video and video-to-video synthesis where the underlying 3D scene is fully editable. To further enhance usability, we develop a scene agent that automatically translates high-level user instructions into the required 3D control signals. Experiments show that LiVER achieves state-of-the-art photorealism and temporal consistency while enabling precise, disentangled control over scene factors, setting a new standard for controllable video generation.
