---
title: "FlowTok: Flowing Seamlessly Across Text and Image Tokens"
canonical_url: "https://www.modelscope.cn/papers/126758"
md_url: "https://www.modelscope.cn/papers/126758.md"
arxiv_id: 2503.10772
published: 2025-03-13
last_updated: 2025-03-13
authors:
  - "Ju He"
  - "Qihang Yu"
  - "Qihao Liu"
  - "Liang-Chieh Chen"
model_name: FlowTok
model_developer: "字节跳动"
domain:
  - "自然语言处理"
  - "计算机视觉"
  - "深度学习"
type:
  - "自然语言处理"
  - "计算机视觉"
  - "深度学习"
  - "Computer Vision and Pattern Recognition (cs.CV)"
arxiv_url: "https://arxiv.org/abs/2503.10772"
pdf_url: "https://arxiv.org/pdf/2503.10772.pdf"
code_link: "https://github.com/bytedance/1d-tokenizer"
---

# FlowTok: Flowing Seamlessly Across Text and Image Tokens

> Bridging different modalities lies at the heart of cross-modality generation. While conventional approaches treat the text modality as a conditioning signal that gradually guides the denoising process from Gaussian noise to the target image modality, we…

「FlowTok: Flowing Seamlessly Across Text and Image Tokens」是 ModelScope 魔搭社区收录的论文，arXiv 2503.10772，作者为 Ju He, Qihang Yu, Qihao Liu et al.，发表于 2025-03-13，属于 自然语言处理、计算机视觉、深度学习 领域。

- **ArXiv**: 2503.10772
- **Published**: 2025-03-13
- **Authors**: Ju He, Qihang Yu, Qihao Liu, Liang-Chieh Chen
- **Model**: FlowTok
- **Developer**: 字节跳动
- **Domain**: 自然语言处理, 计算机视觉, 深度学习
- **ArXiv URL**: https://arxiv.org/abs/2503.10772
- **PDF**: https://arxiv.org/pdf/2503.10772.pdf
- **Code**: https://github.com/bytedance/1d-tokenizer

Source: https://www.modelscope.cn/papers/126758

---

> FlowTok：让文本与图像无缝流转的高效跨模态生成新范式

## 摘要

本文提出了一种名为FlowTok的框架，旨在实现文本和图像模态之间的无缝转换。传统方法通常依赖扩散模型，将文本作为条件信号逐步引导图像生成过程，而FlowTok采用了一种更简单的范式——通过流匹配直接在共享隐空间中演化文本和图像。具体而言，FlowTok将文本和图像分别编码为紧凑的一维隐表示，并对齐其形状以支持直接流匹配。这种方法显著减少了隐空间大小（约3.3倍），消除了复杂条件机制的需求，并提高了训练和推理效率。

FlowTok的核心设计包括：1）使用预训练文本编码器提取初始文本嵌入，并通过轻量级投影器将其映射到低维隐空间；2）基于改进的TA-TiTok方法将图像编码为紧凑的一维隐令牌；3）引入辅助文本对齐损失以保留语义一致性。实验表明，FlowTok不仅在COCO数据集上取得了与现有方法相当的性能，还大幅降低了训练资源需求和推理时间。此外，该框架可自然扩展到图像到文本生成任务，保持了简洁性和高效性。

FlowTok的主要贡献在于提供了一个统一且高效的跨模态生成框架，简化了传统方法的复杂管道，同时实现了最先进的性能。这种设计为未来通用跨模态生成研究奠定了坚实基础。

## Abstract

Bridging different modalities lies at the heart of cross-modality generation. While conventional approaches treat the text modality as a conditioning signal that gradually guides the denoising process from Gaussian noise to the target image modality, we explore a much simpler paradigm-directly evolving between text and image modalities through flow matching. This requires projecting both modalities into a shared latent space, which poses a significant challenge due to their inherently different representations: text is highly semantic and encoded as 1D tokens, whereas images are spatially redundant and represented as 2D latent embeddings. To address this, we introduce FlowTok, a minimal framework that seamlessly flows across text and images by encoding images into a compact 1D token representation. Compared to prior methods, this design reduces the latent space size by 3.3x at an image resolution of 256, eliminating the need for complex conditioning mechanisms or noise scheduling. Moreover, FlowTok naturally extends to image-to-text generation under the same formulation. With its streamlined architecture centered around compact 1D tokens, FlowTok is highly memory-efficient, requires significantly fewer training resources, and achieves much faster sampling speeds-all while delivering performance comparable to state-of-the-art models. Code will be available at this https URL.
