站内快照 · 国内可打开。外网原文可能无法访问。
- 论文公开站arXiv
Learning to Read the Contextual Tokens in Diffusion Transformers
Multimodal Diffusion Transformers (MM-DiTs) jointly process visual and textual representations throughout generation. These models repeatedly update the text tokens through multimodal attention, forming dynamic contextual tokens whose function is not well understood. In this work