Pixel Shuffle

Space-to-depth · Visual token packing

Vision-language models use pixel unshuffle to pack nearby visual features into fewer, richer tokens. This gives the language model a shorter visual sequence to process without simply throwing away the local information.

64×64×1
32×32×4
Camera
Image