Vision-language models use pixel unshuffle to pack nearby visual features into fewer, richer tokens. This gives the language model a shorter visual sequence to process without simply throwing away the local information.