MiniMax Releases VTP-Large-f16d64 Visual Tokenizer With Technical Report and Pretrained Weights
Original titleMiniMaxAI/VTP-Large-f16d64
AISummary
MiniMax released the technical report and pretrained weights for VTP-Large-f16d64, a visual tokenizer that jointly optimizes contrastive, self-supervised, and reconstruction losses.
The model scores 78.2 zero-shot accuracy, 85.7 linear probing, and 0.36 rFID, and its generation performance scales with pretraining compute, parameters, and data. Checkpoint weights were listed as "released very soon" in the source.
Source: MiniMax · new models on Hugging Face · huggingface.coPublished · added here