NVIDIA releases PixelUMM, an encoder-free model for pixel-space image and video tasks
Original titlenvidia/PixelUMM
AISummary
NVIDIA has released PixelUMM, an encoder-free unified multimodal model with 15,199,672,064 parameters that handles text, image, and video understanding and generation directly in pixel space.
It represents images as 16-by-16 RGB pixel patches on a Qwen3-8B language backbone, with iterative denoising for generation. The checkpoint is licensed for non-commercial research or evaluation only, while the source code is under Apache License 2.0.
Source: NVIDIA · new models on Hugging Face · huggingface.coPublished · added here