Skip to content
View original post on X: Intern Large Models· 62/100AI score62/100

Intern Large Models introduces Visual Pretraining learned from visual documents

AISummary

Intern Large Models introduces Visual Pretraining, a pretraining paradigm for foundation models that learns directly from visual documents.

The post says it outperforms text-only pretraining across backbones and benchmarks, and links the arXiv paper 2607.09657 along with Intern-S2-Preview (35B) and Intern-S2-Preview-397B on Hugging Face, the latter presented as a multimodal foundation model trained with this recipe.

Post on XView on X

🔥Introducing Visual Pretraining (VP) 👀: a new, scalable pretraining paradigm for foundation models.
😉It enables models to learn directly from visual documents 📖 and outperforms text-only pretraining across backbones and benchmarks.
📈 Explore the paper and our most capable multimodal foundation model trained with Visual Pretraining as part of its training recipe.
📄 Paper:
https://arxiv.org/abs/2607.09657
🤗 Hugging Face Paper:
https://huggingface.co/papers/2607.09657
🤖 Intern-S2-Preview (35B):
https://huggingface.co/internlm/Intern-S2-Preview
🤖 Intern-S2-Preview-397B:
https://huggingface.co/internlm/Intern-S2-Preview-397B

Source: Intern Large Models · x.comPublished · added here