Read the original: IndexTeam (Bilibili) · new models on Hugging Face·Published · 5d agoAI score29/100
Index-Nailong-2B-FP4 Released as NVFP4 Quantized Translation Model
Original titleIndexTeam/Index-Nailong-2B-FP4
AISummary
IndexTeam has released Index-Nailong-2B-FP4, an official NVFP4 (W4A4) quantization of its Index-Nailong-2B multilingual translation model, which supports 150 languages.
The checkpoint keeps lm_head, embeddings, and MoE router gates in BF16, and a perplexity test on a fixed corpus rose from 3.2806 to 3.4998 (+6.68%), while zh->en and en->zh outputs matched BF16 semantically.
Full FP4 acceleration requires an NVIDIA Blackwell GPU; on Hopper or Ampere, vLLM provides only memory savings, so the FP8 build is recommended.
Source: IndexTeam (Bilibili) · new models on Hugging Face · huggingface.co