Skip to content
Read the original: Jason Wei· Published 36/100AI score36/100

Muse Spark 1.1 beats GPT-5.6 Sol on radiology benchmark RadLE 2.0

Original titleMuse Spark 1.1 outperforms GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam. We don't beat Fable (yet). And Humans are still a lot bet...

AISummary

On Radiology's Last Exam, Meta's Muse Spark 1.1 outperforms OpenAI's GPT-5.6 Sol and Gemini 3.1, but still trails Fable and human experts. The result comes from a post by Jason Wei, with the benchmark context coming from a separate post about RadLE 2.0, an uncertainty-aware radiology diagnosis benchmark.

Read the original x.com

Source: Jason Wei · x.comPublished · added here