Skip to contentSkip to stories

Updated

#xAI

Oct 8

Oct 8Thu
  1. Artificial AnalysisAI score29

    Grok Imagine Video 1.5 Lite nears frontier on three AA-Video-T2V capabilities

    AIArtificial Analysis reports that Grok Imagine Video 1.5 Lite comes closest to the frontier on AA-Video-T2V v2.0 in Multi-Scene & Narrative, Lighting & Materials, and Text Rendering. It is furthest behind in Dialogue & Lip Sync and Human Anatomy. Compared with Grok Imagine Video 1.5, Lite matches it in Physics and trails on the other nine capabilities, by the least in Multi-Scene & Narrative.

  2. Artificial AnalysisAI score28

    Artificial Analysis Pareto frontier: GPT-6 Luna cheapest per task at $0.22

    AIAmong models with a Hallucination-Gated All-Pass Rate above 0%, GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max), and Grok 4.7 (xhigh) set the Pareto frontier for score versus cost per task. GPT-6 Luna (max) is the cheapest at about $0.22 per task, scoring 3.3%, while Grok 4.7 (xhigh) leads at about $9.50 per task and Muse Spark 1.3 (max) costs about $4.20. The three Claude models cost about $18 to $22 per task.

  3. Artificial AnalysisAI score34

    Artificial Analysis compares six hallucination checkers on 20 shared tasks

    AIArtificial Analysis compared six hallucination checkers on the same deliverables from 20 tasks across eight models. GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet. The counts reflect checker behavior rather than establishing accuracy or ruling out self-preference.

Oct 7

Oct 7Wed
  1. Wired · AIAI score60

    Researchers Test GPT-6 Astra Driving a Corolla to In-N-Out

    AIThree Axiom engineers had OpenAI's GPT-6 Astra drive a 2024 Toyota Corolla to an In-N-Out drive-thru through a server linked to cameras and power steering, with a safety driver ready to brake. They also built a parking-lot benchmark, DrivingBench, where Astra completed the course slowly, Claude Fable 5.1 finished 45 percent, and Grok finished 11 percent.

Oct 6

Oct 6Tue
  1. ARC PrizeAI score22

    Grok 4.7 uses more reasoning tokens than Grok 4.6 on ARC-AGI-2

    AIGrok 4.7 used more reasoning tokens on average than Grok 4.6 on ARC-AGI-2 semi-private tasks at medium, high, and xhigh reasoning levels, raising its cost per task. Per test-pair attempt, medium used 136% more tokens, high 125% more, and xhigh 173% more, while low used 27% fewer. A chart compares the two models at xhigh on the 20 public tasks where Grok 4.7 increased token use the most.