Skip to contentSkip to stories
Updated

#xAI

Oct 8

  1. Sherwin WuAI score60

    Harvey LAB-AA v1.1 adds hallucination gate; Grok 4.7 leads at 9.4%

    AISherwin Wu, an OpenAI employee, says the updated Harvey LAB-AA v1.1 benchmark, announced by Artificial Analysis with Harvey, is more useful than the original LAB results. The new Hallucination-Gated All-Pass Rate credits a task only when every rubric criterion passes and no material hallucination appears. Grok 4.7 (xhigh) leads at 9.4%, while GPT-6 Astra (max) at 8.6% has very few material hallucinations.

    Why it matters: The update adds a hallucination gate to a legal benchmark, showing that models with high all-pass rates can rank much lower once material errors count.

  2. Artificial AnalysisAI score62

    GPT-6 Sol (Daybreak Blue) leads Artificial Analysis Cyber Index with trusted access

    AIArtificial Analysis added trusted-access models to its Cyber Index, and GPT-6 Sol (Daybreak Blue, max) now leads the leaderboard. The model is available only through OpenAI's Daybreak program and records no safety blocks, improving 32 points over the publicly available GPT-6 Sol (max). It costs $1.77 per task, below Grok 4.7 (xhigh) at $11.67 per task.

    Why it matters: The post shows how a trusted-access model compares with public models on cyber defense tasks, separating access restrictions from measured capability and cost.

    Image from @ArtificialAnlys's post

Sep 20

  1. xAI News (Grok)AI score72

    xAI releases Grok 4.7, its most capable model for coding and knowledge work

    AIxAI released Grok 4.7, which it calls its most capable model for coding and knowledge work, built on a larger base model than Grok 4.6 and trained with a longer reinforcement learning run. It is priced from $2 per million input tokens and $6 per million output tokens, the same as Grok 4.6, and is available in Cursor, Grok Build, and the Grok API. xAI reports gains on CursorBench 4.0 (46.3%) and AA Briefcase v1.1 (1,657) over Grok 4.6, and says it posts the strongest safety results it has tested on refusals and jailbreak resistance.

    Why it matters: The release pairs a new base model with benchmark tables against named rivals and pricing, letting readers compare its coding and office-work gains against Grok 4.6 and frontier models.

Aug 14

  1. Michael TruellAI score75

    Cursor officially joins SpaceX after acquisition closes

    AICursor's acquisition by SpaceX has officially closed, and Cursor will join the SpaceXAI team. The stated goal is to help make Grok the world's most useful AI and to improve Grok Build, Grok Bot, Grok API, Cursor, and more.

    Why it matters: The acquisition closing ties Cursor's coding tools to Grok's product line, which changes how the two products may be developed and sold together.

  2. Cursor BlogAI score62

    Cursor is acquired by SpaceX, gaining access to its GPU fleet

    AICursor has been acquired by SpaceX, completing a process that began in April when the two companies announced a partnership to accelerate model training. The post says the deal gives Cursor access to what it calls the largest GPU fleet in the world, which it expects to yield more capable models at lower cost. It cites Grok 4.6, released Wednesday, as an early look at what the companies can build together.

    Why it matters: The post confirms a completed acquisition and links it to GPU access and cheaper model serving, which explains why the deal matters for coding tools.

Dec 4, 2025

  1. ARC PrizeAI score62

    ARC Prize 2025 results point to refinement loops as the central AI reasoning trend

    AIARC Prize reports that the top Kaggle entry reached 24% on the ARC-AGI-2 private dataset at $0.20 per task, and that all winning solutions and papers are open source. The top verified commercial model, Opus 4.5 (Thinking, 64k), scored 37.6% at $2.20 per task, while a Poetiq refinement on Gemini 3 Pro reached 54% at $30 per task. The author argues that refinement loops are the main driver of 2025 progress, and says ARC-AGI-3 is planned for early 2026.

    Why it matters: The post links 2025 competition results to a broader argument about refinement loops, showing how benchmark outcomes are being read as evidence of AI reasoning progress.

That’s everything