Skip to contentSkip to stories

Updated

#arXiv

Oct 9

TodayOct 9Fri1 item
  1. X.PINAI score46

    Seed preprint finds DeepSeek V4 long-context retrieval varies by position

    AIA Seed team preprint reports "phase sensitivity" in DeepSeek V4 and V4.1-Flash, where identical information becomes harder to retrieve depending on its position within compressed KV-cache blocks. The compression reduces memory and attention costs, but long-context retrieval accuracy varied by up to 40 percentage points across positions. The authors note that average benchmark scores can hide these recurring weak spots, though the findings concern retrieval specifically rather than all model behavior.

Oct 8

Oct 8Thu

Oct 7

Oct 7Wed
  1. Google ResearchAI score23

    Google Research invites COLM visitors to ContinuousBench walkthrough on DP synthetic data

    AIGoogle Research is hosting a walkthrough at its COLM booth #107 today at 5:00 PM of ContinuousBench, a standardized benchmark for measuring knowledge transfer in differentially private synthetic data. The session, led by Alex Bie, asks whether DP synthetic data preserve actual information or only style. A paper is linked on arXiv.

Oct 6

Oct 6Tue

Oct 1

Oct 1Thu

Jul 22

Jul 22Wed

Feb 25

Feb 25Wed
  1. Quoc LeAI score65

    Aletheia Agent Solves 6 of 10 FirstProof Math Problems Autonomously

    AIGoogle researchers used the Aletheia agent, powered by Gemini 3 Deep Think, to attempt 10 FirstProof challenge problems without modification. The agent operated fully autonomously and solved 6 of the 10 problems, according to the post, with methodology and expert evaluations described in the linked arXiv paper.

    Why it matters: The post gives the autonomous setup and expert-evaluated results for an AI agent on FirstProof math problems, useful for judging how far such systems go on research-level math.

Feb 11

Feb 11Wed