Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 2

Oct 2Fri
  1. TransformerBlogAI score55

    Human oversight may not prevent AI-driven military errors, analysis argues

    AIJoshua Keating argues that keeping a human in the loop on lethal AI decisions is not enough if the humans rely too heavily on AI outputs. He cites a CNN-reported case in which an analyst's AI-assisted report falsely identified a Chinese ship's cargo as nuclear components, nearly prompting a boarding during the Iran war. The piece links this to automation bias and to military AI cases in Gaza and Minab, and warns that AI integration early in a nuclear decision chain is harder to regulate than autonomous launch.

  2. GitHub Blog · AI & MLOfficialAI score23

    Three Skills Developers Need as AI Changes Their Work

    AIAI is changing developer work, and the article recommends three skills: directing AI agents, reviewing AI output instead of trusting the first answer, and using saved time for judgment-heavy problems such as customer needs and tradeoffs. It cites GitHub Copilot's built-in Rubber Duck agent, which uses a second model to critique plans, code, and tests. The author argues that developers remain responsible for outcomes while AI handles more implementation.

  3. a16z NewsBlogAI score32

    The Case for Scaling America's Defense Manufacturing Base Beyond Prototypes

    AIVenture investors have funded defense-tech companies such as SpaceX, Anduril, and Castelion, but the article argues that production capacity in the supplier base is now the bottleneck. Most of America's machine shops and manufacturers are small, with 83% of machine shops employing fewer than 20 people, and 61% of tier-two-and-below defense manufacturers cite tooling, automation, or production-line limits as top expansion barriers.

  4. Thomas WolfXAI score40

    Thomas Wolf questions AI's growth against mathematics' infinite scope

    AIThomas Wolf shares Kevin Buzzard's reflections on whether mathematics is about human understanding and what exponential AI growth means for an infinite field. He quotes Buzzard's view that machines may eventually reach a natural boundary where further progress is not worth the resources, and that humans would then take over from there.

    Image from @Thom_Wolf's post
  5. MIT Technology Review · AINewsAI score62

    AlphaGo's move 37 shows why LLMs do not truly reason, an AlphaGo team member argues

    AIThore Graepel, a core member of the AlphaGo team, argues that current large language models do not truly reason, despite chain-of-thought gains in math and coding. He says they lack an explicit, inspectable epistemic state, keep knowledge and reasoning intertwined in their weights, and often produce post-hoc explanations. He proposes systems that maintain an auditable epistemic state and evaluate each step by how much it resolves uncertainty.

  6. AI Futures ProjectBlogAI score62

    Former OpenAI forecaster urges Senate to curb AI research automation race

    AIDaniel Kokotajlo, who leads the AI Futures Project, testified before a Senate subcommittee on September 30, 2026. He argued that Anthropic and OpenAI are racing toward superintelligence by automating AI research and development, and that his team thinks this could happen as early as 2028. He warned that declining monitorability and models that appear aligned during evaluations make misalignment harder to detect, and he recommended greater industry transparency and redirecting compute away from AI R&D.

  7. indigoXAI score28

    indigo proposes a three-tier Agent usage model for startups

    AIindigo compares AI agent usage to phones, professional computers, and enterprise IT, dividing it into personal, professional, and organizational tiers. The post argues that startups should avoid the consumer tier and focus on professional workflows grounded in personal experience, or enterprise deployment and agent infrastructure.

    Image from @indigox's post
  8. Lucas Beyer (bl16)XAI score45

    Lucas Beyer praises new coding benchmark for finding bugs in repos

    AILucas Beyer calls SWE-sweep a useful new benchmark, where agents must find and fix bugs in a repo checked out at an earlier commit, scored against unit tests from real later bugfixes. He notes two limitations: a model may find valid bugs that don't match the tested ones, and the construction makes training on the test set easy. He advises not overemphasizing small ranking differences once models score highly.

  9. Thomas WolfXAI score44

    Ben Affleck explains fine-tuning open video models for film production

    AIBen Affleck described fine-tuning open video models by freezing base weights and training only the last cinematic layer, using a learning rate of 2e-4. He said his company, InterPositive, built a private model per film from its own dailies after raising money to shoot a controlled dataset over eight months.

    Image from @Thom_Wolf's post
  10. will depueXAI score13

    Will Depue says tweet is not misinterpreted, citing Anthropic employees

    AIWill Depue, who is affiliated with OpenAI, says his tweet is not being misinterpreted and that he has heard the same claim from Anthropic employees about OpenAI. He argues that a later interpretation shifting the framing to US versus China is weak and only marginally better.

  11. Dongxi NLPXAI score27

    LLMs replace condescending engineers by explaining code patiently in many formats

    AIThe author recalls a senior engineer who dismissed a newcomer's question with "oops, forgot," and says LLMs now answer patiently through text, diagrams, videos, and more. The post frames this shift as making dismissive gatekeeping obsolete, building on Andrej Karpathy's tips for turning LLM outputs into easier-to-read formats such as ASD-STE100 writing, diagrams, HTML pages, and generated explainer videos.

  12. jasonXAI score4

    Jason Liu suggests AI may reshape travel booking as it did coding

    AIJason Liu says people claimed the same about coding, suggesting AI could change travel booking similarly. He reacts to a post contrasting booking a flight and hotel in two minutes alone with a ten-minute call where an assistant reviews flight options and sends map screenshots.

  13. Vaibhav (VB) SrivastavXAI score8

    ChatGPT's rapid progress since 2022 illustrates long-term AI growth

    AIVaibhav Srivastav, an OpenAI-affiliated account, posted that people overestimate what they can achieve in one year but underestimate what they can accomplish over a decade. A quoted post recalled how ChatGPT looked in 2022, four years earlier, to frame that point.

  14. TinkerOfficialAI score33

    Tinker praises Fulcrum's cheap, effective style-customization training approach

    AITinker says Fulcrum trains its Echo writing model by building on a base model that already writes well, tailoring both SFT and RL to separate the default LLM voice from authors' voices. The post calls this customization approach both cheap and effective. Fulcrum says Echo beats frontier models at writing tasks such as fiction and technical explanations, at a training cost under $5K.

Oct 1

Oct 1Thu
  1. Yuchen JinXAI score22

    Yuchen Jin wishes AI could generate Andrej Karpathy-style videos

    AIYuchen Jin says no AI today can generate an Andrej Karpathy video from a prompt, and he hopes Karpathy will return, noting he has not uploaded a YouTube video in over a year. The post builds on Karpathy's suggestion that LLMs can produce bespoke explainer videos on arbitrary topics, though that idea is still emerging.

  2. Latent.SpaceXAI score60

    Recursive Language Models explained by MIT's Alex Zhang on coding agents

    AIA Latent.Space podcast episode features MIT researcher Alex Zhang explaining recursive language models (RLMs). He discusses why Claude Code, Codex, and Pi are basically the same, and how RLMs use code, context offloading, and recursive subagents to generalize across tasks. The episode also covers OpenAI's 10,000-agent, 130B-output-token experiment and academia's freedom to pursue ambitious research bets.

    Video from @latentspacepod's post
  3. Ben TossellXAI score10

    Ben Tossell says AI tools fail non-technical users and must improve

    AIBen Tossell argues that AI products are not good enough for semi-technical users like himself, who are neither developers nor casual consumers. He says companies should stop blaming users for poor experiences and make their products better, since current results hurt the broader AI narrative. His reply to his own earlier post calls his experience with dot "💩".

    Image from @bentossell's post
  4. indigoXAI score23

    Rumored fable 5.5 HTML demo shows Superman shifting art styles

    AIA post claims an HTML output generated by the rumored fable 5.5 model shows Superman changing into different artistic styles as he passes through artworks. It also cites a NYT report that Anthropic has been telling religious leaders that Claude has a soul and moral status.

    Video from @indigox's post
  5. Kylie RobisonXAI score6

    Kylie Robison Says Journalists Should Make News About Themselves

    AIJournalist Kylie Robison jokingly argues that it is her duty as a journalist to make this about her, linking to a post on X. The post is brief and offers no substantive news, and the quoted reply from @dwr only comments that "dot" is a useful name for a personal agent if it launches a dot-shaped hardware device.

  6. Harrison ChaseXAI score33

    Harrison Chase Argues Every Agent Harness Needs a Durable Runtime

    AIHarrison Chase argues that every agent harness requires a durable runtime, citing pi-durable as an example alongside deepagents built on LangGraph. The post frames durable execution as a basic requirement for agent systems rather than an optional feature. Pi 1.0 shipped with Pi Durable, which the referenced @pidotdev post invites users to customize.

  7. FireworksOfficialAI score13

    Fireworks details keeping RL rollout and training numerically consistent

    AIRollouts account for most of RL's compute cost, and splitting them from training across separate engines can introduce numerical mismatches. In MoE models, such mismatches can even route tokens to different experts. Fireworks says it co-builds both engines so training stays fast and consistent.

  8. Boris PowerXAI score14

    Boris Power teases the next transportation revolution without naming a product

    AIBoris Power, who is listed as associated with OpenAI, posted only that he is excited for the next transportation revolution, without naming a product, company, or figures. The post links to Bryan Caplan's piece on rideshare economics, which estimates an equilibrium price of about $2 an hour for autonomous vehicles even when they sit empty half the time.

  9. Dongxi NLPXAI score46

    Dongxi jokes about replacing remote consultants with Griffin AI agents

    AIThe author jokes about founding a consulting firm that would use agents for work, Griffin for meetings, and Griffin for interviews to fill remote roles. They then question whether remote engineers and consultancies would still be needed if that became reality. The quoted Tavus post says Griffin passed a video Turing test with 48% of live interlocutors believing it was human.

  10. ClineOfficialAI score34

    Cline reports DeepSeek V4 Pro costs about 30x less than Claude Opus 5

    AICline says two of its largest tasks in the last 30 days each processed 9B tokens, costing about $8,500 on Claude Opus 5 versus about $300 on DeepSeek V4 Pro. The post says that is roughly 30x cheaper for the same token count, letting users run large-horizon work without spending thousands or waiting on limit resets.

  11. Dongxi NLPXAI score42

    arXiv tightens rate limits as AI-assisted research floods submissions

    AIarXiv has introduced stricter rate limiting for all submitters to fairly distribute moderator time. The post links this move to Vibe research, where turning ideas into papers is easier, while noting that standards for judging research output have not kept pace.

  12. Alex HeathXAI score46

    OpenAI's Dots lead ChatGPT's always-on personal agent plans

    AIAlex Heath's podcast with OpenAI's @embirico covers Dots, a new always-on personal agent in ChatGPT that asks permission before acting by default. The episode also discusses Space, a workspace where people and agents collaborate on documents, data, and projects, along with pricing and possible access for free users.

    Video from @alexeheath's post
  13. AnthropicOfficialAI score38

    Harvard physicist builds toolkit to match Claude with science calculations

    AIHarvard physicist Matthew Schwartz argues that LLMs are poorly matched to science when used as human-style collaborators, so he built a toolkit for exact quantitative calculations. Working with Claude, the approach surfaced connections to ecology, population genetics, and a dozen other fields, with domain experts steering it toward interesting questions.

  14. François CholletXAI score62

    Chollet Argues Reasoning Models Differ from Base LLMs by Inductive Program Prediction

    AIFrançois Chollet argues the key difference between base LLMs and modern LRMs is a shift from transductive answer prediction to inductive prediction of the program or reasoning chain behind an answer. He says this enables test-time induction and substantial fluid intelligence in LRMs, which he claims base LLMs largely lack. He cites ARC 1 results: base LLMs remain around 10-15%, while LRMs of the same size or smaller saturated the benchmark in 2025.

  15. Sophia YangXAI score20

    Ember-1 Shows Lower Cost and Faster Reasoning Than Kimi K3

    AIEmber-1, built on the Kimi K3 base model, is reported by @kickingkeys to cost about 20% less and use about 36% fewer reasoning tokens across 124 test prompts. The same tester says it runs about 2x faster, while the source's painting examples show its own expressive style.