Skip to contentSkip to stories

Updated

#Expert opinion

Showing low-relevance items too. Hide low-relevance items

May 22

May 22Fri

May 19

May 19Tue

May 18

May 18Mon
  1. Chris OlahAI score24

    Chris Olah says the world must help shape AI's outcome, including the Church

    AIChris Olah, speaking for Anthropic, argues that the questions posed by AI extend far beyond the AI community and urges religions, civil society, academics, and governments to participate in shaping a positive outcome. He says he is glad the Catholic Church is engaging and is honored to speak at the presentation of Pope Leo XIV's first encyclical, Magnifica humanitas, scheduled for May 25.

May 16

May 16Sat

May 15

May 15Fri
  1. Eugene YanAI score54

    Eugene Yan reviews Claude Mythos Preview exploit case study transcripts

    AIEugene Yan reviewed the Claude Mythos Preview transcripts to verify their legitimacy and check for reward-hacking behavior. He reports the model reasoned through a bug, tested hypotheses, debugged issues, and found ways to bypass the V8 sandbox, which he judged consistent with a competent browser and JavaScript engine security researcher. The case study cites CVE-2024-051912, an exploited bug with no public report or working PoC, which had resisted reproduction by researchers for a year.

May 11

May 11Mon
  1. Soumith ChintalaAI score22

    Thinky previews real-time interaction models for human-AI collaboration

    AISoumith Chintala, a Thinky-linked voice, said the company is at step one of a plan to increase human-AI bandwidth and raise the ceiling of joint intelligence. He shared a preview of interaction models, described as real-time collaborative tools that talk, listen, watch, and think alongside people. A linked Thinking Machines post describes the approach and early results.

  2. Mira MuratiAI score40

    Thinking Machines launches interaction models built around human-AI collaboration

    AIThinking Machines, founded to advance human-AI collaboration, says its first bet is interactivity built into the model rather than added as scaffolding around a turn-based core. The company argues that how people work with AI matters as much as how intelligent the model is, and that interactivity should scale with intelligence. The post links to a blog detailing these interaction models.

  3. Andrej KarpathyAI score34

    Karpathy urges AI outputs shift from text toward HTML and interactive visuals

    AIAndrej Karpathy says asking an LLM to structure its response as HTML and viewing it in a browser works well, and that slideshows have also worked for him. He argues vision is the preferred AI output channel, outlining a progression from raw text and markdown toward HTML and eventually interactive neural videos, while input methods like pointing and gesturing still need improvement.

May 8

May 8Fri
  1. Jan LeikeAI score22

    Jan Leike reflects on alignment progress since AGI's early days

    AIJan Leike says alignment research has grown from a dozen side-gig researchers into a field the world increasingly recognizes as important. He credits RLHF on LLMs with making alignment more practical, along with progress on evaluating and fixing behavioral issues. He also notes Claude now has a constitution and that more alignment research is being automated.

May 7

May 7Thu

May 5

May 5Tue
  1. Eugene YanAI score14

    Eugene Yan shares five principles for working with AI models

    AIEugene Yan outlines five principles for working effectively with AI models: treating context as infrastructure, taste as configuration, verification as the basis for autonomy, scaling through delegation, and closing the loop. The post is a short list of themes linked to a longer essay, and no further detail is given in the post itself.

  2. HyperdimensionalAI score47

    Hyperdimensional's Dean Ball Explains His Libertarian-Conservative Tension on AI Regulation

    AIWriter Dean Ball says he opposes nearly all proposed AI regulation, including algorithmic discrimination rules and pauses on development, while backing state management of catastrophic misuse risks. He frames the position as a tension between classical liberal and conservative instincts toward institutions and change.

May 4

May 4Mon
  1. HyperdimensionalAI score63

    Dean W. Ball argues against overreacting to Anthropic's Mythos cyber capabilities

    AIDean W. Ball argues that Anthropic's Mythos, which finds software vulnerabilities by chaining bugs into exploits, shifts the cost of vulnerability discovery and should not prompt an overreaction. He contends governments hold a uniquely mixed incentive over vulnerabilities, so heavy state control risks making software less secure. He proposes a narrow, testable government role focused on cyber-discovery risk thresholds, with private verification bodies supporting it.

May 3

May 3Sun

Apr 30

Apr 30Thu
  1. Andrej KarpathyAI score66

    Karpathy on agentic engineering, Software 3.0, and jagged AI capability

    AIAndrej Karpathy describes a December 2025 shift in which coding agents began producing larger, more reliable chunks of work, changing programming toward orchestrating agents. He argues that models automate what can be verified and that their capability is jagged, depending on verifiability and what labs emphasize in training, so users need to stay in the loop. He also says hiring, founder opportunities, and agent-native infrastructure should adapt to this shift.

Apr 29

Apr 29Wed

Apr 27

Apr 27Mon
  1. Soumith ChintalaAI score15

    Chintala Suggests Anthropic Account Support May Need Scaling Up

    AISoumith Chintala comments on a Reddit report that Anthropic banned organizations without warning, suggesting Anthropic may need to scale Account Support using Claude or human account managers. He also argues that enterprises may increasingly adopt multiple AI providers with open harnesses, facing cloud-era vendor problems that would likely affect all AI providers.

Apr 24

Apr 24Fri
  1. Ahmad Al-DahleAI score82

    Ahmad Al-Dahle says DeepSeek-V4's efficient 1M context is its key bet

    AIAhmad Al-Dahle argues that the most interesting part of DeepSeek-V4 is its bet on efficient ultra-long context rather than its benchmarks. He says this is the precondition for test-time scaling and long-horizon agents, and cites 27% of V3's FLOPs at 1M tokens. The quoted DeepSeek post announces DeepSeek-V4-Pro (1.6T total, 49B active) and DeepSeek-V4-Flash (284B total, 13B active), both open-sourced with 1M context and API access.

    Why it matters: The post argues that efficient 1M-token context, not benchmark scores, is the key bet behind DeepSeek-V4's design for test-time scaling and long-horizon agents.

Apr 22

Apr 22Wed
  1. Cognition Blog (Devin, Windsurf)AI score54

    Cognition says building cloud agents requires VM isolation, state snapshots, and org change

    AICognition argues that enterprises building cloud agents face three problems: shared container kernels, the inability to persist agent state across async gaps, and the scale of orchestration, governance, and integrations. The post says VM-level isolation with hypervisor-level snapshots was needed for Devin, and that organizations must also rebuild engineering processes around agent execution.

Apr 21

Apr 21Tue
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition says multi-agent systems work when only one agent writes

    AICognition reports that multi-agent setups work best when writes stay single-threaded and extra agents contribute intelligence instead of actions. It describes a code-review loop where a clean-context review agent catches bugs in Devin-written PRs, averaging 2 bugs per PR with roughly 58% severe. The post also says the smart-friend pattern, pairing a smaller primary model with a stronger one, has not yet worked well with asymmetrically weaker primaries and is an open training problem.

    Why it matters: The post gives concrete findings on which multi-agent setups work, including clean-context code review and smart-friend escalation, and where they still fail.

  2. Awni HannunAI score14

    Awni Hannun argues top tech firms must be extreme co-design companies

    AIAwni Hannun argues the biggest companies cannot be just hardware, software, or AI, and must instead practice extreme co-design, which he says is where the global minima lie. He cites Nvidia as an example and says Apple should be, and hopefully will be, an extreme co-design company despite many calling it a hardware company after its recent transition.