Working with this model has been a wild ride.
AIWe’ve come a long way on safety, but we still expect the next capability jump of this scale to be a huge challenge.
Updated
Updated
AIWe’ve come a long way on safety, but we still expect the next capability jump of this scale to be a huge challenge.
AIWe spend much of the 244-page system card and the 60-page risk assessment supplement trying to lay that out.
AIIf we are able to collectively rise to the challenge and confront this risk, it could serve as a blueprint for addressing the even more difficult challenges that lie ahead of us.
AI…internet and world than we had before the advent of AI-powered cyber capabilities.
AIDespite all the resources at my disposal and an incredible support system, I still couldn’t avoid a medical leave. We need better solutions for all patients, now.
AIAndrej Karpathy highlights Farzapedia, a personal wiki Farza built from 2,500 diary entries, Apple Notes, and iMessage conversations, as a good example of his proposed LLM-wiki approach. He argues this file-based memory is explicit, user-owned, interoperable, and usable with any AI tool, letting users control how AI knows them.
AIAndrej Karpathy argues that AI can reverse the historical flow of legibility, letting citizens process dense government data that was once accessible only to specialists. He cites examples such as omnibus bills, budgets, lobbying disclosures, and local council decisions, while acknowledging the same tools could be misused.
AIIt feels short sighted. - Many (if not most) of the best ideas come from junior people with fresh eyes or senior people cross-pollinating across domains - Thanks to AI you can learn much faster
AIAI Futures Project moved Daniel Kokotajlo's Automated Coder median from late 2029 to mid 2028 and Eli's from early 2032 to mid 2030. The main reasons cited are a faster METR time horizon doubling time and the impressive results of Claude Opus 4.6. The authors also say progress in agentic coding has been faster than expected over the past 3 to 5 months.
AIThe best teams build systems that make the right thing easy and the wrong thing hard. The best cultures treat every incident as a systems design question, not a blame assignment
AI…weekend - constantly monitor it for updates - add whatever features you want on top of it all without you ever needing to use your computer
AIAndrej Karpathy reports that an LLM spent four hours strengthening his blog post's argument, then convinced him of the opposite when asked to argue the reverse. He concludes that LLMs are highly capable of arguing almost any direction, which makes them useful for forming opinions if users ask from multiple angles and watch for sycophancy.
AIsomething cool comes out - nobody does anything with it - anthropic finally goes “okay fine we’ll do it” - thing becomes mega popular people fading mcp was remarkably stupid, and you should definitely start building mcp apps.
AIAndrej Karpathy argues that the hardest part of building a deployed app is not the code but the DevOps work of assembling services, API keys, payments, auth, and deployment. He says the goal is for agents to handle this entire lifecycle as code, with agent-native CLI and API access instead of manual web clicking. He calls it a from-scratch redesign that is only now barely technically possible.
AIHamel Husain argues data scientists remain essential as foundation-model APIs let teams ship AI without them, because much of the work lies in evaluation, debugging, and metric design. He says teams often rely on generic off-the-shelf metrics and unverified LLM judges instead of examining their own data. He lists five eval pitfalls, starting with generic metrics, and recommends looking at traces and doing error analysis.
AIA single question from 2 months ago about some topic can keep coming up as some kind of a deep interest of mine with undue mentions in perpetuity. Some kind of trying too hard.
AIJim Fan reposted a report that LiteLLM PyPI release 1.82.8 was compromised and contained a litellm_init.pth file that sends credentials to a remote server and self-replicates. He argues agents make this worse, since files like skills, configs, or PDFs read into context could spread malicious instructions. He concludes that agentic frameworks need guardrails and audited tooling.
AI…cloning directly from humans is the way to break the curse of teleop. 2026 is all about scaling robot learning without robots.
AIHex CEO Barry McCardel posted a graph showing AI agents now create more Hex cells than humans do. Mintlify launched an analytics feature in February to track AI agent traffic to documentation, saying agents may read docs more often than humans. The article argues that writers should consider AI systems as a primary audience.
AICursor says its Composer 2 anniversary marks one year of large model training, with a team of about 40 researchers and engineers now dedicated entirely to software engineering. The company states that every FLOP, token, parameter, and researcher is focused on coding rather than general assistant tasks. Composer 2 is now available in Cursor.
AII'm just going to ask you to do it and it'll only take an hour, tops.
AIEspecially this: “instead of starting with easy sudoku puzzles data and gradually ramping up toward harder ones, the training runs that start with harder puzzles and end with easy puzzles always fare better, and also fare better than mixed difficulty sampling”
AIThe author shares a quoted legal analysis arguing that OpenAI's published Pentagon contract excerpt essentially only permits all lawful use. The analyst notes the excerpt is short, that DoD Directive 3000.09 and other DoD directives referenced in it can be changed by the Department at any time, and that the contract may not guarantee what OpenAI's FAQ implies.
AIA few that I've noticed (no doubt missing many)...
AIIn the future, there will be much more challenging situations of this nature, and it will be critical for the relevant leaders to rise up to the occasion, for fierce competitors to put their differences aside. Good to see that happen today.
AIai allows us to raise our ambitions. but as the technology advances, as a country we should never lower our standards. let freedom ring. 🇺🇸 🇺🇸 🇺🇸
AII don't think they should be. But the profession is changing faster than most people realize. Wrote about what's actually shifting, and what matters next.
AISo much has changed and yet there's still so much to go. Cloud agents and demo-based review are a big leap toward this future.
AIIt's agents proving their work. For products this is demo videos. And we're slowly discovering what this looks like for the rest of software engineering!
AIIt's pretty consistent with what we've seen from interpretability thus far. And it's comparatively actionable in terms of what it suggests for safety.