Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. CNBC · TechnologyAI score50

    US suspends Microsoft, Adobe from green card labor program amid foreign worker crackdown

    AIThe U.S. Department of Labor said it suspended Microsoft and Adobe from its Permanent Labor Certification program, citing multiple active federal investigations. Labor Secretary Keith Sonderling also said no new applications will be accepted for Cognizant, Infosys, Capgemini, Tata, Wipro and HCL. Microsoft said the vast majority of its U.S. employees are Americans and that 80% of its roughly 6,000 H-1B petitions last fiscal year were to extend or change the status of existing employees.

  2. Arena.aiAI score44

    Arena raises $200M Series B led by Lightspeed, launches Alignment Index

    AIArena has secured a $200M Series B, with Lightspeed doubling down on its investment. The company is also launching the Alignment Index, which measures how closely AI behavior aligns with human values in real-world settings. Arena reports annualized revenue above $100M since its Series A, with millions of people helping evaluate frontier models through real-world use.

  3. Tessl BlogAI score29

    One Brain Means Owning Your Organizational Memory

    AILeapfrog, a small team doing high-volume AI visual and production work for fashion and brand clients, is building a "one brain" system that makes company knowledge and client context searchable through natural-language agents. The starter stack described is OpenClaw in a sandbox, a GitHub repository, Obsidian on the local machine, and Telegram as the access point. The system's research structure had roughly 1,200 files at the time of the talk.

  4. Tessl BlogAI score42

    Agent Skills Should Be Treated as Supply Chain Components

    AITessl's talk at AI Native DevCon London argues that agent skills, which can be markdown files with instructions and bundled material, act as supply chain components that can shape agent behavior. The author says reading SKILL.md once is insufficient because risks can sit in supporting files, updates, and workspace trust settings. He identifies the danger as the combination of private context, untrusted content, and external communication, and cites research scanning roughly 4,000 public skills for issues including malware-like behavior.

  5. elvisAI score42

    Voyager: an open harness for creative AI work across video and games

    AIElvis Saravia argues that creative work needs domain-specific agent harnesses rather than coding-oriented ones, and he highlights Voyager as an open harness for video, graphics, and games. According to the quoted post, Voyager lets agents work with local files and drive apps such as Blender, DaVinci Resolve, and Unity, and it is designed to work with models like Opus, Astra, and DeepSeek.

    Video from @omarsar0's post
  6. Tessl BlogAI score52

    Cisco engineer argues agent skills need a context pipeline with evals

    AIJohn Groetzinger, writing in a personal capacity rather than for Cisco, argues that enterprise skills need packaging, evaluation, syncing, and distribution rather than scattered markdown files. He describes using skills to make cheaper models viable, converting curated TAC knowledge-base articles into maintained skills, and rolling out an eval framework across teams. He also describes syncing a repository README to Confluence with a deterministic script.

  7. Artificial AnalysisAI score42

    More output tokens don't guarantee higher scores in AI benchmarks

    AIArtificial Analysis reports that generating more output tokens does not necessarily yield a higher score. GPT-6 Astra (max) scored 8.6% using about 81k output tokens per task, while Grok 4.7 (xhigh) used roughly 180k yet scored lower. Three Claude models produced the most output tokens, about 202k to 562k per task, but scored between 2.8% and 6.4%.

    Image from @ArtificialAnlys's post
  8. Artificial AnalysisAI score28

    Artificial Analysis Pareto frontier: GPT-6 Luna cheapest per task at $0.22

    AIAmong models with a Hallucination-Gated All-Pass Rate above 0%, GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max), and Grok 4.7 (xhigh) set the Pareto frontier for score versus cost per task. GPT-6 Luna (max) is the cheapest at about $0.22 per task, scoring 3.3%, while Grok 4.7 (xhigh) leads at about $9.50 per task and Muse Spark 1.3 (max) costs about $4.20. The three Claude models cost about $18 to $22 per task.

    Image from @ArtificialAnlys's post
  9. Artificial AnalysisAI score34

    Artificial Analysis compares six hallucination checkers on 20 shared tasks

    AIArtificial Analysis compared six hallucination checkers on the same deliverables from 20 tasks across eight models. GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet. The counts reflect checker behavior rather than establishing accuracy or ruling out self-preference.

    Image from @ArtificialAnlys's post
  10. AWS Machine Learning BlogAI score46

    AWS Pays Per Inference for AI Agents with BlockRun and Incarna

    AIAmazon Bedrock AgentCore payments lets AI agents pay for model inference one request at a time, using x402 with USDC on the Base network. Incarna used the service to connect its agents to BlockRun, a pay-as-you-go router serving more than 90 models from more than 15 providers. Spending limits are enforced at the infrastructure layer, outside the model.

  11. 🚨 AI News | TestingCatalogAI score49

    Voyager desktop app lets AI agents work inside creative tools on Mac

    AIVoyager has launched a Mac desktop app that lets AI agents read project files and operate creative tools such as After Effects, DaVinci Resolve, Blender, and Unity. The agents produce editable results for video edits, motion graphics, color grading, 3D scenes, and game prototypes. Built-in and custom skills, plus a memory that learns each user's workflow, are included.

    Video from @testingcatalog's post
  12. OpenAI DevelopersAI score47

    OpenAI expands GPT-6.1 Sol Ultrafast access and EU data residency

    AIOpenAI has made Ultrafast mode for GPT-6.1 Sol available in all supported regions, including US and EU data residency. EU data residency has also been added for GPT-6.1 Sol Fast and GPT-6 Luna Fast. Access to Codex and ChatGPT Work is offered on Pro 500, eligible usage-based Enterprise, and credit-based Edu plans, with Enterprise admins required to enable it.

  13. OpenAI DevelopersAI score62

    OpenAI rolls out Ultrafast for GPT-6.1 Sol in API, Codex, and ChatGPT Work

    AIOpenAI says Ultrafast is rolling out today for GPT-6.1 Sol in the API, Codex, and ChatGPT Work. The company describes it as near-Astra intelligence at up to 8x faster speeds than Sol Standard.

    Why it matters: The post names the access points and a speed comparison to the Sol Standard tier, which helps developers judge whether the faster option fits their workflow.

    Video from @OpenAIDevs's post
  14. DatabricksAI score32

    Databricks' Vibe Data Modeling builds business-specific data models with an agent

    AIDatabricks introduced Vibe Data Modeling, an open-source agent that helps teams build, validate, and evolve business-specific data models. It applies roughly 250 modeling rules while keeping data modelers and business stakeholders involved. Teams can start from 40 industry models as a baseline and iterate toward models that reflect how their business operates.

    Video from @databricks's post
  15. laurenAI score29

    Omarchy seeks feedback on Grok Bot plugins and integrations

    AILauren Tan invites users of Grok Bot on Omarchy and developers building plugins for it to share feedback and feature requests. The post points to the Omarchy plugin catalog and asks what integrations could be supported. Background from DHH says SpaceXAI joined the Omacom Foundation as a Founding Corporate Patron, contributing $1,500,000 in Grok tokens for Omarchy's maintenance and development.

  16. TechCrunch · AIAI score36

    Ben Affleck's AI expertise goes viral as he explains neural networks and fine-tuning

    AIActor Ben Affleck drew attention this week for explaining machine learning concepts, including convolutional neural networks, tensors, and transformers, in several recent interviews. He said he fine-tuned open video models by unfreezing weights and training only the last cinematic layer, using a dataset he built over about eight months for his startup. Affleck said he worries about students and learned helplessness more than Skynet, and predicted AI will be additive to the movie business.

  17. TechCrunch · AIAI score46

    Arena raises $200M at $3.1B valuation, nearly doubling in 10 months

    AIArena, the crowdsourced AI model leaderboard that started as a UC Berkeley research project, raised a $200 million Series B at a $3.1 billion valuation, led by Lightspeed Venture Partners and Khosla Ventures. The company said it reached $100 million in annualized run-rate revenue in June, up from $30 million when it raised its $150 million Series A in January at a $1.7 billion post-money valuation.

  18. TechCrunch · AIAI score58

    OpenAI's annualized revenue reportedly about $20 billion below earlier estimates

    AIOpenAI has reportedly told investors its annualized revenue is approaching $50 billion, about $20 billion below a previously reported $70 billion figure. The Financial Times reports the earlier number came from investor attempts to compare OpenAI with Anthropic, which counts cloud partners' sales differently. OpenAI's IPO has reportedly been pushed to early 2027.

  19. TechCrunch · AIAI score72

    Google launches unified Gemini agent for businesses, consumers to follow

    AIGoogle announced at a Google Cloud event a unified Gemini agent that can plan and complete tasks from a single interface, starting with businesses. The agent has its own Workspace account, connects to systems including Google Workspace, Microsoft 365, Slack, and Jira through MCP, and writes an audit trail attributed to the agent. Google said consumers will get access later, after it addresses security, scale, and performance.

    Why it matters: The source details how the agent takes objectives, connects to business systems, and logs actions, showing how enterprise agent deployment is being structured.

  20. The DecoderAI score80

    Mathematicians call for OpenAI boycott after AI-generated proofs flood the field

    AIThe Association of Historical Mathematicians (AHM) has called for a boycott of OpenAI after the company released more than 700 AI-generated proof files at once. Fields Medalist Terence Tao, who chairs the group, argues that AI solving open problems autonomously reduces seminars, collaborations, and fertile research directions, and that the field should shift its measure of progress toward explanation and community-building.

    Why it matters: The article links the AHM boycott call to Tao's argument that AI-driven proof volume is changing how mathematicians measure progress and whether solutions remain useful.

  21. CNBC · TechnologyAI score38

    Trump's August Disclosure Shows Up to $25 Million in Meta and $5 Million in SpaceX Debt

    AITrump disclosed more than 500 securities transactions in August, including a purchase of up to $25 million in Meta stock and up to $5 million in SpaceX senior unsecured notes. The filing, which reports trades in value ranges, shows total activity of roughly $74.3 million to $273.3 million according to a CNBC analysis. The SpaceX notes were bought two days before Trump signed a national space transportation policy, and the White House says the portfolio is independently managed.

  22. PyTorch BlogAI score62

    NVIDIA Dynamo adds session-level IDs to route and cache agentic inference

    AINVIDIA Dynamo uses a unified session-level identifier to make its inference stack aware of agent sessions, subagents, and their KV cache across turns and tool calls. On SWE-bench, two TP4 MiniMax-M2 replicas on one 8xH100 node gained roughly 12-16% throughput from program-aware scheduling over KV-aware routing alone. The post also describes experimental shared-pool indexing and a proposed KvHint interface for session-aware cache policies in vLLM and SGLang.

    Why it matters: The post explains how session identifiers let an inference stack track agent working sets, with measured throughput gains on SWE-bench and agentic RL rollouts.

  23. KalaAI score34

    Mistral Large 4 and Reflection Beam promise open weights this month

    AIMistral Large 4 and Reflection Beam are previewed now, with Mistral saying weights drop at the end of October and Reflection promising Apache 2.0 weights this month. The post argues that these announced future weights should be treated as a conditional migration dependency, not a current self-hosting option. API previews can be trialed immediately, but they do not prove an unreleased checkpoint will behave the same when downloaded.

  24. TechCrunch · AIAI score65

    OpenAI's math solutions fall short of the field's standards, mathematicians say

    AIOpenAI released hundreds of claimed solutions to hard math problems but did not fully meet guidelines from the Advisory Group on Mathematics and Artificial Intelligence. Only 10 of 719 manuscripts included chain-of-thought releases, and just 42% of proofs were formalized. A Cambridge and King's College paper found discrepancies between a natural language proof and its Lean code for a Navier-Stokes-derived problem.