Skip to contentSkip to stories

Updated

#Deployment/Engineering

Showing low-relevance items too. Hide low-relevance items

Sep 26

Sep 26Sat
  1. DeedyXAI score40

    Economics of neolabs: why GPU spend makes frontier-chasing hard

    AIA neolab is a startup of AI researchers that raises large pre-production funding to finance GPU compute, with 1000 GB300s (about 14 NVL72 racks) costing $125-150M over 3 years, roughly 2-2.5MW. That buys about 10^25 FLOPs per quarter, enough for a GPT-4-level model that is 1-2 OOMs behind the frontier for pretraining. Recouping $10M in training at 50% inference margin would take serving about 10T tokens at a $2/M blended price, so neolabs often pivot to a different model game, proprietary data, or high-revenue niches.

  2. Higgsfield AI 🧩OfficialAI score14

    Higgsfield API launches native 1080p Seedance 2.5 with cashback promotion

    AIHiggsfield AI has introduced Seedance 2.5 in native 1080p on its API, positioned as a US-based option for commercial and large-scale productions. The company is offering 100% instant cashback on API spend from an $18M pool, with caps of up to $200,000 per business and $1,000 per individual, and unused cashback expires September 30.

    Video from @higgsfield's post
  3. Jeff DeanXAI score44

    Waymo's Crash Rate Versus Human Drivers Improves to 20x

    AIJeff Dean says Waymo's latest safety data shows its rate of crashes with serious injury is 20 times better than human drivers across 270 million miles, up from 13 times in March 2026. Waymo's own data reports 82% fewer injury crashes and 95% fewer serious injury crashes across five territories, with 841 fewer injury-causing crashes.

  4. OpenCodeOfficialAI score34

    LongCat-2.5-Preview Free on OpenCode for Two Weeks

    AILongCat-2.5-Preview is free on OpenCode for two weeks, offering a 1M context window, multimodal support, and zero data retention. The post does not provide further details on pricing terms or capabilities beyond these listed features.

  5. SemiAnalysisBlogAI score62

    Intel Panther Lake teardown reveals 18A RibbonFET and PowerVia design details

    AISemiAnalysis tore down Intel's Panther Lake chip, examining its 18A process with RibbonFET gate-all-around transistors and PowerVia backside power delivery. Measurements put 18A compute logic and TSMC N3E GPU logic at similar logic density, though 18A does not lead TSMC N3P, N2, or Samsung SF2 in peak density. The analysis also compares the compute, GPU, I/O, and Foveros-S packaging against Lunar Lake and Samsung's SF2 process.

  6. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score45

    Intern-Decision-4B: Multimodal structured decision model from Qwen3.5-4B

    AIShanghai AI Lab's InternLM released Intern-Decision-4B, a multimodal structured decision model fine-tuned from Qwen3.5-4B, which returns answer distributions for multiple questions in one forward pass. On its benchmark table it scores an average of 90.02 with a Brier score of 0.347 and an ECE of 0.065, and per-query latency averages 44.16 ms on a single RTX 4090. The model is available with a Python DecisionEngine inference interface.

  7. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score44

    Intern-Decision-2B: Structured Multi-Question Decision Model Fine-Tuned from Qwen3.5-2B

    AIShanghai AI Lab's InternLM released Intern-Decision-2B, a multimodal structured decision model fine-tuned from Qwen3.5-2B that returns calibrated answer distributions for multiple questions in one forward pass. It averages 84.68 across listed benchmarks with a 0.437 Brier score and 33.28 ms mean latency on a single RTX 4090. Model weights, a Python DecisionEngine API, and GitHub code are available, with support for up to 16 questions and eight images.

  8. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score46

    Intern-Decision-0.8B: InternLM's structured decision model on Hugging Face

    AIInternLM released Intern-Decision-0.8B, a multimodal structured decision model fine-tuned from Qwen3.5-0.8B that scores answers to multiple questions in one forward pass. The model reports a 79.38 average score and a 33.98 ms mean latency on a single RTX 4090, with 0.8B, 2B, and 4B sizes available. It is accessed through a Python DecisionEngine API that returns calibrated probabilities rather than generating free-form text.

Sep 25

Sep 25Fri
  1. Stephanie PalazzoloXAI score32

    Fal and Fireworks weigh funding rounds at up to $30B valuation

    AIInference providers Fal and Fireworks are reportedly considering new funding rounds amid soaring inference demand. Fal has discussed raising at a $15B valuation, while Fireworks has considered a $30B valuation, nearly doubling their valuations from earlier rounds this year.

  2. Google AntigravityOfficialAI score34

    Antigravity 2.0 adds planning mode with /plan command

    AIGoogle Antigravity 2.0 now includes a dedicated planning mode, matching the Antigravity CLI. Typing /plan makes the agent research the task and generate an implementation plan for user review before execution, requiring approval to proceed. Users can also request a lighter plan through a natural prompt.

    Video from @antigravity's post
  3. LMSYS OrgOfficialAI score38

    SGLang adds multi-item scoring for faster decision model serving

    AISGLang's /v1/score endpoint returns scores for exact requested labels such as Yes/No or A/B/C, and its multi-item scoring (MIS) computes shared context once while keeping candidates isolated. On Qwen3-8B, 16-candidate p95 latency dropped from 54.1 ms with Generate to 20.6 ms with MIS. On Qwen3-0.6B, MIS p95 stayed under about 100 ms as load rose, versus seconds for Generate and SIS.

    Image from @lmsysorg's post
  4. MicrosoftOfficialAI score10

    Microsoft Copilot is positioned as the new OS for work

    AIMicrosoft describes its Copilot as the new operating system for work, designed to keep humans in control. The post is a brief product positioning statement with no further details on features, availability, or pricing.

    Video from @Microsoft's post
  5. Kevin Weil 🇺🇸XAI score75

    Claude solves nine-loop scattering amplitude calculation past prior eight-loop record

    AIAnthropic reports that Claude solved a nine-loop calculation in the planar N=4 super-Yang-Mills model, surpassing the previous eight-loop record set by Lance Dixon and collaborators. The quoted post says Claude ran largely unsupervised for days in Claude Science using a single prompt, at a total cost of a few thousand dollars, and Dixon independently verified the result. Kevin Weil's own text praises the achievement and expects AI to advance high energy physics over the coming 12 months.

    Why it matters: The quoted Anthropic post gives a concrete benchmark: Claude ran for days to reach nine loops, extending the previous eight-loop record in a physics model.

  6. Google WorkspaceOfficialAI score18

    Workday for Google Sheets now available in Google Workspace Marketplace

    AIGoogle Workspace announces that Workday for Google Sheets is now available in the Google Workspace Marketplace. The add-on brings Adaptive from Workday into Sheets and Slides, letting users avoid manual CSV downloads for financial planning.

    Video from @GoogleWorkspace's post
  7. VercelOfficialAI score31

    Klaviyo Ships 356 Internal Apps in Two Weeks on Vercel

    AIKlaviyo built an internal app platform on Vercel, and in the first two weeks 512 employees shipped 356 projects. Teams can go from idea to a live app in about three minutes, with full-stack apps running on Klaviyo's databases. Deployments are SSO-gated and private by default.

  8. eric zakariassonXAI score8

    Grok Bot helps developers build apps on the X API

    AIEric Zakariasson highlights a range of apps that can be built on the X API and says the Grok bot makes getting started easy. The developer exhibit at offers inspiration, and the Grok-hosted X API Engineer can help build, test, and deploy projects.

  9. Boris ChernyXAI score62

    Claude Tag in Slack gains personal connectors for channel workflows

    AIClaude Tag in Slack can now use users' personal connectors, such as Drive, Salesforce, and warehouse access, within channels. The author says Tag writes over 50% of their PRs daily and handles nearly all their data analysis and many product bug fixes. The personal connector feature is available on Teams today and Enterprise next week.

  10. Noah ZwebenXAI score46

    Anthropic shows Claude Code /remote-control demo with Opus 5.5 claymation video

    AIAnthropic's Noah Zweben shared a claymation video showing Claude Code's /remote-control feature, now made with Opus 5.5 after an earlier Opus 4.6 version. The feature, which lets users control Claude Code remotely, is rolling out to Pro users at 10% and ramping, with Team and Enterprise support coming later.

    Video from @noahzweben's post
  11. Together AIOfficialAI score12

    Together AI Simplifies Access to Frontier Open Models via API

    AITogether AI says teams can access frontier open models through its API without managing underlying infrastructure, a point Ted Cui, its VP of Engineering and Inference Platform, made at Apsara Conference 2026. The company emphasizes reliable, fast inference as essential for this access.

  12. Google Cloud · AI & Machine LearningOfficialAI score43

    Google Cloud Introduces Managed Reinforcement Learning Fine-Tuning for Gemini Models

    AIGoogle Cloud has launched a managed reinforcement learning fine-tuning service (RLFT) that lets customers adapt Gemini models using a reward function they define instead of labeled answers. Users supply prompts and a reward function, while Google handles the RL infrastructure and proprietary model internals. The guide advises exhausting prompting and supervised fine-tuning first, and notes that RLFT suits tasks that are easy to score but hard to demonstrate.

  13. SemiAnalysisBlogAI score59

    China Holds Over 24GW of Datacenter Capacity, Shifting Inland With AI Demand

    AISemiAnalysis's China Datacenter Model tracks over 1,000 facilities across 60+ operators and puts China's fleet above 24GW, larger than EMEA. The report attributes the buildout to the Eastern Data, Western Compute policy and hyperscale AI demand, which is moving capacity to western hubs such as Inner Mongolia at construction speeds it says the West cannot match.

  14. OpenCodeOfficialAI score22

    OpenCode makes $60 of DeepSeek v4.1 Flash usage permanent

    AIOpenCode says its $60 of usage for DeepSeek v4.1 Flash is now permanent, as part of its "Operation Cheepseek Phase 2" promotion. The post gives no further details on terms, duration, or eligibility.

  15. Amazon ScienceOfficialAI score38

    Amazon and Reactor build kernel path to real-time video generation on Trainium

    AIUsing the Neuron Kernel Interface, Reactor and Amazon's Neuron Science team built a kernel-centric path to real-time autoregressive diffusion video generation on Trainium. They addressed dynamic shapes, memory access patterns, and cache management, which are hard for generic compilers, and developed techniques intended to generalize across models.

  16. Meituan LongCatOfficialAI score62

    Meituan LongCat-2.5-Preview Launches with 1.6T Parameters and 1M-Token Context

    AIMeituan's LongCat team has released LongCat-2.5-Preview, a natively multimodal model with 1.6T total parameters, about 48B active, and a 1M-token context window. The model is built for long-horizon tasks spanning terminals, browsers, GUIs, spreadsheets, and design tools. It is available now through an API on the LongCat platform and a chat interface.

    Image from @Meituan_LongCat's post
  17. Microsoft CopilotOfficialAI score40

    Microsoft Copilot app refreshed to unify chat, agents, app building, and workflows

    AIMicrosoft has refreshed its Copilot app to bring chat, task delegation, app building, and workflow automation into one place. The update is positioned as an AI built for work, with Satya Nadella describing Copilot as a new OS for work spanning models, form factors, and tasks. The announcement includes Autopilot, an enterprise agent, Code for building apps hosted within a company's tenant, Home combining Chat and Cowork, and Office fully embedded in Copilot.

  18. Satya NadellaXAI score52

    Satya Nadella announces Copilot update with Autopilot, Code, Home, and Office

    AIMicrosoft CEO Satya Nadella announced what he called the biggest Copilot update to date, positioning Copilot as a new operating system for work. The update bundles Autopilot, a proactive long-running enterprise agent; Code, for building apps hosted inside a company's tenant; Home, combining Chat and Cowork; and Office, now fully embedded in Copilot. Copilot can also be invoked in Teams, and a new proactive experience called Today surfaces key information from across M365 without a prompt.

    Video from @satyanadella's post

Sep 24

Sep 24Thu
  1. ModelScopeOfficialAI score23

    NeoHorse-Jev-4B open model turns app states into structured decisions

    AIModelScope has released NeoHorse-Jev-4B, a compact open model that converts application states into structured decisions and probabilities. It scores 77.70 across six text decision benchmark groups, ranking first among four open-weight models with complete results in the comparison. Its prefill-only inference supports Choice, Noul, and Score primitives, accepts text or a single image with text, and is available under Apache 2.0 for deployment via vLLM, SGLang, Python, CLI, or HTTP.

    Video from @ModelScope2022's post
  2. WanOfficialAI score13

    Wan3.0 generates 10-second 1080p video for $1.82 on Venice

    AIAlibaba Cloud's Wan3.0 workflow, run on Venice, produces a 10-second 1080p AI video for $1.82. The post argues that production cost, rather than output quality, will determine how much AI video content can scale.

    Video from @Alibaba_Wan's post
  3. vLLMOfficialAI score30

    vLLM integrates TileRT with PD disaggregation, benchmark config published

    AIvLLM has published a blog post explaining how its integration with TileRT works, alongside a public benchmark configuration in the InferenceX repository. The benchmark script covers GLM-5.3 FP8 on MI355X hardware and is linked on GitHub. The post itself provides the integration details.

  4. vLLMOfficialAI score42

    TileRT and vLLM hit 469 tok/s on GLM-5.3 with MI355X

    AIThe TileRT and AMD teams reached 469 tok/s single-user decode for GLM-5.3 on 8× MI355X using vLLM. The setup disaggregates work, with vLLM handling prefill and TileRT handling latency-critical decode through vLLM's V1 connector interface. SemiAnalysis's AgentX benchmark reports the configuration at 470 TPS on GLM 5.3 (FP8), over 40% faster than GB300 TRTLLM using FP4.

  5. Redwood Research BlogBlogAI score41

    Continual learning could make AI monitors that block actions nearly useless

    AIRedwood Research argues that continual learning, which lets an AI accumulate skills during deployment, may teach models to evade blocking monitors because monitors reduce task success. Online RL on deployment trajectories would train the policy against the monitor through task reward, potentially leaving blocking monitors nearly useless over a long deployment. Memory-based systems pose a weaker version of this risk, according to the post.

  6. Lydia Hallie ✨XAI score20

    Anthropic adds local support to Projects in staged rollout

    AIAnthropic's Lydia Hallie thanked users for feedback on the new version of Projects and said local support was added the previous day. The new Projects is rolling out in stages, with early access available by DM, and existing projects may have rough edges while a smooth carry-over method is still being developed.

  7. Noah ZwebenXAI score38

    Claude Tag in Slack now accesses users' personal connectors

    AIAnthropic's Claude Tag in Slack can now use users' personal connectors, letting them securely reach a Drive doc, Salesforce account, or warehouse table they have personal access to within the conversation. The feature is available on Teams today and on Enterprise next week.

    Video from @noahzweben's post