Skip to contentSkip to stories

Updated

#Deployment/Engineering

Items with an AI score under 20 are hidden. Show low-relevance items

Mar 17

Mar 17Tue
  1. Xiaomi MiMoAI score80

    Xiaomi MiMo-V2-Pro Flagship Model Targets Agent Workloads With 1M Context

    AIXiaomi announced MiMo-V2-Pro, a flagship foundation model for agent workloads with over 1T total parameters, 42B active, and up to 1M-token context. It ranks 8th worldwide and 2nd among Chinese LLMs on the Artificial Analysis Intelligence Index, and its API is publicly available with usage-tiered pricing.

    Why it matters: The post gives benchmark placements, parameter scale, context length, and tiered API pricing, so readers can compare it against Claude and GPT models on concrete terms.

Mar 16

Mar 16Mon

Mar 11

Mar 11Wed
  1. Mistral AI · new models on Hugging FaceAI score62

    Mistral AI releases Leanstral-2603, an open-source Lean 4 proof agent

    AIMistral AI released Leanstral 119B A6B on Hugging Face as an open-source code agent for Lean 4 proof engineering. The model uses 128 experts with 4 active per token, 6.5B activated parameters, a 256k token context window, and accepts text and image input under the Apache 2.0 license. The page also documents vLLM server deployment and Mistral Vibe integration.

    Why it matters: The source specifies Leanstral's 119B MoE architecture, 256k context, Apache 2.0 license, and vLLM setup, showing how the Lean 4 proof agent could be deployed locally.

Mar 10

Mar 10Tue

Mar 9

Mar 9Mon
  1. Black Forest Labs · new models on Hugging FaceAI score39

    Black Forest Labs releases FLUX.2 [klein] 9B-KV with KV-cache for faster multi-reference editing

    AIBlack Forest Labs has released FLUX.2 [klein] 9B-KV, a variant of FLUX.2 [klein] 9B that caches reference-image key-value pairs to speed up multi-reference editing by up to 2.5 times. The 9B flow model, which uses an 8B Qwen3 text embedder and is step-distilled to 4 inference steps, is available for non-commercial use under the FLUX Non-Commercial License and fits in about 29GB VRAM.

Mar 5

Mar 5Thu
  1. Tri DaoAI score62

    FlashAttention-4 paper: attention on Blackwell GPUs nears matmul speed

    AIThe FlashAttention-4 paper is out, reporting that attention on Blackwell GPUs now runs at roughly matmul speed, reaching about 1600 TFLOPs. The forward pass is bottlenecked by exponential computation and the backward pass by shared memory bandwidth, and the redesign uses polynomial exponential emulation, a new online softmax that avoids 90% of softmax rescaling, and 2CTA MMA instructions that let two thread blocks share operands to cut shared memory traffic.

Mar 4

Mar 4Wed
  1. Tri DaoAI score62

    Tri Dao Shares Speculative Speculative Decoding, a Claimed Up-to-2x LLM Inference Speedup

    AITri Dao reposts a quoted post from @tanishqkumar07 introducing Speculative Speculative Decoding (SSD), an LLM inference algorithm claimed to be up to 2x faster than leading inference engines. The quoted post credits collaborators @tri_dao and @avnermay and links to a thread with details. Tri Dao's own text says the approach applies an asynchronous-machines principle seen in GPU kernels to speculative decoding.

  2. Mistral AI · new models on Hugging FaceAI score67

    Mistral Small 4 unifies instruct, reasoning, and coding in one open model

    AIMistral Small 4 combines instruct, reasoning, and Devstral capabilities in one multimodal model with 119B total parameters, 6.5B active per token, and a 256k context window. The source reports a 40% reduction in latency-optimized end-to-end completion time and 3x more requests per second in throughput-optimized setups versus Mistral Small 3. It is released under Apache 2.0 and supports reasoning mode toggling per request.

    Why it matters: The source lists architecture, context length, and mode-switching controls, letting readers compare this release's design with earlier Mistral Small models.

Mar 3

Mar 3Tue

Feb 28

Feb 28Sat
  1. Cognition Blog (Devin, Windsurf)AI score36

    Cognition Previews SWE-1.6, Claims 11% Gain Over SWE-1.5 on SWE-Bench Pro

    AICognition previewed its ongoing SWE-1.6 training run, which scores 11% higher than SWE-1.5 on SWE-Bench Pro and runs at 950 tok/s. The model is post-trained on the same pre-trained model as SWE-1.5, and the company is rolling out early access to a small group of users to gather feedback on behavior such as overthinking and excessive self-verification. The company says training steps now run 6x faster than three months ago, with rollouts in NVFP4 precision.

Feb 26

Feb 26Thu
  1. Cognition Blog (Devin, Windsurf)AI score67

    How Cognition Uses Devin to Build Devin Across Slack, Linear, and Code Review

    AICognition reports merging 659 Devin PRs into its own codebase last week, up from 154 in its best week in 2025. The post describes internal workflows across web, Slack, Linear, CLI, and API, including Devin Review for PR diffs and bug catching, a daily design system audit, automated bug triage on Linear, and DANA for data analysis.

    Why it matters: The post shows concrete workflows for using Devin across Slack, Linear, and code review, with specific usage figures that help teams judge fit for their own engineering processes.

Feb 25

Feb 25Wed

Feb 24

Feb 24Tue
  1. Cognition Blog (Devin, Windsurf)AI score46

    Cognition Launches Cognition for Government to Modernize Federal Software With Devin and Windsurf

    AICognition launched Cognition for Government on February 25, 2026, offering its Devin autonomous software engineering agent and Windsurf AI IDE to modernize U.S. government legacy systems. Devin, available in AWS GovCloud with a FedRAMP High version forthcoming, can complete migrations 5-40x faster than human engineers, while Windsurf is the only FedRAMP High AI IDE and holds DoD IL4/5/6 accreditation.

  2. Replit BlogAI score43

    Replit Pro launches at $100/month as Core drops to $20/month

    AIReplit launched a $100/month Pro plan with Turbo Mode, pooled credits for up to 15 builders, and priority support, while cutting Core from $25 to $20 per month and letting it invite up to 5 collaborators. The Teams plan is being sunset, with Teams users automatically upgraded to Pro at no additional cost for the rest of their term. Economy and Power Modes for Agent are available on all paid plans.

  3. Jim FanAI score62

    NVIDIA's SONIC trains a 42M transformer to control a humanoid robot

    AINVIDIA researchers trained SONIC, a 42M-parameter transformer, to control a humanoid robot's whole body using motion tracking on over 100M mocap frames. After three days of training in simulation, the policy transferred zero-shot to the real G1 robot and reported a 100% success rate across 50 real-world motion sequences. One policy supports VR teleoperation, webcam human video, text prompts, music, and GR00T N1.5 VLA integration with 95% success on mobile tasks, and the code and checkpoints are open-sourced.

Feb 23

Feb 23Mon
  1. Cognition Blog (Devin, Windsurf)AI score46

    Devin 2.2 adds desktop testing, self-review autofix, and 3x faster startup

    AICognition released Devin 2.2, which gives Devin full access to its own Linux desktop so it can launch and test desktop applications, not just browser-based web apps. Devin can also plan, code, review its own output, and fix issues before opening a PR, and it now starts up 3x faster. New users get $10 in free credits, and Desktop support is enabled by default for new sessions as of February 24, 2026.

Feb 22

Feb 22Sun
  1. Artificial IgnoranceAI score62

    Harness engineering emerges as a playbook for managing coding agents

    AIThe article argues that engineers are splitting their work between building a harness of constraints, tools, and documentation for agents and directing the agents' work. It cites OpenAI, Stripe, and Anthropic examples, including architecture guardrails, custom linter messages, AGENTS.md updates, and plan-first execution. The author notes that open problems remain around code maintainability, verification at scale, and adopting these practices in older codebases.

Feb 20

Feb 20Fri
  1. Jim FanAI score75

    DreamDojo: Open-source world model trained on 44K hours of human video

    AIJim Fan announced DreamDojo, an open-source interactive world model that takes robot motor controls and generates future frames in pixels. It is pre-trained on 44K hours of human egocentric video using latent actions, then post-trained onto specific robot hardware, and a real-time version runs at 10 FPS for live teleoperation, policy evaluation, and model-based planning. The author reports a +17% real-world success gain on a fruit packing task, and weights, code, datasets, and the whitepaper are released.

Feb 19

Feb 19Thu
  1. Yi TayAI score78

    Google releases Gemini 3.1 Pro, reporting 77.1% on ARC-AGI-2

    AIGoogle has released Gemini 3.1 Pro, reporting 77.1% on ARC-AGI-2 and more than twice the score of Gemini 3 Pro on that benchmark. The model is rolling out to developers in preview through the Gemini API and Google AI Studio, to enterprises via Vertex AI and Gemini Enterprise, and to consumers in the Gemini app and NotebookLM.

    Why it matters: The post pairs the release with a benchmark table comparing Gemini 3.1 Pro against Gemini 3 Pro, Claude Sonnet 4.6, Claude Opus 4.6, and GPT-5.2 on reasoning and coding tasks.

Feb 13

Feb 13Fri
  1. MiniMax BlogAI score62

    MiniMax details Forge, a scalable agent RL framework behind M2.5

    AIMiniMax describes Forge, its internal reinforcement learning framework for training real-world agents, which was used during the development of MiniMax M2.5. The post explains a Windowed FIFO scheduler, prefix tree merging that the post says yields a 40x training speedup, and CISPO-based training across more than one hundred thousand agent scaffolds and environments.

    Why it matters: The post details how the Forge framework balances throughput, stability, and agent flexibility, with concrete scheduling and prefix-merging methods for training agent RL at scale.

Feb 9

Feb 9Mon
  1. Cognition Blog (Devin, Windsurf)AI score43

    Devin Can Now Autofix Review Comments from Devin Review and Other Bots

    AICognition has configured Devin to automatically autofix incoming review comments from Devin Review and other PR review bots, as well as lint and CI/CD issues. Devin resolves flagged problems and feeds the fixes back into the pull request without human intervention for mechanical fixes. Users can select which bots Devin responds to in Settings > Customization > Autofix settings.

Feb 4

Feb 4Wed
  1. Anthropic EngineeringAI score72

    Anthropic finds container resource limits can shift agentic coding eval scores

    AIAnthropic reports that resource configuration alone can move Terminal-Bench 2.0 scores by up to 6 percentage points, with infra error rates falling from 5.8% under strict enforcement to 0.5% when uncapped. Above about 3x the per-task specs, extra headroom starts letting agents solve tasks they previously could not, so limits can change what the eval measures.

    Why it matters: The source shows how container resource limits shift agentic coding scores, which helps readers interpret small leaderboard gaps and set up evals more consistently.

  2. Anthropic EngineeringAI score75

    Anthropic details how parallel Claude agents built a 100,000-line C compiler

    AINicholas Carlini of Anthropic's Safeguards team describes an agent-team setup where 16 Claude instances worked in parallel on a shared codebase without human intervention to write a Rust-based C compiler. Over nearly 2,000 Claude Code sessions costing about $20,000 in API fees, the team produced a 100,000-line compiler that can build Linux 6.9 on x86, ARM, and RISC-V. The post focuses on harness design, including high-quality tests, lock files for task claiming, GCC as a reference oracle for the kernel, and the limits the project reached.

    Why it matters: The post shows concrete harness design choices for long-running agent teams, including test design, locking, and parallel work division, that readers can adapt to their own autonomous projects.

  3. Guillaume LampleAI score62

    Mistral's Voxtral Realtime streams speech with sub-200ms latency and open weights

    AIVoxtral Realtime is a natively streaming speech model for voice agents and live applications, with latency configurable down to sub-200ms. At 480ms it stays within 1-2% WER of the offline model, and the weights are released under Apache 2.0. The attached FLEURS chart compares word error rates across latency settings for ten languages, including Chinese.

  4. Guillaume LampleAI score62

    Mistral releases Voxtral 2 transcription models with real-time option

    AIMistral announces Voxtral 2 with two transcription models: Voxtral Realtime, released under an Apache 2 license with latency configurable to sub-200 ms, and Voxtral Mini Transcribe 2, which adds speaker diarization, word-level timestamps, and context biasing. The models support 13 languages and are available through the Mistral API, which the post describes as one of the most cost-effective transcription APIs on the market. The attached chart shows word error rates on FLEURS across Italian, Spanish, English, German, Portuguese, French, Russian, Dutch, and Chinese at several latency settings.

Feb 2

Feb 2Mon
  1. Z.ai Release NotesAI score40

    GLM-OCR: Z.ai launches compact OCR model with CogViT and GLM-0.5B encoder-decoder

    AIZ.ai has launched GLM-OCR, a compact, high-performance optical character recognition model built on its self-developed CogViT and GLM-0.5B encoder-decoder architecture. The model uses a dedicated connection layer for cross-modal alignment and CLIP pre-training on billions of image-text pairs for visual semantic understanding and key token extraction. It is designed to stay lightweight for fast inference.

Jan 29

Jan 29Thu
  1. Z.ai (GLM) · new models on Hugging FaceAI score60

    Z.ai releases open-source GLM-OCR multimodal document model

    AIZ.ai has released GLM-OCR, a 0.9B-parameter multimodal OCR model for complex document understanding, under the MIT License. The model scores 94.62 on OmniDocBench V1.5 and supports deployment through vLLM, SGLang, and Ollama, with an official SDK for document parsing.

    Why it matters: The page gives benchmark scores, a 0.9B parameter size, and supported serving frameworks, which help readers weigh OCR deployment options against heavier alternatives.

Jan 27

Jan 27Tue
  1. Cognition Blog (Devin, Windsurf)AI score32

    Cognition opens London office to expand Devin autonomous coding for European businesses

    AICognition is opening a London office to expand rollout of Devin, its autonomous software engineering agent, to leading European businesses. The company says finance has emerged as a clear use case, with Goldman Sachs, Santander, Citi, and BNY among partners using Devin for modernization, migration, security remediation, and codebase documentation.

  2. Cognition Blog (Devin, Windsurf)AI score38

    Cognizant Partners with Cognition to Scale Devin and Windsurf Across Its Engineering Teams

    AICognizant is deploying Cognition's Devin autonomous software engineer and Windsurf agentic IDE across its engineering organization and global client base. Engineers already use Windsurf for agent-assisted coding and are exploring Devin for end-to-end tasks such as code migration, refactoring, testing, and maintenance. Cognition will embed forward-deployed AI engineers to support project selection, engineer enablement, and ROI measurement.