Skip to content

#Deployment/Engineering

Oct 8

TodayOct 8Thu9 items
  1. PandailyAI score57

    Shanghai AI Lab Open-Sources Intern-Decision Small Models for Structured Decisions

    Shanghai AI Lab has open-sourced Intern-Decision, a family of 0.8B, 2B and 4B parameter models that return structured decisions with probabilities instead of free text. The developers self-report that the 4B model averages 90.02% accuracy across seven test suites, ahead of a commercial reference model at 88.74%, with about 44 milliseconds of local latency on a single RTX 4090. Weights are on Hugging Face, and MetaX says the models run on its hardware from launch.

  2. Artificial AnalysisAI score42

    Grok Imagine Video 1.5 Lite ranks #17 in video arena at lower cost

    SpaceXAI's Grok Imagine Video 1.5 Lite ranks #17 on both AA-Video-T2V v2.0 leaderboards, ahead of Google's Veo 3.1 at about a third of its price. It is the fastest model at its quality level in Artificial Analysis benchmarks, with a median of 60.5 seconds for a 10-second 1080p clip, and it costs $0.14 per second at 1080p, 56% of Grok Imagine Video 1.5's $0.25 per second.

  3. Leandro von WerraAI score70

    Carbon-A open model and database predict 566 million gene candidates across 22,617 species

    Carbon-A is an open model that predicts gene locations directly from DNA, and it has been used to annotate genomes from over 22,000 species. The release includes a database of 566 million gene candidates, about 16 times the gene annotations in the RefSeq dataset. Wet-lab RNA experiments supported 239 candidates missing from RefSeq across cats, Syrian hamsters, chickens, and Arabidopsis.

    AIWhy it matters: The source ties an open gene-annotation model to specific wet-lab checks and gene counts, helping readers judge how far its predictions extend beyond well-studied genomes.

  4. Leiphone (雷峰网)AI score62

    Claude Haiku 5.5 gains on computer use but still trails Sonnet 5.5 in terminal coding

    Anthropic released Claude Haiku 5.5, raising its OSWorld 2.1 score from 15.7% to 72.4% and supporting a 1 million token context window. The article notes Haiku 5.5 still scores 39.2% on Terminal-Bench 4.0 against Sonnet 5.5's 70.6%, and that prompts above 100,000 tokens are priced higher, so migration costs need to be measured on real workloads.

  5. OpenRouterAI score44

    Step 5 Preview from @StepFun_ai is live on OpenRouter. Their new flagship for agentic work: a sparse MoE (27B active, 600B total), 1M context, and text, image, and video input. Strong at coding and professional knowledge work, especially finance. Try it: https://openrouter.ai/stepfun/step-5-preview

    Step 5 Preview from @StepFun_ai is live on OpenRouter. Their new flagship for agentic work: a sparse MoE (27B active, 600B total), 1M context, and text, image, and video input. Strong at coding and professional knowledge work, especially finance. Try it: https://openrouter.ai/stepfun/step-5-preview

Oct 7

Oct 7Wed
  1. MarkTechPostAI score67

    Anthropic releases Claude Haiku 5.5, a small model with 1M context

    Anthropic has released Claude Haiku 5.5, its cheapest and fastest small model, priced at $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100K tokens. It keeps a 1M token context window, up to 128K output tokens, and is generally available on the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. Anthropic reports 72.4% on OSWorld 2.1 (offline subset) versus 15.7% for Haiku 4.5, and the article notes that non-default temperature, top_p or top_k values return a 400 error.

  2. AWS Machine Learning BlogAI score56

    Claude Haiku 5.5 becomes available on Amazon Bedrock and Claude Platform on AWS

    Anthropic's Claude Haiku 5.5 is now available on Amazon Bedrock and Claude Platform on AWS. According to Anthropic, it is the fastest and most efficient model in the Claude 5.5 family and costs around 75 percent less than Claude Haiku 4.5 for most tasks. The post also covers pairing it with Claude Opus 5.5 as a subagent layer and provides Boto3, Converse, and Anthropic SDK examples for calling the model.

  3. ThariqAI score67

    Claude Haiku 5.5 returns as a cheaper, faster small model

    Anthropic has released Claude Haiku 5.5, which it describes as the cheapest, fastest, and most capable small model it has released. On average it costs around 75% less to run than Claude Haiku 4.5, and the author says it is 10x cheaper than Haiku 4.5 under 100k tokens. It can be tried with computer use, workflows, and the API.

    This story has a top pick“Anthropic releases Claude Haiku 5.5, scoring 43 on the Intelligence Index”

  4. ClaudeDevsAI score42

    Haiku 5.5 is now available in the Claude Platform and Claude Code. On average, it costs around 75% less to run than Haiku 4.5. It pairs well with Opus 5.5 or Sonnet 5.5 as a subagent. Use it for high-volume, cost-sensitive tasks like summaries, compactions, or database queries.

    Haiku 5.5 is now available in the Claude Platform and Claude Code. On average, it costs around 75% less to run than Haiku 4.5. It pairs well with Opus 5.5 or Sonnet 5.5 as a subagent. Use it for high-volume, cost-sensitive tasks like summaries, compactions, or database queries.

  5. Liquid AIAI score38

    Both Open d1 models run across NVIDIA DGX, RTX, and Jetson, with day-one llama.cpp support to run anywhere. d1-3B single-question latency, measured one request at a time: > NVIDIA RTX 4090: 8 ms > Jetson AGX Thor: 16 ms > Jetson AGX Orin 64 GB: 26 ms > Jetson Orin Nano: 50 ms 4/

    Both Open d1 models run across NVIDIA DGX, RTX, and Jetson, with day-one llama.cpp support to run anywhere. d1-3B single-question latency, measured one request at a time: > NVIDIA RTX 4090: 8 ms > Jetson AGX Thor: 16 ms > Jetson AGX Orin 64 GB: 26 ms > Jetson Orin Nano: 50 ms 4/

  6. vLLMAI score34

    Vela 2.0 brings span-level decisions to routing: safety checks, domain classification, PII spans and unsupported claims in one call. Built by vLLM Semantic Router and KR Labs. Four sizes, 0.3B to 9B. Apache-2.0. https://huggingface.co/collections/vllm-sr/vela-20

    Vela 2.0 brings span-level decisions to routing: safety checks, domain classification, PII spans and unsupported claims in one call. Built by vLLM Semantic Router and KR Labs. Four sizes, 0.3B to 9B. Apache-2.0. https://huggingface.co/collections/vllm-sr/vela-20

Oct 6

Oct 6Tue
  1. Claude Apps Release NotesAI score60

    Claude Haiku 5.5 launches as a fast, low-cost small model, and Max and Team plans gain monthly API credits

    Anthropic launched Claude Haiku 5.5, which it describes as the cheapest, fastest, and most capable small model it has released, aimed at high-volume, cost-sensitive tasks. Max and Team plans now include monthly API credits for running their own apps and agents on the Claude Platform, rolling out over a few days. Users claim the credits by linking a Claude Console organization in Settings > Billing for Max or Organization settings > Billing for Team.

    AIWhy it matters: The notes name a new small model and a credit change for Max and Team plans, with the claim path, which matters for teams budgeting API use.

  2. Gemini API ChangelogAI score58

    Google releases Gemini Nano Banana 2.1 for general availability

    Google has made Gemini Nano Banana 2.1, identified as gemini-nano-banana-2.1, generally available as an image generation and conversational editing model. It improves visual quality, prompt adherence, multi-turn character consistency, and text rendering, and adds panoramic aspect ratios such as 1:4, 4:1, 1:8, and 8:1 at 1K, 2K, and 4K resolutions. The gemini-3.1-flash-image model is deprecated with no shutdown date announced, and developers are told to migrate to the new model.

Oct 5

Oct 5Mon
  1. Liquid AIAI score37

    Liquid AI's d1 decision model adds vision, rivaling GPT-6.1 Sol at lower cost

    Liquid AI released d1 with vision support, accepting images, text, or both as inputs. In tests on six real applications, d1 matched or beat GPT-6.1 Sol on four while costing 19x to 200x less than both GPT-6.1 Sol and Claude Opus 5.5. It returns probabilities for yes/no, choice, or score questions in one forward pass, with text decisions in 200 to 300 ms.

  2. Liquid AI · new models on Hugging FaceAI score67

    Liquid AI releases d1-3B, a 3B multimodal decision model for edge deployment

    Liquid AI has released d1-3B, a 3B parameter multimodal model post-trained to return calibrated, typed answers to yes/no, choice, and score questions in one forward pass. The source reports a Decision Index 0.2.1 score of 48.57, the highest among models under 10B in its table, and 8 ms per decision on an NVIDIA RTX 4090.

    AIWhy it matters: The source gives benchmark scores against named peer models and edge latency figures across several hardware targets, helping readers judge fit for on-device decision pipelines.

Oct 4

Oct 4Sun
  1. Liquid AI BlogAI score70

    Liquid AI releases d1 decision model with image input support

    Liquid AI introduces d1, its first decision model, now accepting both text and images. The company says d1 matches or beats GPT-6.1 Sol on four of six tested applications, at 19x to 200x lower cost and with faster answers on every task. d1 is available on the Liquid AI API and through Vercel and OpenRouter, with text-only support on those two platforms for now.

    AIWhy it matters: The post gives benchmark comparisons against named models along with per-token pricing and latency figures, which makes the cost and speed tradeoff checkable.

Oct 3

Oct 3Sat
  1. IndexTeam (Bilibili) · new models on Hugging FaceAI score22

    Index-Echo-S2ST-9B-FP4 released as NVFP4 quantized speech translation model

    IndexTeam released Index-Echo-S2ST-9B-FP4, an NVFP4 (W4A4) quantization of the Index-Echo-S2ST-9B speech-to-speech translation model, with only its text LLM backbone quantized. Perplexity rose from 3.8218 to 3.9650 (+3.75%) on a fixed corpus, while zh→en and en→zh outputs were semantically equivalent, and full FP4 speedup requires an NVIDIA Blackwell GPU.

  2. IndexTeam (Bilibili) · new models on Hugging FaceAI score27

    Index-Echo-S2ST-2B FP4 Quantized Speech-to-Speech Translation Model Released on Hugging Face

    IndexTeam released Index-Echo-S2ST-2B-FP4, an NVFP4 (W4A4) quantized version of the Index-Echo-S2ST-2B speech-to-speech translation model, with only the text LLM backbone quantized and the audio components kept in BF16. On a fixed corpus, perplexity rose from 5.9332 to 6.4980 (+9.52%), while zh->en and en->zh generations matched the original. Full FP4 acceleration requires an NVIDIA Blackwell GPU, and the model loads via compressed-tensors in vLLM or transformers.

  3. IndexTeam (Bilibili) · new models on Hugging FaceAI score20

    IndexTeam releases NVFP4 quantized Index-Echo-S2TT-9B speech translation model

    IndexTeam published an NVFP4 (W4A4) quantized version of its Index-Echo-S2TT-9B speech-to-text translation model, quantizing only the text LLM backbone while keeping the audio tower and other components in BF16. On an NVIDIA A100, perplexity rose from 3.4155 to 3.5113 (+2.81%), with zh->en and en->zh outputs semantically equivalent under greedy decoding. Full FP4 speedup requires an NVIDIA Blackwell GPU, while older GPUs get only memory reduction.

  4. IndexTeam (Bilibili) · new models on Hugging FaceAI score20

    IndexTeam releases NVFP4 quantized Index-Echo-S2TT-2B speech translation model

    IndexTeam has published an official NVFP4 (W4A4) quantized version of its Index-Echo-S2TT-2B speech-to-text translation model on Hugging Face. Only the text LLM backbone is quantized, while the audio tower, connector, and speech-synthesis components remain in BF16. Perplexity rises 5.80%, from 4.8772 to 5.1599, on a fixed corpus, and full FP4 speedup requires an NVIDIA Blackwell GPU.

  5. IndexTeam (Bilibili) · new models on Hugging FaceAI score22

    Index-Nailong-9B-FP4 NVFP4 quantized translation model released on Hugging Face

    IndexTeam released Index-Nailong-9B-FP4, an official NVFP4 (W4A4) quantization of the Index-Nailong-9B multilingual translation model, which covers 150 languages. In a validation on an NVIDIA A100 against the BF16 checkpoint, perplexity rose 3.10% (2.4339 to 2.5094), and zh-en and en-zh outputs were semantically equivalent. Full FP4 compute acceleration requires an NVIDIA Blackwell GPU, while older GPUs get memory savings only; the FP8 build is recommended for Hopper and Ampere.

  6. IndexTeam (Bilibili) · new models on Hugging FaceAI score29

    Index-Nailong-2B-FP4 Released as NVFP4 Quantized Translation Model

    IndexTeam has released Index-Nailong-2B-FP4, an official NVFP4 (W4A4) quantization of its Index-Nailong-2B multilingual translation model, which supports 150 languages. The checkpoint keeps lm_head, embeddings, and MoE router gates in BF16, and a perplexity test on a fixed corpus rose from 3.2806 to 3.4998 (+6.68%), while zh->en and en->zh outputs matched BF16 semantically. Full FP4 acceleration requires an NVIDIA Blackwell GPU; on Hopper or Ampere, vLLM provides only memory savings, so the FP8 build is recommended.

  7. IndexTeam (Bilibili) · new models on Hugging FaceAI score23

    Index-Homura-9B-FP4 released with NVFP4 quantization for translation model

    IndexTeam released Index-Homura-9B-FP4, an official NVFP4 (W4A4) quantization of the Index-Homura-9B translation model from the Index-Translate family. On a fixed corpus, perplexity rose from 2.5386 in BF16 to 2.6245, a 3.38% increase, and zh->en generations matched the original. Full FP4 compute acceleration requires an NVIDIA Blackwell GPU, while older GPUs get only weight-only memory savings and the FP8 build is recommended for them.

  8. IndexTeam (Bilibili) · new models on Hugging FaceAI score29

    Index-Homura-2B-FP4 released as NVFP4 quantized translation model

    IndexTeam released Index-Homura-2B-FP4, an official NVFP4 (W4A4) quantization of its Index-Homura-2B multilingual translation model, which supports 150 languages. The quantized checkpoint shows a 5.73% perplexity increase over the BF16 original (3.5011 to 3.7017) on a fixed corpus, and its zh-en and en-zh outputs are semantically equivalent under greedy decoding. Full FP4 acceleration requires an NVIDIA Blackwell GPU, while the source recommends the FP8 build for Hopper and Ampere hardware.

Oct 2

Oct 2Fri
  1. OllamaAI score29

    .@Cloudflare's decision models are available on Ollama! Decision models allow you to classify an image, label a bug report or route a support ticket to the right team. Clef (27B): ollama pull clef Clef Flash (9B): ollama pull clef-flash

    .@Cloudflare's decision models are available on Ollama! Decision models allow you to classify an image, label a bug report or route a support ticket to the right team. Clef (27B): ollama pull clef Clef Flash (9B): ollama pull clef-flash

  2. Ai2 (Allen Institute for AI)AI score67

    Ai2 open-sources AstaBrief 8B, a fast open-weights scientific report model

    Ai2 released AstaBrief 8B, a model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. In Asta's Generate a report feature, Fast mode averages 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The model is built on Qwen3-8B with supervised fine-tuning and DPO, and institutions can run its open weights on their own infrastructure.

    AIWhy it matters: The post explains the data filtering and one-pass generation choices behind a fast open-weights report model, showing what worked and what did not.

Oct 1

Oct 1Thu
  1. Harrison ChaseAI score44

    congrats @AravSrinivas! more entrants in the decision model class, and this one is open weights cheap typed answers for the small calls inside a harness (routing, approvals, judging), big model for the rest https://x.com/AravSrinivas/status/2105774153903268288

    congrats @AravSrinivas! more entrants in the decision model class, and this one is open weights cheap typed answers for the small calls inside a harness (routing, approvals, judging), big model for the rest https://x.com/AravSrinivas/status/2105774153903268288

  2. ReplicateAI score47

    FLUX 3 Image is here. The latest from @bfl_ai, generate stunning native 4K imagery, and get maximum control over every pixel. Construct images with hyper-specific layouts using bounding boxes, make multiple targeted edits at once, and stay consistent across edits. For its first week, it's 50% off.

    FLUX 3 Image is here. The latest from @bfl_ai, generate stunning native 4K imagery, and get maximum control over every pixel. Construct images with hyper-specific layouts using bounding boxes, make multiple targeted edits at once, and stay consistent across edits. For its first week, it's 50% off.

Sep 30

Sep 30Wed
  1. Google DeepMindAI score36

    With a 1M token output limit, Argon adds a deeper level of reasoning to tackle long, multi-step problems in one go. Feedback from early testers will help us strengthen our systems before we roll out more broadly to developers, enterprises, and consumers soon. Find out more → https://goo.gle/4rZBiSd

    With a 1M token output limit, Argon adds a deeper level of reasoning to tackle long, multi-step problems in one go. Feedback from early testers will help us strengthen our systems before we roll out more broadly to developers, enterprises, and consumers soon. Find out more → https://goo.gle/4rZBiSd