Skip to contentSkip to stories

Updated

#Tutorial/How-to

Showing low-relevance items too. Hide low-relevance items

Sep 21

Sep 21Mon
  1. Matei ZahariaXAI score32

    Matei Zaharia praises GEPA working with Jev

    AIMatei Zaharia, a prominent AI researcher, said it is very cool that GEPA works on Jev. The post is a short endorsement, linking to background about a test in which GEPA optimized Jev's prompts for extracting suspected adverse drug effects from medical sentences.

  2. Ali GhodsiXAI score14

    Superhuman scales its GEC inference infrastructure, per Ali Ghodsi

    AIDatabricks CEO Ali Ghodsi praises how Superhuman scales its AI infrastructure, linking to a Superhuman blog post on scaling GEC inference. The post itself is only a link with no details provided here, so the specific techniques and results cannot be confirmed from this source.

Sep 20

Sep 20Sun
  1. vLLM BlogOfficialAI score44

    vLLM Reports PD Serving Results for Qwen3.8-2.4T on GB300 NVL72

    AIvLLM achieved 5000 total token throughput per GPU in high-throughput PD serving of Qwen3.8-2.4T on a GB300 NVL72 cluster under an 8K/1K workload. The low-latency scenario reached 180 generated tokens per user, with both results shown on the Pareto frontier. The post also provides srt-slurm recipes and explains the tuning process used to create them.

Sep 19

Sep 19Sat
  1. OpenBMBOfficialAI score34

    OpenBMB's 2B MiniCPM5 powers a local personal news desk

    AIOpenBMB's 2B-parameter MiniCPM5 model runs as a local news desk on an older i5-9400F PC with 16GB RAM and no cloud API. The developer built a system that collects official sources hourly and sends a 24-hour Telegram recap with a lead story and links.

  2. Sebastian RaschkaXAI score36

    Raschka's Inference Scaling Part 1: Sampling for Better Accuracy

    AISebastian Raschka starts a series on inference scaling by modifying text generation with temperature scaling, top-p filtering, and multinomial sampling to produce diverse outputs. He says this enables self-consistency and best-of-N approaches that improve answer accuracy by more than 2x. The video covers chain-of-thought prompting, a MATH-500 evaluation, and accuracy versus compute tradeoffs.

    Video from @rasbt's post

Sep 18

Sep 18Fri
  1. TinkerOfficialAI score31

    Jasper's guide shows how reward tweaks shape search agent behavior

    AIJasper Lu's new blog post walks through training a search agent with GRPO, showing how small reward function changes teach a model to avoid sloppy tool calls, prune unnecessary documents, and balance persistence against token efficiency. The post makes every rollout browsable and releases the code as open source, with the full process from learning rate sweeps to reward shaping documented.

  2. LMSYS OrgOfficialAI score16

    LMSYS releases SGLang SSD expert pack blog post

    AILMSYS Org published a blog post introducing an SGLang SSD expert pack, with the full details available on its website. The post itself gives no further technical specifics, so the summary is limited to the announcement.

  3. LlamaIndex 🦙OfficialAI score16

    LlamaParse Preserves Table Structure in EIA Energy Report Data

    AILlamaIndex launches a Parsed by LlamaParse series, using the EIA's September 2026 Short-Term Energy Outlook to show how table parsing errors can corrupt downstream data. The example value 1,186 in Table 7a means electricity sales to ultimate customers in Q3 2026, in billion kilowatthours, and misparsing its quarter, metric, or unit could flow into dashboards and forecasts. The post says LlamaParse preserves structure and footnote context needed for databases, forecasting, and AI applications.

    Image from @llama_index's post
  4. Google · AI blogOfficialAI score29

    Google co-builds Google Flow tools with two designers for New York Fashion Week runways

    AIGoogle's Envisioning Studio, with Google Labs, co-developed custom Google Flow tools with designers Jane Wade and Sergio Hudson ahead of New York Fashion Week. Wade's Styling Suite let her style runway looks on digital models before producing physical samples, while Hudson's Runway Visualization helped him stage his show within a tight budget. The source says the tools are built with natural language and no coding experience.

  5. MiniMax Design (H3)OfficialAI score18

    Hailuo AI speeds up storyboarding with 3x3 panel-to-video generation

    AIHailuo AI's post suggests that a single 3x3 storyboard image can be converted into a video using the MiniMax H3 Max r2v model at 480p for 15 seconds. The quoted post describes settings with Quality prompt tuning and standard reference strength, and asks for a 2D animation with panel-to-panel cuts while excluding multiple panels and BGM.

  6. Hamel HusainBlogAI score62

    Hamel Husain's FAQ on AI evals: error analysis, judges, and trace review

    AIHamel Husain and Shreya Shankar's FAQ explains AI evals as tests of whether an AI system does what users and the business want. It recommends starting with error analysis on at least 30 traces, then turning recurring failures into binary code-based checks or LLM judges validated against human labels.

Sep 17

Sep 17Thu
  1. Mike KriegerXAI score12

    Anthropic's Krieger builds interactive map of Iron Tangle level

    AIMike Krieger, Anthropic's account owner, used an interactive Claude artifact to visualize the Iron Tangle level from Dungeon Crawler Carl, saying it helped him finally understand the layout. The post shares a link to the artifact, with no further details about its features.

  2. Gemini NotebookOfficialAI score37

    Mariposa Museum exhibit shows town in 1859, a decade after Gold Rush

    AIA Mariposa Museum exhibit photographed by writer Steven Johnson, shared by Gemini Notebook, documents the Sierra Nevada town in 1859, ten years after the Gold Rush began. Johnson says he used the Gemini Notebook mobile app's camera feature to generate a detailed report from photos of the display, which he says was 99% accurate on fact-checking.

  3. JetBrains AI BlogOfficialAI score44

    Building a RAG Pipeline for Semantic Code Search: A Developer Diary

    AIJetBrains describes building Air Context, a RAG pipeline that gives LLM agents semantic code search over real repositories instead of grep. The first installment covers parsing, chunking, and vectorization, arguing that fixed-size line chunks split related code and that structure-aware chunking using language grammar produces better retrieval units.

  4. Z.aiOfficialAI score40

    GLM-5.3 helped build the inference stack serving GLM-5.3-Flash

    AIZ.ai reports that GLM-5.3 helped build and optimize the inference infrastructure for GLM-5.3-Flash. The system went from first successful run to production readiness in under two weeks, with end-to-end throughput tripling over the initial baseline. The team credited dense feedback from local correctness tests, execution traces, microbenchmarks, and end-to-end measurements for enabling targeted hypothesis testing.

Sep 16

Sep 16Wed
  1. Perplexity DevelopersOfficialAI score21

    Perplexity releases a Search SDK cookbook for coding agents

    AIPerplexity has published a new cookbook for its Search SDK, showing how to run focused searches and filter results to official documentation. The recipe extracts relevant passages and produces a source-linked brief that a coding agent can use.

    Video from @perplexitydevs's post
  2. TinkerOfficialAI score32

    Sundial trains Inkling-Small to fix LaTeX errors in under a second

    AISundial fine-tuned Thinking Machines' Inkling-Small with RLVR on 3,978 verified TeX.StackExchange fixes, using rewards for compilation and PDF match and penalties for removed content. The trained model fixes 83.7% of LaTeX errors in under one second at $0.0013 per fix, according to the post. Sundial says it is rolling out the model in its editor, applying fixes as suggestions and rebuilding the PDF.

  3. Google for DevelopersOfficialAI score38

    Three companies use Gemini agentic video understanding to cut token costs

    AIMosaic, Ponder Studio, and Revyl used early access to Google's Gemini Flash models to test agentic video understanding on long footage. Mosaic reports a 97% cut in median token usage and nearly double the ability to handle complex edits, while Ponder Studio reports a 0.967 F1 score and about 72% lower token costs for B-roll selection. Revyl says the approach improved mobile UI bug-catching accuracy by 65%. The capability is available now for video uploads and YouTube videos via the Gemini API.

  4. LlamaIndex 🦙OfficialAI score14

    LlamaIndex webinar on insurance document pipelines with LlamaParse and Extract

    AILlamaIndex Solutions Architect Abrar Mahi will host a webinar on turning insurance documents such as accord forms and policy documents into structured data for underwriting, policy review, and claims. The session covers extracting policy, property, and claims history into a defined schema, verifying values with citations and bounding boxes, and using confidence scores with validation rules to route items to human review.

    Image from @llama_index's post

Sep 15

Sep 15Tue
  1. Noah ZwebenXAI score17

    Anthropic offers Claude Tag office hours for on-call triage feedback

    AIAnthropic is hosting office hours for teams interested in using Claude Tag for on-call work, and it is asking Team or Enterprise plan users to share triage feedback. Claude Tag can start investigating when a Slack alert fires by pulling metrics, diffing deploys, and checking flags to propose a likely cause and fix. Sign-up is through a Google Calendar booking link.

  2. Josh WoodwardXAI score31

    Gemini Notebook adds spoken Q&A and lecture audio notes for students

    AIGoogle's Gemini Notebook now offers live spoken Q&A over class materials in about 100 languages, and lets students record lectures on the go with audio notes saved automatically to a chosen notebook. University students in 140+ countries can also still get a free Google AI Plan for higher limits and access to more Google products.

  3. LlamaIndex 🦙OfficialAI score22

    LlamaIndex Moves Off Stainless for LlamaParse SDK Generation

    AILlamaIndex says Stainless helped it keep LlamaParse SDKs current and pushed it to make the API's names and schemas more consistent. With the Stainless team joining Anthropic, George He and Yong Park explain what worked, what they learned, and why changing SDK generators needs careful handling.

    Image from @llama_index's post
  4. Google · Innovation & AIOfficialAI score52

    Google says its language technology now covers over 300 languages with new speech, data, and on-device tools

    AIGoogle reports that its technologies and products now power everyday interactions in more than 300 languages used by over 7 billion people, about 86% of the global population. The post describes new speech models, including Gemini 3.5 Live Translate and Gemini 3.5 Transcribe, plus the TranslateGemma open translation models trained across 55 languages.

Sep 14

Sep 14Mon
  1. Google Developers BlogOfficialAI score60

    Build zero-trust AI agents that judge intent, not just syntax

    AIPart 2 of the zero-trust agents series moves security checks from agent code to the Gemini Enterprise Agent Platform runtime. Model Armor screens prompts and responses, Semantic Governance Policies judge proposed tool calls against intent and business rules, and Agent Anomaly Detection flags multi-turn drainage that single-turn checks miss. The same Customer Support and Returns Agent from Part 1 is used, with the companion demo open-sourced on GitHub.

    Why it matters: The post walks through a concrete refund agent under four attacks, showing how screening, intent judgment, and anomaly detection each catch what the others miss.

  2. vLLM BlogOfficialAI score62

    How vLLM Speculators trained a DSpark draft model for Kimi K3 on GB300 NVL72

    AIThe vLLM team trained a DSpark speculative decoding draft model for Kimi K3, a 2.8T-parameter model, using the Speculators library on GB300 NVL72 hardware. They added a MooncakeHiddenStatesConnector to stream hidden states from disaggregated vLLM inference nodes to training nodes across multiple machines. The released speculator raises single-stream interactivity from about 110 to about 435 tokens per second per user on math reasoning, with up to about 3.5x higher output throughput under concurrent load.

    Why it matters: The post shows how hidden-state extraction and Mooncake transfers let a 2.8T-parameter model's speculator be trained across multiple nodes, a reusable pattern for similar setups.