Skip to contentSkip to stories

Updated

#RAG

Sep 18

Sep 18Fri
  1. LlamaIndexAI score16

    LlamaParse Preserves Table Structure in EIA Energy Report Data

    AILlamaIndex launches a Parsed by LlamaParse series, using the EIA's September 2026 Short-Term Energy Outlook to show how table parsing errors can corrupt downstream data. The example value 1,186 in Table 7a means electricity sales to ultimate customers in Q3 2026, in billion kilowatthours, and misparsing its quarter, metric, or unit could flow into dashboards and forecasts. The post says LlamaParse preserves structure and footnote context needed for databases, forecasting, and AI applications.

Sep 17

Sep 17Thu

Sep 16

Sep 16Wed
  1. LlamaIndexAI score14

    LlamaIndex webinar on insurance document pipelines with LlamaParse and Extract

    AILlamaIndex Solutions Architect Abrar Mahi will host a webinar on turning insurance documents such as accord forms and policy documents into structured data for underwriting, policy review, and claims. The session covers extracting policy, property, and claims history into a defined schema, verifying values with citations and bounding boxes, and using confidence scores with validation rules to route items to human review.

Sep 14

Sep 14Mon
  1. LlamaIndexAI score29

    LlamaIndex proposes two-pass just-in-time OCR for agent document pipelines

    AILlamaIndex proposes a two-pass just-in-time OCR pattern for agents working through document collections, avoiding parsing every page upfront. LiteParse, an open-source Rust tool supporting 50+ formats, performs a fast layout-aware first pass with bounding boxes, headings, tables, and a per-page complexity flag, processing a full data room in 32 seconds. LlamaParse then parses only the pages needing deeper analysis, returning cell-level tables, bounding boxes, and confidence scores.

Sep 11

Sep 11Fri

Sep 9

Sep 9Wed
  1. LlamaIndexAI score23

    LlamaParse now available as a ChatGPT connector for document parsing

    AILlamaIndex has made LlamaParse available in the ChatGPT plugin directory, following its earlier Claude integration. The connector parses scanned, table-heavy, and chart-filled documents into Markdown, JSON, or HTML, extracts fields into a user-defined schema, searches document collections, and classifies and splits files into sections.

Aug 26

Aug 26Wed
  1. Tencent · new models on Hugging FaceAI score38

    Tencent releases ContextPilot-E4B, a Gemma4-E4B-based checkpoint for proactive context management

    AITencent has published ContextPilot-E4B on Hugging Face, the Gemma4-E4B checkpoint of ContextPilot, a framework that teaches long-horizon language-model agents to plan, maintain long-term memory, and offload less useful context while reasoning and using tools. The checkpoint is intended for research on proactive context management, long-context QA, and deep search, and loading it alone does not execute the context-management tools, which are provided in the ContextPilot repository.

Aug 20

Aug 20Thu
  1. Mistral AIAI score59

    Mistral Agentic Search adds multi-step retrieval for complex enterprise documents

    AIMistral has released Agentic Search, a multi-step retrieval layer available through its Search Toolkit and Libraries. On FinanceBench, the company reports accuracy rising from 26.7% to 86% over one-shot RAG, and on OfficeQA Pro a gain from 6.3% to 51.9%. The system also reports up to 39.6% lower p90 latency and up to one-third lower token use from fewer repeated searches.

Aug 18

Aug 18Tue

Aug 17

Aug 17Mon
  1. Microsoft Foundry BlogAI score62

    Microsoft Foundry adds five Claude agent features to Azure-hosted deployments

    AIMicrosoft Foundry now offers structured outputs, web search, web fetch, MCP connector, and tool search for Claude models on Azure-hosted deployments. Prompts and completions remain within Azure for these deployments, while only usage metadata and safety-flagged content egress to Anthropic. The features were previously available only on Hosted on Anthropic deployments, which required choosing between capability and data-handling commitments.

    Why it matters: The post shows which agent scaffolding now runs on Azure-hosted Claude deployments, which matters for teams needing data residency without rebuilding search, fetch, or tool routing.

Aug 4

Aug 4Tue
  1. Fireworks AI BlogAI score26

    Voyage AI's embedding and reranking models now run natively on Fireworks AI

    AIVoyage AI by MongoDB's full lineup, including the Voyage 4 family, voyage-multimodal-3.5, and rerank-2.5, now runs natively on the Fireworks inference platform. The partnership lets teams run embedding, retrieval, reranking, and generation on one platform and one API. Fireworks says Voyage 4 Large outperforms Voyage 4, Voyage 4 Lite, Gemini Embedding 001, Cohere Embed v4, and OpenAI v3 Large on average retrieval quality.

Jul 28

Jul 28Tue
  1. Fireworks AI BlogAI score46

    Fireworks AI Shows Low-Cost Fine-Tuning Lifts Domain Embedding Retrieval

    AIFireworks AI describes fine-tuning Qwen3-Embedding-8B on private (query, positive) pairs using bidirectional InfoNCE loss through its Training SDK, then serving the model via an OpenAI-compatible embeddings endpoint. The post reports that around 150 training steps was enough, that rank-32 LoRA landed within about one point of full-parameter fine-tuning, and that gains were largest where the base model struggled, while tasks like CoSQA and FiQA2018 showed flat results.

Jun 1

Jun 1Mon
  1. PaddlePaddleAI score36

    PaddleOCR and ERNIE Image now available as official Dify plugins

    AIPaddleOCR and ERNIE Image are now available as official Dify plugins, bringing document parsing and image generation into Dify's agent workflows. PaddleOCR, powered by PP-OCRv5, PP-StructureV3, and PaddleOCR-VL, turns images, scanned PDFs, and multilingual documents into structured data for chunking, vectorization, and RAG, with private or on-prem deployment supported. ERNIE Image offers free generation, a Turbo mode with 8-step inference, and an OpenAI-style API.

May 8

May 8Fri

Apr 4

Apr 4Sat
  1. Andrej KarpathyAI score62

    Andrej Karpathy outlines an LLM-maintained markdown wiki workflow for personal research

    AIKarpathy describes using LLMs to compile raw source documents into a markdown wiki that he views in Obsidian, with the LLM writing and maintaining most of the wiki. He reports that at about 100 articles and 400K words, the LLM agent can answer complex questions directly from the wiki, and he also runs LLM health checks to find inconsistencies and gaps. He shares the underlying idea as an "idea file" that users can give to their own agents to build a customized version.

Apr 2

Apr 2Thu
  1. Andrej KarpathyAI score49

    Karpathy shares an LLM-maintained personal knowledge base workflow

    AIAndrej Karpathy describes using LLMs to compile raw research sources into a markdown wiki of about 100 articles and 400K words, viewed in Obsidian. He says an LLM agent answers complex questions against the wiki without RAG, with outputs filed back to enhance it. He also suggests the workflow could become a product rather than a collection of scripts.

Dec 2, 2025

Dec 2, 2025Tue
  1. Apple · new models on Hugging FaceAI score36

    Apple releases CLaRa-7B-E2E, an end-to-end RAG model with 16x and 128x compression

    AIApple's CLaRa-7B-E2E is a fully end-to-end unified RAG model that jointly optimizes retrieval and generation, with 16x and 128x document compression. It is trained with end-to-end finetuning using differentiable top-k retrieval and a unified language-modeling objective. The model is available on Hugging Face with example end-to-end inference code.

  2. Apple · new models on Hugging FaceAI score36

    Apple releases CLaRa-7B-Instruct for compressed-document retrieval-augmented QA

    AIApple has published CLaRa-7B-Instruct on Hugging Face, an instruction-tuned unified RAG model with built-in semantic document compression at 16× and 128× ratios. The model answers instruction-following questions directly from compressed document representations, and its paper, GitHub repository, and transformers usage example are referenced in the release.

Nov 5, 2025

Nov 5, 2025Wed