Cohere Invites Audience to WMT 2026 Talk and Technical Report
AICohere promotes a video and an October WMT 2026 conference talk featuring Kocmi and the team. The post links to a full technical report on North Small Translate.
Updated
Updated
Showing low-relevance items too. Hide low-relevance items
AICohere promotes a video and an October WMT 2026 conference talk featuring Kocmi and the team. The post links to a full technical report on North Small Translate.
AICohere reports that after the first training step, its model could already translate over 90% of the training documents, which created a data problem. Kocmi describes how the most difficult samples were used to strengthen North Small Translate's capabilities. This post is part 3 of a six-part thread.
AICohere says its supervised fine-tuning extended machine translation with data focused on post-editing, error detection, and terminology. The team also ran multiple reinforcement learning and DPO stages aimed at errors the model showed.

AISebastian Raschka released round six of his Reasoning from Scratch series, introducing Reinforcement Learning with Verifiable Rewards (RLVR) and Group Relative Policy Optimization (GRPO) with an implementation. The video covers accuracy and format rewards, DeepSeek-R1 training, and GRPO versus PPO, then walks through a training loop and evaluates checkpoints on MATH-500.
AIBaseten tested the MetaInfer skills-only approach by having Claude Code build an inference engine, VibeQwen, for Qwen-3.6-35B-A3B in NVFP4 on a single B200. On single-stream text, VibeQwen decoded 90% faster than a tuned vLLM 0.25.1 deployment (1,792 vs. 943 TPS) and cut time to first token from 28 ms to 12 ms, with a 71% throughput gain at concurrency 32. The author notes this was an outcome-focused run that allowed some numerically different outputs as long as accuracy stayed at or above the BF16 baseline.
Why it matters: The post tests a skills-only inference engine method on a real model and states the speed and accuracy constraints used, helping readers judge how far such automated optimization can be trusted.
AIMeta researchers used Muse Spark to uncover a counterexample to a proposed rule about mathematical structures inspired by biology, then developed and refined an alternative characterization. The work is detailed in a linked paper on solvable evolution algebras and a conjecture by Garcia-Martinez and Perez-Rodriguez.
AIResearchers working with Muse Spark connected a number theory idea to a string theory calculation, proving the link holds in more cases than previously known. The work builds on ideas from the 1980s, according to the post.

AIMeta researchers used Muse Spark to prove a clear rule for when a cycle-based relaxation of a hard optimization problem matches the original exactly and when it leaves a gap. The work concerns completed length-three alpha cycles, and the full paper is linked in the post.

AIResearchers disproved a proposed rule about symmetry-describing mathematical structures by finding a single counterexample. Muse Spark generated the search code that located it, and the team verified the result and completed the proof.

AIWith help from Muse Spark, researchers proved that a wave in a laser-inspired model must blow up in finite time under the conditions studied. The result comes from a tug-of-war between one effect squeezing the wave inward and another spreading it out. The paper is titled finite-time blow-up of radial negative-energy solutions for the mass-critical biharmonic nonlinear Schrödinger equation.

AIMathematicians worked with Meta's Muse Spark to answer a question about fitting random points onto the surface of a stretched sphere, known as an ellipsoid. For the setting studied, they proved a sharp cutoff between when an exact fit is likely and when it is unlikely.

AIAI at Meta says mathematicians used Muse Spark 1.1 and Muse Spark 1.2 in Thinking Mode through the standard meta.ai chat interface to find solutions to open problems. The company is sharing six resulting papers, each marking which passages were drafted primarily by humans or AI, with mathematicians guiding the work and a second group reviewing it.
AIGoogle AI linked to a blog post presenting facts about Project Suncatcher, a Google research initiative. The post itself contains no further details beyond the title and link, so specifics cannot be confirmed from this source alone.
AILiquid AI shared a Hugging Face guide to multi-harness reinforcement learning for coding agents, in which a proxy records the token ids and logprobs vLLM samples so training works without changing the harness. Per the quoted post, LFM2.5-2.6B rose from 42% to 54% after training across four harnesses at once, and imitation fine-tuning on 3,189 rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs. The proxy, trainer, tasks, SFT data, training code and seven trained models are described as open.
AIGoogle announces a next-generation federated learning system that uses Trusted Execution Environments to provide verifiable, auditable data anonymization. The system publishes access policies to a public transparency log and is deployed in Gboard, which has launched English and Japanese next-word prediction models with stronger privacy guarantees and improved accuracy. Training time has also sped up significantly because computation moved to the server and is parallelized across many machines.
Why it matters: The post shows how Trusted Execution Environments make federated learning's privacy claims externally verifiable, rather than relying on trust in the server operator.
AIHugging Face and collaborators published a guide to multi-harness RL that trains models through a capture proxy without changing the agent harness. The proxy records the token ids and logprobs vLLM samples, and the source reports LFM2.5-2.6B rising from 42% to 54% after training across four harnesses. Fine-tuning on 3,189 successful rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs, and the capture proxy, trainer, tasks, SFT data, training code, and seven trained models are released openly.
Why it matters: The source gives a concrete method for training models across several agent harnesses, with measured gains and a note that imitation learning underperformed RL.

AICool eval. Simply ask an LLM “Land or Water?” and give it a latitude and longitude coordinate as text. Ask 16,200 times, plot as image. The models know. From compressing the internet.
AIGoogle DeepMind announced SynthID Bio, a family of watermarking methods for AI-generated biological designs. According to the quoted post, the team can embed an imperceptible signature directly into protein sequences without affecting their biological function. Sundar Pichai called it a big step forward for scientific integrity and biosecurity.
AIA Latent.Space podcast episode features MIT researcher Alex Zhang explaining recursive language models (RLMs). He discusses why Claude Code, Codex, and Pi are basically the same, and how RLMs use code, context offloading, and recursive subagents to generalize across tasks. The episode also covers OpenAI's 10,000-agent, 130B-output-token experiment and academia's freedom to pursue ambitious research bets.
AIResearchers Maureen de Seyssel, Jie Chi, and Zakaria Aldeneh found that strengthening language discrimination during pretraining reduces the performance gap between multilingual and monolingual HuBERT speech models. In a controlled English/French setting, phone-ABX error fell from 11.6% to 10.4%, close to the monolingual 10.8%, while lexical sWUGGY scores rose from 52.1% to 56.7%. The gains were largest when language discrimination was introduced in the first training iteration.
AIEpoch AI estimates that high-bandwidth memory shipped in 2025–27 could eventually support about 30–170 million concurrent frontier-model agents once fully deployed and allocated. Using DeepSeek V4 Pro serving benchmarks, the estimate rises to about 1.9 billion concurrent agents. The authors compare the implied API-equivalent spending of $2.6–5.3 trillion per year with projected developer revenue of roughly $1 trillion by end-2027, suggesting demand may lag supply.
Why it matters: The analysis converts HBM shipment data into concurrent agent capacity and compares it with projected API revenue, showing where compute buildout may outpace demand.
AIApple Machine Learning Research reports that discrete diffusion steps match the training distribution only when simultaneously written token positions are conditionally independent given already-fixed tokens. The authors show that per-position distributions cannot determine such dependence, and on the synthetic ScanAndAdd task, confidence-ranked groups of two or more positions were dependent and produced a generated distribution 29 times the sampling-noise floor in total variation.
AIMeta's Jagged Flash Attention kernel, built with TLX on NVIDIA Blackwell B200, outperforms FlashAttention-4 (May 2026 version) on GEM's jagged shapes by about 13% on the forward pass and about 50% on the backward pass. The TLX attention kernel is roughly 3.2K lines of Triton-level code, about 3× shorter than FA4's ~10K-line CuteDSL kernels. The benchmarks use bfloat16 on B200.
AIarXiv has introduced stricter rate limiting for all submitters to fairly distribute moderator time. The post links this move to Vibe research, where turning ideas into papers is easier, while noting that standards for judging research output have not kept pace.
AIHarvard physicist Matthew Schwartz argues that LLMs are poorly matched to science when used as human-style collaborators, so he built a toolkit for exact quantitative calculations. Working with Claude, the approach surfaced connections to ecology, population genetics, and a dozen other fields, with domain experts steering it toward interesting questions.
AIOdyssey introduces PROWL-2, a system in which agents and their world model improve through recursive learning. The post says agents expose errors in imagination, and repairing those errors enables further learning. Odyssey frames this open-ended learning as a critical step toward superintelligence.
AIOdyssey introduces PROWL-2, a recursive loop that delivers up to 91% relative gains over the StarCraft world-model baseline. The post also reports stronger robot coordination in simulation.
AIGoogle Gemma relays a StudentBench study reporting that AI tutors matched expert human tutors on immediate GRE learning gains. The author reports 2,383 students and a cost of 7 cents per AI tutor hour versus $75 for an expert human hour. The post also says the top AI tutor beat expert human tutors on average in 5 of 7 academic topics, and that the data and paper are publicly available.
AINew research finds AI tutoring delivers immediate GRE learning gains comparable to expert human tutors, with Gemma 4 31B performing best. The post presents the findings as a way for educators to scale their impact and make guidance more accessible.
AIGoodfire reports that its method catches more unsafe biological requests than frontier model safeguards while dramatically reducing false positives. The claim is illustrated by a plot in the first tweet of the thread.
AIGoodfire Research developed sequence-aware monitors using protein language model embeddings to flag concerning biological sequences in dual-use AI agent tasks. On a custom benchmark, the monitors outperformed frontier model safeguards with fewer refusals on benign requests, and they held up better against paraphrasing and fragmentation attacks. The paraphrase results rely on in-silico estimates and do not establish whether the redesigned proteins keep biological activity, and the monitors run in milliseconds per sequence.
Why it matters: The post gives a concrete benchmark setup and fragmentation results, showing how sequence embeddings can separate dual-use biology requests that task-based safeguards handle poorly.
AIPrime Intellect reports that after about 100 GRPO steps on Hosted Training, Qwen3.6's reward on held-out problems rose from 0.127 to 0.361, a 2.8x gain. Qwen3.5, trained the same way, reached 0.356, suggesting the method works across model families. Both post-trained models finished well ahead of other open models and narrowed the gap to Claude Opus 4.8, with Qwen3.6 activating only 3B parameters per token.

AIZyphra published a paper and blog explaining how local mixing encodes relative position in global NoPE attention models. The work, titled "How Local Mixing Encodes Relative Position in Global NoPE Attention," is available on arXiv and on Zyphra's website.
AIModels using Kimi Delta Attention show a recency bias, as the mechanism controlling how much earlier information is kept shapes memory. As new information mixes into that memory, nearby words share more information, which gives the model clues about how far apart words are.
AIZyphra derives a mathematical theory showing how local layers produce recency bias that propagates through the residual stream, norms, and MLPs into NoPE attention logits. The authors validate the theory on both randomly initialized and trained networks.

AIIn Zyphra's experiments, models with smaller context windows predicted text better while using less training compute. Zyphra believes the stronger recency bias induced by smaller windows outweighs the reduced compute benefit.

AIHybrid NoPE models combine sliding window attention or recurrent layers, which focus on nearby words, with global attention layers that use no positional encoding (NoPE). The post notes that NoPE layers receive no positional information yet can still learn long-range dependencies, and raises the question of how this works.

AIZyphra says that mixing information within overlapping neighboring windows makes nearby words carry similar information inside a model. This creates a built-in clue about how far apart words are, even before training begins, which the post calls recency bias.

AIZyphra Research explains how language models track word order without explicitly encoding position in attention. The post says local memory layers that read nearby words help global attention layers preserve sequence information. The source is a short teaser thread, so no further technical details are given.
