Generalist shows uncut video of robot model unzipping pencil pouch
AIGeneralist (@GeneralistAI) posted an uncut video of its model performing two tasks back-to-back: unzipping a pencil pouch and then retrieving money from it.
Updated
Updated
Items with an AI score under 20 are hidden. Show low-relevance items
AIGeneralist (@GeneralistAI) posted an uncut video of its model performing two tasks back-to-back: unzipping a pencil pouch and then retrieving money from it.
AIAir's latest release lets users open several projects in one window and run agents across them in parallel, with tasks grouped by project in the sidebar. Markdown files now render as formatted text while editing, with syntax shown only when editing, and Chinese, Japanese, Korean, and other IMEs now work on Windows. The release also adds a Customize screen for keymap, theme, and accent color, and lets users choose the agent and model for Agent Review.
AIUnsloth has released 1-bit quantized versions of Qwen3.8-27B that run on 8GB of RAM while retaining about 77% of BF16 accuracy. The team originally hesitated to publish them but was surprised by how well they performed in internal testing. The release accompanies new Qwen3.8-27B GGUFs that the company says deliver 10% higher accuracy.
AIUnsloth released new Qwen3.8-27B GGUF quantizations built with Unsloth Dynamic v3, which it says gain about 10% top-1% accuracy at the same size. The accuracy was measured with the new Divergence-300 metric, which extends top-1% greedy accuracy to 32 tokens using 300 unseen examples from Terminal Bench and DeepSWE. Unsloth also released 1-bit quants that it says run in 6–8GB, with 8GB RAM cited for running them.
AIUnsloth released new Qwen3.8-27B GGUF quantizations built with Unsloth Dynamic V3, which it says outperform others by more than 10% on Div-300, KLD, and other benchmarks. The release also includes 1-bit quants that retain 77% accuracy and can run on 8GB RAM.

AIGoogle released google/tipsv1-so400m14, the original v1 So400m/14 checkpoint of TIPS, a contrastive vision-language model that produces spatially rich image features aligned with text embeddings. The model has 413M vision parameters and 448M text parameters at 448 resolution, and is licensed under Apache 2.0.
AILiquid AI released 4-bit Q4_0 GGUF checkpoints for LFM2.5-230M, LFM2.5-350M, LFM2.5-1.2B-Instruct, and LFM2.5-2.6B, trained with Quantization-Aware Distillation. The company says the checkpoints recover most accuracy lost to quantization, reaching roughly 97% of their BF16 averages while keeping Q4_0 memory footprint and throughput. Benchmarks compare them against post-training quantized Q4_0 GGUFs and against Q5_K_M, Q4_K_M, and Unsloth's UD-Q4_K_XL.
Why it matters: The post shows how quantization-aware distillation recovers accuracy lost in Q4_0 checkpoints, with throughput measured across four hardware backends for deployment tradeoffs.
AIUnitree Robotics debuted on the Shanghai Stock Exchange STAR Market at 1,100 yuan per share, 629.44% above its 150.8-yuan issue price, with a total market value of 444.9 billion yuan. The company's founder, Wang Xingxing, has been known for taking unconventional positions, including low-cost in-house core components and a skeptical view that data alone will not produce embodied intelligence.
AICursor's update lets cloud agents subscribe to PRs, Slack threads, and scheduled tasks, and wake when something happens. It also adds custom modes that pin a skill in chat, subagents that run on their own virtual machines, and a /goal command for long-lived objectives. Users can also send steering messages while an agent works, with follow-ups applied at the next tool call.
Why it matters: The release lists concrete agent controls such as event subscriptions, custom modes, subagent VMs, and /goal, showing how cloud agents may run longer tasks with less manual steering.
AIOpenRouter announced it is joining forces with Stripe, saying its product, name, mission, and roadmap will remain the same. The company says it processes more than 10 trillion tokens per day from over 400 AI models for a community of over 10 million developers and companies. The transaction is subject to customary closing conditions and is expected to close in the coming weeks.
Why it matters: The announcement states that OpenRouter's product, roadmap, and mission stay unchanged after the Stripe deal, which clarifies what existing developers should expect.
AISequoia Capital argues that companies increasingly should build and own their AI capabilities, citing open-weight models such as Kimi K3 and GLM 5.2 that now approach frontier performance. Owning intelligence can protect margins as inference costs scale, speed up small distilled models, and keep proprietary data in-house, the article says.
AISentence Transformers now supports Answer.AI's ColBERT model, which has only 30M parameters, for local embedding index creation and querying in Python. The update comes with Sentence Transformers v6.0, which adds MultiVectorEncoder for ColBERT-style late interaction models alongside dense, sparse, and reranker models.
AIZhipu GLM has made the GLM-5.3 API generally available, targeting coding, defensive cybersecurity, and long-horizon agentic tasks. Pricing matches GLM-5.2, and access is offered through the official API and partner model gateways.

AIGoogle Labs has opened a waitlist for CC, its experimental AI productivity agent in Gmail, in Australia and New Zealand, and is expanding availability in the US and Canada. CC now helps manage calendars by connecting to Gmail and automatically creating events in a dedicated Google Calendar that update as plans change. Invitations to waitlisted users in the US and Canada begin rolling out today.
AIReplit has launched Free Mode, a new way to use its Agent that lets Core subscribers create up to 30x more on their monthly subscription, with everyday tasks no longer consuming credits. Free Mode is powered by OpenAI's GPT-5.6 Luna and is available to Core and Pro users until they reach usage limits that reset every 5 hours. Core subscribers also receive up to 30 hours per month of chat, and the company is offering the plan for $20 per month.
AIEtched has shipped its first rack to Jane Street, days after raising $700M at a $21B valuation. Tri Dao welcomed the milestone, emphasizing that more inference compute is always welcome.
AIVercel is offering up to $1,000,000 in a public hacker challenge testing its Vercel Sandbox against escapes from the Firecracker microVM and bypasses of the host-side network boundary. Rewards reach $50,000 per report, administered through HackerOne (@Hacker0x01). The company says agents can now exploit vulnerable sandbox boundaries, so it is testing its own defenses in the open.
AIStability AI has released two beta tools for Stable Audio 3.0: a plugin that brings audio generation into digital audio workstations (DAWs) and an upgraded experience at StableAudio.com with more editing options. Both are powered by commercially-safe models, so users own their outputs and can distribute them freely. Some features are experimental, and the company says it will keep iterating in real time.
AICursor's blog describes Continuity, its Git storage system, which stores each push as a write-ahead log entry in S3-compatible object storage. The article contrasts this design with GitHub's earlier Spokes system, which used three-phase commit replication across local disks. Continuity uses stateless replicas that catch up from the log, and the article reports write throughput of up to 120 pushes/s on S3 Standard and over 300 pushes/s on S3 Express One Zone.
Why it matters: The article explains why hosting Git at scale is hard and how Continuity's WAL-based design compares with the earlier Spokes approach, which is useful background for infrastructure work.
AIDaniel Han says Qwen3.8-27B is drawing more usage than any open model Unsloth has released, exceeding the prior most-liked GGUFs, Qwen3.6-35B-A3B at 1.54K likes and DeepSeek-R1 at 1.12K. Unsloth's companion post reports the Qwen3.8-27B GGUF is the #2 trending model on Hugging Face with 2.7M downloads.
AIChip Huyen asks what a good model tiering system looks like, since she is tired of naming specific models per vendor for her agent orchestrator. She wants to instruct the orchestrator by task tier, such as "use models tier ..." for a given kind of task, instead of listing Claude, OpenAI, and other models individually.
AIMicrosoft Foundry now offers structured outputs, web search, web fetch, MCP connector, and tool search for Claude models on Azure-hosted deployments. Prompts and completions remain within Azure for these deployments, while only usage metadata and safety-flagged content egress to Anthropic. The features were previously available only on Hosted on Anthropic deployments, which required choosing between capability and data-handling commitments.
Why it matters: The post shows which agent scaffolding now runs on Azure-hosted Claude deployments, which matters for teams needing data residency without rebuilding search, fetch, or tool routing.
AIReplit now offers black-box pen tests that scan deployed apps over the network and browser, with no access to source code. A Level 3 scan runs them alongside the existing white-box code scan, and the source notes the two catch different kinds of flaws.
Why it matters: The post explains how black-box scans test an app like an outside attacker, showing why source-code review alone misses some exposed doors.
AIMicrosoft says AI is being used to listen to underwater sounds at scale, identifying southern resident orcas in real time and alerting mariners when whales are nearby. The approach aims to help slow vessels and steer clear of the roughly 75 remaining southern resident orcas in the Salish Sea.

AIMark Chen, who is affiliated with OpenAI, said the company signed on for more than 4 GW of capacity with NVIDIA. He called this the scale that frontier training demands. The post quotes a Jensen Huang post, but the source text gives no further detail on terms or timing.
AIOpenAI has agreed to use capacity at the PORTS-Pike Technology Data Center in Pike County, Ohio, alongside SB Energy, NVIDIA, and the U.S. Department of Energy. SB Energy will pay the full cost of grid upgrades and transmission lines, and the project uses closed-loop, air-cooled systems expected to use significantly less water than the historical Portsmouth gaseous diffusion plant.
AIDiG-bench, a 70-game benchmark for discovering hidden rules through interaction, shows Opus 5 and Fable 5 with Claude Code performing best overall, with GPT-5.5 next. Only Opus 5 and Fable 5 beat any Tier 7 tasks, at a 0.2 success rate, while humans reached 100% on the same tests. The authors say the benchmark's games are mostly kept private to avoid training contamination.
AINVIDIA says it is partnering with SB Energy to secure land, power and shell capacity at the PORTS-Pike Technology Campus in Portsmouth, Ohio, to host its compute with OpenAI as tenant. The initial deployment is expected to provide 4.25 gigawatts of AI factory capacity, which NVIDIA estimates could represent about 1.5 million GPUs and $150 billion to $200 billion in revenue per generation. NVIDIA says it is supporting the site for roughly 4 gigawatts over a 20-year term, with its support limited to defined portions of lease and power payments.
AICursor begins rolling out Origin, its code hosting feature, in early beta to all paid plans, excluding enterprise orgs whose admins opt out. Repos can be hosted on Origin, where Origin is the source of truth, or synced from GitHub, where GitHub stays the source of truth and pull requests sync both ways. Vercel, Depot, and Buildkite integrations are already available, and agent-native features are slated to ship soon.
Why it matters: The source specifies how Origin hosts repos alongside GitHub sync, showing how the hosting source of truth differs between the two types of repo.
AIReplit announced enterprise governance updates including more than 50 audit log events that can stream to SIEM tools like Datadog and Splunk. It also launched a beta Admin API for pulling usage, workspace, member, and project data, with workspace settings for company-wide policies and team-level exceptions. Some features are available now, while workspace settings roll out at the end of the week and the Compliance API at the end of August.
AISebastian Raschka walks through building an AI text detector that returns a 0–100 score by fine-tuning a DistilBERT classifier. The project serves as an educational case study in how AI checkers work, their cat-and-mouse limitations, and how a verifier can be used alongside LLMs.
AIAugment rebuilt the Auggie CLI harness as v2, forking the open-source Pi coding harness and moving its context engine into Pi's extension system. On SWE-bench Pro at the same pass rate, Auggie v2 completes a task for $1.27 versus $2.70 for Claude Code, which is 53% cheaper. The gains come mainly from a narrower tool surface, one bash tool plus read, edit, and write, and from codebase retrieval that reduces exploration turns.
Why it matters: The post traces the design trade-offs behind each harness choice and ties them to measured token and cost differences, useful for anyone weighing agent tool surfaces.
AIAndrew Ng's team released an AI Engineering Skills Map, built from analysis of over 10,000 job postings and expert interviews, identifying four priority skills. The skills are building and deploying AI applications, software engineering fundamentals, using coding agents, and shaping the build. Ng says these skills matter for all developers, not only those with the AI Engineer title.
AIAli Ghodsi says Smart Routing on Databricks' AI Gateway lowers costs by about 30% without sacrificing quality. The quoted Databricks post says it matches each coding task to the right model and harness based on task needs, so higher-cost models focus on intelligence while lower-cost models compete on cost and performance.
AICursor has been acquired by SpaceX, completing a process that began in April when the two companies announced a partnership to accelerate model training. The post says the deal gives Cursor access to what it calls the largest GPU fleet in the world, which it expects to yield more capable models at lower cost. It cites Grok 4.6, released Wednesday, as an early look at what the companies can build together.
Why it matters: The post confirms a completed acquisition and links it to GPU access and cheaper model serving, which explains why the deal matters for coding tools.
AIZ.ai announced that an initial group of partners is now offering GLM-5.3-powered services through its official service, with its safeguards and usage policies in place. The company says it will expand partner access through a consistent, responsible process and share updates publicly.
AIZ.ai says GLM-5.3 is available now through GLM Coding Plan and ZCode. API access and open weights will follow in stages after rigorous safety evaluations.
AIAli Ghodsi attributes Databricks' 80% growth at $7B to enterprise AI agents becoming usable, now that Genie Ontology automates the capture of organizational context. He says over 70% of queries on the platform are now generated by Genie agents, and that this usage drives consumption and revenue.
AIDatabricks has made Smart Routing available in Unity AI Gateway to improve coding agent quality and cost. It matches each coding task to a suitable model and harness based on task needs while preserving good cache hit rates. Databricks says this can match frontier quality while cutting task costs by 30% or more.
AIAnthropic says Claude Tag now sends unprompted messages about 45% less often, based on its internal data. The update makes Claude more context-aware across users' work and better at knowing when to stay quiet. Monitoring is included at no extra cost.