Agent & Model Releases
Claude Fable 5.1. Anthropic released Fable 5.1 and Mythos 5.1, claiming better coding/research and up to 45% lower cost for agentic work via cheaper cache reads. The unrestricted Fable 5.1 is available now, though early cost analyses dispute the maximum savings at highest reasoning effort.
- → Claude Fable 5.1 made me a really nice animated pelican
- → Anthropic launches Claude Fable 5.1 and says it’s up to 45 percent cheaper for agentic work
- → Anthropic's Claude Fable 5.1 promises better coding and research at up to 45 percent less
- → Anthropic’s new Fable release is cheaper, less restrictive
OpenAI Astra rollout. OpenAI says its forthcoming Astra model is the first LLM to meet its critical cybersecurity threshold, able to find and exploit unknown flaws autonomously. The company delayed parts of Astra's development after the Hugging Face hack to strengthen safeguards and will limit access to its most advanced cybersecurity capabilities.
New Spark-X2.5 models. Open models Spark-X2.5-4B and 1.7B appeared on Hugging Face with a novel architecture, claiming performance rivaling Qwen 3.5 9B and native 1M context. They don't yet run in mainline llama.cpp but have a custom fork.
Runway Solaris interface model. Runway unveiled Solaris, an 'Interface World Model' that generates UI frames in real time instead of running code, responding to clicks and voice. It remains a research effort, with limitations in text rendering and screen reader support.
Local AI & Hardware
Qwen3.8 local reports. Community reports show Qwen3.8 27B and Flash-Next running across various hardware: 12 tok/s on a 48GB Mac, 280 tok/s decode on dual Epyc R9700s with MXFP4 kernels, and 132 decode tok/s on a single RTX 3090 after custom optimizations. Quant benchmarks suggest UD Q3_K_XL is a good fit for 16GB cards, while some users find Qwen3.8 a good coder but over-eager modifier.
- → Running 104GB Qwen3.8-Flash-Next on 48GB Mac at ~12 tok/s
- → How I got 280 tok/s on Qwen3.8 27B on 2xr9700's and 940k tokens kv cache
- → I pushed Qwen3.8-27B to 2.000 prefill per second and 132 decode per second on A RTX 3090.
- → Kaitchup posted Qwen3.8 27B Benchmarks for quants from Q4 to Q1
- → Everyone is t/s maxing.. 3.8.. but after a week of using it for work I'm tempted to switch back to 3.6
DGX Spark price hikes. Nvidia's DGX Spark and Asus Ascent GX10 saw significant price increases, with the Asus model jumping from $3999 to $5999. Buyers are questioning whether clusters of DGX Sparks are cost-effective versus AMD Epyc systems with more memory bandwidth, though the DGX offers FP4 support and tensor parallelism.
CMP 170HX failures. A user reports 2 of 5 CMP 170HX GPUs died within two weeks and a third had defective tensor cores, arguing current prices don't justify the risk for LLM inference.
INT8 and deceptive quants. Community discussion asks why INT8 W8A8 models aren't more common on RTX 3090s given native INT8 tensor cores, while another user accuses AtomicChat of mislabeling an IQ2_S quant as Q4_K_M.
Intel memory ambitions. Intel hinted it may re-enter the memory business, with its CEO mentioning hiring former SK Hynix executive Seok-Hee Lee and new memory architecture, though details are scarce.
Tools & Frameworks
OpenClaw 2.0 release. OpenClaw 2.0, the open-source personal AI agent, simplifies setup by detecting existing ChatGPT/Claude subscriptions and local models, and rebuilds the browser as primary interface. It also adds shared cloud sessions for collaboration.
Codex local tools. OpenAI's Codex desktop app now includes a full LibreOffice installation and other binaries like Poppler and git in its runtime, enabling document processing skills without external dependencies.
AI software factories. Top AI open source projects like Flue and tldraw are rejecting external PRs, instead using teams of agents to triage, reproduce, fix, and review issues. Vercel's AI SDK project deployed a multi-agent 'software factory' to cut a backlog of over 1,000 issues and 800 PRs.
Terraform for agents. HashiCorp positions HCP Terraform as the governed control plane for AI-driven infrastructure, letting agents author and apply Terraform changes while enforcing policy, identity, and audit controls.
datasette-mcp update. The datasette-mcp plugin released version 0.2, changing SQL query results from arrays to objects to help weaker models map columns correctly.
Research & Safety
BenchMIRT benchmark audit. Hugging Face introduced BenchMIRT, a method to audit LLM benchmarks at individual prompt level, revealing that prompts within a benchmark can measure different abilities and average scores may obscure that.
Gemini agentic video. Google DeepMind launched agentic video understanding across Gemini 3.7/3.6/3.5 Flash-Lite, using native video tools to dynamically search and inspect video segments, improving accuracy and reducing token usage for video analysis.
Claude text watermarking. Anthropic opened its watermark verification API to regulators, media, and fact-checkers, letting approved organizations detect Claude-generated text via a statistical pattern that may survive some editing, as required by the EU AI Act.
Election and bias issues. AlgorithmWatch found Google's election AI Overviews are inconsistent, rely on few sources, and sometimes describe candidates in partisan terms. Separately, Google's AI search gave emergency-call advice based on nationality, prompting fixes.
Industry & Business
AfterQuery's rapid rise. AI training-data startup AfterQuery reportedly raised at a $3.2B valuation, a 10x jump in five months, making it Y Combinator's fastest-ever unicorn. It employs professionals to train models on how to complete tasks like professionals would.
AIR secures $50M. AI security startup AIR came out of stealth with $50M to help companies discover agents and continuously vet the skills, MCP servers, and add-ons they use, blocking unapproved software interactions.
Google courts Hollywood. Google is reportedly seeking licensing deals with major studios to train AI models on copyrighted content, offering large payments, but analysts warn the deals may pose risks to studios' public standing.
Frontier AI focus. Google DeepMind's new chief Koray Kavukcuoglu said frontier AI leadership is the only thing that matters, admitting current models are a bit below frontier but vowing to close the gap, with no concrete updates on Gemini 4 or 3.5 Pro.
Agents in operations. OpenAI's Enterprise Signals shows frontier firms generate 8.3x more output tokens per active user, as startups Basis, Clay, and Exa Labs embed agents into onboarding, account management, and developer integrations.
Apple-OpenAI legal fight. Apple is seeking expedited discovery against OpenAI, alleging the company failed to inspect a former Apple employee's MacBook that contained discussions about destroying forensic data.
Google Pics design tool. Google launched Google Pics, an AI image creation and editing tool for Workspace, powered by Nano Banana, to compete with Canva and Adobe Express with prompt-based design and precise object editing.
Consumer AI agents. Fambot introduced an AI chief of staff for families to manage logistics; John Deere launched a JD AI assistant for farmers using their own data; Amazon added Alexa 'Update Me When' shopping alerts.
That's everything for today - about a 5-minute read.