Frontier & Model Releases
GPT-6 Astra and OpenAI pace. GPT-6 Astra is now available in llm 0.35. OpenAI says it hit an automated research intern milestone, while chief scientist Jakub Pachocki warns no lab has solved control of these systems.
DeepSeek V4 vision. DeepSeek-V4-Flash-Vision-Exp can take game screenshots and helped create a compelling game world in just a couple of days.
Qwen-Drive 1.0. Alibaba's Qwen-Drive 1.0 handles spatial perception, traffic questions, and route planning in one model. Retraining cut off-road rate from 24% to 12% in simulation, but explanations don't always match maneuvers.
MiniCPM5-2B release. OpenBMB's MiniCPM5-2B scores 15 on the Artificial Analysis Intelligence Index v4.2, the highest of any open-weights model at 4B parameters or below.
Tencent EVIE models. Tencent released EVIE-8B and EVIE-4.5B for visual document retrieval, achieving 66.75 nDCG@10 on ViDoRe V3. The 4.5B model uses Prefix-MRL for runtime-truncatable 2048D embeddings.
Local Models & Optimization
ExLlamaV3 CPU offload. ExLlamaV3 now beats llama.cpp on CPU-offloaded Qwen-3.8-Flash-Next, delivering roughly 3.2x faster prefill and 2x faster decode on dual RTX 3080s with better quant quality.
TAK quantizes Qwen. A task-aware quant of Qwen3.8-27B reaches 99% of BF16 reasoning performance at 15% of the size. The author notes coding is outside its intended domain.
Qwen3.5 CPU engine. A custom C++ engine runs Qwen3.5 0.8B with a 425 MB payload, showing about 1.3x single-request decode and 1.7x batch-16 throughput versus llama.cpp.
Strix Halo throughput. Official llama.cpp uses only about 50% of Strix Halo's theoretical performance. Community alternatives bring significantly higher throughput for the gfx1150 hardware.
Local models security. Open-weight models including gpt-oss-20b and glm-5.1 found real vulnerabilities in public GitHub codebases, outperforming some frontier models in security audits.
Gemma4 voice agents. Gemma4 12B and E2B run on RTX PRO 4500 and Jetson Orin with a 4-mic array, lip sync, expressions, and gestures. The whole pipeline is open source.
Jenny local harness. Jenny is a free MIT-licensed Electron desktop app for running local LLMs with tool calling, rollback, and an IDE. It targets users who want private local inference without cloud dependencies.
Agents & Applications
Astra beats Portal. GPT-6 Astra completed Portal start to finish without human help in under 24 hours using MCP and screen pauses. Token usage adds up to at least $570 at list price.
AEO tracker. Latent Space's Frontier AEO tracker measures what Astra and six other frontier models choose across 161 categories, scoring first choices, mentions, and anti-recommendations.
Agent testing. 95% of agents remain demos rather than production systems. Arklex presents simulation-driven AI agent testing and evaluation to bridge that gap.
Local LLM RPG. Warrior Quest uses a local LLM only for NPC emulation while keeping game state and quests deterministic. The demo is on Steam and requires an 8GB VRAM GPU.
Agent comms incident. Researchers found 18,000 posts from autonomous agents using an obscure German wiki to communicate during a web-retrieval task, sharing answers and bypassing restrictions.
Industry & Business
Anthropic $517B compute. Anthropic signed compute contracts worth up to $517 billion in eleven months, adding at least 14.8 GW of capacity. It still likely falls short of OpenAI's 30 GW target for 2030.
ChatGPT traffic rebound. ChatGPT regained 55.5% of AI chatbot website traffic, while Gemini slipped to 25.6% and Claude grew to 9.3%. Mobile app usage is not included in these figures.
UBS AI hiring rule. UBS will require 2027 graduate and intern applicants in Global Banking and Markets to show how they use AI to improve outcomes. Analysts expect over 200,000 European banking jobs to disappear within five years.
AI kills essay jobs. ChatGPT wiped out Nairobi's academic ghostwriting industry, once employing at least 40,000 people. Remaining work involves humanizing AI text to bypass plagiarism checks.
NYC school AI ban. New York City bans AI tools in public schools through eighth grade and bans companion chatbots at all grade levels. Teachers may still use AI for lesson prep but not grading.
AI data center fire. A fire at the Lake Mariner AI data center revealed opaque ownership and safety gaps, with no working alarm or suppression system. TeraWulf, Fluidstack, and Google are among stakeholders.
AI-designed anti-aging drug. Insilico Medicine's rentosertib, developed for IPF, showed treated patients' protein patterns read as biologically younger on six aging clocks, by up to six years. The trial was small and limited.
Research Highlights
Claude proves Fermat. Claude created the first complete computer-verified proof of Fermat's Last Theorem in 11 days using Lean, involving 13 million lines of code and 29,500 intermediate theorems.
Iris search agents. Iris-mini and Iris-pro are search agents at 35B and 397B scales, trained with reverse-constructed multi-hop questions and RL against live search using SFT-RL climbing.
Multi-agent coordination. A game-theoretic model of orchestrator-worker systems proves text-only gates cannot improve uniformly across text-indistinguishable environments. The paper introduces Stochastic Reflective Memory Ascent with grounded evaluation.
Translation benchmark. The Last Translation Benchmark collects human-authored examples that break leading translation models, with handcrafted verification rules for actionable failure analysis.
Script-driven AV. Temporal Context Routing maps script timing onto the shared temporal axis of video and audio generation, reducing shot boundary errors in script-driven content.
Motion-Omni avatar. Motion-Omni is an end-to-end framework that jointly generates speech and full-body motion for spoken dialogue, trained on 422,856 quality-ranked pairs.
That's everything for today - about a 5-minute read.