Model & Agent Releases
Grok 4.6 and Grok Bot. SpaceXAI released Grok 4.6, which ties OpenAI's GPT-5.6 Sol and trails only Claude Opus 5 on the Artificial Analysis index while undercutting frontier prices, and introduced Grok Bot, an always-on AI teammate that works from its own cloud computer to complete assigned tasks.
Qwen 3.8 large and small. Qwen3.8-2.4T-A95B is now available, and a 27B Qwen3.8 page briefly appeared on ModelScope before being pulled, indicating a near-term release. Local-run enthusiasts are already planning hardware splits for the 2.4T model.
DeepSeek V4 Pro update. DeepSeek's V4 Pro 0813 is available via API on OpenRouter, with no official announcement page yet; open weights are likely given prior releases. Its reasoning behavior differs noticeably across low, medium, and high reasoning levels.
Edge vision models. Liquid AI's LFM2.5-VL-3B brings much stronger UI grounding and can run on phones, while Cohere's Apache-2.0 North Micro Vision offers native-resolution image support in a compact 2.4B-parameter package.
Nvidia's trillion-param model. Nvidia is reportedly building Nemotron 4 with at least one trillion parameters and tripling cloud training spend to $28 billion through 2031, aiming to compete with Chinese open-weight models.
Agent Development & Tools
Managed agents and BYOC. LangChain launched Managed Deep Agents as the next step after frameworks, and made LangSmith BYOC generally available on AWS, letting enterprises keep agent traces and data inside their own VPC.
ChatGPT app on Linux. OpenAI released a Linux desktop app bundling ChatGPT, ChatGPT Work, and Codex, but native computer use is not available at launch.
Local Inference & Hardware
RTX PRO 6000 price doubled. Nvidia nearly doubled the MSRP of the 96GB RTX PRO 6000 Blackwell to $16,000, after pre-orders were below $8,000 last year.
DSpark streaming on one GPU. A single RTX PRO 6000 96GB can run DeepSeek V4 Flash 284B at around 31 tokens per second with DSpark's drafter in system RAM, about 15-17% faster than the no-drafter baseline at matched VRAM usage.
Muse Glimmer on Mac. Meta's Muse Glimmer 30B now runs up to about 3.3x faster on Apple Silicon using mlx-dspark, delivering 8-bit quality at near 4-bit speeds on a 48GB Mac.
Gemma 4 KV cache quantization. Quantization-aware training helps Gemma 4 31B retain quality under aggressive KV cache quantization, making large-context inference more practical.
Industry & Business
Cognition raising at $40B. AI coding agent maker Cognition is reportedly in talks to raise at a $40 billion valuation after reaching a $1 billion annualized revenue run rate, up from $492 million three months ago.
Lovable hits $13.3B. Vibe-coding startup Lovable raised $400 million at a $13.3 billion valuation after passing $500 million in annualized revenue, with 60 million hosted projects.
Thrive Holdings enterprise AI. OpenAI-backed Thrive Holdings raised $2 billion at a $12 billion valuation to acquire traditional businesses and implement AI, expanding beyond accounting and IT.
Gemini share drops. Market data from Pangram, OpenRouter, and Similarweb shows Gemini usage declining, with Pangram recording a drop to 1.9% of AI writing while OpenAI holds over 50%.
Claude heads into legal. Anthropic hired legal AI founder Robert Mahari as its first Head of Claude for Legal, building on earlier legal plugins and partnerships.
Policy, Security & Society
LiteLLM supply chain leak. A supply-chain attack on LiteLLM exposed terabytes of credentials from over 2,500 organizations, including Microsoft, Amazon, and Samsung, via malicious PyPI packages.
Claude output watermarks. Anthropic now watermarks Claude outputs to comply with the EU AI Act, upsetting users who worry it will reveal AI use in jobs or school.
Twitch defaults to training. Twitch will train Amazon AI models on streamer content by default, with an opt-out toggle now available; the policy has sparked backlash.
ShieldFont replaces text. Designers released ShieldFont, a font that displays normal text to users but substitutes words in the underlying HTML to poison scrapers.
Hinton, Li, Ng on open AI. At Ai4, Geoffrey Hinton, Fei-Fei Li, and Andrew Ng argued for keeping AI open to avoid a few companies controlling access, despite safety concerns about open weights.
Rare books destroyed. Booksellers fear AI firms are buying rare books and destroying them for training, despite non-destructive scanning alternatives.
Research & Benchmarks
MindTopo topology reasoning. Microsoft Research's MindTopo benchmark shows multimodal models are good at static topological recognition but struggle with interactive planning that requires maintaining relationships over time.
Prompts reconstructed from output. IIT Bombay and Adobe Research developed Previous-Token Prediction, which can reconstruct LLM prompts from output text with near-perfect accuracy without model weights.
That's everything for today - about a 5-minute read.