Open & Local Model Releases
Ornith-1.5 open family. Ornith-1.5 released in 9B dense, 35B MoE, and 397B MoE under MIT with strong agentic and coding benchmarks; community quantizations show the 35B runs at 60 tk/s on a 12GB 4070 Ti and the 9B outperforms earlier small models.
Qwen3.8-27B field reports. Local users report Qwen3.8-27B can autonomously execute 80+ tool calls and handle long agentic chains, but its offline knowledge and strict coding scores are weaker than predecessors, so it works best with clear direction or a copilot.
Qwen3.8 speedups and quants. Community efforts pushed Qwen3.8-27B to 138 tps on an RTX 3090 with DFlash2, released Dynamic V3 quants claiming 10% higher accuracy, a depth-pruned 22.7B Mini-Me, and native FP4 on 2017 V100s matching a 5090.
- → I pushed Qwen3.8-27B limits again... Dflash2 - 134 tps on a RTX 3090
- → DFlash2 speeds Qwen 3.8 27B up to 4 times
- → Introducing Qwen3.8-27B Dynamic v3 Unsloth GGUFs
- → updated unsloth/Qwen3.8-27B-GGUF · Hugging Face
- → Qwen3.8-23B-Mini-Me: A Depth-Pruned Qwen3.8-27B (to ~22.7BB)
- → NVFP4 on VOLTA! Despite being built for Blackwell, I made four 2017 V100s run Qwen 3.8 NVFP4 natively and match my $6000 RTX 5090.
Ling base checkpoints. AntLing open-sourced base and midtrain checkpoints for Ling-3.0-tiny and flash without post-training, and a user pairs Ling tiny as an auxiliary for context compression alongside Qwen3.8.
LFM2.5 QAD quants. Liquid AI released Q4_0 checkpoints trained with quantization-aware distillation, recovering 97% of BF16 accuracy with the same speed and memory.
GLM-5.3 and frontier releases. GLM-5.3 ties Kimi K3 atop open-model rankings with large agentic gains and lower cost, though open weights are delayed two weeks; DeepSeek V4-Pro GA and NVIDIA Nemotron 3.5 Lightning/NeMo Switchyard also shipped.
Agent Tools & Product Updates
Secure sandboxes. Simon Willison tested smolmachines/smolvm as a sandbox for untrusted Python and JavaScript with CPU/RAM limits and no network; Jeremy Morrell argues LLM-authored extensions plus sandbox primitives enable safe extensible software.
Replit Free Mode. Replit introduced Free Mode powered by GPT-5.6 Luna, letting anyone build apps and agents without token-cost constraints.
Calendly AI notes. Calendly entered meeting note-taking with summaries, action items, and an upcoming Callie assistant that uses meeting context to schedule follow-ups.
Meta AI Mac app. Meta launched a Mac app for its AI chatbot with screen sharing, dictation, and Google Workspace integration to compete with Gemini, ChatGPT, and Claude desktop assistants.
WhatsApp on-device scam detection. WhatsApp is testing Scam Alert, an optional on-device model that classifies messages from non-contacts while using confidential federated analytics to preserve privacy.
Google study tools. Google rolled out a Gemini student hub, AI-generated interactive visuals and quizzes in Search, and Deep Research in Gemini Live for back-to-school.
Alexa+ free on Fire TV. Amazon made Alexa+ free on Fire TV without Prime, rolling out conversational search, smart home controls, and AI recommendations automatically.
Industry & Business
Stripe buys OpenRouter. Stripe confirmed a $7.5 billion acquisition of OpenRouter, and founders cited the 'singularity' as a reason to stay private and invest in agent infrastructure.
SpaceX-Cognition report denied. Bloomberg reported SpaceX attempted to acquire Cognition, but CEO Scott Wu denied the story, saying Cognition is not for sale and no talks occurred.
Anthropic revenue lead. Anthropic passed OpenAI on quarterly revenue for the first time, hitting $11.6B with a small operating profit while OpenAI's $6.7B came with deeper losses.
Compute as asset class. Nvidia is working with Apollo, BlackRock, and others on $500B in financing to treat compute as an investable asset; Silicon Data raised $30M to launch CME GPU futures.
Memory and chip supply. DRAM prices are up 500% in 12 months and hyperscalers have locked 2027 supply; China is allowing small H200 batches to ByteDance and Tencent to ease inference capacity.
Data center infrastructure. TerraPower plans a data center project using its Natrium reactor's energy storage, while Relativity Networks raised $22M for hollow-core fiber that transmits data 50% faster.
AI reputation worsens. Pew found 52% of Americans more concerned than excited about AI, and local data center deals are facing public pushback.
Safety, Privacy & Security
OpenAI private safety processing. OpenAI previewed Private Safety Processing for zero-data-retention customers, enabling abuse detection across interactions without exposing content to personnel; it contrasts with Anthropic's 30-day retention policy.
OpenAI fixes Codex and TAC. OpenAI shipped safeguards after Codex deleted real user files through faulty cleanup commands, and confirmed a separate error briefly revoked access to its Trusted Access for Cyber program.
AI-generated ICS exploits. U.S. agencies warned attackers are using AI to build Siemens S7 exploit scripts, cutting the skill and time needed to attack industrial control systems.
AI labs internal controls. Guidelight's first assessment found no AI company fully applies basic internal controls; Anthropic and OpenAI lead with C+, while xAI and Meta score D- and F.
Meta deepfake ads. Meta ran ads for an AI nudification app targeting female politicians despite policies banning sexual material.
Spirit employee data sale. Google won an auction for Spirit Airlines employment data; former flight attendants object that worker confidentiality may not be protected from re-identification.
OpenAI slows development. OpenAI paused some RL training and delayed frontier runs while tightening safeguards, testing whether a voluntary slowdown can work without industry-wide adoption.
Research & Analysis
Claude protein design. Anthropic's Claude models designed minibinders across 15 targets with a 26.8% hit rate, beating typical early drug-discovery results, though independent review is pending.
Vivodyne tissue data. Vivodyne's robotic HIVE labs grow 20 human tissue types and generate causal biological data to address what it sees as AI drug discovery's missing data problem.
Reasoning tokens myth. Research challenges the idea that LLM intermediate tokens are semantically meaningful reasoning; models often produce invalid traces even with correct answers, and corrupted-trace training can still work.
Coding productivity metric. Simon Willison argues lines of debugged production code remain a meaningful productivity metric with coding agents, but reaching that output still takes senior engineering skill.
That's everything for today - about a 5-minute read.