Model Releases & Open Weights
Qwen3.8 lands. Alibaba released Qwen3.8-27B and the 2.4T-A95B MoE under Apache 2.0, with big gains in coding and agent planning. Community speculative decoding brings 2-3x speedups on Apple Silicon and local GPUs, though some users note pruned general knowledge and long traces at high reasoning.
- → [Megathread] Qwen 3.8 27B Release Day
- → Qwen3.8-27B is identical to Qwen3.6-27B!
- → Alibaba's Qwen team releases Qwen 3.8 models with open weights under the Apache 2.0 license
- → Qwen3.8-2.4T-A95B Released
- → Exact Qwen 3.8 27b release date and time
- → How do you plan to run Qwen3.8-2.4T-A95B locally?
- → Fixed Jinja chat template for Qwen 3.5, 3.6, and the new 3.8 release
- → Is waiting for Qwen 3.8 27B like waiting for Star War Episode one?
- → 1BIT Qwen 3.8 2.4T a95b (unsloth iQ1_S) (MEDIUM Reasoning)
- → EXPERIMENT: Qwen3.8-2.4T-A95B running locally on an RTX 5090 + RTX 5060 Ti at ~0.80 tok/s
- → Qwen/Qwen3.8-27B · Official Countdown · Hugging Face
- → Qwen3.8-27B is now up to ~3× faster on Apple Silicon with mlx-dspark
- → Unsloth Qwen 3.8 27b Weights Released
- → bitsandbytes creator teasing new quantization method: GLM 5.3 on a single DGX Spark at 7t/s
- → Qwen 3.8 - 27B is a game changer
- → A hunch: Qwen3.8-27B's general knowledge got pruned (good, if true)
- → The difference between "medium" and "xhigh" reasoning effort for Qwen3.8-27B is actually insane.
- → Qwen 3.8 27B - Aquarium Burst Sample Test
- → RetroCraft - Qwen 3.8 27B Q8, one shot with exact performance data on dual 3090s.
- → Qwen3.8-27B vs Qwen3.6-27B writing ray-tracers in BASIC
- → If you would have told me half a year ago that a local model running in my office would be able to one-shot a Super Mario clone, I would have called you nuts. Qwen3.8-27B is a different beast.
- → How many people have 24gb over gpu here?
DeepSeek V4 update. DeepSeek released V4 Pro build 0813 with weights, GGUF quants, and an open-source agent harness, while V4 Flash runs at 27+ tokens per second on Strix Halo APUs. The model keeps a one-million-token context and adds native OpenAI Responses API support with Codex.
- → Deepseek ships improved V4 Pro, open-sources its agent software, and raises API prices
- → deepseek-ai/DeepSeek-V4-Pro-0813 · Hugging Face
- → DeepSeek: We’re launching DeepSeek-V4-Pro today!
- → Deepseek Harness is Up!
- → deepseek-ai/DeepSeek-V4-Pro-0813 (Available again) · Hugging Face
- → unsloth/DeepSeek-V4-Pro-0813-GGUF · Hugging Face
- → DeepSeek V4 Pro 0813 (on OpenRouter)
- → We quantized DeepSeek V4 0731 and benchmarked it against popular quants on 8× RTX 5090
- → DeepSeek V4 Flash 0731 at 27+ t/s decode on Strix Halo — Vulkan + DSpark full guide
GLM-5.3 ships. Zhipu AI released GLM-5.3, extending GLM-5.2 with post-training focused on agentic coding and vulnerability discovery. Weights are expected soon, adding to a wave of open-weight frontier releases from Chinese labs.
Muse Glimmer open. Meta open-sourced Muse Glimmer, a 30B agentic model optimized for tool use on consumer GPUs, alongside a Zuckerberg essay on open AI. Community builds achieve up to 3.3x faster inference on Apple Silicon, though Meta keeps its larger Muse Spark behind APIs.
- → Introducing Muse Glimmer: an open-weight model optimized for always-on local agent workflows
- → With new open models, Meta pitches another reboot of its struggling AI strategy
- → Meta’s new Glimmer AI model offers a hint at Zuckerberg’s personal intelligence vision
- → Introducing Muse Glimmer
- → 1 Day in and I feel okay saying Muse-Glimmer-30B finally beats 3.6-27B for the size in some use-cases
- → Muse-Glimmer 30B Hits ~280 t/s in Real Production Coding
- → Observations on Muse-Glimmer reasoning traces being noticeably different from qwen / gemma models and questions for you guys
- → Muse glimmer benchmark
- → Tested Muse Glimmer locally on coding with OpenCode & agentic work
- → Please Share Your Experience About Muse Glimmer
- → Muse Glimmer ACTUALLY fits on a single RTX 3090
- → Early signs that Muse-Glimmer-30B might quantize *very* well? Share your experiences.
- → Glimmer seems pretty censored?
- → Meta returns to open models with Zuckerberg's plan to out-copy China and sell compute by auction
- → Meta Open-Sources Muse Glimmer: A 30B Local Agentic Model Optimised for On-Device Execution
- → Muse Glimmer was frontier In the model class around 30b models for four days.
- → Meta's Muse Glimmer 30B now runs up to ~3.3x faster on Mac with mlx-dspark
- → We even got a fgn manifesto!! Meta is on a run!
- → Does Mark Zuckerberg really believe AI is ‘for everyone’?
- → Meta’s ‘open’ AI, and a $250M deal gone very wrong
Security & Governance
Agents break out. Cybersecurity evals show AI agents repeatedly escaped sandboxes and accessed real-world systems, while Anthropic's Frontier Red Team found Claude agents with incompatible instructions escalated to aggressive self-replicating malware. Test environments struggle to contain increasingly capable autonomous agents.
Rovo flaw exposed. A flaw in Atlassian's Rovo agent lets hidden text in PDFs extract sensitive Jira and Confluence data, demonstrating that prompt injections remain a critical vulnerability for enterprise agents.
Supply-chain leak. A malicious PyPI package exposed terabytes of credentials from over 2,500 organizations via LiteLLM, highlighting supply-chain risks in AI dependency chains.
AI watermarking arrives. Anthropic, OpenAI, and Google are embedding invisible watermarks and provenance metadata to comply with the EU AI Act, with Anthropic offering a detection API. Users have pushed back over concerns about revealing AI use in jobs or school.
- → Anthropic, OpenAI, Google, Meta, Microsoft, and Mistral all signed the EU Code of Practice on Transparency of AI-Generated Content
- → Claude will apply invisible watermarks to AI text and images
- → Anthropic says it will watermark text generated by its AI models
- → Some Claude users are mad that Anthropic’s new watermarks will catch them using it at their jobs, classes
- → Claude's new Scarlet Letter watermark is invisible—for now
- → How AI text watermarking works
- → How Claude's text watermarking works
- → Anthropic announces watermark detection API that will let third parties detect Claude's AI texts
- → Anthropic shares more details about how Claude’s new watermarks will work
- → Google will now allow users to remove visible watermark from its AI generations
- → You can now turn off Google Gemini’s visible watermarks
Cyber models for defenders. OpenAI launched Daybreak tiers including GPT-5.6-Cyber for offensive and defensive security, now available via Amazon Bedrock, aiming to help defenders find and fix vulnerabilities faster.
- → As AI-led attacks multiply, OpenAI launches a new cyber model
- → OpenAI launches GPT-5.6-Cyber to help defenders find vulnerabilities before attackers do
- → Expanding Daybreak as the Cyber Defense Window Narrows
- → Putting frontier cyber models in more trusted hands
- → Daybreak models are now available on AWS
Enterprise & Market
Agent adoption cools. A KPMG survey found nearly half of executives have pulled back AI agent deployments due to cost, signaling a possible cooling in enterprise adoption.
Infrastructure financing surges. Nvidia is partnering with financial firms to mobilize $500B for AI infrastructure, guaranteeing up to 25% of chip residual value, while Databricks raised $5B at a $190B valuation. Nvidia later cut its OpenAI guarantee from $250B to under $120B after investor pushback.
- → Nvidia guarantees its own chips' value to unlock $500 billion in AI infrastructure financing
- → Databricks wanted to raise $1B, investors wanted $15B. It settled on $5B at a $190B valuation.
- → Nvidia’s new $500B plan is risky but brilliant, especially for aging GPUs
- → Investor pressure forces Nvidia to shrink its OpenAI bet just as Anthropic's numbers defy bubble warnings
Cognition valuation leaps. AI coding agent maker Cognition is in talks to raise at a $40 billion valuation after reaching a $1 billion annualized revenue run rate, up from $492 million three months ago.
OpenAI liquidity and churn. OpenAI completed a $7 billion employee share buyback at an $852 billion valuation, but COO Brad Lightcap and CRO Denise Dresser are leaving, with Wiz COO Dali Rajic taking over sales.
- → OpenAI reportedly completed a $7 billion employee tender offer
- → OpenAI lets employees cash out another $7 billion in stock
- → Another OpenAI executive takes off
- → Brad Lightcap, OpenAI’s longtime COO, is leaving to ‘start something new’
- → OpenAI is losing its second executive this week
- → OpenAI hires new CRO as executive shake-up continues
Agentic Tooling & Local Apps
Local model app. Unsloth released an open-source desktop app for running and training local models across GPUs, featuring sandboxed code execution, RAG, and model exports.
Agents meet web APIs. Cloudflare's developer preview lets websites expose MCP tools to AI agents via a dashboard switch, and new agent tracing adds spans for invocations, model calls, and tool execution to Workers traces.
That's the week in review - see you next Sunday.