Model Releases & Benchmarks
Gemini 3.8 Flash. Google shipped its third Flash model in six weeks, with improved reasoning, coding, and agentic tool use. Pricing per token is unchanged but the model may consume more tokens; a Cyber variant is available to trusted defenders through the new Fairwind Program.
- → Google releases Gemini 3.8 Flash, its third Flash model in six weeks
- → Introducing Gemini 3.8 Flash and 3.8 Flash Cyber
- → Proactive cyber defense for governments and enterprises
- → Proactive cyber defense for governments and enterprises
- → Gemini 3.8 Flash is Google's third budget model in six weeks while frontier models remain MIA
- → Google says its new Gemini 3.8 Flash model ‘works harder’ but might cost more
- → llm-gemini 0.34
Muse Spark 1.3. Meta's open-weights model now ranks third on AAII, matching GPT-5.6-Sol, and offers over 90% cheaper training if users opt in. Open weights are promised, and users are waiting for smaller variants between Glimmer and Spark.
GLM-5.3 vs DeepSeek-V4. Community comparisons on Mac Studio and DGX Spark find GLM 5.3 Flash faster to conclusions and better at writing and vision, while DeepSeek V4 Flash offers more tokens per second and longer context. Vision support and GGUFs are now available for DeepSeek-V4-Flash-Vision-Exp.
Local AI & Open Source
Qwen3.8 Flash Next tuning. Expert cache PR improves decode from 17 to 25-29 t/s on dual 3090s, and new AP quants plus Q8 N-gram bolting show no speed loss. Users also report plan hallucinations and context corruption, suggesting preview quality.
Perplexity Lily. Perplexity open-sourced its Mac inference server optimized for Qwen 3.6 on Apple silicon, focusing on a single model for best performance.
Local speech tools. VoxGen brings Rust/Vulkan TTS for VoxCPM2 with smooth AMD performance, Microsoft released VibeVoice ASR streaming, and users share Mac Parakeet plus Whisper/Parakeet as a working STT setup for coding.
GLM Minecraft mod. A local GLM 5.3 Flash Q4 on four RTX PRO 6000 cards iteratively built a black hole rifle mod for Minecraft using the Fabric API.
Research & World Models
World Labs Atlas. World Labs unveiled Atlas, a single model that generates, reconstructs, and simulates 3D scenes from just a few photos, using spatial context to anchor every input in 3D space.
H3-World language control. H3-World composes character and camera actions into textual instructions injected via MiniMax-H3’s text pathway, enabling controllable game motion with only 8,000 samples and 0.199% trainable parameters.
Safety, Policy & Legal
Astra safety fears. OpenAI's upcoming Astra uses opaque recurrence, making chain-of-thought harder to monitor, and is rated 'critical' for cyber capabilities. Safety experts warn of a race to the bottom.
DOJ backs fair use. The US government sided with OpenAI in the NYT lawsuit, arguing that training LLMs on copyrighted text is fair use and distinct from outputs, to preserve American AI leadership.
Secret AI safety rules. A lawsuit seeks to force the Trump administration to reveal its secret framework for pre-release frontier AI safety reviews, which currently lacks public or congressional transparency.
NYC student AI ban. New York City banned AI for students through eighth grade and companion chatbots in all grades, with limited high school use and AI literacy classes.
Tumbler Ridge suits. Thirty new lawsuits accuse OpenAI of aiding and abetting the Tumbler Ridge school shooting, alleging the company ignored safety flags about the suspect's ChatGPT conversations.
Business & Enterprise
Palo Alto buys Console. Palo Alto Networks paid $500 million for AI agent startup Console, which automates IT help desk tasks, to add autonomous security outcomes to its Cortex platform.
HiddenLayer raises $100M. AI security startup HiddenLayer raised $100 million as enterprises spend more to protect AI models, agents, and workflows from adversarial attacks.
Wonderful hits $5B. AI OS startup Wonderful raised $550 million, doubling its valuation to $5 billion, expanding from customer service agents to coordinating workflows across 35 countries.
Enterprise agent scaling. Schneider Electric, Vodafone, and monday.com built infrastructure layers for observability, evaluation, and deployment after finding agents easy to prototype but hard to operate.
Alexa scam detection. Amazon's Alexa for Shopping can now compare a message against Amazon's sent records to confirm if it's genuine, helping users spot impersonation scams.
That's everything for today - about a 5-minute read.