Agent & Model Releases
Muse Gadgets. Meta open-sourced firmware and an SDK for building DIY hardware that connects to its Muse agent, such as E Ink displays and HDMI sticks, with a Discord for support.
OpenAI Dot. OpenAI's new Dot agent acts like enterprise software that can also order dinner, running in a virtual machine with access to apps like Blender and GIMP; it's positioned as a coworker rather than a personal shopper.
Unitree humanoid model. Unitree released UnifoLM-WLA-1.0, a 6B model trained on ~2,500 hours of real robot data that handles 64 whole-body and tabletop tasks on the G1, combining optical flow and MMDiT for continuous control.
FrogNano for coding. Microsoft's FrogNano-4B-2609 is an agentic model derived from Qwen3.5-4B, post-trained on 1,500 synthetic SWE tasks with executable rewards to improve repository-level software engineering in a compact form.
Cloudflare Clef. Cloudflare released Clef and Clef-flash decision models that output probabilities over predefined answers, with median latency as low as 39ms, enabling agents to act without a human in the loop.
Tools & Frameworks
MegaCapybara engine. A purpose-built inference engine for RTX 5090 and Qwen3.8 27B claims roughly twice NInfer's decode speed, reaching 500+ t/s single and 2000+ t/s across 12 agents, with dynamic kernels and cache management.
llama.cpp updates. llama.cpp added support for decision models and a CUDA pull request that fuses shared experts to speed up MoE architectures like Qwen 35B A3B.
mlsubgen subtitles. A new local tool generates subtitles in 45 languages entirely on your own machine, detecting per-stretch language, reconciling two speech recognizers with a local LLM, and preferring existing embedded subtitles.
Research & Model Advances
Spotlight memory architecture. Percepta's Spotlight replaces attention with unbounded writable memory that the model indexes itself, separating intelligence from knowledge so capabilities can grow without changing weights.
GPT-6 updates. OpenAI published a GPT-6 family guide and announced DevDay updates including GPT-6.1 Sol, computer use for the Agents API, cloud Codex environments, and a Decisions API.
AstaBrief open-sourced. Hugging Face open-sourced AstaBrief, a small model trained specifically for fast, cited scientific report generation, designed to stay grounded in evidence and verify outputs.
Industry & Business
Apple tightens disk access. Apple is adding controls to macOS Full Disk Access due to AI agent risks, after Meta Muse allegedly read messages without permission; Meta says access is opt-in, and Apple wants very explicit user action.
Data center pushback. Amazon pledged $1B to data center communities and AWS CEO Matt Garman urged an end to moratoriums, claiming foreign disinformation; Microsoft is expanding biomimicry at data centers amid local opposition.
AI becomes 'super intelligence'. The White House got major tech CEOs to sign an AI safety pledge and Trump signed an executive order rebranding AI as 'super intelligence,' while only 2% of consumers are buying AI products.
OpenAI safety shakeup. OpenAI fired three researchers and a fourth departed after an investigation found they leaked confidential information to an outside AI safety organization.
Uber Eats speeds search. Uber rebuilt its search pipeline, cutting end-to-end latency by 50%, using agentic coding to identify and validate optimizations.
Airbnb's inside-out AI. Airbnb CTO Ahmad Al-Dahle is applying an inside-out AI strategy: use generative AI internally to speed product development, then transform the customer experience.
Local AI & Hardware
Strata + Qwen3.8 Next. Users report Strata with Qwen3.8 Next runs at 45-70 t/s on slow DDR4+7900XTX and 150-200 t/s on power-limited RTX 5090 with 96GB DDR5, beating llama.cpp significantly.
iPhone as second GPU. An iPhone 17 Pro Max can accelerate Qwen 3.8 27B prefill by 29-44% on a 24GB MacBook when connected via USB-C and running upper layers on the phone GPU.
Ascend 96GB cards. Two Huawei Atlas 300I Duo cards with 96GB each can run Qwen3.8 Flash-Next at ~30 t/s after software stack tuning, despite not being drop-in CUDA replacements.
DGX Spark and 5090 costs. Nvidia introduced a 64GB DGX Spark while the original 128GB model price rose to $6,950, and buying an RTX 5090 at Micro Center now reportedly requires no-export paperwork.
That's everything for today - about a 5-minute read.