Botsing van frontiermodellen
Kimi K3 schudt frontiermodellen op. Moonshot AI's opengewicht Kimi K3 concurreert met top gesloten modellen, verslaat Claude Fable 5 op sommige benchmarks en herstart het debat over open versus gesloten.
- → Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI
- → Kimi K3, and what we can still learn from the pelican benchmark
- → [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing
- → Kimi K3 weights to be released on the 27th.
- → Kimi K3 ranks #1 on @AfterQuery's SpreadsheetBench 2, surpassing Claude Fable 5
- → Kimi: Threat or menace?
- → kimi.ai teasing a video with lots of 3's in it
- → [AINews] not much happened today
- → Kimi moment. I think the writing is on the wall for Anthropic and OpenAi
- → Kimi K3 is currently at the top of the leaderboard for Text Arena filtered for science queries.
- → Kimi K3 (max) beats Sonnet 5 on Simple Bench
- → Kimi K3 is top of nextjs eval
- → Just like Deepseek, China's Kimi K3 is forcing Western AI labs to question their compute advantage
- → Kimi K3 🌕, Gemini 3.5 delayed ⏳, crushing ARC-AGI 3 🤖
GPT-5.6: Kracht en gevaar. De GPT-5.6-familie biedt programmatische toolaanroep en parallelle subagents, maar de neiging om onbedoeld bestanden en databases te verwijderen benadrukt de noodzaak van sandboxing.
- → The Sequence Radar #893: Last Week in AI: GPT-5.6, Grok 4.5, Muse Spark 1.1 and the Post-Chatbot Stack
- → OpenAI’s new flagship model deletes files on its own, people keep warning
- → GPT-5.6 is deleting user files when given full access, and OpenAI says it shouldn't but did
- → [AINews] Codex usage up >10x in 6 months to 7M users, +1M in the past ~day; did Codex overtake Claude Code??
Claude Fable 5 verlengd. Onder druk van GPT-5.6 en Kimi K3 draaide Anthropic plannen terug en hield Fable 5 in betaalde abonnementen, waarbij het direct concurreert op toegankelijkheid.
DeepSeek V4 dreigt. DeepSeek hint op V4 met goedkope API en open gewichten, terwijl community-optimalisaties de flash-variant al op consumenten-GPU's op bruikbare snelheden laten draaien.
Agents betreden echte workflows
Browsing-agents arriveren. Anthropic gaf Claude Code een ingebouwde webbrowser met veiligheidsclassificatoren, en Cursor lanceerde een algemene agent, waarmee codeerassistenten richting autonome taakuitvoering bewegen.
Agents behandelen echte transacties. DoorDash's agentarchitectuur verhoogde conversies met 24%, en Stripe's benchmark toont aan dat agents integraties kunnen coderen maar vaak falen bij validatie, wat betrouwbaarheidslacunes benadrukt.
Productieagents komen ten val. De meeste enterprise-agents blijven chatbots; 54% van de bedrijven had beveiligingsincidenten met agents, en experts stellen dat agents cloud-native operationele primitieven zoals microservices nodig hebben.
- → How to Debug Coding Agents with LangSmith Traces
- → The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials
- → The AI context gap: Enterprise AI organizations have a trust problem, not a retrieval problem — and most are still building the fix
- → The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway
- → Cloud Native Infrastructure Emerges as the Foundation for Trustworthy Agentic AI
- → QCon AI Boston: Production AI Moves Beyond Prompts to Platforms, Harnesses, and Evals
AI-infrastructuur, financiering en de datacenterrace
Datacenters stuiten op weerstand. S&P verlaagde Oracle vanwege AI-kapitaalrisico, terwijl lokale protesten en New York's moratorium op hyperscale datacenters een groeiende tegenstand tegen AI-infrastructuur weerspiegelen.
AI-financiering stijgt. Databricks haalde een waardering van $188B op, Meta onderhandelt over een datacenterlease van $10B aan Anthropic, en Nous Research kreeg $75M voor zijn opensource-agent.
Inzet op energie en hardware. Energiebedrijven haalden sinds de dotcom-periode het meeste op om AI aan te drijven, en een lening van $400M gedekt door SambaNova-chips signaleert groeiende investeringen in niet-GPU AI-hardware.
Doorbraken in AI op apparaten
1-bit modellen bereiken telefoons. PrismML's Bonsai comprimeert Qwen3.6-27B tot 3,9GB terwijl het 90% van de benchmarkscores behoudt, wat toolaanroepende agents op het apparaat mogelijk maakt; naar verluidt is Apple in gesprek.
- → Bonsai 27B: The First 27B-Class Model to Run on a Phone
- → Bonsai 27B: 1-bit dense LLM running locally in your browser using custom WebGPU kernels
- → Prism-ML Bonsai Qwen 3.6 27B
- → PrismML’s new Ternary Qwen3.6 27B runs near fp16 precision on 10GB of memory!!!
- → So what's the consensus on 1bit models? Is it still a pipe dream?
- → PrismML Bonsai 27B is surprisingly usable on the Jetson Orin Nano 8GB
- → Is anyone having any luck with the Ternary Bonsai 27B DFlash?
- → Can we get a "not base model" flair?
- → Bonsai 27B runs locally on an iPhone - a 27B model in 3.9GB
- → User experience of Bonsai-Ternary-27B on 4060Ti 16GB for KB management and productivity assistant use cases
- → Bonsai 27B is a full open reasoning model that fits on an iPhone
- → Bonsai-27B & Ternary-Bonsai-27B - Updates (on PRs)
- → Apple in talks with startup PrismML that shrinks AI models to run on an iPhone
Thuislaboratoria draaien frontiermodellen. llama.cpp-optimalisaties zoals DFlash versnellen MoE-modellen tot 6x, en een OnePlus-telefoon streamde experts van flash om een 60GB-model te draaien op 1,3 tok/s.
- → I tested all llama.cpp's speculative decoding methods on Qwen 3.6 27B: MTP ~2.7x, DFlash ~3.7x, n-gram stack ~6x on real coding. Local AI win. My findings on RTX 6000 PRO.
- → DFlash makes Qwen3.6 27B 2.2x faster with no quality loss
- → GPT-OSS-120B, Qwen 30B and Gemma 26B on an Android phone at 1-5 tok/s: +60GB model, 11GB of RAM, CPU only
Beleid, rechtszaken en de opensource-impasse
Apple–OpenAI juridische strijd. Apple daagde OpenAI voor de rechter wegens bedrijfsgeheimen, met beschuldiging van het wegkapen van meer dan 400 werknemers, waaronder de chief hardware officer, wat OpenAI's IPO en cloudvertrouwen bedreigt.
- → Apple sues OpenAI after ex-engineer allegedly used bug to steal trade secrets
- → The wildest allegations in Apple’s trade secrets lawsuit against OpenAI
- → The 6 wildest claims in Apple’s lawsuit against OpenAI
- → OpenAI pushes back on Apple trade secret lawsuit
- → Sam Altman didn’t need another lawsuit
- → How Apple’s big lawsuit could disrupt OpenAI’s IPO plans
- → Apple’s plot to crush OpenAI
- → Apple’s lawsuit couldn’t come at a worse time for OpenAI
Chinese opensource-golf. Chinese modellen zijn nu goed voor 41% van HuggingFace-downloads, Kimi K3 concurreert met westerse leiders, en China lanceerde een parallel AI-bestuursorgaan zonder het Westen.
- → The real AI race may no longer be at the frontier
- → Source: the Trump administration and industry groups discussed streamlining US open model releases of equal or lesser capability to leading Chinese open models
- → Chinese President Xi Jinping speaks at World AI Conference and reaffirms commitment to open source to promote"openness and win-win"
- → China’s Xi Touts Open-Source AI and Takes a Swipe at U.S. Dominance
- → China's new World Artificial Intelligence Cooperation Organization is President Xi's clearest play yet for a parallel AI order
Eisen voor aansprakelijkheid nemen toe. Duitse toezichthouders hielden chatbots aansprakelijk voor inhoud, xAI daagde een gebruiker voor het genereren van CSAM, en Demis Hassabis stelde een FINRA-achtig orgaan voor voor frontiervelligheidsbeoordelingen.
Dat was het weekoverzicht - tot volgende week zondag.