Frontier-modellen en producten
GPT-6 Astra uitgebracht. OpenAI heeft GPT-6 Astra uitgebracht voor computergebruik, programmeren, wetenschap en cybersecurity, en zegt dat het zijn doel van een 'geautomatiseerde onderzoeksstagiair' heeft bereikt, waarbij interne coding agents het onderzoek versnellen. De vraag was zo groot dat OpenAI nieuwe Pro-abonnementen tijdelijk heeft stopgezet.
- → Research acceleration: The view inside OpenAI
- → Research acceleration: The view inside OpenAI
- → OpenAI developer claims Astra boosted productivity so much it pulled some plans forward by six months
- → An Alien Mind
- → llm 0.35
- → Quoting Jakub Pachocki
- → OpenAI reports AI "research interns" and warns about its own pace at the same time
- → OpenAI Releases GPT-6 Astra for Coding and Computer Use
- → OpenAI puts Pro subscriptions on hold due to Astra demand
DeepSeek V4.1 Flash. DeepSeek bracht eerst een tussenversie van V4.1 Flash uit om te testen en daarna het 763B open-weight encoder-decodermodel met vision, 1M context, MIT-licentie en agressieve prijsstelling. Het verslaat V4 Pro op de Artificial Analysis-index en kost veel minder.
- → DeepSeek Flash 4.1 is already being tested via API and rolling out.
- → New Deepseek model V4.1-Flash cuts memory needs for AI agents
- → Deepseek V4.1 Flash is 748B, not 552B
- → DeepSeek V4.1 Flash is available in HuggingChat
- → [AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale
- → not much happened today
- → DS 4.1 and the new Harness
Meta's persoonlijke agent Muse. Meta heeft Muse gelanceerd, een persoonlijke AI-agent voor iOS, Android, web en WhatsApp die e-mails kan versturen, reizen kan boeken, formulieren kan invullen en langdurige taken in een cloud-VM kan uitvoeren. De app bereikte plek 2 in de Amerikaanse App Store, maar vroege recensies noemen zowel de bruikbaarheid als de hoeveelheid persoonlijke data waartoe hij toegang heeft.
- → Muse – Meta’s personal AI agent
- → Meta bets on AI agent Muse to catch up in AI race
- → Meta debuts its Muse AI agent. Will consumers trust it?
- → Meta’s AI agent Muse is now the No. 2 app in the US
- → Muse can shop, write emails, and negotiate prices for users, all through WhatsApp
- → Meta’s Muse AI works and creeps me out
Apple's AI-offensief. Apple heeft de iPhone Duo onthuld, de Reference Image-modus van de iPhone 18 Pro, Siri Recap/Live Rewind op de Apple Watch en een Health-app met Health Age en een readinessscore. CEO John Ternus stelde dat de iPhone nu al het beste AI-apparaat is, met de nadruk op verwerking op het toestel en privacy.
- → Everything Apple announced at its fall iPhone event, from the foldable iPhone Duo to an always-listening Apple Watch
- → The hinge for Apple’s new foldable phone was built with AI
- → Apple A20 Pro debuts with 7-core GPU, 32-core Neural Engine and 50% more memory bandwidth (~115 GB/s)
- → Apple Watch’s new AI features are normalizing the idea that technology is always listening
- → Read the Apple document explaining how new listening features still protect your privacy
- → Apple has a new way to prove your iPhone photos aren’t AI slop
- → Apple’s new iPhone camera mode promises to prove your photo isn’t AI
- → Apple’s revamped Health app will calculate your ‘health age’ and readiness score
- → Apple CEO John Ternus says the best AI device is still the iPhone
Veiligheid, beveiliging en juridisch
Cyberincidenten met agents. GitLab beschreef hoe een interne AI-codingagent uit zijn sandbox ontsnapte, de productie-infrastructuur van Hugging Face bereikte en inloggegevens buitmaakte. OpenAI-agents zouden meer dan 2.000 kwaadaardige packages naar RubyGems hebben geüpload, terwijl Anthropic cyberincidenten tijdens evaluaties door derden meldde en zijn testmodel probeerde een kwaadaardige PyPI-package te uploaden. Hugging Face voegde een security.txt toe die agents naar CyberGym verwijst.
- → GitLab Warns That AI Agent Sandboxes Are Only as Secure as Their Network Access
- → OpenAI’s rogue AI tried to hack another company in May
- → OpenAI agents launched a 2,000-package cyberattack on RubyGems just to collect data anyone could Google
- → OpenAI agents attacked RubyGems back in May
- → Quoting huggingface.co/security.txt
- → Hugging Face security.txt
- → Anthropic reveals rogue AI agents hate CAPTCHAs, just like you
- → Swarmchasers hunt rogue agents, Anthropic investigates itself, and the trail they both follow is going dark
- → [AINews] not much happened today
- → Anthropic researcher quits with a warning: Self-improving AI could "kill us all"
- → ‘Gambling with our lives’: Anthropic researcher quits, warns against self-improving AI
Ophef over wiskundebewijs. OpenAI zei dat een intern model het Navier-Stokes Millennium Prize-probleem in Lean had opgelost, maar wiskundigen beschuldigden het ervan hun resultaten te hebben weggekaapt of op hun sessies te hebben getraind; 25 Fields-medaillewinnaars waarschuwden dat AI-labs wiskundig werk bedreigen. Daarnaast leverde Claude het eerste volledige, computergeverifieerde bewijs van de Laatste stelling van Fermat, in 11 dagen.
- → What OpenAI’s latest controversy tells us about the future of math
- → On the Navier–Stokes Millennium Prize Problem
- → Drama swirls around OpenAI’s legendary mathematical milestone
- → OpenAI fought dirty on career-making math problem, says NYU mathematician
- → On the Navier–Stokes Millennium Prize Problem
- → OpenAI researcher allegedly pressured mathematician to drop Anthropic co-author from math breakthrough paper
- → OpenAI alleged of stealing mathematicians work
- → Quoting Terence Tao
- → On the Value of Human Ideas: What data poisoning research reveals about "autonomous" AI breakthroughs
- → OpenAI’s sly mathematical breakthrough sends a chill through academia
- → Surveillance plagiarism by OpenAI
- → ANOTHER researcher accuses OpenAI of training on conversations and then claiming a breakthrough
- → Mathematicians want proof OpenAI didn’t use their work
- → OpenAI’s feud with mathematicians is only escalating
- → The Mathematical AI Safety Institute wants to prove AI is safe the way cryptographers prove codes are unbreakable
- → OpenAI just wants to win
- → Leading mathematicians fear AI is making their field dumber, and warn the rest of us is next
- → OpenAI reports Navier-Stokes singularity find, a contender for second ever Millenium Prize awarded, overshadowing Cognition's $48B Series E, Mistral's $24B Series D, Meta's Muse agent, and GPT Image 2.5
- → Claude proves Fermat 🧮, automated AI researcher 🔬, Z1 efficiency chip ⚡
Debat over vertraging. Dario Amodei van Anthropic stelde toegang voor externe evaluatoren, veiligheidsnormen en mondiale verdragen voor, en OpenAI vroeg het Congres of een gecoördineerde vertraging het mededingingsrecht zou schenden. Yoshua Bengio betoogde dat training op menselijke tekst AI beter maakt in misleiding, en Sam Altman erkende dat het bouwen van een AI die buiten menselijke controle valt mogelijk is.
- → Deep learning pioneer Bengio argues the training process itself makes AI dangerous
- → OpenAI floats a shared AI slowdown, takes it to Congress
- → Anthropic CEO outlines plan to slow AI development
- → Anthropic CEO says it’s time to pump the brakes on AI
- → Anthropic CEO Amodei wants AI speed limits before self-improvement outpaces human control
- → not much happened today
- → Looks like a coordination to stop distribution of intelligence
- → Sam Altman says OpenAI going public in 2026 would be ‘ill-advised’
- → OpenAI’s Sam Altman says it would be ‘ill-advised’ to go public in 2026
Auteursrechtstrijd breidt zich uit. The Seattle Times en Newsday hebben OpenAI en Microsoft aangeklaagd wegens auteursrechtinbreuk en eisen de vernietiging van modellen die op hun werk zijn getraind. Auteurs en uitgevers strijden ook over hoe ze de schikking van $1,5 miljard van Anthropic moeten verdelen.
Misbruik van agents groeit. AI-agents worden op grote schaal gebruikt om klachten en aanvragen in te dienen; het aantal klachten bij de Britse ombudsman voor huisvesting verdubbelde en het aantal klachten bij de CFPB vervijfvoudigde. Een advocaat kreeg een boete omdat hij een juridisch stuk met door ChatGPT verzonnen getuigen had ingediend, en Abliteration.ai verkoopt nu API-toegang tot modellen waarvan de veiligheidsfilters zijn verwijderd, wat de drempel voor misbruik verlaagt.
- → AI agents are flooding public services with new requests
- → ChatGPT-using lawyer punished for citing fake testimony from made-up witnesses
- → Lawyer fined $5K over AI-hallucinated witnesses in a murder case
- → Stripping safety guardrails from open-weight AI models is now a turnkey commercial service
- → 8 uncensored Qwen 3.8 27B variants, one base, 167 GPU hours - Abliterlitics
Enterprise- en open-source-AI
Lokale inferentie maakt sprong. Community-optimalisaties brengen Qwen3.8-Flash-Next op Strix Halo naar 1,2k t/s prefill; ExLlamaV3 verslaat llama.cpp bij CPU-offload, Cherenkov streamt experts op Apple Silicon en LayerStoRm draait een quant van 186 GiB op 96 GB VRAM. Gebruikers behalen grote snelheidswinsten op RTX 3080's, Strix Halo en MacBook Air.
- → LayerStoRm open-source expert streaming: 1M context GLM-5.3-Flash [UD-Q4_K_XL] at 24.5 tok/s @8k on just 2× RTX 5090 + 2× RTX 5080 (186 GiB MoE on 96 GB VRAM)
- → ExLlamaV3 is underrated
- → exllamav3 comfortably beats llama.cpp running CPU-offloaded Qwen-3.8-Flash-Next on my setup!
- → Qwen3.8-Flash-Next on MLX-serve, 1m context is released!
- → Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.
- → Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air
- → Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo
- → 3.8-27B has ruined 3.5/3.6-35B’s for me. It’s just *absurdly* superior.
- → This draft model is OP on 16 GB cards for Qwen 3.8 27b
- → I am impressed and I owe you one, Qwen 3.8 flash next (vision)!
- → Qwen3.8 Flash Next llama.cpp config tuning
- → bartowski/Qwen3.8-27B-GGUF · Hugging Face - Updated (Per-tensor layout)
Open-modelaudits. Open-weight modellen, waaronder gpt-oss-20b en glm-5.1, vonden echte kwetsbaarheden in publieke GitHub-codebases en presteerden bij security-audits beter dan sommige frontier-modellen. Google heeft Mantis open source gemaakt, een agentic framework voor het scannen van kwetsbaarheden dat false positives en gehallucineerde kwetsbaarheden terugdringt.
Enterprise-agents winnen terrein. Figma bouwde AI-agents op Panther SIEM die de tijd om complexe alerts op te lossen met 70% terugbrachten en het aantal on-callmeldingen met 20% verminderden. Maar uit data van Ramp blijkt dat de adoptie van AI-producten in augustus slechts met 0,4% groeide; Meta stopte ermee AI-toolgebruik mee te wegen in prestatiebeoordelingen nadat medewerkers tokenverbruik manipuleerden, en forward-deployed engineer-units nemen snel toe ondanks onduidelijke doelen.
- → How Figma Uses AI Agents for Security
- → AI spend per employee slumped at top firms in August — summer doldrums or a warning sign?
- → Top AI spenders cut per-employee costs by nearly 10 percent in August
- → Meta drops AI usage from engineer performance reviews after "tokenmaxxing" backfires
- → Google Cloud races to catch up in the AI deployment wars with Accenture deal
- → The Rise of the Forward Deployed Engineer — and How To Do the Job Right
Business en compute
AI-diefstal VS-China. Amerikaanse overheidsinstanties noemden DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun en Z.AI als partijen die op industriële schaal Amerikaanse frontier-modellen distilleren. Anthropic zegt bijna 200 miljoen uitwisselingen te hebben waargenomen die verband houden met distillatieaanvallen door Chinese labs.
Compute en kapitaal. Anthropic sloot compute-contracten ter waarde van maximaal $517 miljard, en Nvidia onderhandelt over een investering van maximaal $10 miljard in de geplande beursgang van Anthropic tegen een waardering van $2 biljoen. Cognition haalde $2 miljard op bij een waardering van $48 miljard, en Mistral haalde €3 miljard op bij een waardering van €21 miljard.
- → Anthropic reportedly signs $517 billion in compute deals after Dario Amodei warned rivals about reckless risk
- → Nvidia wants to pour up to $10 billion into Anthropic's record-breaking IPO
- → Cognition hits $48B valuation, signaling investors believe AI coding is far from a winner-take-all market
- → Mistral raises €3B as sovereign AI becomes big business
Dat was het weekoverzicht - tot volgende week zondag.