Modellen & Concurrentie
GPT-5.6 en Work Agent. OpenAI heeft GPT-5.6 uitgebracht in drie redeneringsniveaus en ChatGPT Work gelanceerd voor urenlange autonome taken, maar gebruikslimieten en een verwarrende interface veroorzaakten klachten, wat leidde tot beloftes van oplossingen.
- → OpenAI launches its new family of models with GPT-5.6
- → The new GPT-5.6 family: Luna, Terra, Sol
- → OpenAI rolls out GPT-5.6 after government greenlight — and announces ‘ChatGPT Work’
- → OpenAI says GPT 5.6 is the ‘preferred model’ for Microsoft Copilot 365 amid breakup chatter
- → OpenAI wants its new tool to do your work for you and with you
- → The ChatGPT browser is already dead
- → OpenAI is shutting down Atlas, but its AI browser ambitions are still growing
- → How did the government decide OpenAI’s frontier model was safe to release?
- → OpenAI pairs its GPT-5.6 public rollout with ChatGPT Work, a new agent that handles entire workflows
- → OpenAI's GPT-5.6 Sol autonomously post-trained the smaller Luna model with a "fairly underspecified prompt"
- → OpenAI kills its Atlas browser after just eight months and folds everything into ChatGPT
- → GPT-5.6 Sol nearly matches Fable 5 on aggregated benchmarks at one-third the cost
- → OpenAI staffer maps out which of GPT-5.6 Sol's five reasoning levels fits which task complexity
- → OpenAI admits it "didn't get everything quite right" with ChatGPT Work launch and scrambles to fix UX and costs
Claude Fable 5 presteert het beste op benchmarks. Anthropic's Claude Fable 5 domineerde de Artificial Analysis benchmarks, maar de hoge kosten leidden ertoe dat het bedrijf aanraadde het te gebruiken als planner die taken delegeert aan goedkopere modellen, wat de kosten met ~40% vermindert terwijl het meeste vermogen behouden blijft.
Meta's Agentic Blitz en privacynasleep. Meta lanceerde Muse Spark 1.1, een hoog-contextueel coderingsmodel dat concurrenten onderbiedt op prijs, terwijl de Muse Image-generator werd teruggetrokken nadat het ongeautoriseerd gebruik van Instagram-gezichten mogelijk maakte, wat het debat over toestemming opnieuw aanwakkerde.
- → Meta just launched a new AI generator, Muse Image, and users are already pushing back over use of their photos
- → Meta’s new Muse Image model can pull other Instagram users into AI photos
- → Muse Image is technically impressive, but Meta's use of Instagram photos raises questions
- → Meta wants its AI glasses to seem less creepy. Its AI strategy says otherwise.
- → Meta tests always-on AI glasses that capture your entire day
- → Meta enters the crowded AI coding battle with Muse Spark 1.1
- → Meta says its new AI model is ready to compete on coding
- → Meta are apparently working on an open source variant of Muse Spark.
- → Introducing Muse Spark 1.1
- → Meta's Muse Spark 1.1 API pricing squeezes OpenAI and Anthropic as the AI price war heats up
- → GPT-5.6 🚀, Muse Spark 1.1 ✨, ChatGPT Work 💼
- → Meta removes controversial AI feature on Instagram after backlash
- → Meta turns off the Instagram feature that let users make AI deepfakes of public accounts
Chinese open-sourcegolf. Chinese modellen zijn nu goed voor meer dan 30% van het OpenRouter-verkeer; Tencent's HY3 en GLM‑5.2 draaien op eigen hardware, MiniMax plant een openbare release met 2,7T parameters, en DeepSeek's chipontwerp duidt op verticale integratie te midden van Amerikaanse bezorgdheid.
- → This is what Hy3 is capable of. Mother of god.
- → llama.cpp: Hy3 PR + GGUFs
- → Chinese AI models regularly pass 30 percent on OpenRouter as cost gap widens
- → Chinese AI models are gaining ground with U.S. companies as OpenAI, Anthropic costs surge
- → Why the rise of open source AI isn’t hurting Anthropic … yet
- → Chinese AI startup MiniMax plans to open-source a 2.7 trillion parameter model later this year
- → 4-bit GLM-5.2 (753B MoE) on 4× DGX Spark: 70.8% on Terminal-Bench 2.1 vs 81.0% for the full model
- → Running GLM 5.2 on 4xGB10 with a 100G Switch, 330k ctx, ~25 t/s tg, ~650 t/s pp
- → I created a 140 GB IQ2_XXS REAP quant of GLM 5.2 for coding. Looking for testers.
- → Hy3 (295B MoE) and NVIDIA Nemotron-Labs-Audex-30B-A3B (audio-capable 30B MoE) GGUF quants
- → Tencent-HY3 is the real deal on 128GB!
- → The U.S. tech industry is increasingly anxious about the rising power and competitive price of open-source AI models from China — and whether the Trump administration will respond with yet another executive order | Politico
- → China's DeepSeek developing its own AI chip, sources say
Agent Engineering & Infrastructuur
Bun herschreven door 64 agents. Claude Fable 5 coördineerde 64 parallelle instanties om Bun van Zig naar Rust te herschrijven in 11 dagen, wat resulteerde in meer dan een miljoen regels code en het oplossen van 128 bugs tegen een kostprijs van $165K.
SWE‑Bench Pro audit flopt. OpenAI ontdekte dat ~30% van de SWE‑Bench Pro-taken kapot waren en trok zijn goedkeuring in, wat eerdere zorgen van Artificial Analysis weerspiegelde en de evaluatie van agentische codering ondermijnde.
Modal haalt $355M op. Modal heeft $355 miljoen veiliggesteld om cloudinfrastructuur te bouwen die speciaal is ontworpen voor AI-agents, met focus op sandboxed, snelle iteratieomgevingen voor autonome software.
Cloudflare tijdelijke agents. Cloudflare introduceerde tijdelijke accounts waarmee AI-agents Workers kunnen inzetten zonder permanente inloggegevens, met vervaldata als ze niet worden opgeëist, waardoor agent-gestuurde workflows soepeler verlopen.
Agents op lokale GPU's. Nieuwe kwantiseringstechnieken en kosteneffectieve GPU-opstellingen draaien nu MoE-modellen op consumentenhardware, maar benchmarks tonen aan dat kwantiseringen met lagere bits agentische taken ernstig kunnen verslechteren, wat oproept tot een zorgvuldige selectie van kwant.
- → Qwen3.5 122B is the best?
- → Qwen3.6-27b does not understand software architechure.
- → Qwen 3.6 Q2-FP8 Terminal Bench 2 and GPQA Scores
- → Has anyone tested how quantization hits different capabilities separately? My results are surprising.
- → 2.5x faster Qwen3.6 NVFP4 Unsloth quants
- → Ultra budget 20GB vram with 448GB/s for $100 bucks.
- → I benched quad 5060Tis for code generation with Qwen3.6-27B so you don't have to (it's really good)
Beveiliging, Juridisch & Interpreteerbaarheid
Jacobian Lens brengt modelgeest in kaart. Anthropic's Jacobian Lens onthult een kleine set van verwoordbare representaties die de kernredenering van een model vastleggen, waardoor hallucinatiedetectie, outputsturing en, zoals communitydemo's laten zien, onmiddellijke creatie van schadelijke varianten mogelijk wordt.
- → Anthropic Research - "Verbalizable Representations Form a Global Workspace in Language Models"
- → Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens
- → I tested Anthropic’s new Jacobian Lens on open models, then it turned into a local-model hallucination router
- → Anthropic found a hidden space where Claude puzzles over concepts
- → I created a super harmful model ! :D (by tweaking it's J-Space!!!)
Agent-gestuurde ransomware. Sysdig documenteerde een AI-agent die de technische uitvoering van een ransomware-aanval afhandelde, hoewel een mens de operatie opzette, wat de huidige capaciteiten en beperkingen van agentische malware benadrukt.
AI-misbruikcrises. Een rechtszaak beweert dat xAI's Grok CSAM-afbeeldingen genereerde, terwijl Boko Haram chatbots gebruikt voor aanvalsplanning; de incidenten onderstrepen de dringende behoefte aan veiligheidsmaatregelen nu agentische capaciteiten toenemen.
Apple spant rechtszaak aan tegen OpenAI. Apple heeft een federale rechtszaak aangespannen waarin het OpenAI beschuldigt van het wegkapen van meer dan 400 werknemers om hardwarebedrijfsgeheimen te stelen, mogelijk om een rivaliserende AI-smartphone te bouwen, wat de juridische strijd in de AI-industrie intensiveert.
Dat was het weekoverzicht - tot volgende week zondag.