Modelle & Wettbewerb
GPT-5.6 und Work Agent. OpenAI veröffentlichte GPT-5.6 in drei Denkstufen und startete ChatGPT Work für stundenlange autonome Aufgaben, aber Nutzungsbeschränkungen und eine verwirrende Benutzeroberfläche riefen Beschwerden hervor, was Versprechungen von Korrekturen zur Folge hatte.
- → OpenAI launches its new family of models with GPT-5.6
- → The new GPT-5.6 family: Luna, Terra, Sol
- → OpenAI rolls out GPT-5.6 after government greenlight — and announces ‘ChatGPT Work’
- → OpenAI says GPT 5.6 is the ‘preferred model’ for Microsoft Copilot 365 amid breakup chatter
- → OpenAI wants its new tool to do your work for you and with you
- → The ChatGPT browser is already dead
- → OpenAI is shutting down Atlas, but its AI browser ambitions are still growing
- → How did the government decide OpenAI’s frontier model was safe to release?
- → OpenAI pairs its GPT-5.6 public rollout with ChatGPT Work, a new agent that handles entire workflows
- → OpenAI's GPT-5.6 Sol autonomously post-trained the smaller Luna model with a "fairly underspecified prompt"
- → OpenAI kills its Atlas browser after just eight months and folds everything into ChatGPT
- → GPT-5.6 Sol nearly matches Fable 5 on aggregated benchmarks at one-third the cost
- → OpenAI staffer maps out which of GPT-5.6 Sol's five reasoning levels fits which task complexity
- → OpenAI admits it "didn't get everything quite right" with ChatGPT Work launch and scrambles to fix UX and costs
Claude Fable 5 führt Benchmarks an. Anthropics Claude Fable 5 fegte die Artificial-Analysis-Benchmarks, aber seine hohen Kosten veranlassten das Unternehmen, es als Planer zu empfehlen, der an günstigere Modelle delegiert, was die Ausgaben um etwa 40% senkt und die meisten Fähigkeiten beibehält.
Metas Agentic Blitz und Datenschutz-Fallout. Meta startete Muse Spark 1.1, ein hochkontextuelles Codierungsmodell, das die Konkurrenz preislich unterbietet, während sein Muse-Bildgenerator zurückgezogen wurde, nachdem er die unbefugte Nutzung von Instagram-Gesichtern ermöglichte, was die Debatten über Einwilligung neu entfachte.
- → Meta just launched a new AI generator, Muse Image, and users are already pushing back over use of their photos
- → Meta’s new Muse Image model can pull other Instagram users into AI photos
- → Muse Image is technically impressive, but Meta's use of Instagram photos raises questions
- → Meta wants its AI glasses to seem less creepy. Its AI strategy says otherwise.
- → Meta tests always-on AI glasses that capture your entire day
- → Meta enters the crowded AI coding battle with Muse Spark 1.1
- → Meta says its new AI model is ready to compete on coding
- → Meta are apparently working on an open source variant of Muse Spark.
- → Introducing Muse Spark 1.1
- → Meta's Muse Spark 1.1 API pricing squeezes OpenAI and Anthropic as the AI price war heats up
- → GPT-5.6 🚀, Muse Spark 1.1 ✨, ChatGPT Work 💼
- → Meta removes controversial AI feature on Instagram after backlash
- → Meta turns off the Instagram feature that let users make AI deepfakes of public accounts
Chinesische Open-Source-Welle. Chinesische Modelle machen inzwischen über 30% des OpenRouter-Verkehrs aus; Tencent HY3 und GLM‑5.2 laufen auf heimischer Hardware, MiniMax plant eine offene Veröffentlichung mit 2,7 Billionen Parametern, und DeepSeek's Chipdesign signalisiert vertikale Integration inmitten von US-Bedenken.
- → This is what Hy3 is capable of. Mother of god.
- → llama.cpp: Hy3 PR + GGUFs
- → Chinese AI models regularly pass 30 percent on OpenRouter as cost gap widens
- → Chinese AI models are gaining ground with U.S. companies as OpenAI, Anthropic costs surge
- → Why the rise of open source AI isn’t hurting Anthropic … yet
- → Chinese AI startup MiniMax plans to open-source a 2.7 trillion parameter model later this year
- → 4-bit GLM-5.2 (753B MoE) on 4× DGX Spark: 70.8% on Terminal-Bench 2.1 vs 81.0% for the full model
- → Running GLM 5.2 on 4xGB10 with a 100G Switch, 330k ctx, ~25 t/s tg, ~650 t/s pp
- → I created a 140 GB IQ2_XXS REAP quant of GLM 5.2 for coding. Looking for testers.
- → Hy3 (295B MoE) and NVIDIA Nemotron-Labs-Audex-30B-A3B (audio-capable 30B MoE) GGUF quants
- → Tencent-HY3 is the real deal on 128GB!
- → The U.S. tech industry is increasingly anxious about the rising power and competitive price of open-source AI models from China — and whether the Trump administration will respond with yet another executive order | Politico
- → China's DeepSeek developing its own AI chip, sources say
Agent Engineering & Infrastruktur
Bun von 64 Agenten neu geschrieben. Claude Fable 5 orchestrierte 64 parallele Instanzen, um Bun in 11 Tagen von Zig nach Rust umzuschreiben, wobei über eine Million Codezeilen produziert und 128 Fehler behoben wurden, bei Kosten von 165.000 $.
SWE‑Bench Pro Audit scheitert. OpenAI stellte fest, dass etwa 30% der SWE‑Bench-Pro-Aufgaben fehlerhaft waren, und zog seine Befürwortung zurück, was frühere Bedenken von Artificial Analysis widerspiegelt und die agentische Codebewertung untergräbt.
Modal sammelt 355 Millionen Dollar ein. Modal sicherte sich 355 Millionen Dollar, um eine speziell für KI-Agenten entwickelte Cloud-Infrastruktur aufzubauen, mit Schwerpunkt auf Sandbox-Umgebungen mit schnellen Iterationen für autonome Software.
Cloudflare temporäre Agenten. Cloudflare führte temporäre Konten ein, die es KI-Agenten ermöglichen, Workers ohne dauerhafte Anmeldeinformationen bereitzustellen, mit Ablauf, wenn sie nicht beansprucht werden, was agentengesteuerte Workflows erleichtert.
Agenten erreichen lokale GPUs. Neue Quantisierungstechniken und kosteneffektive GPU-Setups lassen MoE-Modelle nun auf Consumer-Hardware laufen, aber Benchmarks zeigen, dass niederbitige Quantisierungen agentische Aufgaben erheblich beeinträchtigen können, was eine sorgfältige Quant-Auswahl erfordert.
- → Qwen3.5 122B is the best?
- → Qwen3.6-27b does not understand software architechure.
- → Qwen 3.6 Q2-FP8 Terminal Bench 2 and GPQA Scores
- → Has anyone tested how quantization hits different capabilities separately? My results are surprising.
- → 2.5x faster Qwen3.6 NVFP4 Unsloth quants
- → Ultra budget 20GB vram with 448GB/s for $100 bucks.
- → I benched quad 5060Tis for code generation with Qwen3.6-27B so you don't have to (it's really good)
Sicherheit, Recht & Interpretierbarkeit
Jacobian Lens kartiert Modellverstand. Anthropics Jacobian Lens offenbart eine kleine Menge verbalisierbarer Repräsentationen, die das Kernreasoning eines Modells erfassen, was Halluzinationserkennung, Output-Steuerung und, wie Community-Demos zeigen, sofortige Erstellung schädlicher Varianten ermöglicht.
- → Anthropic Research - "Verbalizable Representations Form a Global Workspace in Language Models"
- → Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens
- → I tested Anthropic’s new Jacobian Lens on open models, then it turned into a local-model hallucination router
- → Anthropic found a hidden space where Claude puzzles over concepts
- → I created a super harmful model ! :D (by tweaking it's J-Space!!!)
Agentengesteuerte Ransomware. Sysdig dokumentierte einen KI-Agenten, der die technische Ausführung eines Ransomware-Angriffs übernahm, obwohl ein Mensch den Betrieb einrichtete, was die derzeitigen Fähigkeiten und Grenzen agentischer Malware verdeutlicht.
KI-Missbrauchskrisen. Eine Klage behauptet, dass xAIs Grok CSAM-Bilder generierte, während Boko Haram Chatbots zur Angriffsplanung einsetzt; die Vorfälle unterstreichen die dringende Notwendigkeit von Sicherheitsvorkehrungen, während die agentischen Fähigkeiten zunehmen.
Apple verklagt OpenAI. Apple reichte eine Bundesklage ein, in der es OpenAI beschuldigt, über 400 Mitarbeiter abgeworben zu haben, um Hardware-Geschäftsgeheimnisse zu stehlen, möglicherweise um ein konkurrierendes KI-Smartphone zu bauen, was die Rechtsstreitigkeiten in der KI-Branche verschärft.
Das war die Wochenrückschau - bis nächsten Sonntag.