Modelli e concorrenza
GPT-5.6 e Work Agent. OpenAI ha rilasciato GPT-5.6 in tre livelli di ragionamento e ha lanciato ChatGPT Work per attività autonome di lunga durata, ma i limiti di utilizzo e un'interfaccia confusionaria hanno suscitato lamentele, spingendo a promesse di correzioni.
- → OpenAI launches its new family of models with GPT-5.6
- → The new GPT-5.6 family: Luna, Terra, Sol
- → OpenAI rolls out GPT-5.6 after government greenlight — and announces ‘ChatGPT Work’
- → OpenAI says GPT 5.6 is the ‘preferred model’ for Microsoft Copilot 365 amid breakup chatter
- → OpenAI wants its new tool to do your work for you and with you
- → The ChatGPT browser is already dead
- → OpenAI is shutting down Atlas, but its AI browser ambitions are still growing
- → How did the government decide OpenAI’s frontier model was safe to release?
- → OpenAI pairs its GPT-5.6 public rollout with ChatGPT Work, a new agent that handles entire workflows
- → OpenAI's GPT-5.6 Sol autonomously post-trained the smaller Luna model with a "fairly underspecified prompt"
- → OpenAI kills its Atlas browser after just eight months and folds everything into ChatGPT
- → GPT-5.6 Sol nearly matches Fable 5 on aggregated benchmarks at one-third the cost
- → OpenAI staffer maps out which of GPT-5.6 Sol's five reasoning levels fits which task complexity
- → OpenAI admits it "didn't get everything quite right" with ChatGPT Work launch and scrambles to fix UX and costs
Claude Fable 5 in testa ai benchmark. Claude Fable 5 di Anthropic ha dominato i benchmark di Artificial Analysis, ma il suo costo elevato ha portato l'azienda a raccomandare di usarlo come pianificatore che delega a modelli più economici, riducendo le spese del ~40% mantenendo la maggior parte delle capacità.
Blitz agentico di Meta e ricadute sulla privacy. Meta ha lanciato Muse Spark 1.1, un modello di codifica ad alto contesto che sottoprezza i concorrenti, mentre il suo generatore di immagini Muse è stato ritirato dopo aver permesso l'uso non autorizzato di volti di Instagram, riaccendendo i dibattiti sul consenso.
- → Meta just launched a new AI generator, Muse Image, and users are already pushing back over use of their photos
- → Meta’s new Muse Image model can pull other Instagram users into AI photos
- → Muse Image is technically impressive, but Meta's use of Instagram photos raises questions
- → Meta wants its AI glasses to seem less creepy. Its AI strategy says otherwise.
- → Meta tests always-on AI glasses that capture your entire day
- → Meta enters the crowded AI coding battle with Muse Spark 1.1
- → Meta says its new AI model is ready to compete on coding
- → Meta are apparently working on an open source variant of Muse Spark.
- → Introducing Muse Spark 1.1
- → Meta's Muse Spark 1.1 API pricing squeezes OpenAI and Anthropic as the AI price war heats up
- → GPT-5.6 🚀, Muse Spark 1.1 ✨, ChatGPT Work 💼
- → Meta removes controversial AI feature on Instagram after backlash
- → Meta turns off the Instagram feature that let users make AI deepfakes of public accounts
Onda open-source cinese. I modelli cinesi ora rappresentano oltre il 30% del traffico su OpenRouter; HY3 di Tencent e GLM‑5.2 girano su hardware domestico, MiniMax pianifica un rilascio aperto da 2,7T parametri, e il design dei chip di DeepSeek segnala un'integrazione verticale in mezzo alle preoccupazioni statunitensi.
- → This is what Hy3 is capable of. Mother of god.
- → llama.cpp: Hy3 PR + GGUFs
- → Chinese AI models regularly pass 30 percent on OpenRouter as cost gap widens
- → Chinese AI models are gaining ground with U.S. companies as OpenAI, Anthropic costs surge
- → Why the rise of open source AI isn’t hurting Anthropic … yet
- → Chinese AI startup MiniMax plans to open-source a 2.7 trillion parameter model later this year
- → 4-bit GLM-5.2 (753B MoE) on 4× DGX Spark: 70.8% on Terminal-Bench 2.1 vs 81.0% for the full model
- → Running GLM 5.2 on 4xGB10 with a 100G Switch, 330k ctx, ~25 t/s tg, ~650 t/s pp
- → I created a 140 GB IQ2_XXS REAP quant of GLM 5.2 for coding. Looking for testers.
- → Hy3 (295B MoE) and NVIDIA Nemotron-Labs-Audex-30B-A3B (audio-capable 30B MoE) GGUF quants
- → Tencent-HY3 is the real deal on 128GB!
- → The U.S. tech industry is increasingly anxious about the rising power and competitive price of open-source AI models from China — and whether the Trump administration will respond with yet another executive order | Politico
- → China's DeepSeek developing its own AI chip, sources say
Ingegneria degli agenti e infrastrutture
Bun riscritto da 64 agenti. Claude Fable 5 ha orchestrato 64 istanze parallele per riscrivere Bun da Zig a Rust in 11 giorni, producendo oltre un milione di righe di codice e correggendo 128 bug a un costo di 165.000 dollari.
SWE‑Bench Pro audit fallisce. OpenAI ha scoperto che circa il 30% dei compiti di SWE‑Bench Pro è rotto e ha ritirato la sua approvazione, echeggiando precedenti preoccupazioni di Artificial Analysis e minando la valutazione del coding agentico.
Modal raccoglie 355 milioni di dollari. Modal ha ottenuto 355 milioni di dollari per costruire un'infrastruttura cloud appositamente progettata per agenti AI, concentrandosi su ambienti sandbox e di iterazione rapida per software autonomo.
Cloudflare agenti temporanei. Cloudflare ha introdotto account temporanei che permettono agli agenti AI di distribuire Workers senza credenziali permanenti, con scadenze se non reclamati, semplificando i flussi di lavoro guidati da agenti.
Agenti su GPU locali. Nuove tecniche di quantizzazione e configurazioni GPU economiche ora eseguono modelli MoE su hardware consumer, ma i benchmark mostrano che quantizzazioni a bit più bassi possono degradare gravemente i compiti agentici, sollecitando un'attenta selezione della quantizzazione.
- → Qwen3.5 122B is the best?
- → Qwen3.6-27b does not understand software architechure.
- → Qwen 3.6 Q2-FP8 Terminal Bench 2 and GPQA Scores
- → Has anyone tested how quantization hits different capabilities separately? My results are surprising.
- → 2.5x faster Qwen3.6 NVFP4 Unsloth quants
- → Ultra budget 20GB vram with 448GB/s for $100 bucks.
- → I benched quad 5060Tis for code generation with Qwen3.6-27B so you don't have to (it's really good)
Sicurezza, giuridico e interpretabilità
Jacobian Lens mappa la mente del modello. Jacobian Lens di Anthropic rivela un piccolo insieme di rappresentazioni verbalizzabili che catturano il ragionamento centrale di un modello, consentendo rilevamento di allucinazioni, orientamento dell'output e, come mostrano le demo della community, creazione immediata di varianti dannose.
- → Anthropic Research - "Verbalizable Representations Form a Global Workspace in Language Models"
- → Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens
- → I tested Anthropic’s new Jacobian Lens on open models, then it turned into a local-model hallucination router
- → Anthropic found a hidden space where Claude puzzles over concepts
- → I created a super harmful model ! :D (by tweaking it's J-Space!!!)
Ransomware guidato da agenti. Sysdig ha documentato un agente AI che ha gestito l'esecuzione tecnica di un attacco ransomware, sebbene un umano abbia impostato l'operazione, evidenziando le attuali capacità e limiti del malware agentico.
Crisi di abuso dell'IA. Una causa sostiene che Grok di xAI ha generato immagini CSAM, mentre Boko Haram usa chatbot per la pianificazione di attacchi; gli episodi sottolineano l'urgente necessità di misure di sicurezza mentre le capacità agentiche scalano.
Apple fa causa a OpenAI. Apple ha intentato una causa federale accusando OpenAI di aver sottratto oltre 400 dipendenti per rubare segreti commerciali hardware, potenzialmente per costruire uno smartphone AI rivale, intensificando le battaglie legali nel settore dell'IA.
Questo è il riepilogo della settimana - ci vediamo domenica prossima.