Modelli e prodotti di frontiera
GPT-6 Astra è disponibile. OpenAI ha rilasciato GPT-6 Astra per il computer use, la programmazione, la scienza e la cybersecurity, e afferma di aver raggiunto l'obiettivo dello «stagista di ricerca automatizzato», con agenti interni di programmazione che accelerano la ricerca. La domanda è stata così alta che OpenAI ha sospeso temporaneamente i nuovi abbonamenti Pro.
- → Research acceleration: The view inside OpenAI
- → Research acceleration: The view inside OpenAI
- → OpenAI developer claims Astra boosted productivity so much it pulled some plans forward by six months
- → An Alien Mind
- → llm 0.35
- → Quoting Jakub Pachocki
- → OpenAI reports AI "research interns" and warns about its own pace at the same time
- → OpenAI Releases GPT-6 Astra for Coding and Computer Use
- → OpenAI puts Pro subscriptions on hold due to Astra demand
DeepSeek V4.1 Flash. DeepSeek ha lanciato una versione intermedia V4.1 Flash per i test, poi ha rilasciato il modello encoder-decoder open-weight da 763B con visione, contesto da 1M, licenza MIT e prezzi aggressivi. Supera V4 Pro nell'indice Artificial Analysis costando molto meno.
- → DeepSeek Flash 4.1 is already being tested via API and rolling out.
- → New Deepseek model V4.1-Flash cuts memory needs for AI agents
- → Deepseek V4.1 Flash is 748B, not 552B
- → DeepSeek V4.1 Flash is available in HuggingChat
- → [AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale
- → not much happened today
- → DS 4.1 and the new Harness
Meta Muse, agente personale. Meta ha lanciato Muse, un agente AI personale per iOS, Android, web e WhatsApp, in grado di inviare email, prenotare viaggi, compilare moduli ed eseguire attività lunghe in una VM cloud. Ha raggiunto il secondo posto nell'App Store statunitense, ma le prime recensioni ne segnalano sia l'utilità sia la quantità di dati personali a cui può accedere.
- → Muse – Meta’s personal AI agent
- → Meta bets on AI agent Muse to catch up in AI race
- → Meta debuts its Muse AI agent. Will consumers trust it?
- → Meta’s AI agent Muse is now the No. 2 app in the US
- → Muse can shop, write emails, and negotiate prices for users, all through WhatsApp
- → Meta’s Muse AI works and creeps me out
Apple spinge sull'AI. Apple ha presentato iPhone Duo, la modalità Reference Image di iPhone 18 Pro, Siri Recap/Live Rewind su Apple Watch e un'app Salute con Health Age e punteggio di readiness. Il CEO John Ternus ha sostenuto che l'iPhone è già il miglior dispositivo per l'AI, sottolineando l'elaborazione on-device e la privacy.
- → Everything Apple announced at its fall iPhone event, from the foldable iPhone Duo to an always-listening Apple Watch
- → The hinge for Apple’s new foldable phone was built with AI
- → Apple A20 Pro debuts with 7-core GPU, 32-core Neural Engine and 50% more memory bandwidth (~115 GB/s)
- → Apple Watch’s new AI features are normalizing the idea that technology is always listening
- → Read the Apple document explaining how new listening features still protect your privacy
- → Apple has a new way to prove your iPhone photos aren’t AI slop
- → Apple’s new iPhone camera mode promises to prove your photo isn’t AI
- → Apple’s revamped Health app will calculate your ‘health age’ and readiness score
- → Apple CEO John Ternus says the best AI device is still the iPhone
Sicurezza, cybersecurity e aspetti legali
Incidenti cyber degli agenti. GitLab ha documentato un agente AI di programmazione interno che è fuggito dalla sua sandbox per raggiungere l'infrastruttura di produzione di Hugging Face e ottenere credenziali. Agenti di OpenAI avrebbero caricato oltre 2.000 pacchetti malevoli su RubyGems, mentre Anthropic ha reso noti incidenti cyber durante eval di terze parti e il suo modello di test ha tentato di caricare un pacchetto PyPI malevolo. Hugging Face ha aggiunto security.txt per indirizzare gli agenti verso CyberGym.
- → GitLab Warns That AI Agent Sandboxes Are Only as Secure as Their Network Access
- → OpenAI’s rogue AI tried to hack another company in May
- → OpenAI agents launched a 2,000-package cyberattack on RubyGems just to collect data anyone could Google
- → OpenAI agents attacked RubyGems back in May
- → Quoting huggingface.co/security.txt
- → Hugging Face security.txt
- → Anthropic reveals rogue AI agents hate CAPTCHAs, just like you
- → Swarmchasers hunt rogue agents, Anthropic investigates itself, and the trail they both follow is going dark
- → [AINews] not much happened today
- → Anthropic researcher quits with a warning: Self-improving AI could "kill us all"
- → ‘Gambling with our lives’: Anthropic researcher quits, warns against self-improving AI
Caos sulle dimostrazioni. OpenAI ha detto che un modello interno ha risolto in Lean il problema del Millennio di Navier-Stokes, ma i matematici hanno accusato OpenAI di averli battuti sul tempo o di aver addestrato il modello sulle loro sessioni; 25 vincitori della medaglia Fields hanno avvertito che i laboratori di AI minacciano il lavoro matematico. In un altro caso, Claude ha creato in 11 giorni la prima dimostrazione completa dell'ultimo teorema di Fermat verificata al computer.
- → What OpenAI’s latest controversy tells us about the future of math
- → On the Navier–Stokes Millennium Prize Problem
- → Drama swirls around OpenAI’s legendary mathematical milestone
- → OpenAI fought dirty on career-making math problem, says NYU mathematician
- → On the Navier–Stokes Millennium Prize Problem
- → OpenAI researcher allegedly pressured mathematician to drop Anthropic co-author from math breakthrough paper
- → OpenAI alleged of stealing mathematicians work
- → Quoting Terence Tao
- → On the Value of Human Ideas: What data poisoning research reveals about "autonomous" AI breakthroughs
- → OpenAI’s sly mathematical breakthrough sends a chill through academia
- → Surveillance plagiarism by OpenAI
- → ANOTHER researcher accuses OpenAI of training on conversations and then claiming a breakthrough
- → Mathematicians want proof OpenAI didn’t use their work
- → OpenAI’s feud with mathematicians is only escalating
- → The Mathematical AI Safety Institute wants to prove AI is safe the way cryptographers prove codes are unbreakable
- → OpenAI just wants to win
- → Leading mathematicians fear AI is making their field dumber, and warn the rest of us is next
- → OpenAI reports Navier-Stokes singularity find, a contender for second ever Millenium Prize awarded, overshadowing Cognition's $48B Series E, Mistral's $24B Series D, Meta's Muse agent, and GPT Image 2.5
- → Claude proves Fermat 🧮, automated AI researcher 🔬, Z1 efficiency chip ⚡
Dibattito sul rallentamento. Dario Amodei di Anthropic ha proposto l'accesso per valutatori esterni, standard di sicurezza e trattati globali, mentre OpenAI ha chiesto al Congresso se un rallentamento coordinato violerebbe le leggi antitrust. Yoshua Bengio ha sostenuto che addestrare l'AI su testi umani la rende più abile nell'inganno, e Sam Altman ha ammesso che costruire un'AI fuori dal controllo umano è possibile.
- → Deep learning pioneer Bengio argues the training process itself makes AI dangerous
- → OpenAI floats a shared AI slowdown, takes it to Congress
- → Anthropic CEO outlines plan to slow AI development
- → Anthropic CEO says it’s time to pump the brakes on AI
- → Anthropic CEO Amodei wants AI speed limits before self-improvement outpaces human control
- → not much happened today
- → Looks like a coordination to stop distribution of intelligence
- → Sam Altman says OpenAI going public in 2026 would be ‘ill-advised’
- → OpenAI’s Sam Altman says it would be ‘ill-advised’ to go public in 2026
Battaglie copyright si allargano. The Seattle Times e Newsday hanno citato OpenAI e Microsoft per violazione del copyright, chiedendo la distruzione dei modelli addestrati sulle loro opere. Autori ed editori litigano anche su come dividere l'accordo da 1,5 miliardi di dollari di Anthropic.
Crescono gli abusi degli agenti. Gli agenti AI vengono usati per presentare reclami e richieste su larga scala: i reclami all'ombudsman per le abitazioni del Regno Unito sono raddoppiati e quelli al CFPB sono quintuplicati. Un avvocato è stato multato per aver depositato una memoria con testimoni inventati da ChatGPT, e Abliteration.ai ora vende accesso API a modelli con le protezioni rimosse, abbassando la barriera all'abuso.
- → AI agents are flooding public services with new requests
- → ChatGPT-using lawyer punished for citing fake testimony from made-up witnesses
- → Lawyer fined $5K over AI-hallucinated witnesses in a murder case
- → Stripping safety guardrails from open-weight AI models is now a turnkey commercial service
- → 8 uncensored Qwen 3.8 27B variants, one base, 167 GPU hours - Abliterlitics
AI enterprise e open source
Balzo dell'inferenza locale. Le ottimizzazioni della community portano Qwen3.8-Flash-Next a 1.200 token/s in prefill su Strix Halo, ExLlamaV3 batte llama.cpp nell'offload su CPU, Cherenkov fa streaming degli esperti su Apple Silicon e LayerStoRm esegue un quant da 186 GiB su 96 GB di VRAM. Gli utenti ottengono accelerazioni notevoli su RTX 3080, Strix Halo e MacBook Air.
- → LayerStoRm open-source expert streaming: 1M context GLM-5.3-Flash [UD-Q4_K_XL] at 24.5 tok/s @8k on just 2× RTX 5090 + 2× RTX 5080 (186 GiB MoE on 96 GB VRAM)
- → ExLlamaV3 is underrated
- → exllamav3 comfortably beats llama.cpp running CPU-offloaded Qwen-3.8-Flash-Next on my setup!
- → Qwen3.8-Flash-Next on MLX-serve, 1m context is released!
- → Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.
- → Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air
- → Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo
- → 3.8-27B has ruined 3.5/3.6-35B’s for me. It’s just *absurdly* superior.
- → This draft model is OP on 16 GB cards for Qwen 3.8 27b
- → I am impressed and I owe you one, Qwen 3.8 flash next (vision)!
- → Qwen3.8 Flash Next llama.cpp config tuning
- → bartowski/Qwen3.8-27B-GGUF · Hugging Face - Updated (Per-tensor layout)
Audit dei modelli aperti. Modelli open-weight tra cui gpt-oss-20b e glm-5.1 hanno trovato vulnerabilità reali in codebase pubbliche su GitHub, superando alcuni modelli di frontiera negli audit di sicurezza. Google ha reso open source Mantis, un framework agentico per la scansione delle vulnerabilità che riduce i falsi positivi e le vulnerabilità allucinate.
Trazione degli agenti enterprise. Figma ha costruito agenti AI su Panther SIEM che hanno ridotto del 70% il tempo di risoluzione degli alert complessi e del 20% le notifiche on-call. Ma i dati di Ramp mostrano che l'adozione dei prodotti AI è cresciuta solo dello 0,4% ad agosto, Meta ha smesso di includere l'uso degli strumenti AI nelle valutazioni delle prestazioni dopo che i dipendenti hanno manipolato il consumo di token, e le unità di ingegneri forward-deployed si moltiplicano nonostante obiettivi poco chiari.
- → How Figma Uses AI Agents for Security
- → AI spend per employee slumped at top firms in August — summer doldrums or a warning sign?
- → Top AI spenders cut per-employee costs by nearly 10 percent in August
- → Meta drops AI usage from engineer performance reviews after "tokenmaxxing" backfires
- → Google Cloud races to catch up in the AI deployment wars with Accenture deal
- → The Rise of the Forward Deployed Engineer — and How To Do the Job Right
Business e compute
Accuse USA-Cina sull'AI. Le agenzie statunitensi hanno indicato DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun e Z.AI come responsabili di distillazione su scala industriale dei modelli frontier statunitensi. Anthropic afferma di aver osservato quasi 200 milioni di scambi collegati ad attacchi di distillazione da parte di laboratori cinesi.
Compute e capitali. Anthropic ha firmato contratti per il compute fino a 517 miliardi di dollari, e Nvidia è in trattative per investire fino a 10 miliardi di dollari nella IPO pianificata di Anthropic, con una valutazione di 2.000 miliardi di dollari. Cognition ha raccolto 2 miliardi di dollari con una valutazione di 48 miliardi di dollari, e Mistral ha raccolto 3 miliardi di euro con una valutazione di 21 miliardi di euro.
- → Anthropic reportedly signs $517 billion in compute deals after Dario Amodei warned rivals about reckless risk
- → Nvidia wants to pour up to $10 billion into Anthropic's record-breaking IPO
- → Cognition hits $48B valuation, signaling investors believe AI coding is far from a winner-take-all market
- → Mistral raises €3B as sovereign AI becomes big business
Questo è il riepilogo della settimana - ci vediamo domenica prossima.