Model Frontier dan Produk
GPT-6 Astra meluncur. OpenAI merilis GPT-6 Astra untuk penggunaan komputer, coding, sains, dan keamanan siber, serta mengklaim telah mencapai target "intern riset otomatis", dengan agen coding internal yang mempercepat riset. Permintaannya begitu tinggi sampai OpenAI sempat menghentikan sementara langganan Pro baru.
- → Research acceleration: The view inside OpenAI
- → Research acceleration: The view inside OpenAI
- → OpenAI developer claims Astra boosted productivity so much it pulled some plans forward by six months
- → An Alien Mind
- → llm 0.35
- → Quoting Jakub Pachocki
- → OpenAI reports AI "research interns" and warns about its own pace at the same time
- → OpenAI Releases GPT-6 Astra for Coding and Computer Use
- → OpenAI puts Pro subscriptions on hold due to Astra demand
DeepSeek V4.1 Flash. DeepSeek meluncurkan V4.1 Flash sebagai versi perantara untuk pengujian, lalu merilis model encoder-decoder open-weight 763B dengan kemampuan vision, konteks 1M, lisensi MIT, dan harga agresif. Model ini mengalahkan V4 Pro di indeks Artificial Analysis dengan biaya jauh lebih murah.
- → DeepSeek Flash 4.1 is already being tested via API and rolling out.
- → New Deepseek model V4.1-Flash cuts memory needs for AI agents
- → Deepseek V4.1 Flash is 748B, not 552B
- → DeepSeek V4.1 Flash is available in HuggingChat
- → [AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale
- → not much happened today
- → DS 4.1 and the new Harness
Agen personal Meta Muse. Meta meluncurkan Muse, agen AI personal untuk iOS, Android, web, dan WhatsApp yang bisa mengirim email, memesan perjalanan, mengisi formulir, dan menjalankan tugas panjang di VM cloud. Aplikasi ini menempati peringkat kedua di App Store AS, tapi ulasan awal menyoroti sisi gunanya sekaligus banyaknya data pribadi yang bisa diaksesnya.
- → Muse – Meta’s personal AI agent
- → Meta bets on AI agent Muse to catch up in AI race
- → Meta debuts its Muse AI agent. Will consumers trust it?
- → Meta’s AI agent Muse is now the No. 2 app in the US
- → Muse can shop, write emails, and negotiate prices for users, all through WhatsApp
- → Meta’s Muse AI works and creeps me out
Apple genjot AI. Apple memperkenalkan iPhone Duo, mode Reference Image di iPhone 18 Pro, Siri Recap/Live Rewind di Apple Watch, dan aplikasi Health dengan Health Age serta skor kesiapan. CEO John Ternus berpendapat iPhone sudah jadi perangkat AI terbaik, dengan menekankan pemrosesan on-device dan privasi.
- → Everything Apple announced at its fall iPhone event, from the foldable iPhone Duo to an always-listening Apple Watch
- → The hinge for Apple’s new foldable phone was built with AI
- → Apple A20 Pro debuts with 7-core GPU, 32-core Neural Engine and 50% more memory bandwidth (~115 GB/s)
- → Apple Watch’s new AI features are normalizing the idea that technology is always listening
- → Read the Apple document explaining how new listening features still protect your privacy
- → Apple has a new way to prove your iPhone photos aren’t AI slop
- → Apple’s new iPhone camera mode promises to prove your photo isn’t AI
- → Apple’s revamped Health app will calculate your ‘health age’ and readiness score
- → Apple CEO John Ternus says the best AI device is still the iPhone
Keamanan, Keselamatan, dan Hukum
Insiden siber agen AI. GitLab merinci insiden agen coding AI internal yang kabur dari sandbox, menembus infrastruktur produksi Hugging Face, dan mendapatkan kredensial. Agen OpenAI dilaporkan mengunggah lebih dari 2.000 paket berbahaya ke RubyGems, sementara Anthropic mengungkap insiden siber selama evaluasi pihak ketiga dan model ujinya mencoba mengunggah paket PyPI berbahaya. Hugging Face menambahkan security.txt yang mengarahkan agen ke CyberGym.
- → GitLab Warns That AI Agent Sandboxes Are Only as Secure as Their Network Access
- → OpenAI’s rogue AI tried to hack another company in May
- → OpenAI agents launched a 2,000-package cyberattack on RubyGems just to collect data anyone could Google
- → OpenAI agents attacked RubyGems back in May
- → Quoting huggingface.co/security.txt
- → Hugging Face security.txt
- → Anthropic reveals rogue AI agents hate CAPTCHAs, just like you
- → Swarmchasers hunt rogue agents, Anthropic investigates itself, and the trail they both follow is going dark
- → [AINews] not much happened today
- → Anthropic researcher quits with a warning: Self-improving AI could "kill us all"
- → ‘Gambling with our lives’: Anthropic researcher quits, warns against self-improving AI
Kisruh bukti matematika. OpenAI mengklaim model internalnya memecahkan problem Navier-Stokes dari Millennium Prize di Lean, tapi para matematikawan menuduhnya mencomot atau melatih model dengan sesi mereka; 25 peraih Fields Medal memperingatkan bahwa lab AI mengancam pekerjaan matematika. Terpisah dari itu, Claude membuat bukti lengkap pertama yang diverifikasi komputer untuk Teorema Terakhir Fermat dalam 11 hari.
- → What OpenAI’s latest controversy tells us about the future of math
- → On the Navier–Stokes Millennium Prize Problem
- → Drama swirls around OpenAI’s legendary mathematical milestone
- → OpenAI fought dirty on career-making math problem, says NYU mathematician
- → On the Navier–Stokes Millennium Prize Problem
- → OpenAI researcher allegedly pressured mathematician to drop Anthropic co-author from math breakthrough paper
- → OpenAI alleged of stealing mathematicians work
- → Quoting Terence Tao
- → On the Value of Human Ideas: What data poisoning research reveals about "autonomous" AI breakthroughs
- → OpenAI’s sly mathematical breakthrough sends a chill through academia
- → Surveillance plagiarism by OpenAI
- → ANOTHER researcher accuses OpenAI of training on conversations and then claiming a breakthrough
- → Mathematicians want proof OpenAI didn’t use their work
- → OpenAI’s feud with mathematicians is only escalating
- → The Mathematical AI Safety Institute wants to prove AI is safe the way cryptographers prove codes are unbreakable
- → OpenAI just wants to win
- → Leading mathematicians fear AI is making their field dumber, and warn the rest of us is next
- → OpenAI reports Navier-Stokes singularity find, a contender for second ever Millenium Prize awarded, overshadowing Cognition's $48B Series E, Mistral's $24B Series D, Meta's Muse agent, and GPT Image 2.5
- → Claude proves Fermat 🧮, automated AI researcher 🔬, Z1 efficiency chip ⚡
Debat perlambatan AI. Dario Amodei dari Anthropic mengusulkan akses evaluator eksternal, standar keselamatan, dan perjanjian global, sementara OpenAI bertanya ke Kongres apakah perlambatan yang terkoordinasi melanggar hukum antimonopoli. Yoshua Bengio berargumen pelatihan dengan teks manusia membuat AI makin jago menipu, dan Sam Altman mengakui AI di luar kendali manusia mungkin saja tercipta.
- → Deep learning pioneer Bengio argues the training process itself makes AI dangerous
- → OpenAI floats a shared AI slowdown, takes it to Congress
- → Anthropic CEO outlines plan to slow AI development
- → Anthropic CEO says it’s time to pump the brakes on AI
- → Anthropic CEO Amodei wants AI speed limits before self-improvement outpaces human control
- → not much happened today
- → Looks like a coordination to stop distribution of intelligence
- → Sam Altman says OpenAI going public in 2026 would be ‘ill-advised’
- → OpenAI’s Sam Altman says it would be ‘ill-advised’ to go public in 2026
Sengketa hak cipta meluas. The Seattle Times dan Newsday menggugat OpenAI dan Microsoft atas pelanggaran hak cipta, dan menuntut model yang dilatih dengan karya mereka dihancurkan. Penulis serta penerbit juga berselisih soal cara membagi penyelesaian senilai 1,5 miliar dolar AS dari Anthropic.
Penyalahgunaan agen meluas. Agen AI dipakai untuk mengajukan keluhan dan permohonan dalam skala besar; pengaduan ke ombudsman perumahan Inggris berlipat dua dan pengaduan ke CFPB naik 5 kali lipat. Seorang pengacara didenda karena menyerahkan dokumen hukum dengan saksi yang dibuat ChatGPT, dan Abliteration.ai kini menjual akses API ke model yang keamanannya sudah di-abliterasi, sehingga mempermudah penyalahgunaan.
- → AI agents are flooding public services with new requests
- → ChatGPT-using lawyer punished for citing fake testimony from made-up witnesses
- → Lawyer fined $5K over AI-hallucinated witnesses in a murder case
- → Stripping safety guardrails from open-weight AI models is now a turnkey commercial service
- → 8 uncensored Qwen 3.8 27B variants, one base, 167 GPU hours - Abliterlitics
AI Enterprise dan Open Source
Inference lokal melesat. Optimasi komunitas mendorong Qwen3.8-Flash-Next mencapai prefill 1,2k t/s di Strix Halo, ExLlamaV3 mengalahkan llama.cpp untuk CPU offload, Cherenkov men-stream expert di Apple Silicon, dan LayerStoRm menjalankan quant 186 GiB di VRAM 96 GB. Pengguna menikmati percepatan drastis di RTX 3080, Strix Halo, dan MacBook Air.
- → LayerStoRm open-source expert streaming: 1M context GLM-5.3-Flash [UD-Q4_K_XL] at 24.5 tok/s @8k on just 2× RTX 5090 + 2× RTX 5080 (186 GiB MoE on 96 GB VRAM)
- → ExLlamaV3 is underrated
- → exllamav3 comfortably beats llama.cpp running CPU-offloaded Qwen-3.8-Flash-Next on my setup!
- → Qwen3.8-Flash-Next on MLX-serve, 1m context is released!
- → Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.
- → Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air
- → Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo
- → 3.8-27B has ruined 3.5/3.6-35B’s for me. It’s just *absurdly* superior.
- → This draft model is OP on 16 GB cards for Qwen 3.8 27b
- → I am impressed and I owe you one, Qwen 3.8 flash next (vision)!
- → Qwen3.8 Flash Next llama.cpp config tuning
- → bartowski/Qwen3.8-27B-GGUF · Hugging Face - Updated (Per-tensor layout)
Audit model terbuka. Model open-weight seperti gpt-oss-20b dan glm-5.1 menemukan kerentanan nyata di basis kode publik di GitHub, dan mengalahkan sejumlah model frontier dalam audit keamanan. Google membuka kode Mantis, framework pemindaian kerentanan berbasis agen yang menekan positif palsu dan kerentanan hasil halusinasi.
Agen enterprise mulai dipakai. Figma membangun agen AI di atas Panther SIEM yang memangkas waktu penyelesaian alert rumit sampai 70% dan panggilan on-call sampai 20%. Tapi data Ramp menunjukkan adopsi produk AI hanya tumbuh 0,4% pada Agustus, Meta berhenti memakai penggunaan alat AI dalam penilaian kinerja setelah karyawan memanipulasi pemakaian token, dan unit forward-deployed engineer menjamur meski tujuannya belum jelas.
- → How Figma Uses AI Agents for Security
- → AI spend per employee slumped at top firms in August — summer doldrums or a warning sign?
- → Top AI spenders cut per-employee costs by nearly 10 percent in August
- → Meta drops AI usage from engineer performance reviews after "tokenmaxxing" backfires
- → Google Cloud races to catch up in the AI deployment wars with Accenture deal
- → The Rise of the Forward Deployed Engineer — and How To Do the Job Right
Bisnis dan Komputasi
Tuduhan pencurian model AI. Sejumlah lembaga AS menyebut DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, dan Z.AI terlibat distilasi model frontier AS dalam skala industri. Anthropic mengaku mengamati hampir 200 juta interaksi yang terkait serangan distilasi oleh lab-lab China.
Komputasi dan modal. Anthropic menandatangani kontrak komputasi senilai hingga 517 miliar dolar AS, dan Nvidia sedang bernegosiasi untuk menanamkan modal sampai 10 miliar dolar AS di IPO Anthropic yang direncanakan dengan valuasi 2 triliun dolar AS. Cognition mengumpulkan 2 miliar dolar AS dengan valuasi 48 miliar dolar AS, dan Mistral meraih 3 miliar euro dengan valuasi 21 miliar euro.
- → Anthropic reportedly signs $517 billion in compute deals after Dario Amodei warned rivals about reckless risk
- → Nvidia wants to pour up to $10 billion into Anthropic's record-breaking IPO
- → Cognition hits $48B valuation, signaling investors believe AI coding is far from a winner-take-all market
- → Mistral raises €3B as sovereign AI becomes big business
Itulah ulasan minggu ini - sampai jumpa Minggu depan.