Benturan Model Frontier
Kimi K3 mengguncang frontier. Kimi K3 dengan bobot terbuka dari Moonshot AI bersaing dengan model tertutup teratas, mengalahkan Claude Fable 5 pada beberapa tolok ukur dan memicu kembali perdebatan terbuka vs tertutup.
- → Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI
- → Kimi K3, and what we can still learn from the pelican benchmark
- → [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing
- → Kimi K3 weights to be released on the 27th.
- → Kimi K3 ranks #1 on @AfterQuery's SpreadsheetBench 2, surpassing Claude Fable 5
- → Kimi: Threat or menace?
- → kimi.ai teasing a video with lots of 3's in it
- → [AINews] not much happened today
- → Kimi moment. I think the writing is on the wall for Anthropic and OpenAi
- → Kimi K3 is currently at the top of the leaderboard for Text Arena filtered for science queries.
- → Kimi K3 (max) beats Sonnet 5 on Simple Bench
- → Kimi K3 is top of nextjs eval
- → Just like Deepseek, China's Kimi K3 is forcing Western AI labs to question their compute advantage
- → Kimi K3 🌕, Gemini 3.5 delayed ⏳, crushing ARC-AGI 3 🤖
GPT-5.6: Kekuatan dan Bahaya. Keluarga GPT-5.6 membawa pemanggilan alat secara programatik dan subagen paralel, namun kecenderungannya untuk secara tidak sengaja menghapus file dan database menekankan perlunya sandboxing.
- → The Sequence Radar #893: Last Week in AI: GPT-5.6, Grok 4.5, Muse Spark 1.1 and the Post-Chatbot Stack
- → OpenAI’s new flagship model deletes files on its own, people keep warning
- → GPT-5.6 is deleting user files when given full access, and OpenAI says it shouldn't but did
- → [AINews] Codex usage up >10x in 6 months to 7M users, +1M in the past ~day; did Codex overtake Claude Code??
Claude Fable 5 diperpanjang. Di bawah tekanan dari GPT-5.6 dan Kimi K3, Anthropic membalikkan rencana dan tetap mempertahankan Fable 5 pada paket berbayar, bersaing langsung dalam aksesibilitas.
DeepSeek V4 mengintai. DeepSeek menggoda V4 dengan API berbiaya rendah dan bobot terbuka, sementara optimasi komunitas sudah memungkinkan varian flash-nya berjalan di GPU konsumen dengan kecepatan yang dapat digunakan.
Agen Memasuki Alur Kerja Nyata
Agen penelusuran tiba. Anthropic memberikan Claude Code browser web bawaan dengan pengklasifikasi keamanan, dan Cursor meluncurkan agen umum, menggerakkan asisten kode menuju eksekusi tugas otonom.
Agen menangani transaksi nyata. Arsitektur agen DoorDash meningkatkan konversi sebesar 24%, dan tolok ukur Stripe menunjukkan agen dapat mengodekan integrasi tetapi sering gagal dalam validasi, menyoroti kesenjangan keandalan.
Agen produksi tersandung. Sebagian besar agen perusahaan tetap menjadi chatbot; 54% perusahaan mengalami insiden keamanan agen, dan para ahli berpendapat agen memerlukan primitif operasional cloud-native seperti layanan mikro.
- → How to Debug Coding Agents with LangSmith Traces
- → The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials
- → The AI context gap: Enterprise AI organizations have a trust problem, not a retrieval problem — and most are still building the fix
- → The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway
- → Cloud Native Infrastructure Emerges as the Foundation for Trustworthy Agentic AI
- → QCon AI Boston: Production AI Moves Beyond Prompts to Platforms, Harnesses, and Evals
Infrastruktur AI, Pendanaan, dan Perlombaan Pusat Data
Pusat data menghadapi perlawanan. S&P menurunkan peringkat Oracle atas risiko belanja modal AI, sementara protes lokal dan moratorium New York pada pusat data hiperskala mencerminkan penolakan yang semakin besar terhadap infrastruktur AI.
Pendanaan AI melonjak. Databricks menggalang dana dengan valuasi $188 miliar, Meta sedang menegosiasikan sewa pusat data senilai $10 miliar ke Anthropic, dan Nous Research mendapatkan $75 juta untuk agen open-source-nya.
Taruhan energi dan perangkat keras. Perusahaan energi menggalang dana terbanyak sejak era dot-com untuk menyalakan AI, dan pinjaman $400 juta yang didukung chip SambaNova menandakan investasi yang semakin besar pada perangkat keras AI non-GPU.
Terobosan AI di Perangkat
Model 1-bit mencapai ponsel. Bonsai dari PrismML mengompresi Qwen3.6-27B menjadi 3,9 GB sambil mempertahankan 90% skor tolok ukur, memungkinkan agen pemanggil alat di perangkat; Apple dikabarkan sedang dalam pembicaraan.
- → Bonsai 27B: The First 27B-Class Model to Run on a Phone
- → Bonsai 27B: 1-bit dense LLM running locally in your browser using custom WebGPU kernels
- → Prism-ML Bonsai Qwen 3.6 27B
- → PrismML’s new Ternary Qwen3.6 27B runs near fp16 precision on 10GB of memory!!!
- → So what's the consensus on 1bit models? Is it still a pipe dream?
- → PrismML Bonsai 27B is surprisingly usable on the Jetson Orin Nano 8GB
- → Is anyone having any luck with the Ternary Bonsai 27B DFlash?
- → Can we get a "not base model" flair?
- → Bonsai 27B runs locally on an iPhone - a 27B model in 3.9GB
- → User experience of Bonsai-Ternary-27B on 4060Ti 16GB for KB management and productivity assistant use cases
- → Bonsai 27B is a full open reasoning model that fits on an iPhone
- → Bonsai-27B & Ternary-Bonsai-27B - Updates (on PRs)
- → Apple in talks with startup PrismML that shrinks AI models to run on an iPhone
Laboratorium rumah menjalankan frontier. Optimasi llama.cpp seperti DFlash mempercepat model MoE hingga 6x, dan ponsel OnePlus mengalirkan pakar dari flash untuk menjalankan model 60GB pada 1,3 tok/s.
- → I tested all llama.cpp's speculative decoding methods on Qwen 3.6 27B: MTP ~2.7x, DFlash ~3.7x, n-gram stack ~6x on real coding. Local AI win. My findings on RTX 6000 PRO.
- → DFlash makes Qwen3.6 27B 2.2x faster with no quality loss
- → GPT-OSS-120B, Qwen 30B and Gemma 26B on an Android phone at 1-5 tok/s: +60GB model, 11GB of RAM, CPU only
Kebijakan, Gugatan Hukum, dan Kebuntuan Open-Source
Pertarungan hukum Apple–OpenAI. Apple menggugat OpenAI atas rahasia dagang, menuduh perburuan lebih dari 400 karyawan, termasuk chief hardware officer, mengancam IPO dan kepercayaan cloud OpenAI.
- → Apple sues OpenAI after ex-engineer allegedly used bug to steal trade secrets
- → The wildest allegations in Apple’s trade secrets lawsuit against OpenAI
- → The 6 wildest claims in Apple’s lawsuit against OpenAI
- → OpenAI pushes back on Apple trade secret lawsuit
- → Sam Altman didn’t need another lawsuit
- → How Apple’s big lawsuit could disrupt OpenAI’s IPO plans
- → Apple’s plot to crush OpenAI
- → Apple’s lawsuit couldn’t come at a worse time for OpenAI
Lonjakan open-source Tiongkok. Model Tiongkok kini menyumbang 41% unduhan HuggingFace, Kimi K3 menyaingi pemimpin Barat, dan Tiongkok meluncurkan badan tata kelola AI paralel yang mengecualikan Barat.
- → The real AI race may no longer be at the frontier
- → Source: the Trump administration and industry groups discussed streamlining US open model releases of equal or lesser capability to leading Chinese open models
- → Chinese President Xi Jinping speaks at World AI Conference and reaffirms commitment to open source to promote"openness and win-win"
- → China’s Xi Touts Open-Source AI and Takes a Swipe at U.S. Dominance
- → China's new World Artificial Intelligence Cooperation Organization is President Xi's clearest play yet for a parallel AI order
Tuntutan akuntabilitas meningkat. Regulator Jerman memegang chatbot bertanggung jawab atas konten, xAI menggugat pengguna karena menghasilkan CSAM, dan Demis Hassabis mengusulkan badan mirip FINRA untuk tinjauan keamanan frontier.
Itulah ulasan minggu ini - sampai jumpa Minggu depan.