การปะทะกันของโมเดลแนวหน้า
Kimi K3 สั่นคลอนแนวหน้า. Moonshot AI's open-weight Kimi K3 แข่งขันกับโมเดลปิดชั้นนำ เอาชนะ Claude Fable 5 ในบางเกณฑ์มาตรฐาน และจุดประเด็นการถกเถียงระหว่างโอเพนและปิดอีกครั้ง
- → Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI
- → Kimi K3, and what we can still learn from the pelican benchmark
- → [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing
- → Kimi K3 weights to be released on the 27th.
- → Kimi K3 ranks #1 on @AfterQuery's SpreadsheetBench 2, surpassing Claude Fable 5
- → Kimi: Threat or menace?
- → kimi.ai teasing a video with lots of 3's in it
- → [AINews] not much happened today
- → Kimi moment. I think the writing is on the wall for Anthropic and OpenAi
- → Kimi K3 is currently at the top of the leaderboard for Text Arena filtered for science queries.
- → Kimi K3 (max) beats Sonnet 5 on Simple Bench
- → Kimi K3 is top of nextjs eval
- → Just like Deepseek, China's Kimi K3 is forcing Western AI labs to question their compute advantage
- → Kimi K3 🌕, Gemini 3.5 delayed ⏳, crushing ARC-AGI 3 🤖
GPT-5.6: พลังและอันตราย. ตระกูล GPT-5.6 นำเสนอการเรียกเครื่องมือแบบโปรแกรมและซับเอเจนต์แบบขนาน แต่แนวโน้มที่จะลบไฟล์และฐานข้อมูลโดยไม่ตั้งใจเน้นย้ำถึงความจำเป็นในการแซนด์บ็อกซ์
- → The Sequence Radar #893: Last Week in AI: GPT-5.6, Grok 4.5, Muse Spark 1.1 and the Post-Chatbot Stack
- → OpenAI’s new flagship model deletes files on its own, people keep warning
- → GPT-5.6 is deleting user files when given full access, and OpenAI says it shouldn't but did
- → [AINews] Codex usage up >10x in 6 months to 7M users, +1M in the past ~day; did Codex overtake Claude Code??
Claude Fable 5 ขยายเวลา. ภายใต้แรงกดดันจาก GPT-5.6 และ Kimi K3, Anthropic กลับแผนและคง Fable 5 ไว้ในแผนจ่ายเงิน แข่งขันโดยตรงในเรื่องการเข้าถึง
DeepSeek V4 กำลังใกล้เข้ามา. DeepSeek กำลังเปิดตัว V4 ด้วย API ต้นทุนต่ำและน้ำหนักโอเพน ในขณะที่การเพิ่มประสิทธิภาพของชุมชนทำให้รุ่น flash ทำงานบน GPU ทั่วไปที่ความเร็วที่ใช้งานได้
เอเจนต์เข้าสู่เวิร์กโฟลว์โลกจริง
เอเจนต์ท่องเว็บมาถึงแล้ว. Anthropic มอบเว็บเบราว์เซอร์ในตัวพร้อมตัวจำแนกความปลอดภัยให้ Claude Code และ Cursor เปิดตัวเอเจนต์ทั่วไป ทำให้ผู้ช่วยเขียนโค้ดก้าวไปสู่การดำเนินงานแบบอัตโนมัติ
เอเจนต์จัดการธุรกรรมจริง. สถาปัตยกรรมเอเจนต์ของ DoorDash เพิ่ม Conversion 24% และเกณฑ์มาตรฐานของ Stripe แสดงให้เห็นว่าเอเจนต์สามารถเขียนโค้ดการรวมระบบได้แต่มักล้มเหลวในการตรวจสอบ เน้นย้ำถึงช่องว่างด้านความน่าเชื่อถือ
เอเจนต์ในระบบผลิตสะดุด. เอเจนต์ในองค์กรส่วนใหญ่ยังคงเป็นแชทบอท 54% ของบริษัทมีเหตุการณ์ด้านความปลอดภัยของเอเจนต์ และผู้เชี่ยวชาญโต้แย้งว่าเอเจนต์ต้องการพื้นฐานการดำเนินงานแบบ cloud-native เช่น ไมโครเซอร์วิส
- → How to Debug Coding Agents with LangSmith Traces
- → The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials
- → The AI context gap: Enterprise AI organizations have a trust problem, not a retrieval problem — and most are still building the fix
- → The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway
- → Cloud Native Infrastructure Emerges as the Foundation for Trustworthy Agentic AI
- → QCon AI Boston: Production AI Moves Beyond Prompts to Platforms, Harnesses, and Evals
โครงสร้างพื้นฐาน AI การระดมทุน และการแข่งขัน Data Center
Data Center เผชิญแรงต้าน. S&P ลดอันดับ Oracle เนื่องจากความเสี่ยงด้านรายจ่าย AI ในขณะที่การประท้วงในท้องถิ่นและการระงับชั่วคราวของนิวยอร์กสำหรับ Data Center ขนาดใหญ่สะท้อนถึงการต่อต้านที่เพิ่มขึ้นต่อโครงสร้างพื้นฐาน AI
การระดมทุน AI พุ่งสูง. Databricks ระดมทุนด้วยมูลค่า 188 พันล้านดอลลาร์ Meta กำลังเจรจาเช่า Data Center มูลค่า 10 พันล้านดอลลาร์ให้ Anthropic และ Nous Research ได้รับเงิน 75 ล้านดอลลาร์สำหรับเอเจนต์โอเพนซอร์ส
การเดิมพันด้านพลังงานและฮาร์ดแวร์. บริษัทพลังงานระดมทุนมากที่สุดนับตั้งแต่ยุคดอทคอมเพื่อขับเคลื่อน AI และเงินกู้ 400 ล้านดอลลาร์ที่มีชิป SambaNova ค้ำประกัน ส่งสัญญาณถึงการลงทุนที่เพิ่มขึ้นในฮาร์ดแวร์ AI ที่ไม่ใช่ GPU
ความก้าวหน้า AI บนอุปกรณ์
โมเดล 1 บิตมาถึงโทรศัพท์. Bonsai ของ PrismML บีบอัด Qwen3.6-27B เหลือ 3.9GB ในขณะที่ยังคงคะแนนเกณฑ์มาตรฐานไว้ 90% ทำให้สามารถเรียกเครื่องมือบนอุปกรณ์ได้ มีรายงานว่า Apple กำลังเจรจา
- → Bonsai 27B: The First 27B-Class Model to Run on a Phone
- → Bonsai 27B: 1-bit dense LLM running locally in your browser using custom WebGPU kernels
- → Prism-ML Bonsai Qwen 3.6 27B
- → PrismML’s new Ternary Qwen3.6 27B runs near fp16 precision on 10GB of memory!!!
- → So what's the consensus on 1bit models? Is it still a pipe dream?
- → PrismML Bonsai 27B is surprisingly usable on the Jetson Orin Nano 8GB
- → Is anyone having any luck with the Ternary Bonsai 27B DFlash?
- → Can we get a "not base model" flair?
- → Bonsai 27B runs locally on an iPhone - a 27B model in 3.9GB
- → User experience of Bonsai-Ternary-27B on 4060Ti 16GB for KB management and productivity assistant use cases
- → Bonsai 27B is a full open reasoning model that fits on an iPhone
- → Bonsai-27B & Ternary-Bonsai-27B - Updates (on PRs)
- → Apple in talks with startup PrismML that shrinks AI models to run on an iPhone
แล็ปที่บ้านรันโมเดลแนวหน้า. การเพิ่มประสิทธิภาพของ llama.cpp เช่น DFlash เร่งโมเดล MoE ได้ถึง 6 เท่า และโทรศัพท์ OnePlus สตรีมผู้เชี่ยวชาญจาก flash เพื่อรันโมเดล 60GB ที่ 1.3 tok/s
- → I tested all llama.cpp's speculative decoding methods on Qwen 3.6 27B: MTP ~2.7x, DFlash ~3.7x, n-gram stack ~6x on real coding. Local AI win. My findings on RTX 6000 PRO.
- → DFlash makes Qwen3.6 27B 2.2x faster with no quality loss
- → GPT-OSS-120B, Qwen 30B and Gemma 26B on an Android phone at 1-5 tok/s: +60GB model, 11GB of RAM, CPU only
นโยบาย คดีความ และการเผชิญหน้าระหว่างโอเพนซอร์ส
การต่อสู้ทางกฎหมายระหว่าง Apple-OpenAI. Apple ฟ้อง OpenAI ฐานความลับทางการค้า โดยกล่าวหาว่าลักพาตัวพนักงานกว่า 400 คน รวมถึงหัวหน้าเจ้าหน้าที่ฝ่ายฮาร์ดแวร์ ซึ่งคุกคาม IPO และความไว้วางใจด้านคลาวด์ของ OpenAI
- → Apple sues OpenAI after ex-engineer allegedly used bug to steal trade secrets
- → The wildest allegations in Apple’s trade secrets lawsuit against OpenAI
- → The 6 wildest claims in Apple’s lawsuit against OpenAI
- → OpenAI pushes back on Apple trade secret lawsuit
- → Sam Altman didn’t need another lawsuit
- → How Apple’s big lawsuit could disrupt OpenAI’s IPO plans
- → Apple’s plot to crush OpenAI
- → Apple’s lawsuit couldn’t come at a worse time for OpenAI
โอเพนซอร์สจีนพุ่งสูง. โมเดลจีนตอนนี้คิดเป็น 41% ของการดาวน์โหลดบน HuggingFace, Kimi K3 แข่งขันกับผู้นำตะวันตก และจีนเปิดตัวหน่วยงานกำกับดูแล AI แบบคู่ขนานที่ไม่รวมชาติตะวันตก
- → The real AI race may no longer be at the frontier
- → Source: the Trump administration and industry groups discussed streamlining US open model releases of equal or lesser capability to leading Chinese open models
- → Chinese President Xi Jinping speaks at World AI Conference and reaffirms commitment to open source to promote"openness and win-win"
- → China’s Xi Touts Open-Source AI and Takes a Swipe at U.S. Dominance
- → China's new World Artificial Intelligence Cooperation Organization is President Xi's clearest play yet for a parallel AI order
ความต้องการความรับผิดชอบเพิ่มขึ้น. หน่วยงานกำกับดูแลของเยอรมนีถือว่าแชทบอทต้องรับผิดชอบต่อเนื้อหา xAI ฟ้องผู้ใช้ที่สร้าง CSAM และ Demis Hassabis เสนอองค์กรคล้าย FINRA สำหรับการทบทวนความปลอดภัยแนวหน้า
นี่คือสรุปประจำสัปดาห์ - แล้วพบกันใหม่วันอาทิตย์หน้า