프론티어 모델 대격돌
Kimi K3가 프론티어를 흔들다. Moonshot AI의 오픈 가중치 Kimi K3가 최고 수준의 폐쇄 모델과 경쟁하며 일부 벤치마크에서 Claude Fable 5를 능가해 오픈 대 폐쇄 논쟁에 다시 불을 붙였습니다.
- → Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI
- → Kimi K3, and what we can still learn from the pelican benchmark
- → [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing
- → Kimi K3 weights to be released on the 27th.
- → Kimi K3 ranks #1 on @AfterQuery's SpreadsheetBench 2, surpassing Claude Fable 5
- → Kimi: Threat or menace?
- → kimi.ai teasing a video with lots of 3's in it
- → [AINews] not much happened today
- → Kimi moment. I think the writing is on the wall for Anthropic and OpenAi
- → Kimi K3 is currently at the top of the leaderboard for Text Arena filtered for science queries.
- → Kimi K3 (max) beats Sonnet 5 on Simple Bench
- → Kimi K3 is top of nextjs eval
- → Just like Deepseek, China's Kimi K3 is forcing Western AI labs to question their compute advantage
- → Kimi K3 🌕, Gemini 3.5 delayed ⏳, crushing ARC-AGI 3 🤖
GPT-5.6: 힘과 위험. GPT-5.6 제품군은 프로그래매틱 도구 호출과 병렬 서브에이전트를 제공하지만, 파일과 데이터베이스를 무심코 삭제하는 경향이 있어 샌드박싱의 필요성을 강조합니다.
- → The Sequence Radar #893: Last Week in AI: GPT-5.6, Grok 4.5, Muse Spark 1.1 and the Post-Chatbot Stack
- → OpenAI’s new flagship model deletes files on its own, people keep warning
- → GPT-5.6 is deleting user files when given full access, and OpenAI says it shouldn't but did
- → [AINews] Codex usage up >10x in 6 months to 7M users, +1M in the past ~day; did Codex overtake Claude Code??
Claude Fable 5 연장. GPT-5.6과 Kimi K3의 압박 속에서 Anthropic은 계획을 뒤집고 Fable 5를 유료 요금제에 유지하며 접근성 측면에서 직접 경쟁하게 되었습니다.
DeepSeek V4 등장 임박. DeepSeek은 저비용 API와 오픈 가중치로 V4를 예고하고 있으며, 커뮤니티 최적화를 통해 이미 플래시 변형이 소비자 GPU에서 사용 가능한 속도로 실행되고 있습니다.
에이전트, 실제 워크플로우 진입
브라우징 에이전트 도착. Anthropic은 Claude Code에 안전 분류기를 갖춘 내장 웹 브라우저를 제공했고, Cursor는 일반 에이전트를 출시하여 코딩 어시스턴트를 자율 작업 실행으로 이끌고 있습니다.
에이전트 실제 거래 처리. DoorDash의 에이전트 아키텍처는 전환율을 24% 향상시켰고, Stripe의 벤치마크는 에이전트가 통합을 코딩할 수 있지만 종종 검증에 실패하여 신뢰성 격차를 부각시킵니다.
프로덕션 에이전트 난관. 대부분의 엔터프라이즈 에이전트는 여전히 챗봇이며, 54%의 기업이 에이전트 보안 사고를 겪었고, 전문가들은 에이전트가 마이크로서비스와 같은 클라우드 네이티브 운영 프리미티브가 필요하다고 주장합니다.
- → How to Debug Coding Agents with LangSmith Traces
- → The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials
- → The AI context gap: Enterprise AI organizations have a trust problem, not a retrieval problem — and most are still building the fix
- → The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway
- → Cloud Native Infrastructure Emerges as the Foundation for Trustworthy Agentic AI
- → QCon AI Boston: Production AI Moves Beyond Prompts to Platforms, Harnesses, and Evals
AI 인프라, 자금 조달 및 데이터 센터 경쟁
데이터 센터 저항에 부딪히다. S&P는 AI 자본 지출 위험으로 Oracle의 신용 등급을 하향 조정했으며, 지역 시위와 뉴욕의 하이퍼스케일 데이터 센터 모라토리엄은 AI 인프라에 대한 반발이 커지고 있음을 반영합니다.
AI 자금 조달 급증. Databricks는 1880억 달러의 가치 평가로 자금을 조달했고, Meta는 Anthropic에 100억 달러 규모의 데이터 센터 임대를 협상 중이며, Nous Research는 오픈소스 에이전트를 위해 7500만 달러를 확보했습니다.
에너지 및 하드웨어 투자. 에너지 기업들은 AI 구동을 위해 닷컴 시대 이후 최대 규모의 자금을 조달했으며, SambaNova 칩을 담보로 한 4억 달러 대출은 비GPU AI 하드웨어에 대한 투자 증가를 시사합니다.
온디바이스 AI 혁신
1비트 모델, 휴대폰에 도달. PrismML의 Bonsai는 Qwen3.6-27B를 3.9GB로 압축하면서 벤치마크 점수의 90%를 유지하여 온디바이스 도구 호출 에이전트를 가능하게 합니다. Apple은 협상 중인 것으로 알려졌습니다.
- → Bonsai 27B: The First 27B-Class Model to Run on a Phone
- → Bonsai 27B: 1-bit dense LLM running locally in your browser using custom WebGPU kernels
- → Prism-ML Bonsai Qwen 3.6 27B
- → PrismML’s new Ternary Qwen3.6 27B runs near fp16 precision on 10GB of memory!!!
- → So what's the consensus on 1bit models? Is it still a pipe dream?
- → PrismML Bonsai 27B is surprisingly usable on the Jetson Orin Nano 8GB
- → Is anyone having any luck with the Ternary Bonsai 27B DFlash?
- → Can we get a "not base model" flair?
- → Bonsai 27B runs locally on an iPhone - a 27B model in 3.9GB
- → User experience of Bonsai-Ternary-27B on 4060Ti 16GB for KB management and productivity assistant use cases
- → Bonsai 27B is a full open reasoning model that fits on an iPhone
- → Bonsai-27B & Ternary-Bonsai-27B - Updates (on PRs)
- → Apple in talks with startup PrismML that shrinks AI models to run on an iPhone
홈랩, 프론티어 모델 실행. DFlash와 같은 llama.cpp 최적화는 MoE 모델을 최대 6배 가속하며, OnePlus 휴대폰이 플래시에서 전문가를 스트리밍하여 60GB 모델을 1.3 tok/s로 실행했습니다.
- → I tested all llama.cpp's speculative decoding methods on Qwen 3.6 27B: MTP ~2.7x, DFlash ~3.7x, n-gram stack ~6x on real coding. Local AI win. My findings on RTX 6000 PRO.
- → DFlash makes Qwen3.6 27B 2.2x faster with no quality loss
- → GPT-OSS-120B, Qwen 30B and Gemma 26B on an Android phone at 1-5 tok/s: +60GB model, 11GB of RAM, CPU only
정책, 소송 및 오픈소스 대치
Apple–OpenAI 법적 분쟁. Apple은 OpenAI를 영업 비밀 침해로 고소하며, 최고 하드웨어 책임자를 포함한 400명 이상의 직원을 스카우트했다고 주장, OpenAI의 IPO와 클라우드 신뢰에 위협이 되고 있습니다.
- → Apple sues OpenAI after ex-engineer allegedly used bug to steal trade secrets
- → The wildest allegations in Apple’s trade secrets lawsuit against OpenAI
- → The 6 wildest claims in Apple’s lawsuit against OpenAI
- → OpenAI pushes back on Apple trade secret lawsuit
- → Sam Altman didn’t need another lawsuit
- → How Apple’s big lawsuit could disrupt OpenAI’s IPO plans
- → Apple’s plot to crush OpenAI
- → Apple’s lawsuit couldn’t come at a worse time for OpenAI
중국 오픈소스 급증. 중국 모델이 이제 HuggingFace 다운로드의 41%를 차지하고, Kimi K3는 서구 선두 주자와 경쟁하며, 중국은 서구를 배제한 병렬 AI 거버넌스 기구를 출범시켰습니다.
- → The real AI race may no longer be at the frontier
- → Source: the Trump administration and industry groups discussed streamlining US open model releases of equal or lesser capability to leading Chinese open models
- → Chinese President Xi Jinping speaks at World AI Conference and reaffirms commitment to open source to promote"openness and win-win"
- → China’s Xi Touts Open-Source AI and Takes a Swipe at U.S. Dominance
- → China's new World Artificial Intelligence Cooperation Organization is President Xi's clearest play yet for a parallel AI order
책임 요구 증가. 독일 규제 기관은 챗봇이 콘텐츠에 대해 책임이 있다고 판결했고, xAI는 CSAM을 생성한 사용자를 고소했으며, Demis Hassabis는 프론티어 안전 검토를 위한 FINRA 유사 기구를 제안했습니다.
이번 주 리뷰였습니다 - 다음 일요일에 뵙겠습니다.