模型與競爭
GPT-5.6 與 Work Agent. OpenAI 發布了三個推理層級的 GPT-5.6,並推出了 ChatGPT Work,用於長達數小時的自動化任務,但使用限制和混亂的介面引發了抱怨,促使承諾修復。
- → OpenAI launches its new family of models with GPT-5.6
- → The new GPT-5.6 family: Luna, Terra, Sol
- → OpenAI rolls out GPT-5.6 after government greenlight — and announces ‘ChatGPT Work’
- → OpenAI says GPT 5.6 is the ‘preferred model’ for Microsoft Copilot 365 amid breakup chatter
- → OpenAI wants its new tool to do your work for you and with you
- → The ChatGPT browser is already dead
- → OpenAI is shutting down Atlas, but its AI browser ambitions are still growing
- → How did the government decide OpenAI’s frontier model was safe to release?
- → OpenAI pairs its GPT-5.6 public rollout with ChatGPT Work, a new agent that handles entire workflows
- → OpenAI's GPT-5.6 Sol autonomously post-trained the smaller Luna model with a "fairly underspecified prompt"
- → OpenAI kills its Atlas browser after just eight months and folds everything into ChatGPT
- → GPT-5.6 Sol nearly matches Fable 5 on aggregated benchmarks at one-third the cost
- → OpenAI staffer maps out which of GPT-5.6 Sol's five reasoning levels fits which task complexity
- → OpenAI admits it "didn't get everything quite right" with ChatGPT Work launch and scrambles to fix UX and costs
Claude Fable 5 在基準測試中領先. Anthropic 的 Claude Fable 5 橫掃了 Artificial Analysis 的基準測試,但其高昂的成本導致該公司建議將其用作規劃者,委派給更便宜的模型,從而削減約 40% 的開支,同時保留大部分能力。
Meta 的代理閃電戰與隱私風波. Meta 發布了 Muse Spark 1.1,這是一款高情境編碼模型,在價格上削弱競爭對手,同時其 Muse Image 生成器因允許未經授權使用 Instagram 臉部數據而被下架,重新引發了同意辯論。
- → Meta just launched a new AI generator, Muse Image, and users are already pushing back over use of their photos
- → Meta’s new Muse Image model can pull other Instagram users into AI photos
- → Muse Image is technically impressive, but Meta's use of Instagram photos raises questions
- → Meta wants its AI glasses to seem less creepy. Its AI strategy says otherwise.
- → Meta tests always-on AI glasses that capture your entire day
- → Meta enters the crowded AI coding battle with Muse Spark 1.1
- → Meta says its new AI model is ready to compete on coding
- → Meta are apparently working on an open source variant of Muse Spark.
- → Introducing Muse Spark 1.1
- → Meta's Muse Spark 1.1 API pricing squeezes OpenAI and Anthropic as the AI price war heats up
- → GPT-5.6 🚀, Muse Spark 1.1 ✨, ChatGPT Work 💼
- → Meta removes controversial AI feature on Instagram after backlash
- → Meta turns off the Instagram feature that let users make AI deepfakes of public accounts
中國開源浪潮. 中國模型現佔 OpenRouter 流量的 30% 以上;騰訊的 HY3 和 GLM-5.2 在國產硬體上運行,MiniMax 計劃發布一個 2.7T 參數的開源模型,而 DeepSeek 的晶片設計標誌著垂直整合,引發美國的擔憂。
- → This is what Hy3 is capable of. Mother of god.
- → llama.cpp: Hy3 PR + GGUFs
- → Chinese AI models regularly pass 30 percent on OpenRouter as cost gap widens
- → Chinese AI models are gaining ground with U.S. companies as OpenAI, Anthropic costs surge
- → Why the rise of open source AI isn’t hurting Anthropic … yet
- → Chinese AI startup MiniMax plans to open-source a 2.7 trillion parameter model later this year
- → 4-bit GLM-5.2 (753B MoE) on 4× DGX Spark: 70.8% on Terminal-Bench 2.1 vs 81.0% for the full model
- → Running GLM 5.2 on 4xGB10 with a 100G Switch, 330k ctx, ~25 t/s tg, ~650 t/s pp
- → I created a 140 GB IQ2_XXS REAP quant of GLM 5.2 for coding. Looking for testers.
- → Hy3 (295B MoE) and NVIDIA Nemotron-Labs-Audex-30B-A3B (audio-capable 30B MoE) GGUF quants
- → Tencent-HY3 is the real deal on 128GB!
- → The U.S. tech industry is increasingly anxious about the rising power and competitive price of open-source AI models from China — and whether the Trump administration will respond with yet another executive order | Politico
- → China's DeepSeek developing its own AI chip, sources say
代理工程與基礎設施
Bun 被 64 個代理重寫. Claude Fable 5 協調了 64 個並行實例,在 11 天內將 Bun 從 Zig 重寫為 Rust,生成了超過一百萬行程式碼,修復了 128 個錯誤,花費 165,000 美元。
SWE-Bench Pro 審計失敗. OpenAI 發現約 30% 的 SWE-Bench Pro 任務有問題,並撤回了其認可,呼應了 Artificial Analysis 先前的擔憂,並削弱了代理程式碼評估。
Modal 籌集 3.55 億美元. Modal 獲得了 3.55 億美元,用於建立專為 AI 代理設計的雲端基礎設施,專注於自主軟體的沙盒化、快速迭代環境。
Cloudflare 臨時代理. Cloudflare 推出了臨時帳戶,允許 AI 代理無需永久憑證即可部署 Workers,若未被認領則會過期,從而順暢代理驅動的工作流程。
代理進軍本地 GPU. 新的量化技術和具成本效益的 GPU 設備現在可以在消費級硬體上運行 MoE 模型,但基準測試顯示低位量化可能會嚴重降低代理任務的效能,促使謹慎選擇量化方法。
- → Qwen3.5 122B is the best?
- → Qwen3.6-27b does not understand software architechure.
- → Qwen 3.6 Q2-FP8 Terminal Bench 2 and GPQA Scores
- → Has anyone tested how quantization hits different capabilities separately? My results are surprising.
- → 2.5x faster Qwen3.6 NVFP4 Unsloth quants
- → Ultra budget 20GB vram with 448GB/s for $100 bucks.
- → I benched quad 5060Tis for code generation with Qwen3.6-27B so you don't have to (it's really good)
安全、法律與可解釋性
Jacobian Lens 映射模型思維. Anthropic 的 Jacobian Lens 揭示了一小組可語言化的表徵,這些表徵捕捉了模型的核心推理,從而實現幻覺檢測、輸出控制,而且正如社群演示所示,還能立即創造有害變體。
- → Anthropic Research - "Verbalizable Representations Form a Global Workspace in Language Models"
- → Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens
- → I tested Anthropic’s new Jacobian Lens on open models, then it turned into a local-model hallucination router
- → Anthropic found a hidden space where Claude puzzles over concepts
- → I created a super harmful model ! :D (by tweaking it's J-Space!!!)
代理驅動的勒索軟體. Sysdig 記錄了一個 AI 代理處理勒索軟體攻擊的技術執行,儘管是由人類設置操作,這凸顯了當前代理惡意軟體的能力與限制。
AI 濫用危機. 一起訴訟指控 xAI 的 Grok 生成了 CSAM 圖像,而博科聖地使用聊天機器人進行攻擊策劃;這些事件凸顯了隨著代理能力擴展,迫切需要安全護欄。
Apple 起訴 OpenAI. Apple 提起聯邦訴訟,指控 OpenAI 挖走 400 多名員工以竊取硬體商業機密,可能用於打造競爭對手的 AI 智慧型手機,加劇了 AI 產業的法律戰。
這是本週回顧 - 下週日見。