模型与竞争
GPT-5.6和工作代理. OpenAI发布了三个推理层级的GPT-5.6,并推出了ChatGPT Work,用于长达数小时的自主任务,但使用限制和混乱的界面引发了抱怨,促使承诺修复。
- → OpenAI launches its new family of models with GPT-5.6
- → The new GPT-5.6 family: Luna, Terra, Sol
- → OpenAI rolls out GPT-5.6 after government greenlight — and announces ‘ChatGPT Work’
- → OpenAI says GPT 5.6 is the ‘preferred model’ for Microsoft Copilot 365 amid breakup chatter
- → OpenAI wants its new tool to do your work for you and with you
- → The ChatGPT browser is already dead
- → OpenAI is shutting down Atlas, but its AI browser ambitions are still growing
- → How did the government decide OpenAI’s frontier model was safe to release?
- → OpenAI pairs its GPT-5.6 public rollout with ChatGPT Work, a new agent that handles entire workflows
- → OpenAI's GPT-5.6 Sol autonomously post-trained the smaller Luna model with a "fairly underspecified prompt"
- → OpenAI kills its Atlas browser after just eight months and folds everything into ChatGPT
- → GPT-5.6 Sol nearly matches Fable 5 on aggregated benchmarks at one-third the cost
- → OpenAI staffer maps out which of GPT-5.6 Sol's five reasoning levels fits which task complexity
- → OpenAI admits it "didn't get everything quite right" with ChatGPT Work launch and scrambles to fix UX and costs
Claude Fable 5在基准测试中领先. Anthropic的Claude Fable 5横扫了Artificial Analysis基准测试,但由于成本高昂,该公司建议将其用作规划器,委托给更便宜的模型,从而削减约40%的费用,同时保留大部分能力。
Meta的Agentic Blitz与隐私影响. Meta推出了Muse Spark 1.1,这是一款高上下文编码模型,价格上低于竞争对手,而其Muse Image生成器因允许未经授权使用Instagram人脸而被下架,重新引发了同意辩论。
- → Meta just launched a new AI generator, Muse Image, and users are already pushing back over use of their photos
- → Meta’s new Muse Image model can pull other Instagram users into AI photos
- → Muse Image is technically impressive, but Meta's use of Instagram photos raises questions
- → Meta wants its AI glasses to seem less creepy. Its AI strategy says otherwise.
- → Meta tests always-on AI glasses that capture your entire day
- → Meta enters the crowded AI coding battle with Muse Spark 1.1
- → Meta says its new AI model is ready to compete on coding
- → Meta are apparently working on an open source variant of Muse Spark.
- → Introducing Muse Spark 1.1
- → Meta's Muse Spark 1.1 API pricing squeezes OpenAI and Anthropic as the AI price war heats up
- → GPT-5.6 🚀, Muse Spark 1.1 ✨, ChatGPT Work 💼
- → Meta removes controversial AI feature on Instagram after backlash
- → Meta turns off the Instagram feature that let users make AI deepfakes of public accounts
中国开源浪潮. 中国模型现在占OpenRouter流量的30%以上;腾讯的HY3和GLM-5.2在家用硬件上运行,MiniMax计划发布2.7T参数的开源模型,DeepSeek的芯片设计表明了垂直整合,引发了美国的担忧。
- → This is what Hy3 is capable of. Mother of god.
- → llama.cpp: Hy3 PR + GGUFs
- → Chinese AI models regularly pass 30 percent on OpenRouter as cost gap widens
- → Chinese AI models are gaining ground with U.S. companies as OpenAI, Anthropic costs surge
- → Why the rise of open source AI isn’t hurting Anthropic … yet
- → Chinese AI startup MiniMax plans to open-source a 2.7 trillion parameter model later this year
- → 4-bit GLM-5.2 (753B MoE) on 4× DGX Spark: 70.8% on Terminal-Bench 2.1 vs 81.0% for the full model
- → Running GLM 5.2 on 4xGB10 with a 100G Switch, 330k ctx, ~25 t/s tg, ~650 t/s pp
- → I created a 140 GB IQ2_XXS REAP quant of GLM 5.2 for coding. Looking for testers.
- → Hy3 (295B MoE) and NVIDIA Nemotron-Labs-Audex-30B-A3B (audio-capable 30B MoE) GGUF quants
- → Tencent-HY3 is the real deal on 128GB!
- → The U.S. tech industry is increasingly anxious about the rising power and competitive price of open-source AI models from China — and whether the Trump administration will respond with yet another executive order | Politico
- → China's DeepSeek developing its own AI chip, sources say
代理工程与基础设施
Bun被64个代理重写. Claude Fable 5编排了64个并行实例,在11天内将Bun从Zig重写为Rust,生成了超过一百万行代码,修复了128个错误,花费16.5万美元。
SWE-Bench Pro审计失败. OpenAI发现约30%的SWE-Bench Pro任务有缺陷,因此撤回了认可,呼应了Artificial Analysis早先的担忧,削弱了代理编程评估。
Modal融资3.55亿美元. Modal获得了3.55亿美元,用于构建专为AI代理设计的云基础设施,专注于为自主软件提供沙盒化、快速迭代的环境。
Cloudflare临时代理. Cloudflare推出了临时账户,允许AI代理在没有永久凭证的情况下部署Workers,如果无人认领则会过期,从而简化代理驱动的工作流程。
代理冲击本地GPU. 新的量化技术和经济高效的GPU设备现在可以在消费级硬件上运行MoE模型,但基准测试表明低位量化会严重降低代理任务的表现,敦促谨慎选择量化方案。
- → Qwen3.5 122B is the best?
- → Qwen3.6-27b does not understand software architechure.
- → Qwen 3.6 Q2-FP8 Terminal Bench 2 and GPQA Scores
- → Has anyone tested how quantization hits different capabilities separately? My results are surprising.
- → 2.5x faster Qwen3.6 NVFP4 Unsloth quants
- → Ultra budget 20GB vram with 448GB/s for $100 bucks.
- → I benched quad 5060Tis for code generation with Qwen3.6-27B so you don't have to (it's really good)
安全、法律与可解释性
Jacobian Lens映射模型思维. Anthropic的Jacobian Lens揭示了一小组可言语化的表征,这些表征捕捉了模型的核心推理,实现了幻觉检测、输出引导,并且如社区演示所示,可以立即创建有害变体。
- → Anthropic Research - "Verbalizable Representations Form a Global Workspace in Language Models"
- → Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens
- → I tested Anthropic’s new Jacobian Lens on open models, then it turned into a local-model hallucination router
- → Anthropic found a hidden space where Claude puzzles over concepts
- → I created a super harmful model ! :D (by tweaking it's J-Space!!!)
代理驱动的勒索软件. Sysdig记录了一个AI代理处理勒索软件攻击的技术执行过程,尽管操作是由人类设置的,这突显了当前代理恶意软件的能力和局限。
AI滥用危机. 一起诉讼指控xAI的Grok生成了CSAM图像,而博科圣地利用聊天机器人进行攻击策划;这些事件凸显了随着代理能力的扩展,迫切需要安全护栏。
苹果起诉OpenAI. 苹果提起联邦诉讼,指控OpenAI挖走400多名员工以窃取硬件商业机密,可能用于制造竞争性的AI智能手机,加剧了AI行业的法律斗争。
这是本周回顾 - 下周日见。