モデルと競争
GPT-5.6 と Work Agent. OpenAIは3つの推論段階でGPT-5.6をリリースし、数時間にわたる自律タスク向けのChatGPT Workを開始したが、使用制限と混乱を招くインターフェースが苦情を引き起こし、修正の約束につながった。
- → OpenAI launches its new family of models with GPT-5.6
- → The new GPT-5.6 family: Luna, Terra, Sol
- → OpenAI rolls out GPT-5.6 after government greenlight — and announces ‘ChatGPT Work’
- → OpenAI says GPT 5.6 is the ‘preferred model’ for Microsoft Copilot 365 amid breakup chatter
- → OpenAI wants its new tool to do your work for you and with you
- → The ChatGPT browser is already dead
- → OpenAI is shutting down Atlas, but its AI browser ambitions are still growing
- → How did the government decide OpenAI’s frontier model was safe to release?
- → OpenAI pairs its GPT-5.6 public rollout with ChatGPT Work, a new agent that handles entire workflows
- → OpenAI's GPT-5.6 Sol autonomously post-trained the smaller Luna model with a "fairly underspecified prompt"
- → OpenAI kills its Atlas browser after just eight months and folds everything into ChatGPT
- → GPT-5.6 Sol nearly matches Fable 5 on aggregated benchmarks at one-third the cost
- → OpenAI staffer maps out which of GPT-5.6 Sol's five reasoning levels fits which task complexity
- → OpenAI admits it "didn't get everything quite right" with ChatGPT Work launch and scrambles to fix UX and costs
Claude Fable 5がベンチマークを制覇. AnthropicのClaude Fable 5はArtificial Analysisのベンチマークを制覇したが、その高コストのため同社は、より安価なモデルに委任するプランナーとして使用することを推奨し、ほとんどの機能を維持しながら費用を約40%削減した。
Metaのエージェント攻勢とプライバシー問題. Metaは、競合よりも低価格な高コンテキストコーディングモデルMuse Spark 1.1を立ち上げた一方、Muse ImageジェネレーターはInstagramの顔の無断使用を可能にした後に削除され、同意に関する議論が再燃した。
- → Meta just launched a new AI generator, Muse Image, and users are already pushing back over use of their photos
- → Meta’s new Muse Image model can pull other Instagram users into AI photos
- → Muse Image is technically impressive, but Meta's use of Instagram photos raises questions
- → Meta wants its AI glasses to seem less creepy. Its AI strategy says otherwise.
- → Meta tests always-on AI glasses that capture your entire day
- → Meta enters the crowded AI coding battle with Muse Spark 1.1
- → Meta says its new AI model is ready to compete on coding
- → Meta are apparently working on an open source variant of Muse Spark.
- → Introducing Muse Spark 1.1
- → Meta's Muse Spark 1.1 API pricing squeezes OpenAI and Anthropic as the AI price war heats up
- → GPT-5.6 🚀, Muse Spark 1.1 ✨, ChatGPT Work 💼
- → Meta removes controversial AI feature on Instagram after backlash
- → Meta turns off the Instagram feature that let users make AI deepfakes of public accounts
中国のオープンソースの波. 中国のモデルは現在OpenRouterトラフィックの30%以上を占めている。TencentのHY3とGLM‑5.2は家庭用ハードウェアで動作し、MiniMaxは2.7Tパラメータのオープンリリースを計画、DeepSeekのチップ設計は米国の懸念の中で垂直統合を示唆している。
- → This is what Hy3 is capable of. Mother of god.
- → llama.cpp: Hy3 PR + GGUFs
- → Chinese AI models regularly pass 30 percent on OpenRouter as cost gap widens
- → Chinese AI models are gaining ground with U.S. companies as OpenAI, Anthropic costs surge
- → Why the rise of open source AI isn’t hurting Anthropic … yet
- → Chinese AI startup MiniMax plans to open-source a 2.7 trillion parameter model later this year
- → 4-bit GLM-5.2 (753B MoE) on 4× DGX Spark: 70.8% on Terminal-Bench 2.1 vs 81.0% for the full model
- → Running GLM 5.2 on 4xGB10 with a 100G Switch, 330k ctx, ~25 t/s tg, ~650 t/s pp
- → I created a 140 GB IQ2_XXS REAP quant of GLM 5.2 for coding. Looking for testers.
- → Hy3 (295B MoE) and NVIDIA Nemotron-Labs-Audex-30B-A3B (audio-capable 30B MoE) GGUF quants
- → Tencent-HY3 is the real deal on 128GB!
- → The U.S. tech industry is increasingly anxious about the rising power and competitive price of open-source AI models from China — and whether the Trump administration will respond with yet another executive order | Politico
- → China's DeepSeek developing its own AI chip, sources say
エージェントエンジニアリングとインフラ
Bunが64エージェントによって書き換えられる. Claude Fable 5は64の並列インスタンスを調整して、11日間でBunをZigからRustに書き換え、100万行以上のコードを生成し、128のバグを修正した。費用は16万5000ドル。
SWE‑Bench Pro監査の失敗. OpenAIはSWE‑Bench Proのタスクの約30%が壊れていることを発見し、推奨を撤回した。これはArtificial Analysisの以前の懸念を反映し、エージェントコーディング評価を損なうものだ。
Modalが3億5500万ドルを調達. Modalは、AIエージェント向けに設計されたクラウドインフラを構築するため3億5500万ドルを確保した。自律ソフトウェア向けのサンドボックス型で高速な反復環境に焦点を当てている。
Cloudflareの一時エージェント. Cloudflareは、AIエージェントが永続的な認証情報なしでWorkersをデプロイできる一時アカウントを導入した。未使用の場合は有効期限が切れ、エージェント駆動のワークフローを円滑にする。
エージェントがローカルGPUに到達. 新しい量子化技術とコスト効率の良いGPUリグにより、コンシューマーハードウェアでMoEモデルが動作するようになったが、ベンチマークによると低ビット量子化はエージェントタスクを大幅に劣化させる可能性があり、慎重な量子化選択が求められる。
- → Qwen3.5 122B is the best?
- → Qwen3.6-27b does not understand software architechure.
- → Qwen 3.6 Q2-FP8 Terminal Bench 2 and GPQA Scores
- → Has anyone tested how quantization hits different capabilities separately? My results are surprising.
- → 2.5x faster Qwen3.6 NVFP4 Unsloth quants
- → Ultra budget 20GB vram with 448GB/s for $100 bucks.
- → I benched quad 5060Tis for code generation with Qwen3.6-27B so you don't have to (it's really good)
セキュリティ、法務、解釈可能性
Jacobian Lensがモデルの思考をマッピング. AnthropicのJacobian Lensは、モデルのコア推論を捉える言語化可能な表現の小さなセットを明らかにし、幻覚検出や出力操作を可能にし、コミュニティのデモが示すように、有害なバリアントの即時作成も可能にする。
- → Anthropic Research - "Verbalizable Representations Form a Global Workspace in Language Models"
- → Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens
- → I tested Anthropic’s new Jacobian Lens on open models, then it turned into a local-model hallucination router
- → Anthropic found a hidden space where Claude puzzles over concepts
- → I created a super harmful model ! :D (by tweaking it's J-Space!!!)
エージェント駆動のランサムウェア. Sysdigは、ランサムウェア攻撃の技術的実行を処理したAIエージェントを記録した。人間が操作を設定したが、エージェントマルウェアの現在の能力と限界を浮き彫りにしている。
AI悪用の危機. 訴訟ではxAIのGrokがCSAM画像を生成したと主張され、一方ボコ・ハラムは攻撃計画にチャットボットを使用している。これらの事件は、エージェント機能が拡大する中で安全策の緊急の必要性を強調している。
AppleがOpenAIを提訴. Appleは、OpenAIが400人以上の従業員を引き抜いてハードウェアの営業秘密を盗み、競合するAIスマートフォンを構築する可能性があるとして連邦訴訟を起こし、AI業界の法的戦いを激化させている。
今週のまとめ - また来週の日曜日に。