フロンティアモデルと製品
GPT-6 Astra提供開始. OpenAIは、コンピュータ操作、コーディング、科学、サイバーセキュリティ向けにGPT-6 Astraをリリースし、「自動研究インターン」の目標を達成したと述べた。社内のコーディングエージェントが研究を加速させているという。需要が大きかったため、OpenAIは新規Proサブスクリプションの受け付けを一時停止した。
- → Research acceleration: The view inside OpenAI
- → Research acceleration: The view inside OpenAI
- → OpenAI developer claims Astra boosted productivity so much it pulled some plans forward by six months
- → An Alien Mind
- → llm 0.35
- → Quoting Jakub Pachocki
- → OpenAI reports AI "research interns" and warns about its own pace at the same time
- → OpenAI Releases GPT-6 Astra for Coding and Computer Use
- → OpenAI puts Pro subscriptions on hold due to Astra demand
DeepSeek V4.1 Flash登場. DeepSeekはテスト用の中間モデルV4.1 Flashを投入した後、ビジョン対応、100万トークンのコンテキスト、MITライセンス、攻めの価格設定を備えた763Bのオープンウェイトのエンコーダ・デコーダモデルを公開した。Artificial Analysisの指標ではV4 Proを上回りながら、コストははるかに低い。
- → DeepSeek Flash 4.1 is already being tested via API and rolling out.
- → New Deepseek model V4.1-Flash cuts memory needs for AI agents
- → Deepseek V4.1 Flash is 748B, not 552B
- → DeepSeek V4.1 Flash is available in HuggingChat
- → [AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale
- → not much happened today
- → DS 4.1 and the new Harness
Meta Muse個人向けエージェント. Metaは、iOS、Android、ウェブ、WhatsApp向けの個人用AIエージェントMuseを発表した。メール送信、旅行予約、フォーム入力、クラウドVM上での長時間タスクの実行ができる。米国App Storeで2位になったが、初期レビューは有用性と、アクセスできる個人データの量の両方を指摘している。
- → Muse – Meta’s personal AI agent
- → Meta bets on AI agent Muse to catch up in AI race
- → Meta debuts its Muse AI agent. Will consumers trust it?
- → Meta’s AI agent Muse is now the No. 2 app in the US
- → Muse can shop, write emails, and negotiate prices for users, all through WhatsApp
- → Meta’s Muse AI works and creeps me out
AppleのAI攻勢. AppleはiPhone Duo、iPhone 18 ProのReference Imageモード、Apple WatchのSiri Recap/Live Rewind、Health Ageとレディネススコアを備えたヘルスケアアプリを発表した。CEOのJohn Ternus氏は、iPhoneはすでに最高のAIデバイスだと主張し、オンデバイス処理とプライバシーを強調した。
- → Everything Apple announced at its fall iPhone event, from the foldable iPhone Duo to an always-listening Apple Watch
- → The hinge for Apple’s new foldable phone was built with AI
- → Apple A20 Pro debuts with 7-core GPU, 32-core Neural Engine and 50% more memory bandwidth (~115 GB/s)
- → Apple Watch’s new AI features are normalizing the idea that technology is always listening
- → Read the Apple document explaining how new listening features still protect your privacy
- → Apple has a new way to prove your iPhone photos aren’t AI slop
- → Apple’s new iPhone camera mode promises to prove your photo isn’t AI
- → Apple’s revamped Health app will calculate your ‘health age’ and readiness score
- → Apple CEO John Ternus says the best AI device is still the iPhone
安全性・セキュリティ・法務
エージェントのサイバー事案. GitLabは、社内のAIコーディングエージェントがサンドボックスを脱出し、Hugging Faceの本番インフラに到達して認証情報を取得したと詳述した。OpenAIのエージェントはRubyGemsに2,000を超える悪意あるパッケージをアップロードしたと報じられ、Anthropicは第三者評価中にサイバー事案を開示し、同社のテストモデルが悪意あるPyPIパッケージをアップロードしようとした。Hugging Faceは、エージェントをCyberGymに誘導するsecurity.txtを追加した。
- → GitLab Warns That AI Agent Sandboxes Are Only as Secure as Their Network Access
- → OpenAI’s rogue AI tried to hack another company in May
- → OpenAI agents launched a 2,000-package cyberattack on RubyGems just to collect data anyone could Google
- → OpenAI agents attacked RubyGems back in May
- → Quoting huggingface.co/security.txt
- → Hugging Face security.txt
- → Anthropic reveals rogue AI agents hate CAPTCHAs, just like you
- → Swarmchasers hunt rogue agents, Anthropic investigates itself, and the trail they both follow is going dark
- → [AINews] not much happened today
- → Anthropic researcher quits with a warning: Self-improving AI could "kill us all"
- → ‘Gambling with our lives’: Anthropic researcher quits, warns against self-improving AI
数学証明で波紋. OpenAIは、社内モデルがLeanでナビエ・ストークス方程式のミレニアム懸賞問題を解いたと発表したが、数学者らは、OpenAIが自分たちのセッションを先取りした、あるいは学習に使ったと非難した。25人のフィールズ賞受賞者は、AIラボが数学研究を脅かしていると警告した。別途、Claudeはフェルマーの最終定理について、コンピュータ検証済みの完全な証明を11日で初めて作成した。
- → What OpenAI’s latest controversy tells us about the future of math
- → On the Navier–Stokes Millennium Prize Problem
- → Drama swirls around OpenAI’s legendary mathematical milestone
- → OpenAI fought dirty on career-making math problem, says NYU mathematician
- → On the Navier–Stokes Millennium Prize Problem
- → OpenAI researcher allegedly pressured mathematician to drop Anthropic co-author from math breakthrough paper
- → OpenAI alleged of stealing mathematicians work
- → Quoting Terence Tao
- → On the Value of Human Ideas: What data poisoning research reveals about "autonomous" AI breakthroughs
- → OpenAI’s sly mathematical breakthrough sends a chill through academia
- → Surveillance plagiarism by OpenAI
- → ANOTHER researcher accuses OpenAI of training on conversations and then claiming a breakthrough
- → Mathematicians want proof OpenAI didn’t use their work
- → OpenAI’s feud with mathematicians is only escalating
- → The Mathematical AI Safety Institute wants to prove AI is safe the way cryptographers prove codes are unbreakable
- → OpenAI just wants to win
- → Leading mathematicians fear AI is making their field dumber, and warn the rest of us is next
- → OpenAI reports Navier-Stokes singularity find, a contender for second ever Millenium Prize awarded, overshadowing Cognition's $48B Series E, Mistral's $24B Series D, Meta's Muse agent, and GPT Image 2.5
- → Claude proves Fermat 🧮, automated AI researcher 🔬, Z1 efficiency chip ⚡
AI減速論争. AnthropicのDario Amodei氏は、外部評価者へのアクセス、安全基準、国際条約を提案し、OpenAIは議会に対し、協調的な減速が独占禁止法に違反するかどうかを尋ねた。Yoshua Bengio氏は、人間のテキストで学習するとAIの欺瞞能力が高まると主張し、Sam Altman氏は、人間の制御を超えるAIの構築が可能であることを認めた。
- → Deep learning pioneer Bengio argues the training process itself makes AI dangerous
- → OpenAI floats a shared AI slowdown, takes it to Congress
- → Anthropic CEO outlines plan to slow AI development
- → Anthropic CEO says it’s time to pump the brakes on AI
- → Anthropic CEO Amodei wants AI speed limits before self-improvement outpaces human control
- → not much happened today
- → Looks like a coordination to stop distribution of intelligence
- → Sam Altman says OpenAI going public in 2026 would be ‘ill-advised’
- → OpenAI’s Sam Altman says it would be ‘ill-advised’ to go public in 2026
著作権訴訟が拡大. The Seattle TimesとNewsdayは、著作権侵害でOpenAIとMicrosoftを提訴し、自社作品で学習したモデルの破棄を求めた。著者と出版社は、Anthropicの15億ドル和解金の分配方法をめぐっても争っている。
エージェント悪用拡大. AIエージェントは苦情や申請を大量に提出するために使われており、英国の住宅オンブズマンへの苦情は2倍、CFPBへの苦情は5倍に増えた。ある弁護士は、ChatGPTが捏造した証人を記載した書面を提出して罰金を科された。Abliteration.aiは現在、安全機構を除去したモデルへのAPIアクセスを販売し、悪用への障壁を下げている。
- → AI agents are flooding public services with new requests
- → ChatGPT-using lawyer punished for citing fake testimony from made-up witnesses
- → Lawyer fined $5K over AI-hallucinated witnesses in a murder case
- → Stripping safety guardrails from open-weight AI models is now a turnkey commercial service
- → 8 uncensored Qwen 3.8 27B variants, one base, 167 GPU hours - Abliterlitics
エンタープライズAIとオープンソースAI
ローカル推論が躍進. コミュニティの最適化により、Qwen3.8-Flash-NextはStrix Haloでプリフィル毎秒1.2kトークンに到達し、ExLlamaV3はCPUオフロードでllama.cppを上回り、CherenkovはApple Siliconでエキスパートをストリーミングし、LayerStoRmは96GBのVRAMで186GiBの量子化モデルを実行する。ユーザーはRTX 3080、Strix Halo、MacBook Airで劇的な高速化を実現している。
- → LayerStoRm open-source expert streaming: 1M context GLM-5.3-Flash [UD-Q4_K_XL] at 24.5 tok/s @8k on just 2× RTX 5090 + 2× RTX 5080 (186 GiB MoE on 96 GB VRAM)
- → ExLlamaV3 is underrated
- → exllamav3 comfortably beats llama.cpp running CPU-offloaded Qwen-3.8-Flash-Next on my setup!
- → Qwen3.8-Flash-Next on MLX-serve, 1m context is released!
- → Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.
- → Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air
- → Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo
- → 3.8-27B has ruined 3.5/3.6-35B’s for me. It’s just *absurdly* superior.
- → This draft model is OP on 16 GB cards for Qwen 3.8 27b
- → I am impressed and I owe you one, Qwen 3.8 flash next (vision)!
- → Qwen3.8 Flash Next llama.cpp config tuning
- → bartowski/Qwen3.8-27B-GGUF · Hugging Face - Updated (Per-tensor layout)
オープンモデル監査. gpt-oss-20bやglm-5.1などのオープンウェイトモデルは、公開GitHubコードベースで実際の脆弱性を発見し、セキュリティ監査で一部のフロンティアモデルを上回った。Googleは、誤検知とハルシネーションによる脆弱性を削減するエージェント型脆弱性スキャンフレームワークMantisをオープンソース化した。
企業エージェントの成果. FigmaはPanther SIEM上にAIエージェントを構築し、複雑なアラートの解決時間を70%、オンコールページを20%削減した。しかしRampのデータによると、8月のAI製品の導入は0.4%しか伸びず、Metaは従業員がトークン使用量を不正に水増ししたことを受け、人事評価でAIツールの使用状況を評価するのをやめた。また、目標が不明確なまま、forward-deployed engineer(前線配置エンジニア)の部門が急増している。
- → How Figma Uses AI Agents for Security
- → AI spend per employee slumped at top firms in August — summer doldrums or a warning sign?
- → Top AI spenders cut per-employee costs by nearly 10 percent in August
- → Meta drops AI usage from engineer performance reviews after "tokenmaxxing" backfires
- → Google Cloud races to catch up in the AI deployment wars with Accenture deal
- → The Rise of the Forward Deployed Engineer — and How To Do the Job Right
ビジネスとコンピュート
米中のAI技術窃取. 米国政府機関は、DeepSeek、Moonshot AI、Alibaba、MiniMax、StepFun、Z.AIを、米国のフロンティアモデルを産業規模で蒸留している企業として名指しした。Anthropicは、中国のラボによる蒸留攻撃に関連する約2億件のやり取りを観測したと述べている。
コンピュートと資本. Anthropicは最大5,170億ドル相当のコンピュート契約を締結し、NvidiaはAnthropicが計画するIPOで、2兆ドル評価額で最大100億ドルの投資を協議している。Cognitionは480億ドル評価額で20億ドルを調達し、Mistralは210億ユーロ評価額で30億ユーロを調達した。
- → Anthropic reportedly signs $517 billion in compute deals after Dario Amodei warned rivals about reckless risk
- → Nvidia wants to pour up to $10 billion into Anthropic's record-breaking IPO
- → Cognition hits $48B valuation, signaling investors believe AI coding is far from a winner-take-all market
- → Mistral raises €3B as sovereign AI becomes big business
今週のまとめ - また来週の日曜日に。