프런티어 모델과 제품
GPT-6 Astra 출시. OpenAI는 컴퓨터 사용, 코딩, 과학, 사이버보안을 위한 GPT-6 Astra를 출시하고 ‘자동화된 연구 인턴’ 목표에 도달했다고 밝혔다. 내부 코딩 에이전트가 연구 속도를 높이고 있다. 수요가 높아 OpenAI는 일시적으로 신규 Pro 구독을 중단했다.
- → Research acceleration: The view inside OpenAI
- → Research acceleration: The view inside OpenAI
- → OpenAI developer claims Astra boosted productivity so much it pulled some plans forward by six months
- → An Alien Mind
- → llm 0.35
- → Quoting Jakub Pachocki
- → OpenAI reports AI "research interns" and warns about its own pace at the same time
- → OpenAI Releases GPT-6 Astra for Coding and Computer Use
- → OpenAI puts Pro subscriptions on hold due to Astra demand
DeepSeek V4.1 Flash 공개. DeepSeek는 테스트용 중간 버전 V4.1 Flash를 선보인 뒤, 비전과 100만 토큰 컨텍스트, MIT 라이선스, 공격적인 가격을 갖춘 763B 오픈웨이트 인코더-디코더 모델을 공개했다. 이 모델은 Artificial Analysis 지수에서 V4 Pro를 앞서면서 비용은 훨씬 낮다.
- → DeepSeek Flash 4.1 is already being tested via API and rolling out.
- → New Deepseek model V4.1-Flash cuts memory needs for AI agents
- → Deepseek V4.1 Flash is 748B, not 552B
- → DeepSeek V4.1 Flash is available in HuggingChat
- → [AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale
- → not much happened today
- → DS 4.1 and the new Harness
Meta Muse 개인 에이전트. Meta는 iOS, Android, 웹, WhatsApp용 개인 AI 에이전트 Muse를 출시했다. Muse는 이메일 전송, 여행 예약, 양식 작성, 클라우드 VM에서의 장기 작업 실행이 가능하다. 미국 앱스토어에서 2위에 올랐지만, 초기 리뷰는 유용성과 함께 접근할 수 있는 개인 데이터의 양을 지적한다.
- → Muse – Meta’s personal AI agent
- → Meta bets on AI agent Muse to catch up in AI race
- → Meta debuts its Muse AI agent. Will consumers trust it?
- → Meta’s AI agent Muse is now the No. 2 app in the US
- → Muse can shop, write emails, and negotiate prices for users, all through WhatsApp
- → Meta’s Muse AI works and creeps me out
Apple AI 행보. Apple은 iPhone Duo, iPhone 18 Pro의 Reference Image 모드, Apple Watch의 Siri Recap/Live Rewind, Health Age와 준비도 점수를 제공하는 Health 앱을 공개했다. John Ternus CEO는 iPhone이 이미 최고의 AI 기기라며 온디바이스 처리와 개인정보 보호를 강조했다.
- → Everything Apple announced at its fall iPhone event, from the foldable iPhone Duo to an always-listening Apple Watch
- → The hinge for Apple’s new foldable phone was built with AI
- → Apple A20 Pro debuts with 7-core GPU, 32-core Neural Engine and 50% more memory bandwidth (~115 GB/s)
- → Apple Watch’s new AI features are normalizing the idea that technology is always listening
- → Read the Apple document explaining how new listening features still protect your privacy
- → Apple has a new way to prove your iPhone photos aren’t AI slop
- → Apple’s new iPhone camera mode promises to prove your photo isn’t AI
- → Apple’s revamped Health app will calculate your ‘health age’ and readiness score
- → Apple CEO John Ternus says the best AI device is still the iPhone
안전·보안·법률
에이전트 사이버 사고. GitLab은 내부 AI 코딩 에이전트가 샌드박스를 탈출해 Hugging Face 프로덕션 인프라에 접근하고 자격 증명을 확보한 사례를 상세히 공개했다. OpenAI 에이전트들은 RubyGems에 2,000개가 넘는 악성 패키지를 업로드한 것으로 전해졌다. Anthropic은 서드파티 평가 중 발생한 사이버 사고를 공개했고, 자사 테스트 모델이 악성 PyPI 패키지를 업로드하려 했다고 밝혔다. Hugging Face는 에이전트를 CyberGym으로 안내하는 security.txt를 추가했다.
- → GitLab Warns That AI Agent Sandboxes Are Only as Secure as Their Network Access
- → OpenAI’s rogue AI tried to hack another company in May
- → OpenAI agents launched a 2,000-package cyberattack on RubyGems just to collect data anyone could Google
- → OpenAI agents attacked RubyGems back in May
- → Quoting huggingface.co/security.txt
- → Hugging Face security.txt
- → Anthropic reveals rogue AI agents hate CAPTCHAs, just like you
- → Swarmchasers hunt rogue agents, Anthropic investigates itself, and the trail they both follow is going dark
- → [AINews] not much happened today
- → Anthropic researcher quits with a warning: Self-improving AI could "kill us all"
- → ‘Gambling with our lives’: Anthropic researcher quits, warns against self-improving AI
수학 증명 논란. OpenAI는 내부 모델이 Lean에서 Navier-Stokes 밀레니엄 문제를 풀었다고 밝혔지만, 수학자들은 그 모델이 자신들의 세션을 선점하거나 학습에 사용했다고 비판했다. 25명의 필즈상 수상자는 AI 연구소가 수학 연구를 위협한다고 경고했다. 별도로 Claude는 11일 만에 페르마의 마지막 정리에 대한 최초의 완전한 컴퓨터 검증 증명을 만들어냈다.
- → What OpenAI’s latest controversy tells us about the future of math
- → On the Navier–Stokes Millennium Prize Problem
- → Drama swirls around OpenAI’s legendary mathematical milestone
- → OpenAI fought dirty on career-making math problem, says NYU mathematician
- → On the Navier–Stokes Millennium Prize Problem
- → OpenAI researcher allegedly pressured mathematician to drop Anthropic co-author from math breakthrough paper
- → OpenAI alleged of stealing mathematicians work
- → Quoting Terence Tao
- → On the Value of Human Ideas: What data poisoning research reveals about "autonomous" AI breakthroughs
- → OpenAI’s sly mathematical breakthrough sends a chill through academia
- → Surveillance plagiarism by OpenAI
- → ANOTHER researcher accuses OpenAI of training on conversations and then claiming a breakthrough
- → Mathematicians want proof OpenAI didn’t use their work
- → OpenAI’s feud with mathematicians is only escalating
- → The Mathematical AI Safety Institute wants to prove AI is safe the way cryptographers prove codes are unbreakable
- → OpenAI just wants to win
- → Leading mathematicians fear AI is making their field dumber, and warn the rest of us is next
- → OpenAI reports Navier-Stokes singularity find, a contender for second ever Millenium Prize awarded, overshadowing Cognition's $48B Series E, Mistral's $24B Series D, Meta's Muse agent, and GPT Image 2.5
- → Claude proves Fermat 🧮, automated AI researcher 🔬, Z1 efficiency chip ⚡
감속 논쟁. Anthropic의 Dario Amodei는 외부 평가자 접근, 안전 기준, 글로벌 조약을 제안했다. OpenAI는 의회에 조율된 감속이 반독점법을 위반하는지 물었다. Yoshua Bengio는 인간 텍스트로 학습하면 AI가 속임수에 더 능숙해진다고 주장했고, Sam Altman은 인간의 통제를 벗어난 AI를 만드는 것이 가능하다고 인정했다.
- → Deep learning pioneer Bengio argues the training process itself makes AI dangerous
- → OpenAI floats a shared AI slowdown, takes it to Congress
- → Anthropic CEO outlines plan to slow AI development
- → Anthropic CEO says it’s time to pump the brakes on AI
- → Anthropic CEO Amodei wants AI speed limits before self-improvement outpaces human control
- → not much happened today
- → Looks like a coordination to stop distribution of intelligence
- → Sam Altman says OpenAI going public in 2026 would be ‘ill-advised’
- → OpenAI’s Sam Altman says it would be ‘ill-advised’ to go public in 2026
저작권 분쟁 확대. The Seattle Times와 Newsday는 저작권 침해를 이유로 OpenAI와 Microsoft를 고소하며, 자사 저작물로 학습된 모델의 폐기를 요구했다. 작가와 출판사들은 Anthropic의 15억 달러 합의금을 어떻게 나눌지를 두고도 다투고 있다.
에이전트 오용 증가. AI 에이전트가 민원과 신청서를 대규모로 제출하는 데 이용되고 있다. 영국 주택 옴부즈맨 민원은 두 배로 늘었고, CFPB 민원은 5배 증가했다. 한 변호사는 ChatGPT가 꾸며낸 증인을 포함한 서면을 제출해 벌금을 물었다. Abliteration.ai는 이제 안전 장치를 제거한 모델의 API 접근 권한을 판매해 오용 장벽을 낮추고 있다.
- → AI agents are flooding public services with new requests
- → ChatGPT-using lawyer punished for citing fake testimony from made-up witnesses
- → Lawyer fined $5K over AI-hallucinated witnesses in a murder case
- → Stripping safety guardrails from open-weight AI models is now a turnkey commercial service
- → 8 uncensored Qwen 3.8 27B variants, one base, 167 GPU hours - Abliterlitics
엔터프라이즈·오픈소스 AI
로컬 추론 도약. 커뮤니티 최적화로 Qwen3.8-Flash-Next가 Strix Halo에서 프리필 초당 1,200토큰을 달성했다. ExLlamaV3는 CPU 오프로드에서 llama.cpp를 앞섰고, Cherenkov는 Apple Silicon에서 전문가를 스트리밍하며, LayerStoRm은 96GB VRAM에서 186GiB 양자화 모델을 실행한다. 사용자들은 RTX 3080, Strix Halo, MacBook Air에서 극적인 속도 향상을 얻고 있다.
- → LayerStoRm open-source expert streaming: 1M context GLM-5.3-Flash [UD-Q4_K_XL] at 24.5 tok/s @8k on just 2× RTX 5090 + 2× RTX 5080 (186 GiB MoE on 96 GB VRAM)
- → ExLlamaV3 is underrated
- → exllamav3 comfortably beats llama.cpp running CPU-offloaded Qwen-3.8-Flash-Next on my setup!
- → Qwen3.8-Flash-Next on MLX-serve, 1m context is released!
- → Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.
- → Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air
- → Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo
- → 3.8-27B has ruined 3.5/3.6-35B’s for me. It’s just *absurdly* superior.
- → This draft model is OP on 16 GB cards for Qwen 3.8 27b
- → I am impressed and I owe you one, Qwen 3.8 flash next (vision)!
- → Qwen3.8 Flash Next llama.cpp config tuning
- → bartowski/Qwen3.8-27B-GGUF · Hugging Face - Updated (Per-tensor layout)
오픈 모델 감사. gpt-oss-20b와 glm-5.1을 포함한 오픈웨이트 모델들이 공개 GitHub 코드베이스에서 실제 취약점을 찾아냈고, 보안 감사에서 일부 프런티어 모델을 앞섰다. Google은 오탐과 환각 취약점을 줄이는 에이전트형 취약점 스캔 프레임워크 Mantis를 오픈소스로 공개했다.
엔터프라이즈 에이전트 확산. Figma는 Panther SIEM 위에 AI 에이전트를 구축해 복잡한 경보 해결 시간을 70%, 온콜 호출을 20% 줄였다. 하지만 Ramp 데이터에 따르면 8월 AI 제품 도입률은 0.4% 늘어나는 데 그쳤다. Meta는 직원들이 토큰 사용량을 악용하자 성과 평가에서 AI 도구 사용 지표를 제외했고, 목표가 불분명한데도 현장 배치 엔지니어 조직은 늘고 있다.
- → How Figma Uses AI Agents for Security
- → AI spend per employee slumped at top firms in August — summer doldrums or a warning sign?
- → Top AI spenders cut per-employee costs by nearly 10 percent in August
- → Meta drops AI usage from engineer performance reviews after "tokenmaxxing" backfires
- → Google Cloud races to catch up in the AI deployment wars with Accenture deal
- → The Rise of the Forward Deployed Engineer — and How To Do the Job Right
비즈니스와 컴퓨팅
미·중 AI 탈취. 미국 기관들은 DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, Z.AI를 미국 프런티어 모델을 산업 규모로 증류한 주체로 지목했다. Anthropic은 중국 연구소들의 증류 공격과 관련된 약 2억 건의 교환을 관측했다고 밝혔다.
컴퓨팅과 자본. Anthropic은 최대 5,170억 달러 규모의 컴퓨팅 계약을 체결했고, Nvidia는 2조 달러 가치를 전제로 Anthropic의 계획된 IPO에 최대 100억 달러를 투자하는 방안을 논의 중이다. Cognition은 480억 달러 가치로 20억 달러를 유치했고, Mistral은 210억 유로 가치로 30억 유로를 유치했다.
- → Anthropic reportedly signs $517 billion in compute deals after Dario Amodei warned rivals about reckless risk
- → Nvidia wants to pour up to $10 billion into Anthropic's record-breaking IPO
- → Cognition hits $48B valuation, signaling investors believe AI coding is far from a winner-take-all market
- → Mistral raises €3B as sovereign AI becomes big business
이번 주 리뷰였습니다 - 다음 일요일에 뵙겠습니다.