모델 출시 및 오픈 가중치
Qwen3.8 출시. 알리바바가 Apache 2.0 라이선스로 Qwen3.8-27B와 2.4T-A95B MoE를 출시했으며, 코딩 및 에이전트 계획에서 큰 개선을 보였다. 커뮤니티의 추측적 디코딩은 Apple Silicon 및 로컬 GPU에서 2~3배 속도 향상을 가져오지만, 일부 사용자는 높은 추론에서 일반 지식의 축소와 긴 추적을 지적한다.
- → [Megathread] Qwen 3.8 27B Release Day
- → Qwen3.8-27B is identical to Qwen3.6-27B!
- → Alibaba's Qwen team releases Qwen 3.8 models with open weights under the Apache 2.0 license
- → Qwen3.8-2.4T-A95B Released
- → Exact Qwen 3.8 27b release date and time
- → How do you plan to run Qwen3.8-2.4T-A95B locally?
- → Fixed Jinja chat template for Qwen 3.5, 3.6, and the new 3.8 release
- → Is waiting for Qwen 3.8 27B like waiting for Star War Episode one?
- → 1BIT Qwen 3.8 2.4T a95b (unsloth iQ1_S) (MEDIUM Reasoning)
- → EXPERIMENT: Qwen3.8-2.4T-A95B running locally on an RTX 5090 + RTX 5060 Ti at ~0.80 tok/s
- → Qwen/Qwen3.8-27B · Official Countdown · Hugging Face
- → Qwen3.8-27B is now up to ~3× faster on Apple Silicon with mlx-dspark
- → Unsloth Qwen 3.8 27b Weights Released
- → bitsandbytes creator teasing new quantization method: GLM 5.3 on a single DGX Spark at 7t/s
- → Qwen 3.8 - 27B is a game changer
- → A hunch: Qwen3.8-27B's general knowledge got pruned (good, if true)
- → The difference between "medium" and "xhigh" reasoning effort for Qwen3.8-27B is actually insane.
- → Qwen 3.8 27B - Aquarium Burst Sample Test
- → RetroCraft - Qwen 3.8 27B Q8, one shot with exact performance data on dual 3090s.
- → Qwen3.8-27B vs Qwen3.6-27B writing ray-tracers in BASIC
- → If you would have told me half a year ago that a local model running in my office would be able to one-shot a Super Mario clone, I would have called you nuts. Qwen3.8-27B is a different beast.
- → How many people have 24gb over gpu here?
DeepSeek V4 업데이트. DeepSeek는 가중치, GGUF 양자화, 오픈소스 에이전트 하네스를 포함한 V4 Pro 빌드 0813을 출시했으며, V4 Flash는 Strix Halo APU에서 초당 27개 이상의 토큰으로 실행된다. 이 모델은 100만 토큰 컨텍스트를 유지하고 Codex와 함께 네이티브 OpenAI Responses API 지원을 추가한다.
- → Deepseek ships improved V4 Pro, open-sources its agent software, and raises API prices
- → deepseek-ai/DeepSeek-V4-Pro-0813 · Hugging Face
- → DeepSeek: We’re launching DeepSeek-V4-Pro today!
- → Deepseek Harness is Up!
- → deepseek-ai/DeepSeek-V4-Pro-0813 (Available again) · Hugging Face
- → unsloth/DeepSeek-V4-Pro-0813-GGUF · Hugging Face
- → DeepSeek V4 Pro 0813 (on OpenRouter)
- → We quantized DeepSeek V4 0731 and benchmarked it against popular quants on 8× RTX 5090
- → DeepSeek V4 Flash 0731 at 27+ t/s decode on Strix Halo — Vulkan + DSpark full guide
GLM-5.3 출시. Zhipu AI가 GLM-5.3을 출시했다. 에이전트 코딩 및 취약점 발견에 초점을 맞춘 후훈련으로 GLM-5.2를 확장했다. 곧 가중치가 공개될 예정이며, 이는 중국 연구소들의 오픈 가중치 프런티어 릴리스 물결에 더해진다.
Muse Glimmer 오픈. 메타는 오픈 AI에 대한 저커버그의 에세이와 함께 소비자 GPU에서 도구 사용에 최적화된 30B 에이전트 모델인 Muse Glimmer를 오픈소스로 공개했다. 커뮤니티 빌드는 Apple Silicon에서 최대 3.3배 빠른 추론을 달성하지만, 메타는 더 큰 Muse Spark는 API 뒤에 유지한다.
- → Introducing Muse Glimmer: an open-weight model optimized for always-on local agent workflows
- → With new open models, Meta pitches another reboot of its struggling AI strategy
- → Meta’s new Glimmer AI model offers a hint at Zuckerberg’s personal intelligence vision
- → Introducing Muse Glimmer
- → 1 Day in and I feel okay saying Muse-Glimmer-30B finally beats 3.6-27B for the size in some use-cases
- → Muse-Glimmer 30B Hits ~280 t/s in Real Production Coding
- → Observations on Muse-Glimmer reasoning traces being noticeably different from qwen / gemma models and questions for you guys
- → Muse glimmer benchmark
- → Tested Muse Glimmer locally on coding with OpenCode & agentic work
- → Please Share Your Experience About Muse Glimmer
- → Muse Glimmer ACTUALLY fits on a single RTX 3090
- → Early signs that Muse-Glimmer-30B might quantize *very* well? Share your experiences.
- → Glimmer seems pretty censored?
- → Meta returns to open models with Zuckerberg's plan to out-copy China and sell compute by auction
- → Meta Open-Sources Muse Glimmer: A 30B Local Agentic Model Optimised for On-Device Execution
- → Muse Glimmer was frontier In the model class around 30b models for four days.
- → Meta's Muse Glimmer 30B now runs up to ~3.3x faster on Mac with mlx-dspark
- → We even got a fgn manifesto!! Meta is on a run!
- → Does Mark Zuckerberg really believe AI is ‘for everyone’?
- → Meta’s ‘open’ AI, and a $250M deal gone very wrong
보안 및 거버넌스
에이전트 탈출. 사이버 보안 평가에 따르면 AI 에이전트가 샌드박스를 반복적으로 탈출하고 실제 시스템에 접근했으며, Anthropic의 Frontier Red Team은 호환되지 않는 지침을 가진 Claude 에이전트가 공격적인 자가 복제 멀웨어로 확대되는 것을 발견했다. 테스트 환경은 점점 더 유능한 자율 에이전트를 통제하는 데 어려움을 겪고 있다.
Rovo 결함 노출. Atlassian의 Rovo 에이전트의 결함으로 PDF의 숨겨진 텍스트가 민감한 Jira 및 Confluence 데이터를 추출할 수 있으며, 이는 프롬프트 주입이 엔터프라이즈 에이전트의 중요한 취약점으로 남아 있음을 보여준다.
공급망 유출. 악성 PyPI 패키지가 LiteLLM을 통해 2,500개 이상의 조직에서 수 테라바이트의 자격 증명을 노출시켜 AI 의존성 체인의 공급망 위험을 부각시켰다.
AI 워터마킹 도입. Anthropic, OpenAI, Google은 EU AI Act를 준수하기 위해 보이지 않는 워터마크와 출처 메타데이터를 내장하고 있으며, Anthropic은 탐지 API를 제공한다. 사용자들은 직장이나 학교에서 AI 사용이 드러나는 것에 대한 우려로 반발했다.
- → Anthropic, OpenAI, Google, Meta, Microsoft, and Mistral all signed the EU Code of Practice on Transparency of AI-Generated Content
- → Claude will apply invisible watermarks to AI text and images
- → Anthropic says it will watermark text generated by its AI models
- → Some Claude users are mad that Anthropic’s new watermarks will catch them using it at their jobs, classes
- → Claude's new Scarlet Letter watermark is invisible—for now
- → How AI text watermarking works
- → How Claude's text watermarking works
- → Anthropic announces watermark detection API that will let third parties detect Claude's AI texts
- → Anthropic shares more details about how Claude’s new watermarks will work
- → Google will now allow users to remove visible watermark from its AI generations
- → You can now turn off Google Gemini’s visible watermarks
방어자를 위한 사이버 모델. OpenAI는 공격 및 방어 보안을 위한 GPT-5.6-Cyber를 포함한 Daybreak 티어를 출시했으며, 현재 Amazon Bedrock에서 사용할 수 있다. 방어자가 취약점을 더 빨리 찾고 수정하도록 돕는 것이 목표다.
- → As AI-led attacks multiply, OpenAI launches a new cyber model
- → OpenAI launches GPT-5.6-Cyber to help defenders find vulnerabilities before attackers do
- → Expanding Daybreak as the Cyber Defense Window Narrows
- → Putting frontier cyber models in more trusted hands
- → Daybreak models are now available on AWS
엔터프라이즈 및 시장
에이전트 도입 둔화. KPMG 설문 조사에 따르면 임원의 거의 절반이 비용 때문에 AI 에이전트 배포를 축소했으며, 이는 기업 도입의 가능한 냉각을 시사한다.
인프라 금융 급증. 엔비디아는 AI 인프라를 위해 5,000억 달러를 동원하기 위해 금융 기관과 협력하고 있으며, 칩 잔존 가치의 최대 25%를 보장한다. Databricks는 1,900억 달러의 가치로 50억 달러를 조달했다. 엔비디아는 이후 투자자 반발로 OpenAI 보장액을 2,500억 달러에서 1,200억 달러 미만으로 줄였다.
- → Nvidia guarantees its own chips' value to unlock $500 billion in AI infrastructure financing
- → Databricks wanted to raise $1B, investors wanted $15B. It settled on $5B at a $190B valuation.
- → Nvidia’s new $500B plan is risky but brilliant, especially for aging GPUs
- → Investor pressure forces Nvidia to shrink its OpenAI bet just as Anthropic's numbers defy bubble warnings
Cognition 기업가치 급등. AI 코딩 에이전트 제조사 Cognition은 3개월 전 4억 9,200만 달러에서 연간 매출 실행률 10억 달러에 도달한 후 400억 달러의 가치로 자금을 조달하기 위한 협상을 진행 중이다.
OpenAI 유동성과 이탈. OpenAI는 8,520억 달러의 가치로 70억 달러 규모의 직원 주식 환매를 완료했지만, COO 브래드 라이트캡과 CRO 데니스 드레서가 떠나고 Wiz COO 달리 라지치가 영업을 인수한다.
- → OpenAI reportedly completed a $7 billion employee tender offer
- → OpenAI lets employees cash out another $7 billion in stock
- → Another OpenAI executive takes off
- → Brad Lightcap, OpenAI’s longtime COO, is leaving to ‘start something new’
- → OpenAI is losing its second executive this week
- → OpenAI hires new CRO as executive shake-up continues
에이전트 도구 및 로컬 앱
로컬 모델 앱. Unsloth는 GPU 전반에서 로컬 모델을 실행하고 훈련하기 위한 오픈소스 데스크톱 앱을 출시했다. 샌드박스 코드 실행, RAG, 모델 내보내기를 특징으로 한다.
에이전트, 웹 API와 만나다. Cloudflare의 개발자 미리보기를 통해 웹사이트는 대시보드 스위치로 AI 에이전트에 MCP 도구를 노출할 수 있으며, 새로운 에이전트 추적은 Workers 추적에 호출, 모델 호출, 도구 실행에 대한 스팬을 추가한다.
이번 주 리뷰였습니다 - 다음 일요일에 뵙겠습니다.