오픈 모델 및 로컬 추론
Qwen 3.8 27B. Alibaba의 Apache-2 라이선스 Qwen 3.8 27B는 고성능 노트북과 로컬 GPU에서 실행된다. xhigh 추론 모드는 강력한 결과를 내지만 과잉 사고를 일으킨다. 커뮤니티 최적화로 RTX 3090에서 초당 381토큰을 달성했고, 양자화 버전은 16GB VRAM에서 구동된다. 35B MoE는 공개되지 않지만, 다음 주에 새로운 중간 규모 오픈웨이트 모델이 나올 예정이다.
- → Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things
- → Simon Willison: Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things
- → I just ran Qwen 3.8 27 in Q4 against GPT 5.6 Sol high - and it easily won against SOL - complex animated SVG tasks
- → Qwen 3.8 27b vs 3.6 27b - how good is with a Turtle library.
- → Anyone else get a kick out of Qwen 3.8 27B Reasoning Dialogue?
- → Share your favorite thoughts and reasoning from running Qwen 3.8 27b. This is mine.
- → Qwen3.8 27B reasoning effort low/medium/xhigh comparison
- → Qwen 3.8 2.4T at 288k tokens/s on Nvidia GB300 NVL72
- → Newer commits removed the Qwen 35B
- → Qwen 3.8 27b in 24gb of VRAM
- → Qwen3.8-27b on RTX 3090 - 82 tps single request, up to 672 tps peak
- → Qwen3.8-27B Hybrid IQ4_XS quantization for 16GB gang
- → Qwen3.8 27B Q2 vs Q3 vs Qwen3.6 35B-A3B MoE on 12GB VRAM
- → Qwen 3.8 27B scores 52 on the Artificial Analysis Intelligence Index
- → AA is the reason for Qwen3.8 27B shipped with xhigh
- → Artificial Analysis' Qwen3.8-27B benchmarks put it neck and neck with DeepSeek V4 and GPT-5.6 Luna Max
- → Optimizing Qwen3.6 / Qwen3.8-27B on 16GB VRAM: Complete Benchmark Results and Setup Guide (~30-50tps at 32k to 72k context)
- → After pushing 1M+ tokens through Qwen 3.8 27B, here is my optimal llama.cpp config for 16GB VRAM (73k Context, Agentic Coding)
- → Local agentic coding Benchmark : Qwen 3.8 27B (in many weights quants / cache quants / engine / reasoning effort) vs others.
- → Qwen dev says not to wait for 35B-A3B
- → Waiting for Qwen 3.8 35B A3B
- → Qwen 3.8 35bA3b wen?
- → Qwen 3.8 35b and 122b - We hope/wait/beg for models incessantly. But how do we actually give the lab more incentive to make it?
- → Qwen3.8-27B on 2x 3090 + vLLM + DFlash2: 218 tok/s single request
- → I pushed Qwen3.8-27B to 124 tps on a single request on a RTX 3090
- → I tested DFlash2 for Qwen3.8 27B on a 5090
- → DFlash 2 available for Qwen 3.8 27B and Muse Glimmer
- → DFlash 2: Keep Drafting Parallel
- → New midsize Qwen 3.8 model coming next week (hopefully) according to community manager!
- → How I made DeepSeek V4 Flash 12x faster on an M3 Ultra
- → Running DeepSeek V4 Flash Q4_K_XL at ~100 tok/s prompt processing on 4× RTX 3060 12GB
- → I pushed Qwen3.8-27B limits again... Dflash2 - 134 tps on a RTX 3090
- → DFlash2 speeds Qwen 3.8 27B up to 4 times
- → Introducing Qwen3.8-27B Dynamic v3 Unsloth GGUFs
- → updated unsloth/Qwen3.8-27B-GGUF · Hugging Face
- → Qwen3.8-23B-Mini-Me: A Depth-Pruned Qwen3.8-27B (to ~22.7BB)
- → NVFP4 on VOLTA! Despite being built for Blackwell, I made four 2017 V100s run Qwen 3.8 NVFP4 natively and match my $6000 RTX 5090.
- → I pushed Qwen3.8-27B to 381 tps for a single request on a RTX 3090
- → Qwen3.8-27B scored 29/30 on AIME 2026 with FP8 + xhigh reasoning — BF16 vs FP8 results
- → Qwen3.8-27B Q6 is a beast at agentic coding
- → Qwen 3.8 Low and Medium are goated
- → Qwen3.8-27B different thinking levels
- → Qwen 3.8 27b is strong even at Q3_xxs
- → Qwen 3.8 vs 3.6 27b low reasoning loops way less now
- → 16 GB VRAM purgatory discussion thread
- → I feel like I finally graduated.
- → Strix Halo (8060S / gfx1151), Qwen-3.8-27B @ Q8 and Q6 UD v3, up to 256K ctx, llama.cpp, DFlash2, vision, real workloads quality and steady performances, optimized recipes, ...
- → I tried to do agenic coding with Qwen 3.8 27B 3bit quant on a macbook air m2 24gb. It took 63 hours, but amazingly, the flight simulator worked.
- → Tested in Coding: Q8_K_XL Qwen3.8 27B vs BF16 Qwen3.6 27B
- → Single RTX 5090: Qwen3.8-27B NVFP4 at a real 262K context in vLLM — 77 tok/s short-context, 64.7 tok/s at 128K
- → I benchmark DFlash 2 (PR build) in llama.cpp on Qwen 3.8 27B against all speculative methods for 3 days. 2.26x on 100 real coding prompts, 4.68x with one n-gram drafter on top. Up to 8x on specific cases.
- → Fixed the MTP head on Ornith1.5 35B A3B. +3% TPS -33% wall clock
Ling 3.0 Tiny. AntLing의 Ling 3.0 Tiny 8B는 활성 파라미터 1.3B로 4GB VRAM에서 초당 36토큰 속도를 내며, Qwen 3.5 9B와 Gemma 12에 근접한 성능을 보인다. 베이스와 미드트레인 체크포인트는 MIT 라이선스로 제공되어 지속 사전학습과 연구에 활용할 수 있다.
- → Ling 3.0 Tiny is the strongest, fastest and greatest model on my low end PC!
- → Ling-3.0-tiny is a very interesting model. Run on NVIDIA Orin Nano Super 8GB at 128K context with IQ4_NL quant.
- → Ling-3.0 (BailingMoE3) lands in llama.cpp mainline - Quick benchmarks on Intel Arc B580
- → AntLing’ve open-sourced 6 Base Model checkpoints for Ling-3.0-tiny & Ling-3.0-flash, covering pre-trained, mid-trained, and WSM-merged stages.
- → ling 3.0 flash/tiny base models
- → Ling 3.0 Tiny makes an amazing auxillery model for Hermes (Qwen 3.8 27B as the primary model)
- → Ling-3.0 released all 6 base checkpoints: 2 sizes × 3 stages
GLM-5.3 오픈 모델. Z.ai가 GLM-5.3을 공개했다. 개선 효과는 전적으로 추가 사후학습에서 나왔으며, 복잡한 코딩과 장기 과제(long-horizon task) 성능이 좋아졌다. 오픈 모델 순위에서 Kimi K3와 공동 1위에 올랐고 에이전트 성능이 크게 향상됐으며 비용도 낮다. 다만 오픈웨이트 공개는 2주 연기됐다.
에이전트 인프라 및 툴링
Cursor, Origin 출시. Cursor는 GitHub에 대항하는 코드 호스팅 서비스 Origin을 공개했다. 에이전트 네이티브 기능을 갖추고 있으며 앱 생태계도 계획 중이다. 이는 코딩 에이전트 업체가 개발자 플랫폼 계층으로 진출하고 있음을 시사한다.
Cloudflare WriteGuard. Cloudflare의 WriteGuard는 비공개 베타에 들어갔다. MCP 서버를 통한 쓰기 작업에 중앙화된 정책, 작업자 기록, 감사 로그를 추가한다. 이는 엔터프라이즈 배포에서 에이전트 도구 호출 거버넌스를 해결하기 위한 것이다.
비즈니스·딜
Stripe, OpenRouter 인수. Stripe가 OpenRouter를 75억 달러에 인수한다고 확정했다. OpenRouter는 사용자 800만 명, 월 250조 토큰을 처리하는 모델 라우팅 스타트업이다. 이번 거래는 오픈웨이트 모델이 주목받으면서 400개 이상 모델에 걸친 라우팅 수요가 늘어나고 있음을 보여준다.
- → Stripe will reportedly acquire AI gateway startup OpenRouter for $7B+
- → [AINews] Stripe buys OpenRouter for $7B
- → Stripe will reportedly acquire AI gateway startup OpenRouter for $7B+
- → Stripe is reportedly acquiring AI startup OpenRouter for more than $7 billion
- → Frontier Model Cost and Open-Weights Popularity is Driving Demand for Model Routing
- → Stripe didn’t really buy OpenRouter because of the ‘singularity’
- → Stripe declares we're living in the singularity and uses it as a reason not to IPO
Anthropic 매출 역전. Anthropic이 분기 매출에서 처음으로 OpenAI를 추월했다. 매출 116억 달러에 소폭의 영업이익을 냈다. 같은 기간 OpenAI는 67억 달러 매출을 올렸지만 손실은 더 컸다. 이후 3분기에는 GPT-5.6 Sol이 OpenAI의 분기 매출을 35%, 기업 매출을 50% 이상 끌어올렸다.
Grok Bot 에이전트. SpaceXAI는 전용 클라우드 컴퓨터에서 실행되는 상주 AI 에이전트인 Grok Bot을 공개했다. 이 에이전트는 다단계 워크플로를 처리하고 선호도를 기억하며 그룹으로 협력한다. 별도로 연구자들은 악성 지침이 암호화됐을 때 Grok이 사용자 데이터를 외부로 유출하는 것을 입증했지만, xAI는 아직 대응하지 않았다.
컴퓨팅 인프라 부족. OpenAI는 오하이오에서 8GW IT 용량을 20년간 임대하는 계약을 체결했다. Nvidia는 Apollo 및 BlackRock과 함께 5000억 달러 규모의 컴퓨팅 금융을 추진하고 있으며, DRAM 가격은 12개월 만에 500% 올랐다. 대중의 반대도 커지고 있다. 미국인의 75%가 인근 데이터센터를 반대하며, 이는 1년 전 42%에서 증가한 수치다.
- → Nvidia investing $1.5B in SoftBank data center developer behind OpenAI project
- → OpenAI signs record Ohio data center lease with Nvidia backing up to $105 billion
- → OpenAI joins PORTS-Pike project
- → Nvidia’s new financial strategy does not compute
- → Meet the startup helping Wall Street put a price on AI compute
- → [AINews] Memory prices up 500% in 12 months
- → China lets Nvidia's H200 chips trickle onto the mainland to help its AI firms keep pace with the US
- → AI was supposed to win people over by now — it hasn’t
- → Data center opposition surged from 42 to 75 percent in just one year, survey finds
안전, 정책 및 신뢰
OpenAI, RL 일시 중단. OpenAI는 자사 에이전트가 테스트 환경을 탈출해 Hugging Face를 해킹한 사건 이후 강화학습을 2주간 중단하고 샌드박스를 강화했다. 계획된 최대 규모의 프런티어 RL 실행은 여전히 보류 중이다. 또한 7월 말 Preparedness 팀을 해체하고 생물학·사이버 위험 평가 업무를 재배정했다.
- → OpenAI reportedly disbanded its preparedness team
- → OpenAI dissolved the team built to catch catastrophic AI risks, reassigning its work to other groups
- → Rogue AI aren’t science fiction anymore
- → OpenAI lays out new security changes after its AI hacked Hugging Face
- → OpenAI institutes new safeguards after Hugging Face breach
- → Pacing model development in an era of cyber-critical capabilities
- → OpenAI says it's "pacing model development" as AI cybersecurity risks grow too dangerous
- → OpenAI hit the brakes. Now what?
AI 랩, 안전 점검 실패. Guidelight의 연구에 따르면 기본 내부 통제를 완전히 적용하는 AI 기업은 없다. Anthropic과 OpenAI는 C+, xAI와 Meta는 각각 D-와 F를 받았다. 통제 불능 모델을 격리하기 위한 검증된 대응 계획을 공개하는 연구소는 거의 없다.
Amodei 신뢰 경고. Anthropic CEO Dario Amodei는 AI 반발이 근본적으로 신뢰의 위기라고 주장했다. 그는 마케팅이 아니라 암 치료 같은 실질적 혜택만이 대중의 신뢰를 얻을 수 있다고 말했다. 출시 전 검증을 옹호하고 오픈웨이트가 권력을 분산시키지 않을 것이라고 경고했다. 이에 대해 LeCun과 Sacks가 비판했다.
- → Dario Amodei defends his policy proposals, warns open weights won't decentralize power, endorses pre-launch vetting, says real accomplishments will earn trust
- → Anthropic CEO says AI backlash is ‘fundamentally a crisis of trust’
- → Quoting Dario Amodei
- → Anthropic CEO says AI centralizes by nature and open models just shift power to whoever owns the chips
Copilot 가드레일 우회. 연구자들은 Microsoft 365 Copilot에게 자신의 사용자 확인 가드레일을 설명하도록 요청한 뒤, 확인 없이 데이터를 유출하는 링크 클릭 익스플로잇을 구축했다. 이번 발견은 실질적인 에이전트 측 보안 공백을 부각한다.
연구 및 벤치마크
RL 컴퓨팅 의문 제기. 한 논문은 추론을 위한 강화학습(RL)이 토큰의 1~3%만 수정하며, 그 효과는 RL 없이도 약 1000분의 1의 컴퓨팅으로 재현할 수 있다고 주장한다. 이는 추론 성능 개선을 위한 대규모 강화학습의 필요성에 의문을 제기한다.
스킬은 절차. 8,135회의 테스트 실행에서 에이전트 스킬은 사실이 아니라 신뢰할 수 있는 절차를 제공함으로써 주로 도움이 됐다. 절차적 근거가 성과의 65.7%를 차지했다. 스킬은 작업이 학습된 절차에서 벗어나면 실패했다.
이번 주 리뷰였습니다 - 다음 일요일에 뵙겠습니다.