모델 및 경쟁
GPT-5.6 및 Work Agent. OpenAI가 세 가지 추론 계층으로 GPT-5.6을 출시하고 수시간 동안 지속되는 자율 작업을 위한 ChatGPT Work를 출시했지만, 사용 제한과 혼란스러운 인터페이스로 인해 불만이 제기되어 수정 약속이 이어졌습니다.
- → OpenAI launches its new family of models with GPT-5.6
- → The new GPT-5.6 family: Luna, Terra, Sol
- → OpenAI rolls out GPT-5.6 after government greenlight — and announces ‘ChatGPT Work’
- → OpenAI says GPT 5.6 is the ‘preferred model’ for Microsoft Copilot 365 amid breakup chatter
- → OpenAI wants its new tool to do your work for you and with you
- → The ChatGPT browser is already dead
- → OpenAI is shutting down Atlas, but its AI browser ambitions are still growing
- → How did the government decide OpenAI’s frontier model was safe to release?
- → OpenAI pairs its GPT-5.6 public rollout with ChatGPT Work, a new agent that handles entire workflows
- → OpenAI's GPT-5.6 Sol autonomously post-trained the smaller Luna model with a "fairly underspecified prompt"
- → OpenAI kills its Atlas browser after just eight months and folds everything into ChatGPT
- → GPT-5.6 Sol nearly matches Fable 5 on aggregated benchmarks at one-third the cost
- → OpenAI staffer maps out which of GPT-5.6 Sol's five reasoning levels fits which task complexity
- → OpenAI admits it "didn't get everything quite right" with ChatGPT Work launch and scrambles to fix UX and costs
Claude Fable 5, 벤치마크 선두. Anthropic의 Claude Fable 5가 Artificial Analysis 벤치마크를 휩쓸었으나, 높은 비용으로 인해 회사는 이를 더 저렴한 모델에 위임하는 플래너로 사용할 것을 권장하여 비용을 약 40% 절감하면서도 대부분의 성능을 유지했습니다.
Meta의 에이전트 공세 및 프라이버시 파장. Meta가 경쟁사보다 저렴한 가격의 고맥락 코딩 모델인 Muse Spark 1.1을 출시했지만, Muse 이미지 생성기가 무단 인스타그램 얼굴 사용을 가능하게 한 후 철수되어 동의 논쟁이 재점화되었습니다.
- → Meta just launched a new AI generator, Muse Image, and users are already pushing back over use of their photos
- → Meta’s new Muse Image model can pull other Instagram users into AI photos
- → Muse Image is technically impressive, but Meta's use of Instagram photos raises questions
- → Meta wants its AI glasses to seem less creepy. Its AI strategy says otherwise.
- → Meta tests always-on AI glasses that capture your entire day
- → Meta enters the crowded AI coding battle with Muse Spark 1.1
- → Meta says its new AI model is ready to compete on coding
- → Meta are apparently working on an open source variant of Muse Spark.
- → Introducing Muse Spark 1.1
- → Meta's Muse Spark 1.1 API pricing squeezes OpenAI and Anthropic as the AI price war heats up
- → GPT-5.6 🚀, Muse Spark 1.1 ✨, ChatGPT Work 💼
- → Meta removes controversial AI feature on Instagram after backlash
- → Meta turns off the Instagram feature that let users make AI deepfakes of public accounts
중국 오픈소스 물결. 중국 모델이 이제 OpenRouter 트래픽의 30% 이상을 차지합니다. Tencent의 HY3와 GLM‑5.2는 자체 하드웨어에서 실행되며, MiniMax는 2.7T 파라미터의 오픈 릴리즈를 계획하고, DeepSeek의 칩 설계는 미국의 우려 속에서 수직적 통합을 시사합니다.
- → This is what Hy3 is capable of. Mother of god.
- → llama.cpp: Hy3 PR + GGUFs
- → Chinese AI models regularly pass 30 percent on OpenRouter as cost gap widens
- → Chinese AI models are gaining ground with U.S. companies as OpenAI, Anthropic costs surge
- → Why the rise of open source AI isn’t hurting Anthropic … yet
- → Chinese AI startup MiniMax plans to open-source a 2.7 trillion parameter model later this year
- → 4-bit GLM-5.2 (753B MoE) on 4× DGX Spark: 70.8% on Terminal-Bench 2.1 vs 81.0% for the full model
- → Running GLM 5.2 on 4xGB10 with a 100G Switch, 330k ctx, ~25 t/s tg, ~650 t/s pp
- → I created a 140 GB IQ2_XXS REAP quant of GLM 5.2 for coding. Looking for testers.
- → Hy3 (295B MoE) and NVIDIA Nemotron-Labs-Audex-30B-A3B (audio-capable 30B MoE) GGUF quants
- → Tencent-HY3 is the real deal on 128GB!
- → The U.S. tech industry is increasingly anxious about the rising power and competitive price of open-source AI models from China — and whether the Trump administration will respond with yet another executive order | Politico
- → China's DeepSeek developing its own AI chip, sources say
에이전트 엔지니어링 및 인프라
64개 에이전트가 Bun 재작성. Claude Fable 5가 64개의 병렬 인스턴스를 조율하여 11일 만에 Bun을 Zig에서 Rust로 재작성했으며, 백만 줄 이상의 코드를 생성하고 128개의 버그를 수정하는 데 165,000달러가 소요되었습니다.
SWE‑Bench Pro 감사 실패. OpenAI가 SWE‑Bench Pro 작업의 약 30%가 손상되었음을 발견하고 지지를 철회했으며, 이는 Artificial Analysis의 이전 우려를 반영하고 에이전트 코딩 평가를 훼손합니다.
Modal, 3억 5500만 달러 조달. Modal이 AI 에이전트 전용 클라우드 인프라를 구축하기 위해 3억 5500만 달러를 확보했으며, 자율 소프트웨어를 위한 샌드박스 처리된 빠른 반복 환경에 중점을 둡니다.
Cloudflare 임시 에이전트. Cloudflare가 AI 에이전트가 영구 자격 증명 없이 Workers를 배포할 수 있는 임시 계정을 도입했으며, 청구되지 않으면 만료되어 에이전트 주도 워크플로우를 원활하게 합니다.
에이전트, 로컬 GPU 공략. 새로운 양자화 기술과 비용 효율적인 GPU 장비로 이제 소비자 하드웨어에서 MoE 모델을 실행할 수 있지만, 벤치마크에 따르면 낮은 비트 양자화가 에이전트 작업을 심각하게 저하시킬 수 있어 신중한 양자화 선택이 필요합니다.
- → Qwen3.5 122B is the best?
- → Qwen3.6-27b does not understand software architechure.
- → Qwen 3.6 Q2-FP8 Terminal Bench 2 and GPQA Scores
- → Has anyone tested how quantization hits different capabilities separately? My results are surprising.
- → 2.5x faster Qwen3.6 NVFP4 Unsloth quants
- → Ultra budget 20GB vram with 448GB/s for $100 bucks.
- → I benched quad 5060Tis for code generation with Qwen3.6-27B so you don't have to (it's really good)
보안, 법률 및 해석 가능성
Jacobian Lens, 모델 마인드 매핑. Anthropic의 Jacobian Lens는 모델의 핵심 추론을 포착하는 소수의 언어화 가능한 표현을 밝혀내어, 환각 탐지, 출력 조정, 그리고 커뮤니티 데모가 보여주듯 유해한 변종의 즉각적인 생성을 가능하게 합니다.
- → Anthropic Research - "Verbalizable Representations Form a Global Workspace in Language Models"
- → Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens
- → I tested Anthropic’s new Jacobian Lens on open models, then it turned into a local-model hallucination router
- → Anthropic found a hidden space where Claude puzzles over concepts
- → I created a super harmful model ! :D (by tweaking it's J-Space!!!)
에이전트 기반 랜섬웨어. Sysdig가 랜섬웨어 공격의 기술적 실행을 처리한 AI 에이전트를 기록했지만, 인간이 작업을 설정했으며, 이는 에이전트 기반 멀웨어의 현재 능력과 한계를 강조합니다.
AI 오용 위기. 소송은 xAI의 Grok이 CSAM 이미지를 생성했다고 주장하는 한편, 보코 하람이 공격 계획에 챗봇을 사용합니다. 이러한 사건들은 에이전트 능력이 확장됨에 따라 안전 가드레일의 긴급한 필요성을 강조합니다.
Apple, OpenAI 고소. Apple이 OpenAI가 400명 이상의 직원을 스카우트하여 하드웨어 영업 비밀을 훔치고 잠재적으로 경쟁 AI 스마트폰을 구축했다고 비난하는 연방 소송을 제기하여 AI 업계의 법적 분쟁을 심화시켰습니다.
이번 주 리뷰였습니다 - 다음 일요일에 뵙겠습니다.