โมเดลและการแข่งขัน
GPT-5.6 และ Work Agent. OpenAI เปิดตัว GPT-5.6 ในสามระดับการให้เหตุผล และเปิดตัว ChatGPT Work สำหรับงานอัตโนมัติที่ใช้เวลาหลายชั่วโมง แต่ข้อจำกัดในการใช้งานและอินเทอร์เฟซที่สับสนทำให้เกิดข้อร้องเรียน นำไปสู่สัญญาว่าจะแก้ไข
- → OpenAI launches its new family of models with GPT-5.6
- → The new GPT-5.6 family: Luna, Terra, Sol
- → OpenAI rolls out GPT-5.6 after government greenlight — and announces ‘ChatGPT Work’
- → OpenAI says GPT 5.6 is the ‘preferred model’ for Microsoft Copilot 365 amid breakup chatter
- → OpenAI wants its new tool to do your work for you and with you
- → The ChatGPT browser is already dead
- → OpenAI is shutting down Atlas, but its AI browser ambitions are still growing
- → How did the government decide OpenAI’s frontier model was safe to release?
- → OpenAI pairs its GPT-5.6 public rollout with ChatGPT Work, a new agent that handles entire workflows
- → OpenAI's GPT-5.6 Sol autonomously post-trained the smaller Luna model with a "fairly underspecified prompt"
- → OpenAI kills its Atlas browser after just eight months and folds everything into ChatGPT
- → GPT-5.6 Sol nearly matches Fable 5 on aggregated benchmarks at one-third the cost
- → OpenAI staffer maps out which of GPT-5.6 Sol's five reasoning levels fits which task complexity
- → OpenAI admits it "didn't get everything quite right" with ChatGPT Work launch and scrambles to fix UX and costs
Claude Fable 5 ครองอันดับใน Benchmark. Claude Fable 5 ของ Anthropic กวาดคะแนนใน Benchmark ของ Artificial Analysis แต่ค่าใช้จ่ายสูงทำให้บริษัทแนะนำให้ใช้เป็นตัววางแผนที่มอบหมายงานให้โมเดลราคาถูกกว่า ลดค่าใช้จ่ายประมาณ 40% ในขณะที่ยังคงความสามารถส่วนใหญ่ไว้
การโจมตีของ Meta และผลกระทบด้านความเป็นส่วนตัว. Meta เปิดตัว Muse Spark 1.1 โมเดลการเขียนโค้ดที่มีบริบทสูง ซึ่งต่ำกว่าราคาคู่แข่ง ขณะที่เครื่องกำเนิดภาพ Muse ถูกถอนออกหลังจากเปิดให้ใช้ใบหน้า Instagram โดยไม่ได้รับอนุญาต จุดประเด็นถกเถียงเรื่องความยินยอมอีกครั้ง
- → Meta just launched a new AI generator, Muse Image, and users are already pushing back over use of their photos
- → Meta’s new Muse Image model can pull other Instagram users into AI photos
- → Muse Image is technically impressive, but Meta's use of Instagram photos raises questions
- → Meta wants its AI glasses to seem less creepy. Its AI strategy says otherwise.
- → Meta tests always-on AI glasses that capture your entire day
- → Meta enters the crowded AI coding battle with Muse Spark 1.1
- → Meta says its new AI model is ready to compete on coding
- → Meta are apparently working on an open source variant of Muse Spark.
- → Introducing Muse Spark 1.1
- → Meta's Muse Spark 1.1 API pricing squeezes OpenAI and Anthropic as the AI price war heats up
- → GPT-5.6 🚀, Muse Spark 1.1 ✨, ChatGPT Work 💼
- → Meta removes controversial AI feature on Instagram after backlash
- → Meta turns off the Instagram feature that let users make AI deepfakes of public accounts
คลื่นโอเพนซอร์สของจีน. ขณะนี้โมเดลจีนมีสัดส่วนมากกว่า 30% ของปริมาณการใช้งาน OpenRouter; HY3 และ GLM-5.2 ของ Tencent ทำงานบนฮาร์ดแวร์ภายในประเทศ, MiniMax วางแผนปล่อยโอเพนซอร์สที่มีพารามิเตอร์ 2.7T, และการออกแบบชิปของ DeepSeek ส่งสัญญาณถึงการบูรณาการแนวตั้งท่ามกลางความกังวลของสหรัฐฯ
- → This is what Hy3 is capable of. Mother of god.
- → llama.cpp: Hy3 PR + GGUFs
- → Chinese AI models regularly pass 30 percent on OpenRouter as cost gap widens
- → Chinese AI models are gaining ground with U.S. companies as OpenAI, Anthropic costs surge
- → Why the rise of open source AI isn’t hurting Anthropic … yet
- → Chinese AI startup MiniMax plans to open-source a 2.7 trillion parameter model later this year
- → 4-bit GLM-5.2 (753B MoE) on 4× DGX Spark: 70.8% on Terminal-Bench 2.1 vs 81.0% for the full model
- → Running GLM 5.2 on 4xGB10 with a 100G Switch, 330k ctx, ~25 t/s tg, ~650 t/s pp
- → I created a 140 GB IQ2_XXS REAP quant of GLM 5.2 for coding. Looking for testers.
- → Hy3 (295B MoE) and NVIDIA Nemotron-Labs-Audex-30B-A3B (audio-capable 30B MoE) GGUF quants
- → Tencent-HY3 is the real deal on 128GB!
- → The U.S. tech industry is increasingly anxious about the rising power and competitive price of open-source AI models from China — and whether the Trump administration will respond with yet another executive order | Politico
- → China's DeepSeek developing its own AI chip, sources say
วิศวกรรมเอเจนต์และโครงสร้างพื้นฐาน
Bun ถูกเขียนใหม่โดย 64 เอเจนต์. Claude Fable 5 จัดการ 64 อินสแตนซ์แบบขนานเพื่อเขียน Bun ใหม่จาก Zig เป็น Rust ใน 11 วัน สร้างโค้ดมากกว่าล้านบรรทัดและแก้ไข 128 บัก ด้วยต้นทุน $165K
การตรวจสอบ SWE‑Bench Pro ล้มเหลว. OpenAI พบว่างานใน SWE‑Bench Pro ประมาณ 30% ทำงานไม่ได้และถอนการรับรอง สะท้อนความกังวลก่อนหน้านี้จาก Artificial Analysis และบั่นทอนการประเมินการเขียนโค้ดแบบเอเจนต์
Modal ระดมทุน $355M. Modal ได้รับเงิน $355 ล้านเพื่อสร้างโครงสร้างพื้นฐานคลาวด์ที่ออกแบบมาเฉพาะสำหรับเอเจนต์ AI โดยเน้นสภาพแวดล้อมแบบแซนด์บ็อกซ์ที่ทำซ้ำได้เร็วสำหรับซอฟต์แวร์อัตโนมัติ
Cloudflare เอเจนต์ชั่วคราว. Cloudflare เปิดตัวบัญชีชั่วคราวที่ให้เอเจนต์ AI ปรับใช้ Workers โดยไม่มีข้อมูลรับรองถาวร โดยจะหมดอายุหากไม่มีการเรียกร้อง ทำให้เวิร์กโฟลว์ที่ขับเคลื่อนโดยเอเจนต์ราบรื่นขึ้น
เอเจนต์ทำงานบน GPU ในเครื่อง. เทคนิคการหาปริมาณใหม่และชุด GPU ที่คุ้มค่าต้นทุนสามารถรันโมเดล MoE บนฮาร์ดแวร์ผู้บริโภคได้ แต่ Benchmark แสดงว่าการหาปริมาณบิตต่ำสามารถทำให้งานแบบเอเจนต์เสื่อมสภาพอย่างรุนแรง จึงควรเลือกการหาปริมาณอย่างระมัดระวัง
- → Qwen3.5 122B is the best?
- → Qwen3.6-27b does not understand software architechure.
- → Qwen 3.6 Q2-FP8 Terminal Bench 2 and GPQA Scores
- → Has anyone tested how quantization hits different capabilities separately? My results are surprising.
- → 2.5x faster Qwen3.6 NVFP4 Unsloth quants
- → Ultra budget 20GB vram with 448GB/s for $100 bucks.
- → I benched quad 5060Tis for code generation with Qwen3.6-27B so you don't have to (it's really good)
ความปลอดภัย กฎหมาย และการตีความ
Jacobian Lens แผนที่ความคิดของโมเดล. Jacobian Lens ของ Anthropic เผยชุดการแสดงผลที่สามารถอธิบายเป็นคำพูดได้จำนวนน้อยที่จับการให้เหตุผลหลักของโมเดล ทำให้สามารถตรวจจับภาพหลอน ควบคุมผลลัพธ์ และดังที่การสาธิตของชุมชนแสดง การสร้างตัวแปรที่เป็นอันตรายได้ทันที
- → Anthropic Research - "Verbalizable Representations Form a Global Workspace in Language Models"
- → Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens
- → I tested Anthropic’s new Jacobian Lens on open models, then it turned into a local-model hallucination router
- → Anthropic found a hidden space where Claude puzzles over concepts
- → I created a super harmful model ! :D (by tweaking it's J-Space!!!)
แรนซัมแวร์ที่ขับเคลื่อนโดยเอเจนต์. Sysdig บันทึกเอเจนต์ AI ที่จัดการการดำเนินการทางเทคนิคของการโจมตีแรนซัมแวร์ แม้ว่ามนุษย์จะเป็นผู้ตั้งค่าการดำเนินการ โดยเน้นความสามารถและข้อจำกัดในปัจจุบันของมัลแวร์แบบเอเจนต์
วิกฤตการใช้ AI ในทางที่ผิด. คดีความกล่าวหาว่า Grok ของ xAI สร้างภาพ CSAM ขณะที่ Boko Haram ใช้แชทบอทเพื่อวางแผนการโจมตี เหตุการณ์เหล่านี้เน้นย้ำถึงความจำเป็นเร่งด่วนสำหรับมาตรการป้องกันความปลอดภัยเมื่อความสามารถแบบเอเจนต์ขยายตัว
Apple ฟ้อง OpenAI. Apple ยื่นฟ้องในศาลรัฐบาลกลางกล่าวหาว่า OpenAI ล่อลวงพนักงานกว่า 400 คนเพื่อขโมยความลับทางการค้าด้านฮาร์ดแวร์ ซึ่งอาจเป็นการสร้างสมาร์ทโฟน AI ที่เป็นคู่แข่ง ทำให้การต่อสู้ทางกฎหมายในอุตสาหกรรม AI รุนแรงขึ้น
นี่คือสรุปประจำสัปดาห์ - แล้วพบกันใหม่วันอาทิตย์หน้า