نماذج ومنافسة
GPT-5.6 وWork Agent. أطلقت OpenAI إصدار GPT-5.6 في ثلاث طبقات استدلال وأطلقت ChatGPT Work للمهام المستقلة التي تستمر لساعات، لكن حدود الاستخدام وواجهة مربكة أثارت شكاوى، مما دفع إلى وعود بالإصلاحات.
- → OpenAI launches its new family of models with GPT-5.6
- → The new GPT-5.6 family: Luna, Terra, Sol
- → OpenAI rolls out GPT-5.6 after government greenlight — and announces ‘ChatGPT Work’
- → OpenAI says GPT 5.6 is the ‘preferred model’ for Microsoft Copilot 365 amid breakup chatter
- → OpenAI wants its new tool to do your work for you and with you
- → The ChatGPT browser is already dead
- → OpenAI is shutting down Atlas, but its AI browser ambitions are still growing
- → How did the government decide OpenAI’s frontier model was safe to release?
- → OpenAI pairs its GPT-5.6 public rollout with ChatGPT Work, a new agent that handles entire workflows
- → OpenAI's GPT-5.6 Sol autonomously post-trained the smaller Luna model with a "fairly underspecified prompt"
- → OpenAI kills its Atlas browser after just eight months and folds everything into ChatGPT
- → GPT-5.6 Sol nearly matches Fable 5 on aggregated benchmarks at one-third the cost
- → OpenAI staffer maps out which of GPT-5.6 Sol's five reasoning levels fits which task complexity
- → OpenAI admits it "didn't get everything quite right" with ChatGPT Work launch and scrambles to fix UX and costs
Claude Fable 5 يتصدر المعايير. اجتاح Claude Fable 5 من Anthropic معايير Artificial Analysis، لكن تكلفته العالية دفعت الشركة للتوصية باستخدامه كمخطط يفوض النماذج الأرخص، مما يقلل النفقات بنحو 40% مع الاحتفاظ بمعظم القدرات.
هجوم Meta الوكيل وتداعيات الخصوصية. أطلقت Meta نموذج Muse Spark 1.1، وهو نموذج برمجة عالي السياق يخفض أسعار المنافسين، بينما تم سحب مولد الصور Muse بعد تمكين استخدام الوجوه غير المصرح به على Instagram، مما أعاد إشعال نقاشات الموافقة.
- → Meta just launched a new AI generator, Muse Image, and users are already pushing back over use of their photos
- → Meta’s new Muse Image model can pull other Instagram users into AI photos
- → Muse Image is technically impressive, but Meta's use of Instagram photos raises questions
- → Meta wants its AI glasses to seem less creepy. Its AI strategy says otherwise.
- → Meta tests always-on AI glasses that capture your entire day
- → Meta enters the crowded AI coding battle with Muse Spark 1.1
- → Meta says its new AI model is ready to compete on coding
- → Meta are apparently working on an open source variant of Muse Spark.
- → Introducing Muse Spark 1.1
- → Meta's Muse Spark 1.1 API pricing squeezes OpenAI and Anthropic as the AI price war heats up
- → GPT-5.6 🚀, Muse Spark 1.1 ✨, ChatGPT Work 💼
- → Meta removes controversial AI feature on Instagram after backlash
- → Meta turns off the Instagram feature that let users make AI deepfakes of public accounts
موجة المصادر المفتوحة الصينية. تمثل النماذج الصينية الآن أكثر من 30% من حركة مرور OpenRouter؛ نموذجا HY3 من Tencent وGLM-5.2 يعملان على أجهزة منزلية، وتخطط MiniMax لإصدار مفتوح بـ 2.7 تريليون معلمة، ويشير تصميم رقاقة DeepSeek إلى التكامل الرأسي وسط مخاوف أمريكية.
- → This is what Hy3 is capable of. Mother of god.
- → llama.cpp: Hy3 PR + GGUFs
- → Chinese AI models regularly pass 30 percent on OpenRouter as cost gap widens
- → Chinese AI models are gaining ground with U.S. companies as OpenAI, Anthropic costs surge
- → Why the rise of open source AI isn’t hurting Anthropic … yet
- → Chinese AI startup MiniMax plans to open-source a 2.7 trillion parameter model later this year
- → 4-bit GLM-5.2 (753B MoE) on 4× DGX Spark: 70.8% on Terminal-Bench 2.1 vs 81.0% for the full model
- → Running GLM 5.2 on 4xGB10 with a 100G Switch, 330k ctx, ~25 t/s tg, ~650 t/s pp
- → I created a 140 GB IQ2_XXS REAP quant of GLM 5.2 for coding. Looking for testers.
- → Hy3 (295B MoE) and NVIDIA Nemotron-Labs-Audex-30B-A3B (audio-capable 30B MoE) GGUF quants
- → Tencent-HY3 is the real deal on 128GB!
- → The U.S. tech industry is increasingly anxious about the rising power and competitive price of open-source AI models from China — and whether the Trump administration will respond with yet another executive order | Politico
- → China's DeepSeek developing its own AI chip, sources say
هندسة الوكلاء والبنية التحتية
إعادة كتابة Bun بواسطة 64 وكيلاً. نظم Claude Fable 5 64 مثيلاً متوازياً لإعادة كتابة Bun من Zig إلى Rust في 11 يوماً، منتجاً أكثر من مليون سطر من الكود وإصلاح 128 خطأ بتكلفة 165 ألف دولار.
فشل تدقيق SWE‑Bench Pro. وجدت OpenAI أن حوالي 30% من مهام SWE‑Bench Pro معطلة وسحبت تأييدها، مرددة المخاوف السابقة من Artificial Analysis وتقويض تقييم البرمجة الوكيلة.
Modal تجمع 355 مليون دولار. حصلت Modal على 355 مليون دولار لبناء بنية تحتية سحابية مصممة خصيصاً لوكلاء الذكاء الاصطناعي، مع التركيز على بيئات معزولة سريعة التكرار للبرمجيات المستقلة.
وكلاء مؤقتون من Cloudflare. قدمت Cloudflare حسابات مؤقتة تتيح لوكلاء الذكاء الاصطناعي نشر Workers بدون بيانات اعتماد دائمة، مع انتهاء صلاحية إذا لم يتم المطالبة بها، مما يسهل سير العمل المدعوم بالوكلاء.
الوكلاء يصلون إلى وحدات معالجة الرسوم المحلية. تقنيات التكميم الجديدة وأجهزة GPU الفعالة من حيث التكلفة تدير الآن نماذج MoE على أجهزة المستهلك، لكن المعايير تظهر أن التكميم منخفض البت يمكن أن يقلل بشدة من المهام الوكيلة، مما يحث على اختيار دقيق للتكميم.
- → Qwen3.5 122B is the best?
- → Qwen3.6-27b does not understand software architechure.
- → Qwen 3.6 Q2-FP8 Terminal Bench 2 and GPQA Scores
- → Has anyone tested how quantization hits different capabilities separately? My results are surprising.
- → 2.5x faster Qwen3.6 NVFP4 Unsloth quants
- → Ultra budget 20GB vram with 448GB/s for $100 bucks.
- → I benched quad 5060Tis for code generation with Qwen3.6-27B so you don't have to (it's really good)
الأمن والقانون وقابلية التفسير
عدسة Jacobian ترسم عقل النموذج. تكشف عدسة Jacobian من Anthropic عن مجموعة صغيرة من التمثيلات القابلة للتعبير اللفظي التي تلتقط الاستدلال الأساسي للنموذج، مما يتيح اكتشاف الهلوسة، توجيه المخرجات، وكما تظهر عروض المجتمع، الإنشاء الفوري للمتغيرات الضارة.
- → Anthropic Research - "Verbalizable Representations Form a Global Workspace in Language Models"
- → Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens
- → I tested Anthropic’s new Jacobian Lens on open models, then it turned into a local-model hallucination router
- → Anthropic found a hidden space where Claude puzzles over concepts
- → I created a super harmful model ! :D (by tweaking it's J-Space!!!)
فدية مدفوعة بواسطة وكيل. وثقت Sysdig وكيل ذكاء اصطناعي تولى التنفيذ التقني لهجوم فدية، على الرغم من أن إنساناً أعد العملية، مما يسلط الضوء على القدرات الحالية وحدود البرمجيات الخبيثة الوكيلة.
أزمات إساءة استخدام الذكاء الاصطناعي. يدعي دعوى قضائية أن Grok من xAI أنتج صور استغلال الأطفال (CSAM)، بينما تستخدم بوكو حرام روبوتات المحادثة لتخطيط الهجمات؛ تؤكد الحوادث على الحاجة الملحة لحواجز الأمان مع توسع القدرات الوكيلة.
Apple تقاضي OpenAI. رفعت Apple دعوى قضائية اتحادية تتهم OpenAI باختطاف أكثر من 400 موظف لسرقة أسرار تجارية للأجهزة، بهدف بناء هاتف ذكي منافس للذكاء الاصطناعي، مما يزيد حدة المعارك القانونية في صناعة الذكاء الاصطناعي.
هذا ملخص الأسبوع - نراكم الأحد القادم.