ফ্রন্টিয়ার মডেলের সংঘাত
Kimi K3 ফ্রন্টিয়ারকে নাড়িয়ে দিয়েছে. Moonshot AI-এর ওপেন-ওয়েট Kimi K3 শীর্ষ ক্লোজড মডেলের সাথে প্রতিযোগিতা করে, কিছু বেঞ্চমার্কে Claude Fable 5-কে পরাজিত করে এবং ওপেন বনাম ক্লোজড বিতর্ক পুনরায় জাগিয়ে তোলে।
- → Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI
- → Kimi K3, and what we can still learn from the pelican benchmark
- → [AINews] Kimi K3 2.8T-A50B: the largest open model ever released; Opus 4.8-class at Sonnet 5 pricing
- → Kimi K3 weights to be released on the 27th.
- → Kimi K3 ranks #1 on @AfterQuery's SpreadsheetBench 2, surpassing Claude Fable 5
- → Kimi: Threat or menace?
- → kimi.ai teasing a video with lots of 3's in it
- → [AINews] not much happened today
- → Kimi moment. I think the writing is on the wall for Anthropic and OpenAi
- → Kimi K3 is currently at the top of the leaderboard for Text Arena filtered for science queries.
- → Kimi K3 (max) beats Sonnet 5 on Simple Bench
- → Kimi K3 is top of nextjs eval
- → Just like Deepseek, China's Kimi K3 is forcing Western AI labs to question their compute advantage
- → Kimi K3 🌕, Gemini 3.5 delayed ⏳, crushing ARC-AGI 3 🤖
GPT-5.6: শক্তি এবং বিপদ. GPT-5.6 পরিবার প্রোগ্রাম্যাটিক টুল কলিং এবং প্যারালাল সাবএজেন্ট নিয়ে আসে, কিন্তু ফাইল এবং ডেটাবেস অনিচ্ছাকৃতভাবে মুছে ফেলার প্রবণতা স্যান্ডবক্সিংয়ের প্রয়োজনীয়তা জোর দেয়।
- → The Sequence Radar #893: Last Week in AI: GPT-5.6, Grok 4.5, Muse Spark 1.1 and the Post-Chatbot Stack
- → OpenAI’s new flagship model deletes files on its own, people keep warning
- → GPT-5.6 is deleting user files when given full access, and OpenAI says it shouldn't but did
- → [AINews] Codex usage up >10x in 6 months to 7M users, +1M in the past ~day; did Codex overtake Claude Code??
Claude Fable 5 বাড়ানো হয়েছে. GPT-5.6 এবং Kimi K3-এর চাপে, Anthropic পরিকল্পনা পরিবর্তন করে Fable 5-কে পেইড প্ল্যানে রেখেছে, সরাসরি অ্যাক্সেসিবিলিটিতে প্রতিযোগিতা করে।
DeepSeek V4 উদীয়মান. DeepSeek V4-কে কম খরচের API এবং ওপেন ওয়েট দিয়ে টিজ করছে, যখন কমিউনিটি অপটিমাইজেশন ইতিমধ্যেই এর ফ্ল্যাশ ভেরিয়েন্টকে ভোক্তা GPU-তে ব্যবহারযোগ্য গতিতে চালাতে দেয়।
এজেন্টরা বাস্তব-বিশ্বের ওয়ার্কফ্লোতে প্রবেশ করে
ব্রাউজিং এজেন্ট এসেছে. Anthropic Claude Code-কে নিরাপত্তা ক্লাসিফায়ার সহ একটি বিল্ট-ইন ওয়েব ব্রাউজার দিয়েছে, এবং Cursor একটি সাধারণ এজেন্ট চালু করেছে, কোডিং সহায়কদের স্বায়ত্তশাসিত কাজ সম্পাদনের দিকে নিয়ে যাচ্ছে।
এজেন্টরা বাস্তব লেনদেন পরিচালনা করে. DoorDash-এর এজেন্ট আর্কিটেকচার রূপান্তর ২৪% বাড়িয়েছে, এবং Stripe-এর বেঞ্চমার্ক দেখায় যে এজেন্টরা ইন্টিগ্রেশন কোড করতে পারে কিন্তু প্রায়শই ভ্যালিডেশনে ব্যর্থ হয়, নির্ভরযোগ্যতার ফাঁক উজ্জ্বল করে।
প্রোডাকশন এজেন্টরা হোঁচট খায়. বেশিরভাগ এন্টারপ্রাইজ এজেন্ট চ্যাটবট থাকে; ৫৪% ফার্মের এজেন্ট নিরাপত্তা ঘটনা ছিল, এবং বিশেষজ্ঞরা যুক্তি দেন যে এজেন্টদের মাইক্রোসার্ভিসের মতো ক্লাউড-নেটিভ অপারেশনাল প্রিমিটিভ প্রয়োজন।
- → How to Debug Coding Agents with LangSmith Traces
- → The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials
- → The AI context gap: Enterprise AI organizations have a trust problem, not a retrieval problem — and most are still building the fix
- → The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway
- → Cloud Native Infrastructure Emerges as the Foundation for Trustworthy Agentic AI
- → QCon AI Boston: Production AI Moves Beyond Prompts to Platforms, Harnesses, and Evals
এআই অবকাঠামো, তহবিল এবং ডেটা সেন্টার প্রতিযোগিতা
ডেটা সেন্টার প্রতিরোধের মুখে. S&P Oracle-কে AI ক্যাপেক্স ঝুঁকির জন্য ডাউনগ্রেড করেছে, যখন স্থানীয় প্রতিবাদ এবং নিউ ইয়র্কের হাইপারস্কেল ডেটা সেন্টারে স্থগিতাদেশ AI অবকাঠামোর বিরুদ্ধে বর্ধিত প্রতিক্রিয়া প্রতিফলিত করে।
AI তহবিল বেড়েছে. Databricks $188B মূল্যায়নে তহবিল সংগ্রহ করেছে, Meta Anthropic-এর সাথে $10B ডেটা সেন্টার লিজ নিয়ে আলোচনা করছে, এবং Nous Research তার ওপেন-সোর্স এজেন্টের জন্য $75M সুরক্ষিত করেছে।
শক্তি এবং হার্ডওয়্যার বাজি. শক্তি সংস্থাগুলি ডট-কম যুগের পর সবচেয়ে বেশি তহবিল সংগ্রহ করেছে AI চালানোর জন্য, এবং SambaNova চিপ দ্বারা সমর্থিত $400M ঋণ নন-GPU AI হার্ডওয়্যারে বর্ধিত বিনিয়োগের ইঙ্গিত দেয়।
অন-ডিভাইস এআই অগ্রগতি
1-বিট মডেল ফোনে পৌঁছেছে. PrismML-এর Bonsai Qwen3.6-27B-কে 3.9GB-তে সংকুচিত করে, বেঞ্চমার্ক স্কোরের 90% ধরে রেখে, অন-ডিভাইস টুল-কলিং এজেন্ট সক্ষম করে; Apple কথিত আলোচনায় রয়েছে।
- → Bonsai 27B: The First 27B-Class Model to Run on a Phone
- → Bonsai 27B: 1-bit dense LLM running locally in your browser using custom WebGPU kernels
- → Prism-ML Bonsai Qwen 3.6 27B
- → PrismML’s new Ternary Qwen3.6 27B runs near fp16 precision on 10GB of memory!!!
- → So what's the consensus on 1bit models? Is it still a pipe dream?
- → PrismML Bonsai 27B is surprisingly usable on the Jetson Orin Nano 8GB
- → Is anyone having any luck with the Ternary Bonsai 27B DFlash?
- → Can we get a "not base model" flair?
- → Bonsai 27B runs locally on an iPhone - a 27B model in 3.9GB
- → User experience of Bonsai-Ternary-27B on 4060Ti 16GB for KB management and productivity assistant use cases
- → Bonsai 27B is a full open reasoning model that fits on an iPhone
- → Bonsai-27B & Ternary-Bonsai-27B - Updates (on PRs)
- → Apple in talks with startup PrismML that shrinks AI models to run on an iPhone
হোম ল্যাব ফ্রন্টিয়ার চালায়. DFlash-এর মতো llama.cpp অপটিমাইজেশন MoE মডেলকে 6x পর্যন্ত ত্বরান্বিত করে, এবং একটি OnePlus ফোন ফ্ল্যাশ থেকে বিশেষজ্ঞ স্ট্রিম করে 60GB মডেল 1.3 tok/s-এ চালায়।
- → I tested all llama.cpp's speculative decoding methods on Qwen 3.6 27B: MTP ~2.7x, DFlash ~3.7x, n-gram stack ~6x on real coding. Local AI win. My findings on RTX 6000 PRO.
- → DFlash makes Qwen3.6 27B 2.2x faster with no quality loss
- → GPT-OSS-120B, Qwen 30B and Gemma 26B on an Android phone at 1-5 tok/s: +60GB model, 11GB of RAM, CPU only
নীতি, মামলা এবং ওপেন-সোর্স দ্বন্দ্ব
Apple–OpenAI আইনি লড়াই. Apple OpenAI-কে ট্রেড সিক্রেটের জন্য মামলা করেছে, প্রধান হার্ডওয়্যার অফিসার সহ 400-এর বেশি কর্মচারী শিকার করার অভিযোগ, OpenAI-এর IPO এবং ক্লাউড বিশ্বাসকে হুমকির মুখে ফেলেছে।
- → Apple sues OpenAI after ex-engineer allegedly used bug to steal trade secrets
- → The wildest allegations in Apple’s trade secrets lawsuit against OpenAI
- → The 6 wildest claims in Apple’s lawsuit against OpenAI
- → OpenAI pushes back on Apple trade secret lawsuit
- → Sam Altman didn’t need another lawsuit
- → How Apple’s big lawsuit could disrupt OpenAI’s IPO plans
- → Apple’s plot to crush OpenAI
- → Apple’s lawsuit couldn’t come at a worse time for OpenAI
চীনা ওপেন-সোর্স উত্থান. চীনা মডেল এখন HuggingFace ডাউনলোডের 41% অংশ, Kimi K3 পশ্চিমা নেতাদের প্রতিদ্বন্দ্বিতা করে, এবং চীন পশ্চিম বাদ দিয়ে একটি সমান্তরাল AI গভর্নেন্স বডি চালু করেছে।
- → The real AI race may no longer be at the frontier
- → Source: the Trump administration and industry groups discussed streamlining US open model releases of equal or lesser capability to leading Chinese open models
- → Chinese President Xi Jinping speaks at World AI Conference and reaffirms commitment to open source to promote"openness and win-win"
- → China’s Xi Touts Open-Source AI and Takes a Swipe at U.S. Dominance
- → China's new World Artificial Intelligence Cooperation Organization is President Xi's clearest play yet for a parallel AI order
জবাবদিহিতার দাবি বাড়ছে. জার্মান নিয়ন্ত্রকরা চ্যাটবটকে বিষয়বস্তুর জন্য দায়ী করেছে, xAI একজন ব্যবহারকারীকে CSAM উৎপাদনের জন্য মামলা করেছে, এবং Demis Hassabis ফ্রন্টিয়ার সেফটি পর্যালোচনার জন্য FINRA-সদৃশ একটি সংস্থার প্রস্তাব করেছে।
এটি সপ্তাহের পর্যালোচনা - আগামী রবিবার দেখা হবে।