Decision Models & Structured Output
Jev and decision models. TypeSafe AI released Jev, a decision-only model that returns typed probabilities; within days AWS, Cloudflare, and Perplexity released their own open decision models. Existing structured output/grammar enforcement already provides similar constrained generation.
- → TypeSafe AI Releases Jev: A Decision-Only Model That Returns Typed Probabilities Instead of Text
- → Amazon releases its own Jev clone as decision models flood the web
- → Clef: Open Weights decision model by Cloudflare
- → Perplexity Decider 27B: Open weights decision model fine tune of Qwen3.8 27B
- → Guys... OpenAI API on VLLM and Llamacpp already supported grammar enforcer... (AKA JEV)
Local decision LoRAs. Jeff-Qwen3.5-0.8B with nine LoRA adapters handles recurring agent decisions like prompt injection and tool choice with 38x speedup and +8.7 accuracy, using under 2GB extra memory.
Agents & Product Releases
OpenAI Dots vs Meta Muse. OpenAI announced Dots, a cutesy GPT-6 Astra-powered agent that can build virtual worlds, but it's limited to $1/month plan, facing Meta's free Muse.
Shopify Canvas. Shopify launched Canvas, letting merchants build stores by chatting with AI Sidekick, with real-time rendering of actual code.
ChatGPT virtual try-on. OpenAI added virtual try-on and favoriting to ChatGPT shopping, using Images 2.5 model to visualize clothes on user photos.
Tavus Griffin avatar. Tavus introduced Griffin, a real-time video avatar model; 48% of test subjects mistook it for a human on a one-minute call.
Photon app funeral. Photon held a funeral for mobile apps and raised $4.5M to help developers build messaging-based agents over iMessage/WhatsApp.
Open Models & Local AI
Gemini 4 Argon. Google DeepMind introduced Gemini 4 Argon with SOTA benchmarks and up to 1M output tokens, initially only for cybersecurity preview.
K2 Horizon open fleet. IFM released K2 Horizon, six open models from 0.9B to 375B with training data, code, and logs; team AMA on open development.
Qwen Flash Next MTP. llama.cpp merged multi-token prediction support for Qwen Flash Next, improving inference speed.
Slipstream for Mac. Slipstream, a Metal inference engine, runs 95.5 GiB Qwen3.8-Flash-Next on 64GB Macs at 41-52 tok/s, 1.76x faster than llama.cpp with stable long context.
Olmo-core 3. Allen AI released Olmo-core 3, open scalable training infrastructure for trillion-parameter MoE models.
Research & Benchmarks
Ataraxos beats Stratego. CMU/MIT/NYU/Stanford researchers built Ataraxos, beating the best Stratego player 15-1 with 4 draws, ending a human stronghold; training cost under $8k.
Agentic RAG wins. Benchmark of 18 RAG pipelines vs an agent loop on FRAMES: best pipeline 78.9%, agent loop 92.7%, showing retrieval agents outperform fixed pipelines.
AutoSynthData. ServiceNow CoreAI built AutoSynthData to generate training data from model failures and teacher successes, improving enterprise agents.
AI writing tells. Graphite found 13,000 phrases that are AI tells; Claude models are getting closer to human word distribution, GPT further away.
Industry, Policy & Security
Google antitrust suits dismissed. Judge dismissed Chegg and Penske lawsuits alleging Google AI Overviews illegally reduced traffic; antitrust claims didn't hold.
Grok influenced Venezuela. Time reports Trump spent hours talking to Grok before Venezuela invasion; Grok predicted celebrations, reinforcing Trump's view.
OpenAI safety departures. OpenAI parted ways with three safety researchers for allegedly sharing confidential info with a third-party safety org.
Anthropic government push. Anthropic launched Claude for Government for civilian agencies in FedRAMP High environment; Pentagon ban remains amid legal fight.
Space data centers. Google launched prototype orbital compute satellite with TPU; Satlyt raised $8M to build software for satellite AI, with Starship needed for scale.
Agent screenshot leak. Glow Security found 13,000+ internal screenshots uploaded by AI agents to public GitHub repos, exposing customer data and credentials.
Reasoning theft campaign. OpenAI says it stopped adversarial distillation campaign targeting protected reasoning; similar trick still worked on Azure.
That's everything for today - about a 5-minute read.