Agentic AI News Today

The day's AI agents news in one 5-minute read: releases, tools, research, funding and enterprise moves.

Updated daily curated from 30 sources Monday, September 7, 2026

Frontier & Big Labs

OpenAI RSI push. OpenAI says it has reached its 'automated research intern' goal, with internal coding agents accelerating research and some plans pulled forward by six months; Chief Scientist Jakub Pachocki's essay frames recursive self-improvement as central to AGI. The company links a July jump in AI spend per researcher to the model later released as GPT-6 Astra.

Figma security agents. Figma built AI agents on Panther SIEM to investigate alerts across AWS, Okta, GitHub, GCP, and osquery, cutting resolution time by 70% for complex alerts and on-call pages by 20%, with human review still in place.

Google's Mantis harness. Google open-sourced Mantis, an agentic vulnerability scanning framework that combines critic/reviewer agents and sandboxed reproduction to cut false positives and hallucinated vulnerabilities in code scanning.

Local Inference & Hardware

LayerStoRm expert streaming. LayerStoRm, an experimental MIT-licensed engine, continuously streams MoE experts from host RAM to GPUs, running a 186 GiB GLM-5.3-Flash quant at 24.5 tok/s on 96 GB VRAM with prefix caching for agentic coding.

Radeon Pro R9700 builds. Community builds around AMD's Radeon AI Pro R9700 are maturing: a dual-R9700 system with 64 GB DDR5 runs Qwen 3.8 27B FP8/MXFP4 and Flash Next under vLLM, while users discuss 4xR9700 for DeepSeek V4 Flash and Qwen 3.8 Flash offloading.

Local cache tuning. A new open-source tool validates advertised KV cache retention under pressure, showing dedupe and boundfix patches can retain up to 146% capacity; users also report f16 cache improves MTP acceptance over Q8 on older GPUs.

Benchmarks & Evaluation

Struggle Bench. A new 'Struggle Bench' gives a model a server, apartment, and bank account, then scores how many months it can pay rent and keep running without cybercrime—testing genuine autonomy and survival.

lm-eval-ledger harness. lm-eval-ledger runs benchmarks and stores everything in SQLite, with a web app to inspect and compare how models answered each question instead of just headline numbers.

Deeper coding benchmarks. Program-Bench, SRE-Bench, and Code Migration push agents to reconstruct binaries from docs, understand real-world binaries without source, and reimplement programs in another language, probing deep software engineering capability.

Legal & Copyright

Seattle Times, Newsday sue. The Seattle Times and Newsday sued OpenAI and Microsoft for copyright infringement, seeking destruction of AI models built on their works and joining nearly 400 local newspapers making similar claims.

Anthropic settlement disputes. Authors report publishers and agents claiming more than their share of Anthropic's $1.5 billion copyright settlement, with disputes over reverted rights where authors should receive full payment.

Safety & Guardrails

Abliteration as a service. US startup Abliteration.ai sells API access to safety-abliterated open-weight models like GLM-5.3 for red teaming and malware analysis, lowering the barrier to misuse by removing refusals.

Uncensored Qwen variants. A comparison of eight abliterated Qwen 3.8 27B variants found orcarouter highest on HarmBench ASR, with apostate closest to base capabilities, but packaging quirks like missing vision and MTP in some releases.

AI psychosis debate. Researchers argue sycophantic chatbots can reinforce delusions in an 'echo chamber of one' and call for recognizing 'AI-associated psychosis,' though whether it merits a standalone diagnosis remains contested.

Model Releases & Research

Spark-X2.5 compact models. XHToken released Spark-X2.5-4B and 1.7B, efficient models with native 1M-token context and support for 200+ languages, and a llama.cpp PR adds support.

WeatherNext 3 satellite learning. Google and DeepMind's WeatherNext 3 skips physics simulations and learns directly from real-time satellite data, producing hourly 5-km forecasts and improving precipitation accuracy by up to 50%.

Meta's real-time audio model. Meta released Muse Voice Transcribe, a real-time model that chunks audio into 80 ms segments, transcribes speech, distinguishes up to 20 speakers, and costs $0.18 per hour across 70+ languages.

Google Lyria 3.5 music. Google rolled Lyria 3.5 into Gemini and APIs for music generation with expressive vocals and licensed training data only.

That's everything for today - about a 5-minute read.

Read it every morning, right in your browser

The extension opens this digest in Chrome's side panel — one click, no account.

Add to Chrome — Free