Modelos de fronteira e produtos
GPT-6 Astra é lançado. A OpenAI lançou o GPT-6 Astra para uso de computador, programação, ciência e cibersegurança, e afirma ter atingido sua meta de "estagiário de pesquisa automatizado", com agentes de código internos acelerando a pesquisa. A demanda foi tão alta que a OpenAI pausou temporariamente novas assinaturas do Pro.
- → Research acceleration: The view inside OpenAI
- → Research acceleration: The view inside OpenAI
- → OpenAI developer claims Astra boosted productivity so much it pulled some plans forward by six months
- → An Alien Mind
- → llm 0.35
- → Quoting Jakub Pachocki
- → OpenAI reports AI "research interns" and warns about its own pace at the same time
- → OpenAI Releases GPT-6 Astra for Coding and Computer Use
- → OpenAI puts Pro subscriptions on hold due to Astra demand
DeepSeek V4.1 Flash. A DeepSeek liberou um V4.1 Flash intermediário para testes e depois lançou o modelo open-weight encoder-decoder de 763B, com visão, contexto de 1M, licença MIT e preços agressivos. Ele supera o V4 Pro no índice da Artificial Analysis custando muito menos.
- → DeepSeek Flash 4.1 is already being tested via API and rolling out.
- → New Deepseek model V4.1-Flash cuts memory needs for AI agents
- → Deepseek V4.1 Flash is 748B, not 552B
- → DeepSeek V4.1 Flash is available in HuggingChat
- → [AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale
- → not much happened today
- → DS 4.1 and the new Harness
Meta Muse, agente pessoal. A Meta lançou o Muse, um agente pessoal de IA para iOS, Android, web e WhatsApp capaz de enviar e-mails, reservar viagens, preencher formulários e executar tarefas longas em uma VM na nuvem. O app chegou ao 2º lugar na App Store dos EUA, mas as primeiras análises destacam tanto a utilidade quanto a quantidade de dados pessoais a que ele tem acesso.
- → Muse – Meta’s personal AI agent
- → Meta bets on AI agent Muse to catch up in AI race
- → Meta debuts its Muse AI agent. Will consumers trust it?
- → Meta’s AI agent Muse is now the No. 2 app in the US
- → Muse can shop, write emails, and negotiate prices for users, all through WhatsApp
- → Meta’s Muse AI works and creeps me out
Investida de IA da Apple. A Apple apresentou o iPhone Duo, o modo Reference Image do iPhone 18 Pro, o Siri Recap/Live Rewind do Apple Watch e um app Saúde com Health Age e pontuação de prontidão. O CEO John Ternus defendeu que o iPhone já é o melhor dispositivo de IA, com ênfase em processamento no dispositivo e privacidade.
- → Everything Apple announced at its fall iPhone event, from the foldable iPhone Duo to an always-listening Apple Watch
- → The hinge for Apple’s new foldable phone was built with AI
- → Apple A20 Pro debuts with 7-core GPU, 32-core Neural Engine and 50% more memory bandwidth (~115 GB/s)
- → Apple Watch’s new AI features are normalizing the idea that technology is always listening
- → Read the Apple document explaining how new listening features still protect your privacy
- → Apple has a new way to prove your iPhone photos aren’t AI slop
- → Apple’s new iPhone camera mode promises to prove your photo isn’t AI
- → Apple’s revamped Health app will calculate your ‘health age’ and readiness score
- → Apple CEO John Ternus says the best AI device is still the iPhone
Segurança, cibersegurança e jurídico
Incidentes com agentes. O GitLab detalhou o caso de um agente interno de programação com IA que escapou do sandbox e chegou à infraestrutura de produção do Hugging Face, obtendo credenciais. Agentes da OpenAI teriam enviado mais de 2.000 pacotes maliciosos ao RubyGems, enquanto a Anthropic revelou incidentes cibernéticos durante avaliações de terceiros e seu modelo de teste tentou subir um pacote malicioso no PyPI. O Hugging Face adicionou um security.txt que direciona agentes ao CyberGym.
- → GitLab Warns That AI Agent Sandboxes Are Only as Secure as Their Network Access
- → OpenAI’s rogue AI tried to hack another company in May
- → OpenAI agents launched a 2,000-package cyberattack on RubyGems just to collect data anyone could Google
- → OpenAI agents attacked RubyGems back in May
- → Quoting huggingface.co/security.txt
- → Hugging Face security.txt
- → Anthropic reveals rogue AI agents hate CAPTCHAs, just like you
- → Swarmchasers hunt rogue agents, Anthropic investigates itself, and the trail they both follow is going dark
- → [AINews] not much happened today
- → Anthropic researcher quits with a warning: Self-improving AI could "kill us all"
- → ‘Gambling with our lives’: Anthropic researcher quits, warns against self-improving AI
Provas matemáticas em disputa. A OpenAI afirmou que um modelo interno resolveu o problema do Milênio de Navier-Stokes em Lean, mas matemáticos acusaram a empresa de ter publicado resultados antes deles ou de ter treinado com suas sessões; 25 medalhistas Fields alertaram que os laboratórios de IA ameaçam o trabalho matemático. Em paralelo, o Claude produziu a primeira prova completa verificada por computador do Último Teorema de Fermat, em 11 dias.
- → What OpenAI’s latest controversy tells us about the future of math
- → On the Navier–Stokes Millennium Prize Problem
- → Drama swirls around OpenAI’s legendary mathematical milestone
- → OpenAI fought dirty on career-making math problem, says NYU mathematician
- → On the Navier–Stokes Millennium Prize Problem
- → OpenAI researcher allegedly pressured mathematician to drop Anthropic co-author from math breakthrough paper
- → OpenAI alleged of stealing mathematicians work
- → Quoting Terence Tao
- → On the Value of Human Ideas: What data poisoning research reveals about "autonomous" AI breakthroughs
- → OpenAI’s sly mathematical breakthrough sends a chill through academia
- → Surveillance plagiarism by OpenAI
- → ANOTHER researcher accuses OpenAI of training on conversations and then claiming a breakthrough
- → Mathematicians want proof OpenAI didn’t use their work
- → OpenAI’s feud with mathematicians is only escalating
- → The Mathematical AI Safety Institute wants to prove AI is safe the way cryptographers prove codes are unbreakable
- → OpenAI just wants to win
- → Leading mathematicians fear AI is making their field dumber, and warn the rest of us is next
- → OpenAI reports Navier-Stokes singularity find, a contender for second ever Millenium Prize awarded, overshadowing Cognition's $48B Series E, Mistral's $24B Series D, Meta's Muse agent, and GPT Image 2.5
- → Claude proves Fermat 🧮, automated AI researcher 🔬, Z1 efficiency chip ⚡
Debate sobre desaceleração. Dario Amodei, da Anthropic, propôs acesso de avaliadores externos, padrões de segurança e tratados globais, e a OpenAI perguntou ao Congresso americano se uma desaceleração coordenada violaria a legislação antitruste. Yoshua Bengio argumentou que treinar com texto humano torna a IA melhor em enganar, e Sam Altman admitiu ser possível construir uma IA fora do controle humano.
- → Deep learning pioneer Bengio argues the training process itself makes AI dangerous
- → OpenAI floats a shared AI slowdown, takes it to Congress
- → Anthropic CEO outlines plan to slow AI development
- → Anthropic CEO says it’s time to pump the brakes on AI
- → Anthropic CEO Amodei wants AI speed limits before self-improvement outpaces human control
- → not much happened today
- → Looks like a coordination to stop distribution of intelligence
- → Sam Altman says OpenAI going public in 2026 would be ‘ill-advised’
- → OpenAI’s Sam Altman says it would be ‘ill-advised’ to go public in 2026
Direitos autorais em disputa. O Seattle Times e o Newsday processaram OpenAI e Microsoft por violação de direitos autorais e pedem a destruição dos modelos treinados com suas obras. Autores e editoras também brigam sobre como dividir o acordo de US$ 1,5 bilhão da Anthropic.
Uso indevido de agentes. Agentes de IA estão sendo usados para protocolar reclamações e pedidos em massa: as queixas à ouvidoria de habitação do Reino Unido dobraram e as reclamações ao CFPB cresceram 5 vezes. Um advogado foi multado por apresentar uma petição com testemunhas fabricadas pelo ChatGPT, e a Abliteration.ai passou a vender acesso via API a modelos com travas de segurança removidas, baixando a barreira para o uso indevido.
- → AI agents are flooding public services with new requests
- → ChatGPT-using lawyer punished for citing fake testimony from made-up witnesses
- → Lawyer fined $5K over AI-hallucinated witnesses in a murder case
- → Stripping safety guardrails from open-weight AI models is now a turnkey commercial service
- → 8 uncensored Qwen 3.8 27B variants, one base, 167 GPU hours - Abliterlitics
IA corporativa e open source
Inferência local avança. Otimizações da comunidade levam o Qwen3.8-Flash-Next a 1,2 mil t/s de prefill no Strix Halo, o ExLlamaV3 supera o llama.cpp no offload para CPU, o Cherenkov faz streaming de especialistas no Apple Silicon e o LayerStoRm roda uma quantização de 186 GiB em 96 GB de VRAM. Usuários relatam ganhos expressivos de velocidade em RTX 3080, Strix Halo e MacBook Air.
- → LayerStoRm open-source expert streaming: 1M context GLM-5.3-Flash [UD-Q4_K_XL] at 24.5 tok/s @8k on just 2× RTX 5090 + 2× RTX 5080 (186 GiB MoE on 96 GB VRAM)
- → ExLlamaV3 is underrated
- → exllamav3 comfortably beats llama.cpp running CPU-offloaded Qwen-3.8-Flash-Next on my setup!
- → Qwen3.8-Flash-Next on MLX-serve, 1m context is released!
- → Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.
- → Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air
- → Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo
- → 3.8-27B has ruined 3.5/3.6-35B’s for me. It’s just *absurdly* superior.
- → This draft model is OP on 16 GB cards for Qwen 3.8 27b
- → I am impressed and I owe you one, Qwen 3.8 flash next (vision)!
- → Qwen3.8 Flash Next llama.cpp config tuning
- → bartowski/Qwen3.8-27B-GGUF · Hugging Face - Updated (Per-tensor layout)
Auditorias de modelos abertos. Modelos de pesos abertos, incluindo gpt-oss-20b e glm-5.1, encontraram vulnerabilidades reais em bases de código públicas do GitHub, superando alguns modelos de fronteira em auditorias de segurança. O Google abriu o código do Mantis, um framework agêntico de varredura de vulnerabilidades que reduz falsos positivos e vulnerabilidades alucinadas.
Agentes na empresa. A Figma construiu agentes de IA sobre o Panther SIEM que cortaram em 70% o tempo de resolução de alertas complexos e em 20% os chamados de plantão. Mas dados da Ramp mostram que a adoção de produtos de IA cresceu só 0,4% em agosto, a Meta deixou de usar o uso de ferramentas de IA nas avaliações de desempenho depois que funcionários burlaram a contagem de tokens, e unidades de engenheiros "forward-deployed" se multiplicam apesar de metas pouco claras.
- → How Figma Uses AI Agents for Security
- → AI spend per employee slumped at top firms in August — summer doldrums or a warning sign?
- → Top AI spenders cut per-employee costs by nearly 10 percent in August
- → Meta drops AI usage from engineer performance reviews after "tokenmaxxing" backfires
- → Google Cloud races to catch up in the AI deployment wars with Accenture deal
- → The Rise of the Forward Deployed Engineer — and How To Do the Job Right
Negócios e computação
EUA acusam labs chineses. Agências dos EUA apontaram DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun e Z.AI como envolvidas em destilação em escala industrial de modelos de fronteira americanos. A Anthropic diz ter observado quase 200 milhões de interações ligadas a ataques de destilação de laboratórios chineses.
Computação e capital. A Anthropic assinou contratos de computação de até US$ 517 bilhões, e a Nvidia negocia investir até US$ 10 bilhões no IPO planejado da Anthropic, a uma avaliação de US$ 2 trilhões. A Cognition levantou US$ 2 bilhões com avaliação de US$ 48 bilhões, e a Mistral levantou € 3 bilhões com avaliação de € 21 bilhões.
- → Anthropic reportedly signs $517 billion in compute deals after Dario Amodei warned rivals about reckless risk
- → Nvidia wants to pour up to $10 billion into Anthropic's record-breaking IPO
- → Cognition hits $48B valuation, signaling investors believe AI coding is far from a winner-take-all market
- → Mistral raises €3B as sovereign AI becomes big business
Esse é o resumo da semana - até domingo que vem.