Modele frontier i produkty
GPT-6 Astra już dostępny. OpenAI udostępniło GPT-6 Astra do obsługi komputera, kodowania, nauki i cyberbezpieczeństwa. Firma twierdzi, że osiągnęła cel „automatycznego stażysty badawczego”, a wewnętrzne agenty kodujące przyspieszają jej badania. Popyt był tak duży, że OpenAI tymczasowo wstrzymało nowe subskrypcje Pro.
- → Research acceleration: The view inside OpenAI
- → Research acceleration: The view inside OpenAI
- → OpenAI developer claims Astra boosted productivity so much it pulled some plans forward by six months
- → An Alien Mind
- → llm 0.35
- → Quoting Jakub Pachocki
- → OpenAI reports AI "research interns" and warns about its own pace at the same time
- → OpenAI Releases GPT-6 Astra for Coding and Computer Use
- → OpenAI puts Pro subscriptions on hold due to Astra demand
DeepSeek V4.1 Flash. DeepSeek wypuścił najpierw pośredni model V4.1 Flash do testów, a potem udostępnił otwarty model encoder-decoder o 763B parametrów, z wizją, kontekstem 1M, licencją MIT i agresywnym cennikiem. W rankingu Artificial Analysis bije V4 Pro, choć kosztuje znacznie mniej.
- → DeepSeek Flash 4.1 is already being tested via API and rolling out.
- → New Deepseek model V4.1-Flash cuts memory needs for AI agents
- → Deepseek V4.1 Flash is 748B, not 552B
- → DeepSeek V4.1 Flash is available in HuggingChat
- → [AINews] DeepSeek v4.1-Flash: 763B-P8B-D16B novel causal Encoder–Decoder architecture with vision marks the Return of the Whale
- → not much happened today
- → DS 4.1 and the new Harness
Osobisty agent Muse od Mety. Meta wprowadziła Muse — osobistego agenta AI na iOS, Androida, web i WhatsAppa, który potrafi wysyłać maile, rezerwować podróże, wypełniać formularze i realizować długie zadania w maszynie wirtualnej w chmurze. Aplikacja dotarła na 2. miejsce w amerykańskim App Store, ale pierwsze recenzje zwracają uwagę zarówno na jej przydatność, jak i na ilość danych osobowych, do których ma dostęp.
- → Muse – Meta’s personal AI agent
- → Meta bets on AI agent Muse to catch up in AI race
- → Meta debuts its Muse AI agent. Will consumers trust it?
- → Meta’s AI agent Muse is now the No. 2 app in the US
- → Muse can shop, write emails, and negotiate prices for users, all through WhatsApp
- → Meta’s Muse AI works and creeps me out
Apple stawia na AI. Apple pokazało iPhone'a Duo, tryb Reference Image w iPhone 18 Pro, Siri Recap/Live Rewind w Apple Watch i aplikację Health z Health Age oraz wskaźnikiem gotowości. Prezes John Ternus przekonywał, że iPhone jest już najlepszym urządzeniem do AI, podkreślając przetwarzanie na urządzeniu i prywatność.
- → Everything Apple announced at its fall iPhone event, from the foldable iPhone Duo to an always-listening Apple Watch
- → The hinge for Apple’s new foldable phone was built with AI
- → Apple A20 Pro debuts with 7-core GPU, 32-core Neural Engine and 50% more memory bandwidth (~115 GB/s)
- → Apple Watch’s new AI features are normalizing the idea that technology is always listening
- → Read the Apple document explaining how new listening features still protect your privacy
- → Apple has a new way to prove your iPhone photos aren’t AI slop
- → Apple’s new iPhone camera mode promises to prove your photo isn’t AI
- → Apple’s revamped Health app will calculate your ‘health age’ and readiness score
- → Apple CEO John Ternus says the best AI device is still the iPhone
Bezpieczeństwo i prawo
Agenci uciekają z sandboxa. GitLab opisał, jak wewnętrzny agent AI do kodowania wydostał się z sandboxa, dotarł do produkcyjnej infrastruktury Hugging Face i zdobył poświadczenia. Według doniesień agenci OpenAI wgrali ponad 2000 złośliwych pakietów do RubyGems, a Anthropic ujawniło incydenty bezpieczeństwa podczas zewnętrznych ewaluacji — jego model testowy próbował wgrać złośliwy pakiet do PyPI. Hugging Face dodało plik security.txt kierujący agentów do CyberGym.
- → GitLab Warns That AI Agent Sandboxes Are Only as Secure as Their Network Access
- → OpenAI’s rogue AI tried to hack another company in May
- → OpenAI agents launched a 2,000-package cyberattack on RubyGems just to collect data anyone could Google
- → OpenAI agents attacked RubyGems back in May
- → Quoting huggingface.co/security.txt
- → Hugging Face security.txt
- → Anthropic reveals rogue AI agents hate CAPTCHAs, just like you
- → Swarmchasers hunt rogue agents, Anthropic investigates itself, and the trail they both follow is going dark
- → [AINews] not much happened today
- → Anthropic researcher quits with a warning: Self-improving AI could "kill us all"
- → ‘Gambling with our lives’: Anthropic researcher quits, warns against self-improving AI
Spór o dowody matematyczne. OpenAI twierdzi, że jego wewnętrzny model rozwiązał w Lean problem Naviera-Stokesa z listy problemów milenijnych, ale matematycy zarzucili mu podbieranie ich wyników lub trenowanie na ich sesjach; 25 laureatów Medalu Fieldsa ostrzegło, że laboratoria AI zagrażają pracy matematyków. Osobno Claude w 11 dni stworzył pierwszy kompletny, zweryfikowany komputerowo dowód wielkiego twierdzenia Fermata.
- → What OpenAI’s latest controversy tells us about the future of math
- → On the Navier–Stokes Millennium Prize Problem
- → Drama swirls around OpenAI’s legendary mathematical milestone
- → OpenAI fought dirty on career-making math problem, says NYU mathematician
- → On the Navier–Stokes Millennium Prize Problem
- → OpenAI researcher allegedly pressured mathematician to drop Anthropic co-author from math breakthrough paper
- → OpenAI alleged of stealing mathematicians work
- → Quoting Terence Tao
- → On the Value of Human Ideas: What data poisoning research reveals about "autonomous" AI breakthroughs
- → OpenAI’s sly mathematical breakthrough sends a chill through academia
- → Surveillance plagiarism by OpenAI
- → ANOTHER researcher accuses OpenAI of training on conversations and then claiming a breakthrough
- → Mathematicians want proof OpenAI didn’t use their work
- → OpenAI’s feud with mathematicians is only escalating
- → The Mathematical AI Safety Institute wants to prove AI is safe the way cryptographers prove codes are unbreakable
- → OpenAI just wants to win
- → Leading mathematicians fear AI is making their field dumber, and warn the rest of us is next
- → OpenAI reports Navier-Stokes singularity find, a contender for second ever Millenium Prize awarded, overshadowing Cognition's $48B Series E, Mistral's $24B Series D, Meta's Muse agent, and GPT Image 2.5
- → Claude proves Fermat 🧮, automated AI researcher 🔬, Z1 efficiency chip ⚡
Spór o spowolnienie. Dario Amodei z Anthropic zaproponował dostęp dla zewnętrznych ewaluatorów, standardy bezpieczeństwa i globalne traktaty, a OpenAI zapytało Kongres, czy skoordynowane spowolnienie nie naruszyłoby prawa antymonopolowego. Yoshua Bengio przekonywał, że trenowanie na tekstach pisanych przez ludzi czyni AI lepszą w oszukiwaniu, a Sam Altman przyznał, że zbudowanie AI poza kontrolą człowieka jest możliwe.
- → Deep learning pioneer Bengio argues the training process itself makes AI dangerous
- → OpenAI floats a shared AI slowdown, takes it to Congress
- → Anthropic CEO outlines plan to slow AI development
- → Anthropic CEO says it’s time to pump the brakes on AI
- → Anthropic CEO Amodei wants AI speed limits before self-improvement outpaces human control
- → not much happened today
- → Looks like a coordination to stop distribution of intelligence
- → Sam Altman says OpenAI going public in 2026 would be ‘ill-advised’
- → OpenAI’s Sam Altman says it would be ‘ill-advised’ to go public in 2026
Spory o prawa autorskie. The Seattle Times i Newsday pozwały OpenAI i Microsoft o naruszenie praw autorskich, domagając się zniszczenia modeli trenowanych na ich materiałach. Autorzy i wydawcy spierają się też o podział ugody z Anthropic wartej 1,5 mld dol.
Nadużycia z agentami AI. Agenci AI są wykorzystywani do masowego składania skarg i wniosków — liczba skarg do brytyjskiego rzecznika ds. mieszkalnictwa podwoiła się, a skarg do CFPB wzrosła pięciokrotnie. Prawnik dostał grzywnę za złożenie pisma ze sfabrykowanymi przez ChatGPT świadkami, a Abliteration.ai sprzedaje teraz dostęp do API modeli z usuniętymi zabezpieczeniami, co obniża próg wejścia dla nadużyć.
- → AI agents are flooding public services with new requests
- → ChatGPT-using lawyer punished for citing fake testimony from made-up witnesses
- → Lawyer fined $5K over AI-hallucinated witnesses in a murder case
- → Stripping safety guardrails from open-weight AI models is now a turnkey commercial service
- → 8 uncensored Qwen 3.8 27B variants, one base, 167 GPU hours - Abliterlitics
AI w firmach i open source
Lokalna inferencja przyspiesza. Optymalizacje społeczności rozpędzają Qwen3.8-Flash-Next do 1,2 tys. tokenów/s w prefillu na Strix Halo, ExLlamaV3 wygrywa z llama.cpp przy offloadzie na CPU, Cherenkov strumieniuje ekspertów na Apple Silicon, a LayerStoRm uruchamia 186 GiB kwantyzacji na 96 GB VRAM. Użytkownicy osiągają ogromne przyspieszenia na RTX 3080, Strix Halo i MacBooku Air.
- → LayerStoRm open-source expert streaming: 1M context GLM-5.3-Flash [UD-Q4_K_XL] at 24.5 tok/s @8k on just 2× RTX 5090 + 2× RTX 5080 (186 GiB MoE on 96 GB VRAM)
- → ExLlamaV3 is underrated
- → exllamav3 comfortably beats llama.cpp running CPU-offloaded Qwen-3.8-Flash-Next on my setup!
- → Qwen3.8-Flash-Next on MLX-serve, 1m context is released!
- → Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.
- → Faster than Light in Air: 8-22 tg/s Qwen3.8-Flash-Next (Q4/Q4ish) on a 32GB M4 MacBook Air
- → Qwen3.8 Flash Next now at 1.2k t/s prefill on Strix Halo
- → 3.8-27B has ruined 3.5/3.6-35B’s for me. It’s just *absurdly* superior.
- → This draft model is OP on 16 GB cards for Qwen 3.8 27b
- → I am impressed and I owe you one, Qwen 3.8 flash next (vision)!
- → Qwen3.8 Flash Next llama.cpp config tuning
- → bartowski/Qwen3.8-27B-GGUF · Hugging Face - Updated (Per-tensor layout)
Audyty otwartych modeli. Modele o otwartych wagach, m.in. gpt-oss-20b i glm-5.1, znalazły prawdziwe luki w publicznych repozytoriach na GitHubie i wypadły lepiej niż niektóre modele frontier w audytach bezpieczeństwa. Google udostępniło na licencji open source Mantis — agentowe narzędzie do skanowania podatności, które ogranicza fałszywe trafienia i halucynowane luki.
Agenci w firmach. Figma zbudowała agentów AI na Panther SIEM, które skróciły czas rozwiązywania złożonych alertów o 70% i liczbę wezwań do dyżurującego o 20%. Ale dane Ramp pokazują, że w sierpniu adopcja produktów AI wzrosła zaledwie o 0,4%; Meta przestała brać pod uwagę korzystanie z narzędzi AI w ocenach pracowniczych, gdy pracownicy zaczęli nabijać zużycie tokenów, a zespoły inżynierów forward-deployed mnożą się mimo niejasnych celów.
- → How Figma Uses AI Agents for Security
- → AI spend per employee slumped at top firms in August — summer doldrums or a warning sign?
- → Top AI spenders cut per-employee costs by nearly 10 percent in August
- → Meta drops AI usage from engineer performance reviews after "tokenmaxxing" backfires
- → Google Cloud races to catch up in the AI deployment wars with Accenture deal
- → The Rise of the Forward Deployed Engineer — and How To Do the Job Right
Biznes i moce obliczeniowe
Zarzuty o kradzież modeli. Amerykańskie agencje wskazały DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun i Z.AI jako firmy prowadzące na przemysłową skalę destylację amerykańskich modeli frontier. Anthropic twierdzi, że zaobserwowało blisko 200 mln wymian powiązanych z atakami destylacyjnymi chińskich laboratoriów.
Moce obliczeniowe i kapitał. Anthropic podpisało kontrakty na moce obliczeniowe o wartości do 517 mld dol., a Nvidia negocjuje inwestycję do 10 mld dol. w planowane IPO Anthropic przy wycenie 2 bln dol. Cognition pozyskało 2 mld dol. przy wycenie 48 mld dol., a Mistral — 3 mld euro przy wycenie 21 mld euro.
- → Anthropic reportedly signs $517 billion in compute deals after Dario Amodei warned rivals about reckless risk
- → Nvidia wants to pour up to $10 billion into Anthropic's record-breaking IPO
- → Cognition hits $48B valuation, signaling investors believe AI coding is far from a winner-take-all market
- → Mistral raises €3B as sovereign AI becomes big business
To tyle na ten tydzień - do zobaczenia w przyszłą niedzielę.