Google Research proposes RRSI, using regularized recursive self-improvement to prevent agent harnesses from overfitting to training tasks.
2026-09-23
— The large-model price war has begun, and the shadow of security and military AI surfaced on the same day.
OpenAI releases GPT-6 Sol and Luna, priced at half of the previous generation; Anthropic releases Claude Opus 5.5, matching Fable 5.1 performance at 40% lower cost. The two companies launched on the same day, marking the entry of frontier model competition into a cost-effectiveness phase. On the security front, ShinyHunters claims to have breached the FBI and stolen all employee data; the Pentagon admits that over-reliance on AI led to the mistaken bombing of an Iranian school.
Schlagzeilen
OpenAI releases GPT-6 Sol and Luna, prices halvedMehrere Quellen ×5
OpenAI launches GPT-6 Sol and GPT-6 Luna as lightweight complements to GPT-6 Astra, focused on cost efficiency. Sol targets complex tasks such as coding, while Luna targets high-volume paperwork; both are priced at half of comparable GPT-5.6 models for input/output. Why it matters: for developers building applications on the API, this means inference costs drop substantially while maintaining near-frontier performance, especially benefiting agents and batch-processing scenarios.
The comment section generally acknowledges the price cuts and high cost-effectiveness, but some believe the performance gains are limited and the release timing appears aimed at competitors.
Anthropic releases Claude Opus 5.5, matching Fable 5.1 performance at 40% lower costMehrere Quellen ×3
Anthropic releases Claude Opus 5.5, claiming it reaches Claude Fable 5.1 levels on most tasks, runs at 40% lower cost than Opus 5, and cuts output token pricing from $25 per million to $20 per million. The model is Anthropic's first release since it called for slowing the frontier race, and it underwent external evaluation by Frontier Design and METR. Why it matters: a flagship-tier model cutting prices and improving inference efficiency directly lowers API costs for complex coding and knowledge-work workloads, while also reinforcing the safety-testing narrative.
Most acknowledge that Opus 5.5 is more natural, cheaper, and strong in performance, but some point to the version-number jump, still-high pricing, and skepticism toward benchmarks.
ShinyHunters claims to have breached the FBI, stealing all employee and applicant data
The hacker group ShinyHunters claims to have breached multiple FBI-related services and stolen data on all FBI employees and applicants, including names, home addresses, phone numbers, and spouse information. 404 Media verified that some of the data matches public records. Why it matters: if true, this is one of the most serious data breaches in the history of a U.S. law enforcement agency, potentially triggering large-scale counterintelligence risks and once again exposing security shortcomings in government systems.
The comment section generally believes this breach exposes serious negligence in U.S. government cybersecurity, but some doubt the authenticity of the data and say the actual leaked contents are needed for confirmation.
Pentagon admits over-reliance on AI led to mistaken bombing of Iranian school
A Pentagon investigation found that intelligence gaps, outdated imagery, and over-reliance on AI together led to a missile strike on a school in Minab, Iran, on February 28, 2026, killing 123 children. Why it matters: this is a typical case of AI participating in a military kill chain and causing major civilian casualties, raising sharp accountability questions about the boundaries of automated decision-making in defense systems.
Comments generally believe AI is just a scapegoat, that the real responsibility lies with the people who decided to use AI and the military, and that human accountability must be pursued; but some also believe the technology provider Palantir is equally culpable.
GPT-6 Astra cracks Enigma message unsolved since 2005
On September 15, 2026, Carter Leffer used GPT-6 Astra to crack the July 10, 1941 German Army Enigma message MVUEH, which had remained unbroken since 2005. Why it matters: it demonstrates LLMs' assistive capability in cryptanalysis, but the community is skeptical of its authenticity, believing it may involve training-data leakage or hype.
Comments generally question the real value of GPT-6 cracking the Enigma message, believing it may be hype or training-data leakage, but some believe it demonstrates the potential of human-machine collaboration.
Jeden Morgen ein Tech-Digest, für dich kuratiert
Das Web zeigt das große Ganze; Abonnenten bekommen ihr eigenes — nach deinen Interessen kuratiert, dein privates RSS integriert, mit Community-Stimmen, jeden Morgen zugestellt. Dauerhaft kostenlos.
73 Ausgaben erschienen · täglich 150+ Meldungen auf 30 gesiebt
KI-News
The RoboDawn interface transfers VLM intelligence to robot control, exploring generalization from the digital to the physical world.
Tencent ARC's WorldCrafter uses implicit 3D-aware memory to achieve long-horizon consistent video world models.
D-RAC proposes a retrieval-aware chunking method for enterprise documents, improving RAG quality through PDF normalization and multimodal Markdown conversion.
GPT-6 improves prompt caching, with higher cache hit rates and up to 90% discounts on cached input tokens.
Dev & Open Source
Drop is a rootless Linux sandbox with gVisor support that can isolate third-party programs and coding agents.
llm 0.36 adds gpt-6-sol and gpt-6-luna models, and supports declaring models that do not support conversation.
The llm-anthropic 0.29 plugin adds Claude Opus 5.5 support.
The llm-typesafe 0.1a0 plugin supports TypeSafe AI's Jev decision model, outputting structured yes/no or choice answers.
An experiment explores whether gzip can serve as a language model, generating text based on the compression-prediction equivalence.
Community-Themen
Alibaba officially announces Qwen 4 at the Apsara Conference, with the LocalLLaMA community focused on its open-source weights and local deployment potential.
Xiaomi MiMo-V2.6-Pro 1T-A42B becomes the new open-source weight leader, with training costs of only $3 million.
MIT Tech Review warns readers to be wary of this summer's AI hype, including events such as vulnerability discoveries, math breakthroughs, and resignation warnings.
Cisco Talos releases the open-source framework CAIRN to identify malware and hacking tools powered by AI chatbots.
Microsoft leads the takedown of the AI-assisted fraud platform EvilTokens, which had compromised 12,000 Microsoft accounts.
GitHub Trending
Star agent-substrate / substrate Agent Substrate: the core system
Star dream-num / univer The Office Harness for AI Agents — Spreadsheets, Docs, Slides, Canvas, Relational Tables, and PDF in one runtime.
Sponsor Star davila7 / claude-code-templates CLI tool for configuring and monitoring Claude Code
Star google / ax Google's open agentic orchestration runtime
Star mvt-project / mvt MVT (Mobile Verification Toolkit) helps with conducting forensics of mobile devices in order to find signs of a potential compromise.
Star superdesigndev / treg OpenRouter for agent tools. Join community here: https://discord.gg/6mQYYfFMAn
Star browser-use / video-use Edit videos with coding agents
Weitere Fundstücke(62 weitere)
I thought we were supposed to be slowing down the frontier.
We talked to Google’s Oscar winning “Giganerd” about automating science, solving climate change, and how future generations can contribute to science in the age of superintelligent AI
Researchers say another looming threat hangs over some of America's most important critical infrastructure.
We present onPanda, an interactive tool for efficiently annotating LLM alignment data and agent trajectories. onPanda adopts token-level correction as its core interaction: while reading a model response, the annotator locates the first inappropriate token and either picks a substitute from the model's candidate tokens or types the correct text via free-form editing. The system then truncates everything after that position and continues generation from the corrected prefix, repeating this locate
Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We introduce a continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience. The memory widens by acquiring
9月22日,在2026杭州云栖大会企业级Agent实践峰会上,基元律动联合创始人兼CTO韩凯发表演讲《从Harness到RSI飞轮》。
This post is written by a human and I'd appreciate it if you treated it as such. Thanks. So, I've been noticing a pretty clear interest in developing as good a coding and agentic tool-calling model as possible, especially at smaller sizes, sub-50 gigs. However, I'm finding that at least for my use of AI, if I really want to move away from big providers, I am going to require a model that has better world knowledge than the current offerings. Qwen 3.8 27B is a truly fantastic model for tons and t
Repository-level software engineering (SWE) comprises heterogeneous task categories, whose progress under pooled agentic reinforcement learning can be uneven: gains in some categories coincide with regressions in others, while aggregate resolution obscures these changes. Motivated by this category see-saw, we develop a category-aware expert-training and policy-integration framework. Executable task construction and SWE Labeler, an evidence-grounded multi-axis labeling system, organize the traini
Agentic memory is becoming essential for long-horizon AI agents, yet many existing systems rely on autoregressive LLMs to control how memories are organized, retrieved, and used, placing expensive generation on the critical path of memory operations. We introduce \method, a new agentic memory architecture inspired by System-One/System-Two cognition. System One captures fast, lightweight decision-making, whereas System Two performs slower, deliberative reasoning. Jev-Mem brings this division of l
Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the best harness varies across domains, instances, and models, a general-purpose agent must either settle for a suboptimal shared harness or route among an ever-growing set of specialized ones. We therefore study agent harness distillation: using a domain- or instance-optimized harness as training-time gui
Sparse on-policy distillation (OPD) allocates teacher supervision to a small subset of tokens in student-generated trajectories. However, useful teacher guidance can yield a noisy update when its gradient is estimated from a sampled next token. We study this estimation problem at a fixed prefix in information geometry and propose an information-efficiency ratio (IER) based on a signal-to-noise decomposition. IER characterizes relative gradient estimation error under an optimal scalar baseline. A
Skills can improve the performance of Large Language Model (LLM) agents by providing task-specific procedural guidance, while skill optimization further improves their effectiveness through iterative refinement. However, existing skill optimization methods typically represent skills as unstructured natural-language instructions, creating two key challenges: 1) Unstructured skills often lack explicit workflow-level guidance and contain substantial redundancy, making them difficult for LLMs to exe
Building general-purpose agents for industrial deployment requires integrating multiple capabilities, each typically acquired at a distinct stage of training. Yet there is currently no well-established recipe for Agent Continual Learning (ACL), with little understanding of the trade-offs among existing integration paradigms. To address this gap, we introduce ACLArena, a framework for comprehensively studying, analyzing, and evaluating ACL. We first build a sequential training pipeline and conduc
A 421M-parameter model just played Flappy Bird on my desktop CPU (OpenVINO int8) Running on my Intel Core i7 12th gen CPU Converted laya system one model to OpenVINO and quantized to int8
AntLing open sourced the Ming-Image-0.1-Design family: • Ming-Image-0.1-Design, 6B • Ming-Image-0.1-Design-Layer, 6B • Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard
After turning a string of spectacular mathematical results into a reputational crisis, OpenAI is consulting human mathematicians to help it figure out a less disastrous path forward. On Monday, the company announced a new independent panel of mathematicians tasked with advising it and other AI companies on their interactions with mathematical research and the wider […]
GPT‑6 Astra allowed Parallel’s agents to research and synthesize labor-market data in half the time and at half the cost vs. prior models.
Hugging Face CEO:「太棒了」
Manipulation policies must know which objects matter and where they are, yet the pretrained backbones that current robot foundation models build on, from language in vision-language-action models (VLAs) to video generation in world-action models (WAMs), do not directly require this metric grounding, leaving it to be learned implicitly from robot demonstrations. We propose Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding. GAM can be conditioned usin
Business intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in software such as Power BI and Tableau. In traditional BI workflows, users need to prepare data by (1) identifying relevant tables, (2) performing data transformations, and (3) building join relationships, before they can (4) answer their business questions. These steps can be complex and time-consuming, making BI challenging. Given the strong capabilities of large language models (
Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requires grounding the conversation in evolving visual context. We present Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction. Each model serves as a complete co
• Ming-Image-0.1-Design, 6B • Ming-Image-0.1-Design-Layer, 6B • Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard.
better make no mistakes I've been running Qwen 3.8 Flash Next and it's a great driver for Hermes and Pi. I told it to add CUDA_DISABLE_PERF_BOOST=1 to reduce my server's idle power draw
Assign work to AI agents, like any teammate Discussion | Link
Don't know how they did it, but for under 10GB model, the results are astonishing. I am running it on Unsloth Studio. They just released the update, so if you are not seeing the option, I recommend updating your Unsloth Studio. Cheers!
Rabbit, the company behind the underwhelming R1 device, is rolling out a standalone AI agent that you don't need its hardware to use, as reported earlier by Wired. The startup says its new OS3 "agentic operating system" runs in the cloud but operates locally across Windows, Mac, and Linux devices. According to Rabbit, you can […]
British Columbia sues OpenAI, demands Tumbler Ridge shooter’s ChatGPT logs.
Qualcomm said that its new top chip can run 30B mixture-of-expert model locally.
In the hope of uncovering new details about ancient life, researchers have developed a large language model that fills in the gaps in papyrus fragments.
Two years after trying to sidestep mobile apps with dedicated AI hardware, Rabbit is launching OS3, a cross-platform agent that lives on the screens you already use.
The English hospitals recovered patient care data only.
Autonomy-1 will have a small, transformer-based AI model taking charge of a space probe.
Visual FoxPro stopped at version 9 in 2007. A surprising amount of it is still running, in 32 bits, because rewriting a 20-year-old business app is how you lose the business. A customer wanted to keep milking their app for the foreseeable future, so here it is: the same language on a new runtime (Rust, compiled to wasm, checked against the real vfp9.exe), tables no longer stopped at 2 GB, the old 32-bit .fll add-ins still loading, and lambdas, JSON and an HTTP server bolted on for good measure.
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Don’t be fooled by this summer of AI hype —Timnit Gebru, executive director of the Distributed AI Research Institute (DAIR), and Emily M. Bender, professor of linguistics at the University of…
For the people that are unaware or haven't seen the news yet.
Human video datasets offer an abundant and diverse source of interaction data that can complement expensive real-robot data. To bridge the human-to-robot embodiment gap, existing approaches either robotize videos in task-matched settings or address observation and action alignment separately at scale. In this work, we systematically examine whether robotized human videos can serve as an effective and scalable source of supervision for VLA pretraining. To this end, we develop a robotization pipel
OpenAI outlines priorities and principles for rigorous, secure, and independent third-party AI safety assessments of frontier models and safeguards.
Existing approaches to persona simulation with Large Language Models (LLMs) mostly rely on shallow character descriptions that fail to sustain coherent character behavior across extended interactions. We introduce Deep Persona, a psychologically grounded, three-layered architecture that organizes personas into hierarchical levels of observable expression, latent beliefs, and core motivational drives, for constructing highly convincing role-playing agents. Governed by the principles of scripted d
Basically the title.We did not get a new moe model with qwen 3.8 and Alibaba did not announce any small moe models on apsara.I know we might get an announcement later but ngl I kinda lost hope
The deterministic memory layer for AI Discussion | Link
We need studies on conversational AI (CAI) at scale to understand human behavior and shape CAI design. However, fragmented reporting of systems and study configurations hinders replication, extension, and knowledge accumulation. We present Gricea, an open-science platform representing studies as configurable, deployable research artifacts that researchers can run, inspect, share, and reuse. Informed by a formative analysis of prior CAI research, Gricea couples study procedures, participant-facin
I have been successfully hosting Vaultwarden for the last year or so and have had no major issues to write home about. It serves myself and my mum, but I'm intending to expand that to other family. What I'm interested in is understanding if I might be better of just using the official Bitwarden for self hosting instead of Vaultwarden. I'm concerned that if the Vaultwarden maintainer is let go from Bitwarden or changes priorities or whatever, then I'm at risk. One of the initial reasons for choos