Google Research 提出 RRSI,用正则化递归自改进防止 agent harness 在训练任务上过拟合。
2026-09-23
— 大模型价格战开打,安全与军事 AI 的阴影同日浮现。
OpenAI 发布 GPT-6 Sol 和 Luna,价格较前代减半;Anthropic 发布 Claude Opus 5.5,性能对标 Fable 5.1 且成本降 40%。两家公司同日发布,标志前沿模型竞争进入性价比阶段。安全方面,ShinyHunters 声称入侵 FBI 并窃取全部员工数据;五角大楼承认过度依赖 AI 导致伊朗学校误炸。
头条
OpenAI 发布 GPT-6 Sol 和 Luna,价格减半多源事件 ×5
OpenAI 推出 GPT-6 Sol 和 GPT-6 Luna,作为 GPT-6 Astra 的轻量级补充,主打成本效率。Sol 面向复杂任务如编码,Luna 面向高容量文书工作;两者输入/输出价格均为 GPT-5.6 同档模型的一半。为什么重要:对于以 API 构建应用的开发者,这意味着在保持接近前沿性能的同时,推理成本大幅下降,尤其利好 agent 和批量处理场景。
评论区普遍认可降价与高性价比,但也有人认为性能提升有限、发布时机疑似针对竞品。
Anthropic 发布 Claude Opus 5.5,性能对标 Fable 5.1 且成本降 40%多源事件 ×3
Anthropic 发布 Claude Opus 5.5,宣称在多数任务上达到 Claude Fable 5.1 水平,运行成本比 Opus 5 低 40%,输出 token 定价从 25 美元/百万降至 20 美元/百万。该模型是 Anthropic 自呼吁放缓前沿竞赛以来的首个发布,经 Frontier Design 和 METR 外部评估。为什么重要:旗舰级模型降价并提升推理效率,直接降低复杂编码与知识工作负载的 API 开销,同时强化了安全测试叙事。
多数人认可 Opus 5.5 更自然、更便宜且性能强,但也有人认为版本号跳跃、定价仍高,且对基准测试持怀疑态度。
每天早晨,一份为你精选的科技日报
网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。
已发布 73 期 · 每天筛过 150+ 条只留值得读的 30 条
AI 动态
RoboDawn 接口将 VLM 智能迁移到机器人控制,探索数字到物理世界的泛化。
腾讯 ARC 的 WorldCrafter 用隐式 3D 感知记忆实现长时程一致的视频世界模型。
D-RAC 提出面向企业文档的检索感知分块方法,通过 PDF 归一化和多模态 Markdown 转换提升 RAG 质量。
GPT-6 改进 prompt caching,缓存命中率更高,缓存输入 token 折扣最高 90%。
开发与开源
Drop 是一个无 root 的 Linux 沙箱,支持 gVisor,可隔离第三方程序和 coding agent。
llm 0.36 新增 gpt-6-sol 和 gpt-6-luna 模型,并支持声明不支持对话的模型。
llm-anthropic 0.29 插件加入 Claude Opus 5.5 支持。
llm-typesafe 0.1a0 插件支持 TypeSafe AI 的 Jev 决策模型,可输出结构化 yes/no 或 choice 答案。
实验探讨 gzip 能否作为语言模型,基于压缩-预测等价性生成文本。
社区热议
阿里在 Apsara 大会正式宣布 Qwen 4,LocalLLaMA 社区关注其开源权重与本地部署潜力。
小米 MiMo-V2.6-Pro 1T-A42B 成为新的开源权重榜首,训练成本仅 300 万美元。
MIT Tech Review 提醒读者警惕今夏 AI 炒作,包括漏洞发现、数学突破和离职警告等事件。
Cisco Talos 发布开源框架 CAIRN,用于识别由 AI 聊天机器人驱动的恶意软件和黑客工具。
微软牵头捣毁 AI 辅助诈骗平台 EvilTokens,该平台已攻陷 12,000 个微软账户。
GitHub Trending
Star agent-substrate / substrate Agent Substrate: the core system
Star dream-num / univer The Office Harness for AI Agents — Spreadsheets, Docs, Slides, Canvas, Relational Tables, and PDF in one runtime.
Sponsor Star davila7 / claude-code-templates CLI tool for configuring and monitoring Claude Code
Star google / ax Google's open agentic orchestration runtime
Star mvt-project / mvt MVT (Mobile Verification Toolkit) helps with conducting forensics of mobile devices in order to find signs of a potential compromise.
Star superdesigndev / treg OpenRouter for agent tools. Join community here: https://discord.gg/6mQYYfFMAn
Star browser-use / video-use Edit videos with coding agents
更多值得一看(内容池 62 条)
I thought we were supposed to be slowing down the frontier.
We talked to Google’s Oscar winning “Giganerd” about automating science, solving climate change, and how future generations can contribute to science in the age of superintelligent AI
Researchers say another looming threat hangs over some of America's most important critical infrastructure.
We present onPanda, an interactive tool for efficiently annotating LLM alignment data and agent trajectories. onPanda adopts token-level correction as its core interaction: while reading a model response, the annotator locates the first inappropriate token and either picks a substitute from the model's candidate tokens or types the correct text via free-form editing. The system then truncates everything after that position and continues generation from the corrected prefix, repeating this locate
Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We introduce a continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience. The memory widens by acquiring
9月22日,在2026杭州云栖大会企业级Agent实践峰会上,基元律动联合创始人兼CTO韩凯发表演讲《从Harness到RSI飞轮》。
This post is written by a human and I'd appreciate it if you treated it as such. Thanks. So, I've been noticing a pretty clear interest in developing as good a coding and agentic tool-calling model as possible, especially at smaller sizes, sub-50 gigs. However, I'm finding that at least for my use of AI, if I really want to move away from big providers, I am going to require a model that has better world knowledge than the current offerings. Qwen 3.8 27B is a truly fantastic model for tons and t
Repository-level software engineering (SWE) comprises heterogeneous task categories, whose progress under pooled agentic reinforcement learning can be uneven: gains in some categories coincide with regressions in others, while aggregate resolution obscures these changes. Motivated by this category see-saw, we develop a category-aware expert-training and policy-integration framework. Executable task construction and SWE Labeler, an evidence-grounded multi-axis labeling system, organize the traini
Agentic memory is becoming essential for long-horizon AI agents, yet many existing systems rely on autoregressive LLMs to control how memories are organized, retrieved, and used, placing expensive generation on the critical path of memory operations. We introduce \method, a new agentic memory architecture inspired by System-One/System-Two cognition. System One captures fast, lightweight decision-making, whereas System Two performs slower, deliberative reasoning. Jev-Mem brings this division of l
Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the best harness varies across domains, instances, and models, a general-purpose agent must either settle for a suboptimal shared harness or route among an ever-growing set of specialized ones. We therefore study agent harness distillation: using a domain- or instance-optimized harness as training-time gui
Sparse on-policy distillation (OPD) allocates teacher supervision to a small subset of tokens in student-generated trajectories. However, useful teacher guidance can yield a noisy update when its gradient is estimated from a sampled next token. We study this estimation problem at a fixed prefix in information geometry and propose an information-efficiency ratio (IER) based on a signal-to-noise decomposition. IER characterizes relative gradient estimation error under an optimal scalar baseline. A
Skills can improve the performance of Large Language Model (LLM) agents by providing task-specific procedural guidance, while skill optimization further improves their effectiveness through iterative refinement. However, existing skill optimization methods typically represent skills as unstructured natural-language instructions, creating two key challenges: 1) Unstructured skills often lack explicit workflow-level guidance and contain substantial redundancy, making them difficult for LLMs to exe
Building general-purpose agents for industrial deployment requires integrating multiple capabilities, each typically acquired at a distinct stage of training. Yet there is currently no well-established recipe for Agent Continual Learning (ACL), with little understanding of the trade-offs among existing integration paradigms. To address this gap, we introduce ACLArena, a framework for comprehensively studying, analyzing, and evaluating ACL. We first build a sequential training pipeline and conduc
A 421M-parameter model just played Flappy Bird on my desktop CPU (OpenVINO int8) Running on my Intel Core i7 12th gen CPU Converted laya system one model to OpenVINO and quantized to int8
AntLing open sourced the Ming-Image-0.1-Design family: • Ming-Image-0.1-Design, 6B • Ming-Image-0.1-Design-Layer, 6B • Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard
After turning a string of spectacular mathematical results into a reputational crisis, OpenAI is consulting human mathematicians to help it figure out a less disastrous path forward. On Monday, the company announced a new independent panel of mathematicians tasked with advising it and other AI companies on their interactions with mathematical research and the wider […]
GPT‑6 Astra allowed Parallel’s agents to research and synthesize labor-market data in half the time and at half the cost vs. prior models.
Hugging Face CEO:「太棒了」
Manipulation policies must know which objects matter and where they are, yet the pretrained backbones that current robot foundation models build on, from language in vision-language-action models (VLAs) to video generation in world-action models (WAMs), do not directly require this metric grounding, leaving it to be learned implicitly from robot demonstrations. We propose Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding. GAM can be conditioned usin
Business intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in software such as Power BI and Tableau. In traditional BI workflows, users need to prepare data by (1) identifying relevant tables, (2) performing data transformations, and (3) building join relationships, before they can (4) answer their business questions. These steps can be complex and time-consuming, making BI challenging. Given the strong capabilities of large language models (
Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requires grounding the conversation in evolving visual context. We present Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction. Each model serves as a complete co
• Ming-Image-0.1-Design, 6B • Ming-Image-0.1-Design-Layer, 6B • Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard.
better make no mistakes I've been running Qwen 3.8 Flash Next and it's a great driver for Hermes and Pi. I told it to add CUDA_DISABLE_PERF_BOOST=1 to reduce my server's idle power draw
Assign work to AI agents, like any teammate Discussion | Link
Don't know how they did it, but for under 10GB model, the results are astonishing. I am running it on Unsloth Studio. They just released the update, so if you are not seeing the option, I recommend updating your Unsloth Studio. Cheers!
Rabbit, the company behind the underwhelming R1 device, is rolling out a standalone AI agent that you don't need its hardware to use, as reported earlier by Wired. The startup says its new OS3 "agentic operating system" runs in the cloud but operates locally across Windows, Mac, and Linux devices. According to Rabbit, you can […]
British Columbia sues OpenAI, demands Tumbler Ridge shooter’s ChatGPT logs.
Qualcomm said that its new top chip can run 30B mixture-of-expert model locally.
In the hope of uncovering new details about ancient life, researchers have developed a large language model that fills in the gaps in papyrus fragments.
Two years after trying to sidestep mobile apps with dedicated AI hardware, Rabbit is launching OS3, a cross-platform agent that lives on the screens you already use.
The English hospitals recovered patient care data only.
Autonomy-1 will have a small, transformer-based AI model taking charge of a space probe.
Visual FoxPro stopped at version 9 in 2007. A surprising amount of it is still running, in 32 bits, because rewriting a 20-year-old business app is how you lose the business. A customer wanted to keep milking their app for the foreseeable future, so here it is: the same language on a new runtime (Rust, compiled to wasm, checked against the real vfp9.exe), tables no longer stopped at 2 GB, the old 32-bit .fll add-ins still loading, and lambdas, JSON and an HTTP server bolted on for good measure.
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Don’t be fooled by this summer of AI hype —Timnit Gebru, executive director of the Distributed AI Research Institute (DAIR), and Emily M. Bender, professor of linguistics at the University of…
For the people that are unaware or haven't seen the news yet.
Human video datasets offer an abundant and diverse source of interaction data that can complement expensive real-robot data. To bridge the human-to-robot embodiment gap, existing approaches either robotize videos in task-matched settings or address observation and action alignment separately at scale. In this work, we systematically examine whether robotized human videos can serve as an effective and scalable source of supervision for VLA pretraining. To this end, we develop a robotization pipel
OpenAI outlines priorities and principles for rigorous, secure, and independent third-party AI safety assessments of frontier models and safeguards.
Existing approaches to persona simulation with Large Language Models (LLMs) mostly rely on shallow character descriptions that fail to sustain coherent character behavior across extended interactions. We introduce Deep Persona, a psychologically grounded, three-layered architecture that organizes personas into hierarchical levels of observable expression, latent beliefs, and core motivational drives, for constructing highly convincing role-playing agents. Governed by the principles of scripted d
Basically the title.We did not get a new moe model with qwen 3.8 and Alibaba did not announce any small moe models on apsara.I know we might get an announcement later but ngl I kinda lost hope
The deterministic memory layer for AI Discussion | Link
We need studies on conversational AI (CAI) at scale to understand human behavior and shape CAI design. However, fragmented reporting of systems and study configurations hinders replication, extension, and knowledge accumulation. We present Gricea, an open-science platform representing studies as configurable, deployable research artifacts that researchers can run, inspect, share, and reuse. Informed by a formative analysis of prior CAI research, Gricea couples study procedures, participant-facin
I have been successfully hosting Vaultwarden for the last year or so and have had no major issues to write home about. It serves myself and my mum, but I'm intending to expand that to other family. What I'm interested in is understanding if I might be better of just using the official Bitwarden for self hosting instead of Vaultwarden. I'm concerned that if the Vaultwarden maintainer is let go from Bitwarden or changes priorities or whatever, then I'm at risk. One of the initial reasons for choos