DawnSift
订阅日报
周五 · 科技日报 · 第 89 期

2026-10-09

— 今天的主线:AI 从“会写”走向“会干活”,但可靠性与成本仍是两道硬门槛。

今日 TL;DR

谷歌发布 Gemini Agent,正式进入办公智能体混战,支持调用 Claude 等多模型。JetBrains 开源 12B MoE 编码模型 Mellum2.1,SWE-bench Verified 从 2.0 跃升至 47.0。Anthropic 推出免费开源安全扫描服务 OSS Scanner。安全方面,韩国银行攻击事件显示单人即可组合开源渗透工具与多个 LLM 完成入侵。

You give it objectives, not instructions.(给它目标,而非逐步指令。)—— Google Cloud CEO Thomas Kurian 谈 Gemini Agent 的设计原则

头条

1

谷歌发布 Gemini Agent:通用办公智能体,支持调用 Claude

Google Cloud 发布 Gemini Agent,一个可长期运行、跨应用规划并执行任务的通用办公 Agent,拥有独立企业身份(邮箱、日历、账号),并能根据任务自动选择底层模型,甚至调用 Anthropic 的 Claude。为什么重要:办公 Agent 赛道竞争白热化,多模型路由与子 Agent 协作正成为企业级 AI 的标准架构,对后端集成与权限模型提出新要求。

2

JetBrains 开源 Mellum2.1:12B MoE 编码模型,SWE-bench 从 2.0 到 47.0

JetBrains 发布 Mellum2.1,Apache 2.0 许可的 12B 混合专家思考模型,每 token 激活 2.5B 参数,131,072-token 上下文。升级几乎全部来自真实软件环境中的强化学习(RL),SWE-bench Verified 得分从 2.0 提升至 47.0。为什么重要:小参数、可自托管的编码 Agent 模型正在逼近大模型能力,GGUF 构建从 7.0 GB 起,可直接用于 llama.cpp、Ollama 和 LM Studio。

3

韩国银行攻击事件:单人组合开源渗透工具与多个 LLM 完成入侵

CrowdStrike 报告显示,上周韩国多家大型银行遭受的网络攻击可能由单人完成,攻击者组合使用开源 AI 渗透工具 ARTEX、DeepSeek v4.1-Flash、GLM-5.3、Grok 4.6 和 Claude Code。为什么重要:LLM 正在显著降低复杂攻击的门槛,安全团队需要重新评估威胁模型,尤其是针对具备自主工具调用能力的编码 Agent 的防护。

4

Anthropic 推出免费开源安全扫描服务 OSS Scanner

Anthropic 发布 OSS Scanner,为开源项目提供由最强模型(包括 Mythos)驱动的周期性安全漏洞扫描,完全免费但报告无人工审核。为什么重要:模型生成的漏洞报告能加速开源项目发现安全问题,但缺乏人工复核意味着误报与漏报风险并存,开发者需自行验证。

5

Goodfire 推出“由内而外”的 AI Agent 监控方案,成本大幅降低

Goodfire 发布新型监控器,不再让第二个 AI 通读 Agent 的全部输出,而是直接观察模型内部工作状态,仅在异常时触发备份审查,成本显著低于传统方案。为什么重要:长时间运行的 Agent 产生的文本量巨大,传统外部审查成本高昂;内部可解释性监控为 Agent 安全治理提供了新的工程路径。

每天早晨,一份为你精选的科技日报

网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。

已发布 89 期 · 每天筛过 150+ 条只留值得读的 30 条

AI 动态

开发与开源

Whistle: Speech to Text in 16.9 MB

Whistle 发布 16.9 MB 语音识别模型,CPU 运行、无依赖,支持 7 种语言转录与词级时间戳。

普遍认可英语识别准确且体积小、CPU可跑,但也有人认为非英语及嘈杂场景效果差、不如Parakeet。

ttok 0.4Simon Willison2 min开发工具AI

Simon Willison 更新 ttok 0.4,修复 Click 警告并新增 --list-models 命令。

社区热议

“Math 2.0” will need to value mathematical progress more holistically

陶哲轩发文呼吁数学界更整体地衡量 AI 时代的数学进展,评论区普遍认同但担忧数学职业存续。

评论区普遍认同AI将深刻改变数学,需更整体地衡量进展;但也有人认为未来模型会远超人类,数学职业或难存续。

GitHub Trending

Star boykopovar / AnyPS5 Tool for automatic PS5 executables porting to Linux and Windows

Sponsor Star cathrynlavery / diagram-design Editorial diagram design for Claude Code, Codex, GitHub Copilot, Factory Droid, and Pi. 42 diagram types. Self-contained HTML + SVG. No shadows. No Mermaid slop.

morluto/rea★ 27636

Star morluto / rea Reverse engineer anything with agents, from app behavior down to native binaries.

Sponsor Star mattpocock / skills Skills for Real Engineers. Straight from my .agents directory.

Sponsor Star thedotmack / claude-mem Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More

Star EpicGames / raddebugger A native, user-mode, multi-process, graphical debugger.

Star anthropics / knowledge-work-plugins Open source repository of plugins primarily intended for knowledge workers to use in Claude Cowork

Star storytold / artcraft ArtCraft is an intentional crafting engine for artists, designers, and filmmakers

更多值得一看(内容池 68 条)
Gains and Collapse in On-Policy Distillation:A Reinforcement Learning Perspective

On-policy distillation (OPD) has become an important approach to language model post-training. However, despite its performance gains, OPD can also collapse into excessively long and repetitive generation, and the mechanism underlying these divergent outcomes remains poorly understood. We explain these outcomes through a reinforcement learning perspective: the teacher implicitly rewards student behaviors, even those it rarely exhibits itself. From this perspective, our experiments show that OPD

DLoop: Looped Speculative Decoding

Speculative decoding accelerates autoregressive generation in large language models. In each drafting stage, a lightweight draft model proposes tokens that the target model subsequently verifies. With increasingly capable draft models, we find that the target model frequently accepts all tokens produced in a drafting stage. A verification nevertheless follows each drafting stage, resulting in unnecessary target-model forward passes even when drafting could have continued. Adaptive draft length m

Humanize: Judgement Engineering for Agentic Coding

arXiv:2610.08900v1 Announce Type: new Abstract: Agentic coding makes code generation cheap, but reliable completion remains difficult: the agent that writes the code is a weak judge of whether it is done. We present Humanize, a multi-agent orchestration workflow for agentic coding built around judgement engineering: explicit, mechanically enforced decisions at the boundaries between planning, implementation, review, and learning. A human approves a plan contract, a builder agent implements it in

Perplexity AI Releases pplx-embed-v2-late: A 0.6B Edge Model and a 9B Model Scoring 92.4% on MADQA

Perplexity's pplx-embed-v2-late comes in 2 sizes: a 0.6B model built to run on edge devices, and a 9B model for building high-quality indexes. Its best score is 92.4% on MADQA, and its weakest is 61.2% on ViDoRe v3 Markdown. Both are MIT-licensed and ready to self-host. The post Perplexity AI Releases pplx-embed-v2-late: A 0.6B Edge Model and a 9B Model Scoring 92.4% on MADQA appeared first on MarkTechPost .

Agent Plasticity: Measuring Self-Improvement Through Experience

arXiv:2610.08902v1 Announce Type: new Abstract: AI agents increasingly operate in environments where they can diagnose failures and improve through experience, yet existing evaluations largely measure what an agent can do at a fixed point in time rather than how effectively it learns. Evaluating self-improvement requires answering three questions: does future performance improve and generalize beyond the interactions that enabled learning; how efficiently are new capabilities acquired; and where

Sequential Probabilistic Uncertainty Estimation for Parallel Multi-Agent Reasoning Systems

arXiv:2610.08901v1 Announce Type: new Abstract: LLM-based multi-agent systems (MAS) have attracted growing attention for improving reasoning through interaction among multiple agents. In this work, we focus on parallel multi-agent reasoning systems, where several agents solve the same problem over multiple rounds and aggregate their outputs into a final answer. Despite their strong reasoning performance, uncertainty estimation for such systems remains underexplored: the reliability of a MAS depe

From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery

Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controll

Strata rewrote their Github history to wipe evidence of Claude-authoring

Just noticed this today when I went to run the built-in "UPDATE" script and git failed because there was no common ancestor. Looked into why, and apparently every historical commit has been re-written to strip the "Co-Authored by Claude" text from the descriptions. Personally I think that's pretty gross. I'm struggling to think of any reason to do this other than an intention to be dishonest about the origins of the project.

Internalizing Agent Experience into Diffusion Model Weights via On-Policy Context Distillation

Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continually construct and revise prompts, thereby eliciting better images. These gains, however, remain external to the diffusion model and are realized only while the full harness runs. We propose Diffusion On-Policy Context Distillation (D-OPCD), which treats the agent-imp

AgentGarten: Code Worlds for Evolving Agents

Interactive virtual worlds allow agents to learn through exploration and interaction. What agents can learn is bounded by the environments they practice in, which must be faithful, with consistent state, rules, and dynamics, and realistic, with observations that follow the real-world visual distributions. Achieving both across diverse worlds remains a bottleneck. We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time intera

recently saw a bunch of open decision models pop out of nowhere in the last two weeks (laya, liquid's d1, cloudflare's clef-flash, interfaze's lev), so I wanted to see how far apart they actually are on the same GPU(yes, model size is a huge factor, but still isn't the only factor). all four had the same task of reading nine wikipedia articles about centipedes (9,534 words) word by word and flag every word that names a centipede. one /v1/systemone call per word, the next word goes out the second

[audio.cpp] Recent updates you might have missed: Higgs Audio TTS use 48% less VRAM (< 6GB), HTDemucs 2.2× faster, PocketTTS 2.2× faster on CPU, and WebUI generation history feature

Hi all, a bunch of performance improvements have been landed in audio.cpp. The biggest highlight is Higgs Audio TTS, which now runs with around 6 GB VRAM , a 48% reduction in peak memory usage compared to the previous implementation. Thanks to We also made some models significantly faster, especially HTDemucs on GPU and PocketTTS on CPU. No compromises in parity and correctness. Here's a summary of the improvements: Model Peak memory reduction Speedup Higgs Audio TTS 48% VRAM 1.01–1.09× CUDA ACE

RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments

General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces. Its 84 tasks span manipulation, mobile manipulation, locomotion, driving, and aerial control, with explicit interaction budgets and exec

PhysEvo: Astra Can Act, Let It

Astra can act, yet reliable manipulation depends on the system through which it observes and controls the world. We introduce PhysEvo, a framework for physical recursive self-improvement (RSI) around a single frozen model. A task agent executes robot tasks; a meta-agent uses the resulting trajectories to diagnose failures, revise tools and skills, and test corrections. The meta-agent can also improve its own diagnostic tools, so retained revisions support both later action and later self-improve

Nvidia’s erroneous paper accepted as ICML’s spotlight [D]

Here is the story: Nvidia has published this work (with source code available) called dreamDojo which is a world model for robotics based of their prior work Cosmos 2.5 which is cited about 100 times and got ICML’s spotlight. Authors are very well known and respected in the field with too many peer reviewed papers already published. The work doesn’t have much novelty (which I don’t care) but it is yet another foundation model. The gist is that they collected about 44k hours of human data (data a

Introducing the Anthropic Cyber Mission

Today we’re launching the Anthropic Cyber Mission, a long-term commitment to securing the systems everyone depends on. The Cyber Mission is a new effort to support defenders with tools, research, and resources to secure their software and systems. We’re starting with two areas: Critical infrastructure: Starting with securing the operational technology behind power grids, water systems, and transportation networks, and protecting government systems. Today, we’re introducing the Critical Infrastru

An Empirical Study of Agent Skills' Downstream Utility

arXiv:2610.08875v1 Announce Type: new Abstract: Agent Skills package procedural guidance and resources for reuse, but a relevant Skill does not necessarily improve task performance. Existing studies characterize Skill content and evaluate downstream performance, yet provide limited explanations of how utility depends on content, execution configuration, and multi-Skill organization. We conduct an empirical study on 87 SkillsBench tasks, defining downstream utility as the pass-rate difference fro

Inverting Multi-Vector Visual Document Indices

Prevailing multi-vector visual document retrievers store each page as about a thousand patch vectors, often in vector databases run by a third party. Since no one can read a page from its vectors, this index is easily treated as less sensitive than the page. However, because the index keeps one vector per patch in raster order, and each vector is computed by a vision-language model pre-trained to read documents, we hypothesize that whoever runs or breaches the store can reproduce a page from its

How Fragile Is On-Device Language Model Safety? Localizing Safety-Critical Parameters for Sparse Fault Analysis

arXiv:2610.09000v1 Announce Type: new Abstract: As small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern. We investigate whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated within a sparse subset of parameters, creating a reduced fault surface for targeted analysis. We study two complementary localizati

AdaGuard: Enhancing Safety and Policy Compliance with Reasoning-Enabled LLM-As-A-Judge Guardrails

arXiv:2610.08923v1 Announce Type: new Abstract: Enterprise generative AI applications require robust safety mechanisms that can accommodate diverse risk postures, evolving policies, and varying latency constraints. Current guardrail solutions often suffer from rigidity, relying on fixed policy sets and offering limited transparency or reasoning flexibility. We present Adaguard, an adaptive LLM-as-a-Judge framework designed to address these challenges through dynamic policy enforcement and adapti

On-Policy Distillation with Negative-Policy Rollouts

On-policy distillation (OPD) has been widely studied as a post-training method in which a student model obtains token-level supervision from a stronger teacher on its own rollouts. Recent studies have improved OPD through alternative distillation reward formulations and teacher configurations, while the objective of distillation remains centered on mimicking the teacher. However, when a stronger teacher has limited distributional overlap with the student, such positive guidance can provide insuf

MIMESIS: Learning User Simulators as Training Environments for Interactive Agents

Training and evaluating interactive language agents typically requires rich user interactions, yet collecting human feedback is expensive and difficult to scale. Simulated users offer a scalable alternative, but they must both resemble real user behavior and provide useful learning experiences for agents. In contrast, most agent-training frameworks rely on off-the-shelf assistant LLMs, whose helpfulness can make them overly cooperative, explicit, and behaviorally homogeneous compared with real u

Verify Less, Evolve More: Training Idea-Level Critics for Verification-Efficient ML Evolving Agents

arXiv:2610.08993v1 Announce Type: new Abstract: As large language models become more powerful, self-evolving agents are able to tackle challenging tasks including AI for machine learning (AI4ML). In AI4ML, while empirical verification is available, it often requires computationally costly model training and evaluation, limiting the speed and scale of agent evolution. Yet verification efficiency remains under-explored, and frontier models provide only limited gains when used directly as idea sele

CADFather: Autonomous CAD Reconstruction through Coordinated Tool Use

Reconstructing an editable CAD model from a 3D shape remains a challenging engineering task. Existing methods can propose CAD operations, but no single source of proposals works equally well across different part geometries and stages of reconstruction. We introduce CADFather, an autonomous agentic system that coordinates complementary tools to recover parametric CAD programs from 3D meshes. A vision-language assistant inspects renders of the target and intermediate reconstructions, then decides

Improving Proactive AI Assistance with Hierarchical Procedural Understanding

Proactive AI assistants continuously observe a user's activity and decide whether to provide new guidance or remain silent. They should provide appropriate guidance for the task, determine when to provide the next guidance based on task progress, and adjust the guidance level to the user's expertise and needs. Supporting these capabilities requires training and evaluation data that reflect procedural structure and capture how guidance should adapt to task progress and user needs. However, existi

Thank you :) Swift Models hit 2.2 million+ downloads / Early Access to New Models, Free Compute for Researchers

Hey everyone, Jovan from UkisAI (Swift Qwen) here! For those who don't know us, UkisAI is a small lab making tiny frontier LLMs, tools and datasets (+doing it open-source!). I'm one of the guys running it aka I train the models and post on Reddit. Our first open-source release is Swift, a series of reasoning-efficient LLMs. It is proof of how penalizing pathological overthinking patterns inside of various LLMs can bring their token usage down -58.3% and speed x1.95 without losing accuracy if RL-

Can any open-weight models handle a decomp/recomp project yet?

Opus5.5, Sol6.1, Fable, and Astra have all proven they can and the scene has exploded this past week. Part of that is from the tools and feedback loops maturing though. Are any open weight models (at all, so including K3, GLM5.3, Qwen3.8-Max, and Mimo-2.6) able to do this? Can the larger models this sub regularly runs (GLM 5.3-Flash, Qwen 3.8-Next-Flash, V4.1-deepseek Flash..) handle a simpler one (GBA and PSP having smaller roms and mature pipelines)?

UniWAM: Unified World-Action Model

Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semantic understanding and reasoning under distribution shifts. We introduce UniWAM, a unified architecture that integrates a physical reasoner, a world generator, and an action predicto

Anthropic changes usage policy to ban model abuse and election interference

Anthropic's updated usage policy explicitly prohibits users from repeatedly abusing Claude in extreme cases, though ordinary frustration and criticism are still allowed. The new rules also address election interference, deceptive campaigns, weapons software, and surveillance.

Tetris3D: 3D Scene Generation With Objects That Fit Together

We propose Tetris3D, a generative framework for single-image 3D scene reconstruction that recovers objects which are physically and geometrically coherent as a scene. Existing methods often generate objects independently or couple them implicitly, providing limited guidance for ensuring fine-grained spatial compatibility between neighboring objects that interact with one another. To address this, we explicitly condition the generation of each object on the geometry of surrounding objects and the

Computer programming is, fundamentally, about two things: Problem-solving using computers Learning to control complexity while solving these problems I have a hard time imagining a future where knowing how to solve problems with computers and how to control the complexity of those solutions is less valuable than it is today, so I think it will continue to be a viable career even with the advent of AI tools. — Carson Gross Tags: computer-science , carson-gross , careers , ai

VIEScore2: Unified Image Evaluation with Spatially Grounded Explanations

Existing synthetic image evaluators typically provide only a scalar quality score and do not identify the image regions that support it. We introduce VIEScore2, a unified evaluator for image generation and editing tasks with optional conditioning images. VIEScore2 represents an image as an N x N grid and jointly predicts quality scores and defect locations in a single model pass. Its text-native grid representation provides a common interface for heterogeneous spatial supervision and enables dir

Accelerating Floating-Point Satisfiability Solving via Gradient Normalization

arXiv:2610.08808v1 Announce Type: new Abstract: Satisfiability Modulo Theories (SMT) solvers are foundational to software verification, program analysis, and compiler testing, particularly over the theory of Quantifier-Free Floating-Point (QF_FP). While recent optimization-based SMT solvers have successfully applied gradient descent to continuous relaxations of logical formulas, they are fundamentally bottlenecked by gradient domination, a phenomenon where a small subset of difficult clauses hij

2026 Usage Policy update

Each year, Anthropic updates its Usage Policy in response to the evolving capabilities of our models, and the feedback we’ve received from our customers. We're publishing a new version of the policy today. In this post, we summarize the changes we’ve made. Most of the updates in the latest version are intended to clarify existing rules. In the year since our last refresh, Claude has taken on longer, more independent work. This update provides new examples that show how our rules apply to Claude’

Q-Learning with Scalar Adjoint Matching

Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian p

Learning to Report Unsafe Tasks in a Multi-Agent Game

arXiv:2610.09002v1 Announce Type: new Abstract: When agents share a reward for completed tasks, reporting unsafe work can reduce the reporter's reward by stopping a task. Audits can make reporting optimal without ensuring that further training teaches a silent team to report. We study this learning problem in a game where any witness can stop a task by reporting. With $k$ witnesses per task sharing a policy and drawing independently, the expected-reward derivative with respect to their shared si

Can AI Agents Make Open-Ended Scientific Discovery? Evidence from Station

arXiv:2610.08927v1 Announce Type: new Abstract: Recent AI systems have made rapid progress in scientific discovery when given well-defined metrics, but whether they can autonomously undertake open-ended scientific discovery remains unclear. We investigate AI's ability to tackle open-ended tasks in Station, an open-world environment in which multiple agents simulate a scientific ecosystem. To tackle challenges specific to open-ended tasks, we propose augmenting Station with two mechanisms: a Supe

A self-learning scientific agent for X-ray diffraction

A central challenge for scientific agents is to turn analytical experience into reusable expertise grounded in physical evidence. Here we introduce Gan Jiang, a self-learning agent for powder X-ray diffraction built on a diffraction-analysis ecosystem we developed: XMatcher, XQueryer, XDecomposer and WPEM. Together, these engines span phase identification, multiphase decomposition and physics-constrained whole-pattern modelling. Gan Jiang converts analytical experience into executable skills by

jevman: AI decision models play Pac-Man

The other week I posted about Jev vs. Kev compared and since then, OpenAI released the decisions endpoint, Cloudflare released Clef and many here asked about Laya as well. This time we compared six popular decision models by making them play Pac-Man: kev 1.13, Kev 4B, Clef, Clef Flash, GPT-6 Luna and Laya. Since they respond within ms it works for them to play the game in real time. We published a leaderboard and the repo is open-source, so anyone can run their own decision model, like your own

Anti-Patterns in Software Blogging

Anti-Patterns in Software Blogging Some excellent writing advice from Michael Lynch. Michael warns against "meandering intros", misjudging your reader's existing knowledge, assuming they'll read your previous posts, and excessive formality. He also warns against overreliance on links as an excuse not to explain terminology. This one hurt! I do this all the time, but I have a nagging suspicion that almost nobody ever clicks on them. (In a Lobste.rs comment Michael clarifies that "My rule of thumb

每天早晨,一份为你精选的科技日报