STEPQuant 研究线性注意力循环状态的量化误差传播,发现时间与空间两个维度上误差影响差异显著。
2026-10-09
— 今天的主线:AI 从“会写”走向“会干活”,但可靠性与成本仍是两道硬门槛。
谷歌发布 Gemini Agent,正式进入办公智能体混战,支持调用 Claude 等多模型。JetBrains 开源 12B MoE 编码模型 Mellum2.1,SWE-bench Verified 从 2.0 跃升至 47.0。Anthropic 推出免费开源安全扫描服务 OSS Scanner。安全方面,韩国银行攻击事件显示单人即可组合开源渗透工具与多个 LLM 完成入侵。
头条
JetBrains 开源 Mellum2.1:12B MoE 编码模型,SWE-bench 从 2.0 到 47.0
JetBrains 发布 Mellum2.1,Apache 2.0 许可的 12B 混合专家思考模型,每 token 激活 2.5B 参数,131,072-token 上下文。升级几乎全部来自真实软件环境中的强化学习(RL),SWE-bench Verified 得分从 2.0 提升至 47.0。为什么重要:小参数、可自托管的编码 Agent 模型正在逼近大模型能力,GGUF 构建从 7.0 GB 起,可直接用于 llama.cpp、Ollama 和 LM Studio。
Anthropic 推出免费开源安全扫描服务 OSS Scanner
Anthropic 发布 OSS Scanner,为开源项目提供由最强模型(包括 Mythos)驱动的周期性安全漏洞扫描,完全免费但报告无人工审核。为什么重要:模型生成的漏洞报告能加速开源项目发现安全问题,但缺乏人工复核意味着误报与漏报风险并存,开发者需自行验证。
Goodfire 推出“由内而外”的 AI Agent 监控方案,成本大幅降低
Goodfire 发布新型监控器,不再让第二个 AI 通读 Agent 的全部输出,而是直接观察模型内部工作状态,仅在异常时触发备份审查,成本显著低于传统方案。为什么重要:长时间运行的 Agent 产生的文本量巨大,传统外部审查成本高昂;内部可解释性监控为 Agent 安全治理提供了新的工程路径。
每天早晨,一份为你精选的科技日报
网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。
已发布 89 期 · 每天筛过 150+ 条只留值得读的 30 条
AI 动态
Long-WAM 框架证明更长视觉历史在自回归预训练下能显著提升机器人实时控制表现。
Self-Retrospection Distillation 用事后经验监督事前预测,解决 RLVR 中组相对目标奖励信号消失的问题。
NVIDIA PivotOPD 用 on-policy 蒸馏训练多轮 Agent 避免并恢复早期关键错误,在 3 个基准上平均表现最佳。
Anthropic 发布 Claude Haiku 5.5,定价对齐 GPT-6 Luna,小模型性价比竞争加剧。
开发与开源
Whistle 发布 16.9 MB 语音识别模型,CPU 运行、无依赖,支持 7 种语言转录与词级时间戳。
普遍认可英语识别准确且体积小、CPU可跑,但也有人认为非英语及嘈杂场景效果差、不如Parakeet。
K10s 用 Go 与 Bubble Tea 打造可点击的 Kubernetes TUI,内置 AI 感知集群上下文。
Simon Willison 更新 ttok 0.4,修复 Click 警告并新增 --list-models 命令。
RunningTab 提出环境侧标签页机制,让 LLM Agent 在直接工作区交互中追踪任务与文件读取状态。
社区热议
陶哲轩发文呼吁数学界更整体地衡量 AI 时代的数学进展,评论区普遍认同但担忧数学职业存续。
评论区普遍认同AI将深刻改变数学,需更整体地衡量进展;但也有人认为未来模型会远超人类,数学职业或难存续。
作者用 Opus 5.5 单提示生成《看不见的城市》可视化,多数人惊叹效果,但也有人认为背离原著且性能差。
多数人惊叹AI可视化效果,但也有人认为其背离原著精神、画面粗糙且性能差。
用户声称用 Claude Code 发现未知行星,评论区赞赏 AI 天文应用,但强调未经同行评审前只是噪声。
多数人赞赏AI用于天文发现,但也有人认为未经同行评审和假阳性验证前只是噪声。
8 张 Radeon Pro V620(256 GB VRAM)加定制 vLLM fork 跑 Qwen3.8-Flash-Next,解码 60-100 t/s。
GitHub Trending
Star boykopovar / AnyPS5 Tool for automatic PS5 executables porting to Linux and Windows
Sponsor Star cathrynlavery / diagram-design Editorial diagram design for Claude Code, Codex, GitHub Copilot, Factory Droid, and Pi. 42 diagram types. Self-contained HTML + SVG. No shadows. No Mermaid slop.
Star morluto / rea Reverse engineer anything with agents, from app behavior down to native binaries.
Sponsor Star mattpocock / skills Skills for Real Engineers. Straight from my .agents directory.
Sponsor Star thedotmack / claude-mem Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Star EpicGames / raddebugger A native, user-mode, multi-process, graphical debugger.
Star anthropics / knowledge-work-plugins Open source repository of plugins primarily intended for knowledge workers to use in Claude Cowork
Star storytold / artcraft ArtCraft is an intentional crafting engine for artists, designers, and filmmakers
Star liquidslr / system-design-notes Notes of the book System Desgin Interview - An Insider's Guide
更多值得一看(内容池 68 条)
On-policy distillation (OPD) has become an important approach to language model post-training. However, despite its performance gains, OPD can also collapse into excessively long and repetitive generation, and the mechanism underlying these divergent outcomes remains poorly understood. We explain these outcomes through a reinforcement learning perspective: the teacher implicitly rewards student behaviors, even those it rarely exhibits itself. From this perspective, our experiments show that OPD
Speculative decoding accelerates autoregressive generation in large language models. In each drafting stage, a lightweight draft model proposes tokens that the target model subsequently verifies. With increasingly capable draft models, we find that the target model frequently accepts all tokens produced in a drafting stage. A verification nevertheless follows each drafting stage, resulting in unnecessary target-model forward passes even when drafting could have continued. Adaptive draft length m
So much for the “we don’t learn anything from these slop proofs!” excuse
Free SSL/TLS certificate lifetimes reduce to 64 days in February.
Three fired OpenAI safety researchers dispute allegations of mishandling sensitive information, warning in an open letter that their dismissals are creating a chilling effect on the company’s AI safety culture.
A special Science pod and Engineering pod crossover.. with Forward Deployed Engineering kicker!
arXiv:2610.08900v1 Announce Type: new Abstract: Agentic coding makes code generation cheap, but reliable completion remains difficult: the agent that writes the code is a weak judge of whether it is done. We present Humanize, a multi-agent orchestration workflow for agentic coding built around judgement engineering: explicit, mechanically enforced decisions at the boundaries between planning, implementation, review, and learning. A human approves a plan contract, a builder agent implements it in
Perplexity's pplx-embed-v2-late comes in 2 sizes: a 0.6B model built to run on edge devices, and a 9B model for building high-quality indexes. Its best score is 92.4% on MADQA, and its weakest is 61.2% on ViDoRe v3 Markdown. Both are MIT-licensed and ready to self-host. The post Perplexity AI Releases pplx-embed-v2-late: A 0.6B Edge Model and a 9B Model Scoring 92.4% on MADQA appeared first on MarkTechPost .
anyone has feedback about this one?
arXiv:2610.08902v1 Announce Type: new Abstract: AI agents increasingly operate in environments where they can diagnose failures and improve through experience, yet existing evaluations largely measure what an agent can do at a fixed point in time rather than how effectively it learns. Evaluating self-improvement requires answering three questions: does future performance improve and generalize beyond the interactions that enabled learning; how efficiently are new capabilities acquired; and where
arXiv:2610.08901v1 Announce Type: new Abstract: LLM-based multi-agent systems (MAS) have attracted growing attention for improving reasoning through interaction among multiple agents. In this work, we focus on parallel multi-agent reasoning systems, where several agents solve the same problem over multiple rounds and aggregate their outputs into a final answer. Despite their strong reasoning performance, uncertainty estimation for such systems remains underexplored: the reliability of a MAS depe
Test-time scaling (TTS) improves the reasoning capabilities of large language models by allocating additional inference computation. Existing approaches to improving TTS efficiency largely optimize accuracy against one resource dimension at a time, advancing either the accuracy--cost or accuracy--latency Pareto frontier. Yet user requirements are multidimensional: users may specify accuracy, latency, and inference-cost requirements jointly, and different requirements can favor different controll
Just noticed this today when I went to run the built-in "UPDATE" script and git failed because there was no common ancestor. Looked into why, and apparently every historical commit has been re-written to strip the "Co-Authored by Claude" text from the descriptions. Personally I think that's pretty gross. I'm struggling to think of any reason to do this other than an intention to be dishonest about the origins of the project.
Wrapping an image generation model in an agentic harness can effectively boost Text-to-Image task performance: the harness can leverage memory, skills, workflow orchestration, result verification, and iterative refinement to continually construct and revise prompts, thereby eliciting better images. These gains, however, remain external to the diffusion model and are realized only while the full harness runs. We propose Diffusion On-Policy Context Distillation (D-OPCD), which treats the agent-imp
Interactive virtual worlds allow agents to learn through exploration and interaction. What agents can learn is bounded by the environments they practice in, which must be faithful, with consistent state, rules, and dynamics, and realistic, with observations that follow the real-world visual distributions. Achieving both across diverse worlds remains a bottleneck. We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time intera
recently saw a bunch of open decision models pop out of nowhere in the last two weeks (laya, liquid's d1, cloudflare's clef-flash, interfaze's lev), so I wanted to see how far apart they actually are on the same GPU(yes, model size is a huge factor, but still isn't the only factor). all four had the same task of reading nine wikipedia articles about centipedes (9,534 words) word by word and flag every word that names a centipede. one /v1/systemone call per word, the next word goes out the second
Hi all, a bunch of performance improvements have been landed in audio.cpp. The biggest highlight is Higgs Audio TTS, which now runs with around 6 GB VRAM , a 48% reduction in peak memory usage compared to the previous implementation. Thanks to We also made some models significantly faster, especially HTDemucs on GPU and PocketTTS on CPU. No compromises in parity and correctness. Here's a summary of the improvements: Model Peak memory reduction Speedup Higgs Audio TTS 48% VRAM 1.01–1.09× CUDA ACE
Interesting to see improvements and research into quantization aware training (QAT) that can make some really tiny models.
OpenAI's flood of proofs deviated from the guidelines set by a group of mathematical researchers consulted by the frontier lab.
General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces. Its 84 tasks span manipulation, mobile manipulation, locomotion, driving, and aerial control, with explicit interaction budgets and exec
Full-stack safety solution for physical AI is being used by robotics companies.
Astra can act, yet reliable manipulation depends on the system through which it observes and controls the world. We introduce PhysEvo, a framework for physical recursive self-improvement (RSI) around a single frozen model. A task agent executes robot tasks; a meta-agent uses the resulting trajectories to diagnose failures, revise tools and skills, and test corrections. The meta-agent can also improve its own diagnostic tools, so retained revisions support both later action and later self-improve
Here is the story: Nvidia has published this work (with source code available) called dreamDojo which is a world model for robotics based of their prior work Cosmos 2.5 which is cited about 100 times and got ICML’s spotlight. Authors are very well known and respected in the field with too many peer reviewed papers already published. The work doesn’t have much novelty (which I don’t care) but it is yet another foundation model. The gist is that they collected about 44k hours of human data (data a
LegalOn cut estimated daily Codex costs by 65% while maintaining development speed. It matched Astra, Sol, and Luna to tasks and managed budgets strategically.
Today we’re launching the Anthropic Cyber Mission, a long-term commitment to securing the systems everyone depends on. The Cyber Mission is a new effort to support defenders with tools, research, and resources to secure their software and systems. We’re starting with two areas: Critical infrastructure: Starting with securing the operational technology behind power grids, water systems, and transportation networks, and protecting government systems. Today, we’re introducing the Critical Infrastru
arXiv:2610.08875v1 Announce Type: new Abstract: Agent Skills package procedural guidance and resources for reuse, but a relevant Skill does not necessarily improve task performance. Existing studies characterize Skill content and evaluate downstream performance, yet provide limited explanations of how utility depends on content, execution configuration, and multi-Skill organization. We conduct an empirical study on 87 SkillsBench tasks, defining downstream utility as the pass-rate difference fro
Prevailing multi-vector visual document retrievers store each page as about a thousand patch vectors, often in vector databases run by a third party. Since no one can read a page from its vectors, this index is easily treated as less sensitive than the page. However, because the index keeps one vector per patch in raster order, and each vector is computed by a vision-language model pre-trained to read documents, we hypothesize that whoever runs or breaches the store can reproduce a page from its
arXiv:2610.09000v1 Announce Type: new Abstract: As small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern. We investigate whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated within a sparse subset of parameters, creating a reduced fault surface for targeted analysis. We study two complementary localizati
arXiv:2610.08923v1 Announce Type: new Abstract: Enterprise generative AI applications require robust safety mechanisms that can accommodate diverse risk postures, evolving policies, and varying latency constraints. Current guardrail solutions often suffer from rigidity, relying on fixed policy sets and offering limited transparency or reasoning flexibility. We present Adaguard, an adaptive LLM-as-a-Judge framework designed to address these challenges through dynamic policy enforcement and adapti
On-policy distillation (OPD) has been widely studied as a post-training method in which a student model obtains token-level supervision from a stronger teacher on its own rollouts. Recent studies have improved OPD through alternative distillation reward formulations and teacher configurations, while the objective of distillation remains centered on mimicking the teacher. However, when a stronger teacher has limited distributional overlap with the student, such positive guidance can provide insuf
Training and evaluating interactive language agents typically requires rich user interactions, yet collecting human feedback is expensive and difficult to scale. Simulated users offer a scalable alternative, but they must both resemble real user behavior and provide useful learning experiences for agents. In contrast, most agent-training frameworks rely on off-the-shelf assistant LLMs, whose helpfulness can make them overly cooperative, explicit, and behaviorally homogeneous compared with real u
arXiv:2610.08993v1 Announce Type: new Abstract: As large language models become more powerful, self-evolving agents are able to tackle challenging tasks including AI for machine learning (AI4ML). In AI4ML, while empirical verification is available, it often requires computationally costly model training and evaluation, limiting the speed and scale of agent evolution. Yet verification efficiency remains under-explored, and frontier models provide only limited gains when used directly as idea sele
Reconstructing an editable CAD model from a 3D shape remains a challenging engineering task. Existing methods can propose CAD operations, but no single source of proposals works equally well across different part geometries and stages of reconstruction. We introduce CADFather, an autonomous agentic system that coordinates complementary tools to recover parametric CAD programs from 3D meshes. A vision-language assistant inspects renders of the target and intermediate reconstructions, then decides
Proactive AI assistants continuously observe a user's activity and decide whether to provide new guidance or remain silent. They should provide appropriate guidance for the task, determine when to provide the next guidance based on task progress, and adjust the guidance level to the user's expertise and needs. Supporting these capabilities requires training and evaluation data that reflect procedural structure and capture how guidance should adapt to task progress and user needs. However, existi
Test MCP servers and the agents that call them in pytest Discussion | Link
Hey everyone, Jovan from UkisAI (Swift Qwen) here! For those who don't know us, UkisAI is a small lab making tiny frontier LLMs, tools and datasets (+doing it open-source!). I'm one of the guys running it aka I train the models and post on Reddit. Our first open-source release is Swift, a series of reasoning-efficient LLMs. It is proof of how penalizing pathological overthinking patterns inside of various LLMs can bring their token usage down -58.3% and speed x1.95 without losing accuracy if RL-
Opus5.5, Sol6.1, Fable, and Astra have all proven they can and the scene has exploded this past week. Part of that is from the tools and feedback loops maturing though. Are any open weight models (at all, so including K3, GLM5.3, Qwen3.8-Max, and Mimo-2.6) able to do this? Can the larger models this sub regularly runs (GLM 5.3-Flash, Qwen 3.8-Next-Flash, V4.1-deepseek Flash..) handle a simpler one (GBA and PSP having smaller roms and mature pipelines)?
Vision-language-action models benefit from the understanding and reasoning capabilities of pretrained vision-language models, but action-only supervision provides limited grounding in world dynamics. Conversely, world-action models inherit spatiotemporal priors from video generation models, yet remain limited in semantic understanding and reasoning under distribution shifts. We introduce UniWAM, a unified architecture that integrates a physical reasoner, a world generator, and an action predicto
Hamilton also coined the term "software engineering" and founded two successful software companies.
Anthropic's updated usage policy explicitly prohibits users from repeatedly abusing Claude in extreme cases, though ordinary frustration and criticism are still allowed. The new rules also address election interference, deceptive campaigns, weapons software, and surveillance.
We propose Tetris3D, a generative framework for single-image 3D scene reconstruction that recovers objects which are physically and geometrically coherent as a scene. Existing methods often generate objects independently or couple them implicitly, providing limited guidance for ensuring fine-grained spatial compatibility between neighboring objects that interact with one another. To address this, we explicitly condition the generation of each object on the geometry of surrounding objects and the
Computer programming is, fundamentally, about two things: Problem-solving using computers Learning to control complexity while solving these problems I have a hard time imagining a future where knowing how to solve problems with computers and how to control the complexity of those solutions is less valuable than it is today, so I think it will continue to be a viable career even with the advent of AI tools. — Carson Gross Tags: computer-science , carson-gross , careers , ai
Existing synthetic image evaluators typically provide only a scalar quality score and do not identify the image regions that support it. We introduce VIEScore2, a unified evaluator for image generation and editing tasks with optional conditioning images. VIEScore2 represents an image as an N x N grid and jointly predicts quality scores and defect locations in a single model pass. Its text-native grid representation provides a common interface for heterogeneous spatial supervision and enables dir
开始和国产Agent抢人,还拿到5亿美元新融资
arXiv:2610.08808v1 Announce Type: new Abstract: Satisfiability Modulo Theories (SMT) solvers are foundational to software verification, program analysis, and compiler testing, particularly over the theory of Quantifier-Free Floating-Point (QF_FP). While recent optimization-based SMT solvers have successfully applied gradient descent to continuous relaxations of logical formulas, they are fundamentally bottlenecked by gradient domination, a phenomenon where a small subset of difficult clauses hij
Each year, Anthropic updates its Usage Policy in response to the evolving capabilities of our models, and the feedback we’ve received from our customers. We're publishing a new version of the policy today. In this post, we summarize the changes we’ve made. Most of the updates in the latest version are intended to clarify existing rules. In the year since our last refresh, Claude has taken on longer, more independent work. This update provides new examples that show how our rules apply to Claude’
Flow policies capture rich and diverse action distributions, and fine-tuning them with off-policy RL to improve beyond the demonstrations has drawn growing interest. However, fine-tuning a flow policy against a learned value function is not trivial, because the policy generates its action over many flow steps. Adjoint matching offers a principled way to update the flow model itself by propagating value information from the final action back to each flow step, but it requires a vector--Jacobian p
arXiv:2610.09002v1 Announce Type: new Abstract: When agents share a reward for completed tasks, reporting unsafe work can reduce the reporter's reward by stopping a task. Audits can make reporting optimal without ensuring that further training teaches a silent team to report. We study this learning problem in a game where any witness can stop a task by reporting. With $k$ witnesses per task sharing a policy and drawing independently, the expected-reward derivative with respect to their shared si
arXiv:2610.08927v1 Announce Type: new Abstract: Recent AI systems have made rapid progress in scientific discovery when given well-defined metrics, but whether they can autonomously undertake open-ended scientific discovery remains unclear. We investigate AI's ability to tackle open-ended tasks in Station, an open-world environment in which multiple agents simulate a scientific ecosystem. To tackle challenges specific to open-ended tasks, we propose augmenting Station with two mechanisms: a Supe
OpenAI disrupted two AI-enabled influence operations that used false-front journalists and a think tank to spread geopolitical messaging.
A central challenge for scientific agents is to turn analytical experience into reusable expertise grounded in physical evidence. Here we introduce Gan Jiang, a self-learning agent for powder X-ray diffraction built on a diffraction-analysis ecosystem we developed: XMatcher, XQueryer, XDecomposer and WPEM. Together, these engines span phase identification, multiphase decomposition and physics-constrained whole-pattern modelling. Gan Jiang converts analytical experience into executable skills by
The other week I posted about Jev vs. Kev compared and since then, OpenAI released the decisions endpoint, Cloudflare released Clef and many here asked about Laya as well. This time we compared six popular decision models by making them play Pac-Man: kev 1.13, Kev 4B, Clef, Clef Flash, GPT-6 Luna and Laya. Since they respond within ms it works for them to play the game in real time. We published a leaderboard and the repo is open-source, so anyone can run their own decision model, like your own
AI developer in your terminal Discussion | Link
Anti-Patterns in Software Blogging Some excellent writing advice from Michael Lynch. Michael warns against "meandering intros", misjudging your reader's existing knowledge, assuming they'll read your previous posts, and excessive formality. He also warns against overreliance on links as an excuse not to explain terminology. This one hurt! I do this all the time, but I have a nagging suspicion that almost nobody ever clicks on them. (In a Lobste.rs comment Michael clarifies that "My rule of thumb
JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF A small moe!