DawnSift
订阅日报
周四 · 科技日报 · 第 88 期

2026-10-08

— 模型价格战与 agent 可靠性,今天同时按下加速键。

今日 TL;DR

Anthropic 发布 Claude Haiku 5.5,以 $0.10/M 输入 token 的价格切入高吞吐场景;OpenAI 向所有用户开放 GPT-6 与 Intelligent UI,交互式可视化成为新默认。Liquid AI 开源 d1 决策模型系列,主打零输出 token 的实时决策。同时,多项研究聚焦 agent 可靠性:从工具调用证据链到跨 tokenizer 蒸馏,以及万亿参数 RL 的权重同步优化。

Maybe AGI is here after all.

头条

1

Anthropic 发布 Claude Haiku 5.5:1M 上下文,输入 $0.10/M token多源事件 ×3

Anthropic 推出 Claude Haiku 5.5,定位为迄今最便宜、最快的小模型,支持 1M token 上下文与最高 128K 输出,输入定价 $0.10/M token,较 Haiku 4.5 在 100K token 以内提示词场景下降约 90%。同时 Sonnet 5.5 的 cache reads 价格减半。 为什么重要:对于高吞吐、成本敏感的工程任务(摘要、压缩、分类、子代理调用),这个价位直接改变了小模型选型的经济账;1M 上下文与可调 effort 参数也让它更适合作为编码工作流中的 subagent。

多数人认可性能提升与 Max 订阅 API 额度,但也有人担忧 10 万 token 后涨价及网络安全限制。

2

OpenAI 向所有用户开放 GPT-6 与 Intelligent UI

OpenAI 将 GPT-6 推广至每周 12 亿 ChatGPT 用户,并引入 Intelligent UI 能力:模型可组合文本、图表、表单、可点击按钮等交互式元素来回答问题。 为什么重要:这标志着 LLM 输出从纯文本向可交互界面的范式迁移,对前端生成、数据可视化、即时工具构建等场景有直接启发;但过度可视化也可能影响文本可导出性与可审计性。

评论区普遍认可交互式 UI 是自然演进,但也有人认为其可能过度可视化、降低文本可导出性,并担忧准确性与依赖风险。

3

Liquid AI 开源 d1 决策模型:零输出 token 的实时决策多源事件 ×3

Liquid AI 发布 Open d1 系列:d1-3B 支持文本与图像,d1-omni-600M 支持文本+图像或文本+音频,均不生成文本,而是在单次前向传播中返回校准后的类型化答案,零输出 token。d1-3B 在 Decision Index 0.2.1 上得分 48.57,超过所有 4B 和 9B 模型。 为什么重要:决策模型绕开 token 生成,直接输出结构化结果,在边缘设备(Jetson AGX Thor 上 16ms)和实时控制场景中可能比传统 LLM 更高效;llama.cpp 首日支持也降低了本地部署门槛。

4

研究聚焦 agent 可靠性:从证据链到工具调用失败模式多源事件 ×3

多篇论文同时关注 agent 的可靠性问题:CheckerBench 用 300 个来自 297 个 CVE 的任务评估长程 agent 能否合成静态分析 checker;From Evidence to Action 研究工具调用 agent 在证据到行动链上的断裂点;UNREAL 提出用单一模型统一检索与长上下文证据选择。 为什么重要:agent 从 demo 走向生产的关键瓶颈正是可靠性——正确结果不等于有证据支撑的行动,这些工作为评估和修复 agent 失败模式提供了可复现的基准与方法。

5

Chrome 155 正式支持 JPEG XL 解码

Chrome 从 155 版本开始支持 JPEG XL (.jxl) 图像格式解码,提供比 JPEG 高 30-50% 的压缩率、无损压缩、内置 HDR 支持及无损 JPEG 转码。实现采用 Rust 以保证内存安全。 为什么重要:对于 Web 开发者,JPEG XL 在高保真与无损场景下是 AVIF 之外的重要补充,尤其适合摄影图像与渐进式解码需求;Rust 实现也再次印证浏览器对内存安全语言的投入。

评论区普遍欢迎 Chrome 重新支持 JPEG XL,认为这是重要进展,但也有人认为 AVIF 在低码率下更优,且生态支持仍不完善。

每天早晨,一份为你精选的科技日报

网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。

已发布 88 期 · 每天筛过 150+ 条只留值得读的 30 条

AI 动态

开发与开源

Docker 发布 docker-agent CLI 插件,用 YAML 声明式配置构建多 agent 协作,无需写代码。

评论区普遍质疑Docker Agent定位模糊、与Docker关联不明,认为其像跟风的agent框架,但也有人认为它对安全可复现的容器化开发流程有潜在价值。

社区热议

Meta and Microsoft take steps to reduce employee usage of Claude AI

Meta 与微软减少内部使用 Claude AI,转向自研编码工具;评论区认为主因是成本控制与内部模型自用,而非 Claude 质量下降。

评论区普遍认为此举是出于成本控制和推动内部模型自用,而非质疑AI价值;但也有人认为这更多是数据治理和接口调整,不代表Claude质量差。

GitHub Trending

morluto/rea★ 15576

Star morluto / rea Reverse engineer anything with agents, from app behavior down to native binaries.

Sponsor Star mattpocock / skills Skills for Real Engineers. Straight from my .agents directory.

Star boykopovar / AnyPS5 Tool for automatic PS5 executables porting to Linux and Windows

Star ayghri / i-have-adhd A skill to stop your coding agent from burying the answer. ADHD-friendly output.

Sponsor Star cathrynlavery / diagram-design Editorial diagram design for Claude Code, Codex, GitHub Copilot, Factory Droid, and Pi. 42 diagram types. Self-contained HTML + SVG. No shadows. No Mermaid slop.

Star addyosmani / agent-skills Production-grade engineering skills for AI coding agents.

Star EpicGames / raddebugger A native, user-mode, multi-process, graphical debugger.

Sponsor Star thedotmack / claude-mem Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More

Star manaflow-ai / cmux Open source Ghostty-based macOS terminal with vertical tabs and notifications for AI coding agents. Built for multitasking, organization, and programmability.

trycua/cua★ 28773

Sponsor Star trycua / cua Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.

更多值得一看(内容池 49 条)
DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks

LLM agents often lack the operational knowledge to act reliably in new environments, as they must discover specific tool behaviors or environment conventions on their own. Without memory of past attempts, they repeat the same mistakes across tasks, leading to more task failures and longer trajectories. To address this, agentic systems typically rely on human-written guidelines or on procedural memory built from training tasks and an oracle verifier, both of which require prior knowledge of the e

AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model

Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a t

Image-text retrieval with EmbeddingGemma 2's vision tower, running in the browser on WebGPU

EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in TypeScript, and made a small photo gallery where search runs entirely on your GPU in the browser. ruNNtime also supports plenty of other vision-like models, and you can play with them in the interactive docs source:

AGO AI Quality Gate: Evidence-First Release Decisions for Retrieval-Augmented Generation

Enterprises adopting retrieval-augmented generation (RAG) face a recurring operational decision: promote, revise, or block a system version. The evidence is incomplete and the metrics come from fallible LLM judges. We report on AGO AI Quality Gate (AGO), an evidence-first quality-gate framework deployed in industrial RAG assessment engagements. AGO integrates four key components: a four-state decision model that treats missing data and judge errors as explicit outcomes; layered scoring combining

DiffGate: Difficulty-Gated Teacher Guidance for On-Policy Distillation

On-policy distillation (OPD) has emerged as a widely used paradigm for post-training large language models, reducing the train--test mismatch of conventional distillation by supervising the student on its own generated trajectories. However, existing OPD objectives remain largely token-local and outcome-agnostic, optimizing teacher--student agreement at each prefix despite reasoning quality being determined at the trajectory level. Reinforcement learning with verifiable rewards (RLVR), particula

SlimWise: Decoupling Expert Pruning Across Prefill and Decode for Efficient MoE Serving

Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughput benefit. We present SlimWise, a serving framework that tailors the expert pool to each inference phase. SlimWise performs prefill with the full model and decode with a pruned mo

Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding

On-policy reinforcement learning has become a central paradigm for improving the reasoning abilities of large language models. However, its effectiveness is often limited by reward sparsity: when a model fails to discover correct trajectories for difficult problems, the optimization process receives little useful signal and may stagnate. Existing approaches mitigate this issue by incorporating off-policy demonstrations, expert traces, or model-generated solutions, but they typically require the

Questioning the Questions: Sustaining Self-Evolution in Reasoning Models

Self-evolving reasoning models learn from their own generated questions, yet repeated self-training can lead to performance collapse. In this paper, we investigate why performance deteriorates over successive rounds and how to sustain self-evolution. Our analysis identifies two recurring quality problems in self-generated questions: invalid questions and repeated variants of the same mathematical questions. First, invalid questions become more prevalent across rounds, and answer-consistency filt

TRIAGE: Direction-Aware Mismatch Stabilization of Native NVFP4 Reinforcement Learning

Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sampler execution can destabilize policy optimization. In this paper, we characterize the interaction between mismatch and the policy-gradient direction, distinguishing locally amplifying from contracting update contributions that mismatch magnitude alone cannot identify. In native NVFP4 runs, we observe an early imbalance between the two amplifying re

Understanding and Enhancing Backdoor Persistency in LLM Agent Post-Training

Developers can build LLM agents by adapting third-party models through benign post-training. We study a supply-chain threat in which an attacker supplies a model with a backdoor: hidden behavior that produces malicious outputs when a particular input pattern appears. Focusing on software-engineering agents, we ask whether such backdoors survive the developer's supervised fine-tuning (SFT) and subsequent task-level reinforcement learning (RL). We observe that benign SFT substantially reduces atta

MiniCorp: The Last Mile of the AI Agent Firm

The last mile toward enterprise AGI is a company that runs itself. Training and adapting such agents require longitudinal enterprise data, which remain scarce, costly to acquire, and often restricted by privacy constraints. Historical archives are also frequently incomplete and record only what actually happened. They cannot show the outcomes of alternative decisions. We introduce MiniCorp, an office simulator for studying how agents can collectively run a company while generating enterprise dat

ReikaProduct Hunt1 min开发工具AI

A coding agent CLI designed around small local models first Discussion | Link

A Developer’s Guide to Laya: Zero-Shot Decisions and Calibration

Explore a comprehensive coding guide to Laya, the open-source zero-shot decision engine. Learn how to implement typed decisions, fit custom temperatures, and build reliable abstention gates using real-world CLINC150 banking data. The post A Developer’s Guide to Laya: Zero-Shot Decisions and Calibration appeared first on MarkTechPost .

Towards In-Parameter Memory Augmentation for Large Language Models

Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge acquired after pretraining, e.g., domain facts, user preferences, documents, and interaction experience. In-context learning (ICL) and ICL-based agent harness remain flexible, but they consume context capacity and incur repeated discretized encoding cost that grows with context length. In-parameter memory offers a complementary substrate: reusable memory information is represented in model param

HLA: Expressive Hybrid Linear Attention via Chunk-Wise Dynamic Mixing

Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, but this compression can make selective access to sparse and distant information difficult. Existing chunk-based extensions increase memory capacity, yet learned chunk-mixing coefficients may remain fixed with respect to input content and therefore cannot adapt historical access to each query. We introduce Hybrid Linear Attention (HLA), a query-dependent chunk-level attention mec

Harness-Aware Distillation for Small Language Model Agents

Language model agents are deployed with a harness, the software around the model that manages its context, tools, and feedback. When such an agent is distilled into a smaller one, the harness stays in place, so the student mainly needs the teacher-specific abilities that the harness cannot provide, such as acting correctly on harness information. Standard distillation, however, imitates the teacher's full outputs and treats the harness as part of the input. We propose Harness-Aware Distillation

OPD Before RL: Warm-Starting Rubric-Based RL with On-Policy Distillation

Many useful language-model tasks cannot be evaluated by exact outcome verification. Rubric-based reinforcement learning (RL) addresses this issue by scoring open-ended responses against explicit criteria. However, because the reward is assigned after the complete response, the training signal does not directly identify which individual decisions contributed to the final score. We propose a two-stage training framework that uses rubrics first as privileged teacher context for dense token-level su

Microsoft is giving Copilot more control over Windows and your files

At today's Windows and Surface event, Microsoft showed off an upgrade to its Copilot AI system that will give it access to local files on your PC and the ability to take actions across the OS. It's part of an idea Microsoft is calling "Hybrid Intelligence," where apps and tools rely on a mix of […]

Hey Pocket-ID Dev-Team, Hey stonith404 , I just wanted to say thank you for your work. After a couple of months of intensive use, I have to say: Pocket ID is the service that makes using my self-hosted apps so much easier and more comfortable. I came across Pocket ID while searching for an encrypted file transfer service, and I've been following its development ever since. In the beginning, I didn't have much trust that a young developer could build secure and reliable software. Back then, I did

The New ChatGPT Is More Show Than Tell

OpenAI is updating ChatGPT for all users with an “Intelligent UI” that’s more visual—generating interactive elements as part of the chatbot’s outputs.

Taming VLAs under Robot Execution Errors: Self-Compensation and Stress Testing

Vision-language-action (VLA) policies often fail when a robot's executed motion deviates from their commanded action. Such execution errors arise from the robot's mechanics and operating conditions, such as wear and payload changes. We propose self-compensating VLA, a deployment-time adaptation method that enables a VLA policy to pre-compensate for the robot's execution errors when generating commands. Without task rewards or labels, it updates the policy online using the residual between the ac

Trained a ~20K LM (probably smallest) that can still write stories

I’ve been pushing TinyStories-style models downward in size, and this is the smallest one so far: MacroStories — 19,969 parameters, 81 KB FP32 For scale: → ~50× smaller than the 1M TinyStories model → ~3,000× smaller than AlexNet → 32-dim hidden state → 378-token vocabulary → one decoder block, recurrently applied 4 times with shared weights It’s obviously not a general-purpose LM, but within its constrained story distribution it can maintain a 100–300 word narrative with a goal, problem, releva

Adaptive Latent Capacity for World Models

We introduce Adaptive LeWorldModel (ALeWM), a world model based on a joint-embedding predictive architecture (JEPA) that learns to concentrate predictive information in compact prefixes of a wide latent representation. To encourage this ordering, ALeWM learns a sequence-conditioned distribution over prefix lengths and trains the predictor to estimate the full next embedding from a sampled input prefix. As standard anti-collapse objectives encourage variation across latent coordinates and do not

Sherpa: Teaching LLMs to Teach Adaptively

Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks like. However, these signals are often not grounded in individual student learning outcomes, where effective teaching strategies can vary substantially across learners. To address t

I set up borg via borgmatic like a year ago. 3-2-1 strategy. Confirmed it was backing up and did a quick extract test. That was it. That was a year ago. Well I set a vm in Proxmox and because I was still learning I didn’t set up some directories correctly. Later I installed Immich but apparently installed it under a directory owned by Nextcloud. I never updated Nextcloud because it was local and then I decided to make it available remotely via a reverse proxy and all that fun stuff. So I wanted

A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action Policies

A safe action is not necessarily a viable one. A frozen vision-language-action (VLA) policy can favor a locally admissible move that leaves no policy-supported route to safe task completion. We call this the feasibility-likelihood gap: likelihood ranks the next move, while feasibility depends on the futures it leaves open. To bring those futures into the decision, we derive the exact next-block marginal of the history-conditioned policy-environment trajectory law restricted to safe task completi

每天早晨,一份为你精选的科技日报