DawnSift
订阅日报
周三 · 科技日报 · 第 66 期

2026-09-16

— 开源模型追到只剩 4 个月差距,而安全事件提醒我们:信任前先扫描。

今日 TL;DR

Google 发布 Gemini 3.8 Live 系列语音模型,瞄准生产级语音 agent;TypeSafe 推出 System One Models,主打结构化决策的速度与成本;Mozilla 报告显示开源模型与前沿闭源模型差距缩至 4.4 个月;安全方面,Strix 用 25 分钟拿到 Baseten 生产 GitHub 管理员权限,CrofAI 被曝为 OpenRouter 套壳后清空线上存在。

Models have been superhuman at chat for years, so where is all the automation?

头条

1

Google 发布 Gemini 3.8 Live 与 Extended Thinking,主打生产级语音 agent多源事件 ×4

Google 推出 Gemini 3.8 Live 和 Gemini 3.8 Live Extended Thinking 两款原生语音到语音模型,前者面向规模与成本效率,后者面向高复杂度多步推理;两者均支持后台工具调用、实时视觉输入和 97 种语言切换,Extended Thinking 在 Artificial Analysis 语音质量指数上以 82.6 分排名第一。为什么重要:语音 agent 从演示走向生产的关键瓶颈是推理与工具执行不打断对话流,这两款模型直接针对该缺口,且已通过 Gemini Live API 和 AI Studio 开放。

评论区对语音体验多有认可,但也有人认为演示翻车、版本推送混乱且扩展思考名不副实。

2

TypeSafe 发布 System One Models:面向结构化决策的新模型类别

TypeSafe AI 创始人 Diogo Almeida(前 OpenAI ChatGPT 研究团队成员)宣布推出 System One Model,定位为快速、结构化决策的前沿模型类别,并配套发布 Jev 工具,瞄准分类等结构化任务的速度与成本优势。为什么重要:如果该模型在结构化决策上确实以更低成本达到前沿水平,将改变 agent 流水线中大量重复性判断任务的模型选型逻辑。

社区普遍认可其在结构化任务上的速度与成本优势,但也有人认为缺乏基准证明、与 LLM 对比不公平。

3

Mozilla 报告:开源模型与前沿闭源模型差距缩至 4.4 个月

Mozilla《State of Open Source AI》报告显示,美国科技公司前沿 AI 模型与中国最佳开源权重模型之间的性能差距已缩小到 4.4 个月,Moonshot AI 的 Kimi K3 在 Artificial Analysis Intelligence Index 上的综合得分仅落后 3 分。为什么重要:这意味着多数常规工作负载用开源模型即可满足,前沿闭源模型只在极窄的高难度场景下才值得 5 倍成本溢价,直接影响企业的模型采购与部署策略。

4

Strix 安全团队 25 分钟拿到 Baseten 生产 GitHub 管理员权限

安全公司 Strix 在评估是否将数据托管给 Baseten 时,用自研自主黑客 agent Strix 扫描其外部暴露面,约 25 分钟后获得一个对 Baseten 内部仓库具有仓库级管理员权限的实时 GitHub token。为什么重要:这再次印证了第三方服务供应链安全的脆弱性——即使是估值 130 亿美元、被大量严肃公司依赖的基础设施提供商,也可能存在快速可被利用的凭证泄露路径。

5

CrofAI 被曝为 OpenRouter 套壳,以 20 倍加价路由到更小模型后清空线上存在

自称'全球最便宜推理服务商'的 CrofAI 被曝光实为 OpenRouter 包装,将请求路由到更小、更便宜的模型并加价最高 20 倍;面对欺诈指控,CrofAI 先否认、后改口,3 小时后清空全部线上存在。为什么重要:这是追逐廉价 token 的典型教训——推理服务商若无透明的模型路由与定价机制,'便宜'可能只是套壳加价的伪装。

社区普遍视其为追逐廉价 token 的警示故事。

每天早晨,一份为你精选的科技日报

网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。

已发布 66 期 · 每天筛过 150+ 条只留值得读的 30 条

AI 动态

Atria Dawn: The Dawn of Agentic Superintelligence

Atria Dawn Preview:面向科研与工程工作流的 agentic 基础语言模型,在 16 个基准上与前沿 agent 竞争。

🤖Atria Dawn Preview is a foundation agentic language model trained through verified tool interactions that achieves strong benchmark results and demonstrates a shift toward human-AI project-level collaboration in scientific research.

ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search

ZGCM-1:7B 全开源基础模型,通过内部推理与外部工具调用耦合,在 256K 上下文上实现高效训练。

🤖ZGCM-1 is a 7B open foundation model that combines internal reasoning with external tool use, trained via efficient architecture-system co-design, progressive long-context scaling, and autonomous agent workflows to achieve strong reasoning and efficiency.

Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction

Grouped Value Attention 通过按需重建 key 减少 KV cache,以更小持久缓存达到接近 GQA 的精度。

🤖Grouped Value Attention reduces transformer KV cache size by storing grouped values and reconstructing keys via a learned linear map, achieving near-GQA accuracy with a smaller persistent cache.

开发与开源

Java 27 GA 发布,包含 9 个 JEP:G1 成为所有环境默认 GC、TLS 1.3 后量子混合密钥交换、紧凑对象头默认启用等。

评论区普遍认可 Java 持续更新,但也有人认为 Oracle 节奏过快、企业仍用旧版,爱恨交织。

社区热议

25 years of mass surveillance is enough

Schneier 撰文称 25 年大规模监控已远超反恐初衷,成为执法常规工具;评论区普遍认为监控不会停止且将加剧。

评论区普遍认为大规模监控不会停止且将加剧,但也有人认为应转向本地化替代方案或争取数字权利保护。

US confirms for first time it has deployed space weapons

美国首次确认在轨部署太空武器;评论区普遍批评此举引发军备竞赛与轨道碎片危机。

评论区普遍批评美国部署太空武器会引发军备竞赛和轨道碎片危机,但也有人认为太空军事化早已存在,此次披露可能只是政治噱头。

GitHub Trending

Star alibaba / open-code-review Fast, efficient, battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines + LLM Agent, precise line-level comments, built-in multi-language ruleset (NPE, thread-safety, XSS, SQL injection), OpenAI & Anthropic compatible.

Star JustVugg / colibri Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

Sponsor Star ever-co / ever-gauzy Ever® Gauzy™ - Open Business Management Platform (ERP/CRM/HRM/ATS/PM) - https://gauzy.co

Star debpalash / VoiceStudio VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

Sponsor Star Homebrew / BrewUI 📺 Homebrew's official macOS GUI

Star melgarafael / DeskcommCRM Open-source AI sales OS — self-hosted CRM with native AI agents + WhatsApp (WAHA). Open alternative to Kommo, Octadesk & Intercom for any business that sells by chat. MCP-ready, multi-tenant, LGPD.

Sponsor Star danny-avila / LibreChat Enhanced ChatGPT Clone: Features Agents, MCP, Skills, DeepSeek, Anthropic, AWS, OpenAI, Responses API, Azure, Groq, o1, GPT-5, Mistral, OpenRouter, Vertex AI, Gemini, Artifacts, AI model switching, message search, Code Interpreter, langchain, DALL-E-3, OpenAPI Actions, Functions, Secure Multi-User Auth, Presets, open-source for self-hosting. Active

Star pacifio / atlas Source control for agents. Use multiple coding agents, track their changes and query them in one place

更多值得一看(内容池 61 条)
Meta Introduces ZGateway: A Stateless Proxy Tier That Unifies ZippyDB Traffic and Handles Over 1 Billion Operations Per Second

Meta engineering team introduced ZGateway, a proxy tier that now sits between client applications and ZippyDB, the Meta’s most widely used key value store. ZippyDB backs product metadata, counters, and configuration at billions of operations per second. ZGateway started as a fix for connection sprawl across more than a million client hosts and grew into […] The post Meta Introduces ZGateway: A Stateless Proxy Tier That Unifies ZippyDB Traffic and Handles Over 1 Billion Operations Per Second appe

When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis

Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to stop. This test-time strategy makes it difficult to measure how agent performance scales. We study open-ended tasks that provide continuous scores for intermediate submissions, making progress observable throughout long trajectories. We propose Elo-per-token analysis, which tracks the best solution found at each token budget and uses a Bradley-Terr

I trained a 44M parameter quantized LLM from scratch on 45B tokens. It ships in 19.8 MB and runs at ~1,900 tok/s on CPU. [P]

Three weeks back , i posted SHADOW-250M here. It got 360 upvotes, 293 on r/LocalLLaMA and 94 GitHub stars. Thank you. That model was 60 MB, ran around 400 tok/s on CPU and could retrieve records from an archive on disk. What it couldn’t do reliably was reason over what it retrieved or compute. So I built a smaller one to experiment with those two problems. SHADOW-50M is actually 44M parameters, trained from scratch on 45B tokens. 19.8 MB complete model, ~1,900 tok/s on laptop CPU, ~41 MB RAM, te

Token Efficient Task Execution via Application Behavior Modeling for Web Agents

arXiv:2609.13491v1 Announce Type: new Abstract: The strong performance of AI Agents across an impressive variety of tasks is driving an unprecedented investment in agentic infrastructures, however the cost of processing tokens is fast increasing. Web agents automate the execution of web-application tasks described in natural language, by analyzing the web-application's user interface (UI) and interacting with it. This work introduces OdoBot, a novel web-agent architecture that completes tasks at

Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents

arXiv:2609.13543v1 Announce Type: new Abstract: LLM agents are predominantly benchmarked on short, single-task trajectories, yet real deployments run for hours under contention, surfacing a different class of failures. We use the Clinical Environment Simulator (CES), in which an agent manages an entire emergency-department shift under continuous time and resource pressure, as a testbed: long-horizon execution failures manifest measurably in a single rollout under structured, multi-dimensional gr

Agent-net Open Sources Webagent: A Go Harness That Turns Any Website into a Guarded AI Agent

Agent-net, the team building an agent-to-agent marketplace where AI agents discover, trust, and pay each other, has released Webagent, an open source harness for standing up public-facing business agents. So, basically you give it your website, get an agent, and let it talk to other agents. Instead of writing orchestration code, a business fills in […] The post Agent-net Open Sources Webagent: A Go Harness That Turns Any Website into a Guarded AI Agent appeared first on MarkTechPost .

LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents

Diffusion large language models (dLLMs) achieve high decoding efficiency through block-parallel, arbitrary-order generation, making them attractive for latency-sensitive applications. GUI agents represent a natural testbed for this paradigm, as they must repeatedly perceive screen states and emit structured, spatially grounded actions in real time. However, whether dLLMs can be extended into capable multimodal GUI agents while preserving their parallel decoding advantage remains an open question

HazardAuditor: From Executable Threats to Safer Computer-Use Agents

Computer-use agents increasingly interact with browsers, terminals, file systems, and external services, introducing safety risks that emerge through runtime behavior rather than generated content alone. Existing guard models target static prompts and responses and are poorly suited to agent execution; existing executable safety platforms produce evaluation verdicts rather than the normalized supervision a guard model needs to learn across heterogeneous agent frameworks. We introduce HazardAudit

Voodoo Dynamic Quant - Now MIT Licensed

Two months ago I announced I had found a new dynamic quant method called Voodoo Quant which was SOTA for the most aggressive quant levels on some smaller Qwen3.5 GGUF models. I kept the methodology private at the time, but I've seen too many requests for dyn quants for various models lately, so I decided to give my method to the community since I don't have the time to scale this into something that could do it justice. Hopefully it will also inspire some researchers to find out more about it an

Cut Qwen3.8-27B Reasoning Tokens by 40% -- 3.8 'ThinkingCap' benchmarked!

EDIT: Sorry for the unclear title. This model is UkisAI's Swift-Qwen3.8-27B, not a new version of BottleCap AI's 3.6-ThinkingCap. All credit goes to UkisAI for making great fine-tune, and I made this post to celebrate their work. I meant no disrespect by mentioning another model in the title. I doubt I'm in the minority here when I say I love Qwen models, but the overthinking is a major timekiller. It was bad in 3.6-27B, and it's worse in 3.8. I know there are some who say, "well that's how it a

ByteShape Qwen 3.8 27B: To KL Diverge or Not to KL Diverge, Part 2: Metric Boogaloo

Hey r/LocalLLaMA , We’ve released our full ShapeLearn GGUFs for Qwen 3.8 27B. Blog / Download models TL;DR 3.84 bpw (GPU-5) reaches 99.63% of BF16’s aggregate score of 8 benchmarks, being the most accurate quant we’ve evaluated; 3.23 bpw (GPU-4) reaches 98.72%. These average BF16-normalized scores across instruct and thinking benchmarks. All five new models sit on the measured quality/speed-bpw frontier across six GPUs. In this model’s case, lower BPW translates directly to TPS. Comparisons incl

If you have a 3090, or other 30xx for local LLMs, I have something for you

I have a custom fork of llama.cpp designed around the ampere architecture specifically (though many of the upgrades also translate to faster performance of blackwell + lovelace). The recommended config supports 90+ TPS (for agentic/coding, at temp 1; greedy will of course be faster) through 100K tokens, with context of up to 240K. If you want the repo, it is here: I recommend running with this quant, which is ~ 4 K M quality but considerably faster (technically, a 3 K XL upgrade) If you want the

Generalized Agent Iteration: One Formal Framework for Iterative Policy Improvement and Recursive Self-Improvement

arXiv:2609.13406v1 Announce Type: new Abstract: When we speak of recursive self-improvement (RSI), are we speaking of a phenomenon, a mechanism, or a prospect? Towards autonomous and evolving intelligence, RSI is being claimed at many scales, while no single framework that formally describes these emerging instances exists. Its counterpart in the classical realm, iterative policy improvement, is characterized by generalized policy iteration (GPI), a framework of broad applicability with well-und

Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?

arXiv:2609.13436v1 Announce Type: new Abstract: Large Language Model (LLM) agents offer a promising path toward autonomously managing long-term physical tasks without human intervention. However, physical tasks require agents to continuously observe the environment, make consequential actions, and remain effective as the environment changes. Existing approaches either require substantial data and retraining, or primarily focus on agents operating in the virtual world. In this work, we explore th

Discovery Foundation Models: Toward Open-Ended Discovery Intelligence

Foundation models have progressed from learning and reasoning over existing knowledge, to increasingly learning through action, tool use, and outcome feedback. We argue that the next frontier is a further transition: from solving and acting within problems specified by humans to participating in the process by which new problems, representations, explanations, and knowledge are created. We refer to this capability as Discovery Intelligence. We formulate Discovery Foundation Models (DFMs) as gene

BVB: Benchmarking Agentic Video Understanding via Programmatic Reconstruction in Blender

Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Yet video understanding benchmarks still evaluate models mainly through question answering. If an agent truly understands a video, it can reconstruct it programmatically. We introduce BVB, Blender-VideoBench, a benchmark that tests this ability by asking agents to reconstruct real-world videos as animated Blender scenes. To ensure fair comparison, each agent programs the reconst

AlayaVista: Streaming World Modeling from Panoramic States to Perspective Video

Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity observations with low latency. Existing approaches face a representation trade-off: perspective models operate on local views and must preserve off-screen content over long rollouts, whereas broader spatial coverage is typically obtained by synthesizing full-sphere videos or constructing explicit 3D representations. Motivated by the complementary roles of global context and selectiv

Pelican-Sim 1.0: A General World Model Simulator for Embodied Intelligence

In this technical report, we propose Pelican-Sim 1.0, a general world model simulator for embodied intelligence that predicts future observations from visual context and robot actions to support downstream learning and decision making. The model incorporates four key design features: (1) Unified action representation: a 28-dimensional action value space covering most mainstream embodiments, keeping one model valid across heterogeneous devices. (2) Action-visual injection: URDF- and camera-render

Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support an interpretation before an utterance or sound event is complete. If that interpretation enters memory as a fact, later reasoning can keep relaying it even after audio contradicts it. We call this failure premature cross-modal commitment. We propose Omni-Streaming Thinking (OST), which generates structured outputs that include evidence observed s

Apple Foundation Models: local AI natively on MacOS 27

Maybe some of you know but I didn’t see any post about this. Apple just made available their AFM model on MacOS 27 natively. Just run fm chat in a terminal. Disclaimer: I’m an open weight person. I prefer open models and ecosystem, but I’ll still open the discussion. Did you test them? Build using them? Are these models good? I feel like this is still a huge step in the direction of local AI that a company like Apple does this and release hardware optimized models. So what do you think?

Broken Windows, Abstractions and the Cost of Always Keeping Things Simple

The patterns already in a codebase have a huge influence on what gets added next, even when everyone can see those patterns are starting to break down. This article discusses this phenomenon through the lens of YAGNI and the Broken Window Theory.

LabAgent: Customize Any Research Hubs for Scientific Discoveries Using AI Agents

arXiv:2609.13437v1 Announce Type: new Abstract: Scientific research is a continuous process that emphasizes inheritance. Methods developed by predecessors are often expanded upon by new researchers to explore more novel and in-depth scientific questions. However, the change of lab staff, such as student graduation, leads to a lack of personnel capable of replicating methods. Methods that have been developed with significant effort and resources cannot be continued. To address these limitations,

Roundtables: Could AI really kill us all?

Listen to the session or watch below Employees at the world’s leading AI labs are saying there’s a real possibility that advanced AI could destroy humanity. Are they right? Or is this more scaremongering and hype? Watch a conversation unpacking AI extinction fears: where they come from, whether they hold any water, and, if so,…

Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents

arXiv:2609.13422v1 Announce Type: new Abstract: LLM judges are increasingly used to evaluate and improve AI-generated outputs, yet their reliability for complex professional work remains unclear. We study this problem through Vibe Patenting, an end-to-end patent-drafting testbed for AI-agent evaluation. A separately-invoked LLM judge evaluates generated patent drafts and provides structured feedback for iterative revision. Across multiple inventions and drafting-agent configurations, judge-guide

AI models need more data about biology, and OpenAI is paying to create it

Last year Ruxandra Teslo, a policy analyst who focuses on clinical trials, posted an idea for supercharging medical AI systems: Use data from failed biotech companies. By bidding at their bankruptcy proceedings, she proposed, it might be possible to obtain detailed regulatory filings, manufacturing strategies, and safety data—types of information usually considered trade secrets. She…

Lets be honest, 80% of the services we run our labs are just tools to deploy, monitor, log, backup or do other server-related stuff. Cool for nerds, but nothing to write home about honestly. The really useful and "fun" services, e.g. Vaultwarden, Paperless, Jellyfin etc. are rather sparse I want to know whats your favorite "fun" service one might not know about yet. I'll start: Recently came across RECLIP and I throughly enjoy it. Used to work with halfassed browser extensions and sketchy websit

Think Before You Link: Rarity, Reasoning, and Retrieval in Multilingual Entity Linking

Multimodal entity linking grounds entity mentions in text and images to knowledge-base entries. These systems degrade on rare entities, but prior work measures rarity primarily through popularity-based metrics such as pageviews. We broaden this view using knowledge-graph structural metrics that capture how well an entity is documented and connected. These metrics identify many rare entities that popularity metrics miss. Across the resulting rare-entity slices, state-of-the-art accuracy drops by

每天早晨,一份为你精选的科技日报