DawnSift
订阅日报
周二 · 科技日报 · 第 44 期

2026-08-25

— 今天的主线:AI 正在从模型能力转向系统与信任的博弈。

今日 TL;DR

OpenAI 对 GPT-5.6 系列大幅降价并接入 Kiro 开发代理,价格战与 agent 落地同步加速。小米发布 Xring O3 芯片,单核追平苹果、多核更强,但功耗与能效仍存疑。微软 Paint 被曝在本地生成图片中静默嵌入含 GUID 的隐形水印,隐私争议升温。IPFS 核心维护团队 Shipyard 宣布因资金中断将于 9 月底停止相关工作。

We see a future where intelligence is a utility like electricity or water and people buy it from us on a meter and use it for whatever they want to use it for.

头条

1

OpenAI 大幅下调 GPT-5.6 系列价格,并接入 Kiro 开发代理

OpenAI 将 GPT-5.6 Sol 的短上下文输入价格降至 $4.00/百万 token,缓存输入 $0.40,输出 $20.00;促销价至少持续到 2026 年 11 月 21 日。同时 GPT-5.6 全系(Sol、Terra、Luna)已接入软件工程代理 Kiro,用于规划、构建、评审和测试。为什么重要:旗舰模型降价直接降低 agent 工作流的 token 成本,而 Kiro 集成意味着 OpenAI 正在把最新模型推向 AI-native 软件开发生命周期,对依赖 LLM 做代码生成的团队有实际成本与工具链影响。

社区普遍欢迎降价,认为价格战利好开源模型和用户;但也有人指出降价不影响订阅用户,且模型命名混乱。

2

小米发布 Xring O3 芯片:单核追平苹果,多核大幅领先

小米发布 Xring O3 处理器,其 C1-Ultra 大核在单线程任务上大致追平苹果核心,多线程执行则明显更快;芯片总缓存达 44MB,超过多数笔记本 CPU,并支持 SME2 矩阵扩展与 SVE2 数据并行。为什么重要:这是中国厂商在消费级 CPU 单核性能上首次逼近苹果,对移动端 AI 推理和本地 LLM 部署有直接意义;但功耗与能效比仍是关键未知数。

评论区普遍认可性能亮眼,但质疑功耗与能效比,认为单核仅追平苹果去年产品、多核靠核心数取胜;也有人认为竞争利好消费者。

3

微软 Paint 在本地生成图片中静默嵌入含 GUID 的隐形水印

安全研究员发现,Microsoft Paint 和 Photos 在本地 AI 生成图片时,会将提示词发送到远程服务器进行内容审核,服务器返回一个 GUID,该 GUID 被嵌入到本地生成的图片中作为隐形水印;Copilot+ PC 上图片生成虽在本地完成,但提示词审核仍为远程。为什么重要:这直接关系到本地 AI 功能的隐私边界——用户以为完全离线的操作,实际上仍会向微软服务器泄露提示词,并在输出中留下可追踪标识,对匿名性和数据主权构成实质影响。

评论区普遍担忧隐私侵犯与匿名性受损,认为静默嵌入水印不可接受;但也有人认为这是对抗深度伪造的合理措施。

4

IPFS 核心维护团队 Shipyard 宣布因资金中断停止相关工作

Protocol Labs 通知 Shipyard 不再续签资金支持,Shipyard 将于 2026 年 9 月 30 日结束所有 IPFS 相关的工程、维护和基础设施运营。为什么重要:IPFS 是去中心化存储与内容寻址的关键基础设施,核心维护团队的退出将直接影响生态系统的稳定性和后续演进,对依赖 IPFS 做数据分发的开发者是不容忽视的风险信号。

评论区普遍惋惜,认为项目理念好但执行和生态支持不足;也有人认为这仅是个别团队调整,项目并未终结。

5

seL4 在 AArch64 上完成安全隔离的形式化证明

Proofcraft 宣布在 AArch64 架构上完成了 seL4 微内核的机密性证明,加上此前已完成的功能正确性与完整性证明,seL4 实现代码在 AArch64 上强制安全隔离的数学证明已全部完成。为什么重要:这是操作系统内核形式化验证的里程碑,意味着运行在 seL4 上的应用无法在未授权情况下获取信息,对高安全场景(如关键基础设施、可信执行环境)的开发者提供了可验证的隔离保证。

每天早晨,一份为你精选的科技日报

网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。

已发布 44 期 · 每天筛过 150+ 条只留值得读的 30 条

AI 动态

OmniScientist: An Omni-Modal Omni-Discipline AI Scientist

OmniScientist 提出端到端全模态 AI 科学家,直接从异构原始证据进行多学科研究,代码已开源。

🤖OmniScientist is an end-to-end omni-modal AI scientist that performs multidisciplinary research directly from heterogeneous raw evidence using autonomous agents and lifecycle-wide perception, improving evidence-grounded discovery across diverse scientific modalities.

开发与开源

社区热议

Coding expertise is going to collapse from AI reliance

《Coding expertise is going to collapse from AI reliance》引发热议,多数评论认同 AI 依赖正在削弱编程技能,但也有人视其为技术演进必然。

评论普遍认同AI依赖正削弱编程技能,但也有人认为这是技术演进的必然,类似计算器或编译器的影响。

GitHub Trending

基于官方 DeepSeek Harness 打造的 Electron 桌面端,深度适配 macOS 和 Windows,提供最佳的,开箱即用的体验。

Strip multi-vendor AI provenance marks: Unicode text hygiene, statistical rewrite hooks, and C2PA/metadata from PNG/JPEG/SVG/PDF/DOCX/HTML/MD

Self-evolving Context Database for AI Agents. Unify Agent Memory, Knowledge RAG and Skills.

更多值得一看(内容池 54 条)
Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye

Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now AI is accelerating some types of progress but not others:…A nice METR study lays out where acceleration is showing up…Here’s a little analysis from METR which […]

[Paper] ToMoE: Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning

ToMoE : Converting Dense Large Language Models to Mixture-of-Experts through Dynamic Structural Pruning Large Language Models (LLMs) have demonstrated remarkable abilities in tackling a wide range of complex tasks. However, their huge computational and memory costs raise significant challenges in deploying these models on resource-constrained devices or efficiently serving them. Prior approaches have attempted to alleviate these problems by permanently removing less important model structures, y

TielCoder's 22 GB 4-bit quant matches Opus4.6 medium on recent real life coding issues, surpassing KAT-Coder and Nail as strongest and fastest MoE picks.

Qwen3.8-27B is amazing, but it’s slow. A stronger 35B-A3B Mixture of Experts-coder that can run and solve real codebase issues fast (even on constrained hardware) is a valuable addition to the arsenal. This one is the strongest and most consistent 35B-A3B I’ve benchmarked, on both correctness and speed, in addition to being the fastest to fix out of all the 35B-A3B models when you throw them at real codebases. On top of Ornith-1.5’s fine tune, TielCoder uses a code-weighted imatrix for dynamic q

deepseek-v4-flash-0731 - surprisingly usable

I just finished building my (relatively) low rent local inference machine: * Epyc 7663 * 256GB ECC DDR4-3200 * 1x RTX 5090 32GB Yeah I realize it's weird to throw a 5090 and 256GB of anything together and call it low end, but relative to ~151GB of weights it is. I'm running UD-Q8_K_XL and getting 23.8-24.6 tokens/sec, with pp ranging from 60 on the first prompt to 385 near the last (no doubt lots of caching) on tasks using 100-128k total context. It was slower with DFlash so I took that out. It

JetBrains local AI (using Qwen3.6 27B)

Sounds quite interesting, a big IDE provider optimizing for local AI with their coding harness. Especially that they picked Qwen3.6 over Qwen3.8 because of the thinking needs. Haven't read the full article yet, but sounds really cool.

Instinct’s powerful AI assistant is raising privacy and security concerns

Early testers are raving about what Instinct can do, but some say the AI assistant’s sweeping access, broad terms and ability to act on users’ behalf come with uncomfortable trade-offs.

Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts

Mixture-of-Experts (MoE) architectures significantly expand model capacity without a proportional increase in computational cost. However, optimizing their hyperparameters---particularly the learning rate---at extreme scales of both model size and token budget via sweeping remains computationally prohibitive. In this paper, we propose a compute-efficient, two-step hyperparameter transfer framework that estimates optimal learning rates for training large MoE models by transferring them across sca

Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory

arXiv:2608.20397v1 Announce Type: new Abstract: Agentic large language models (LLMs) on the Model Context Protocol (MCP) re-encode verbose tool schemas every turn, so prefill - quadratic in sequence length - dominates time-to-first-token (TTFT) as the tool registry grows. Nexus's primary lever is to decouple routing from the schema-prefill cost: an INT8 semantic lookaside buffer (SLB) with a calibrated cross-encoder margin gate selects tools by retrieval, and arguments are generated over a compr

When Retrieval Fails Before It Begins: Structurally Indirect Prerequisite Eviction as a Retention Failure in Agentic Memory

arXiv:2608.20400v1 Announce Type: new Abstract: Agentic memory under a fixed budget involves two stages: retention and retrieval. Existing retrieval-centered paradigms implicitly assume necessary evidence survives eviction, but we challenge this by isolating a pre-retrieval failure mode: structurally indirect prerequisite eviction, in which upstream blocks weakly aligned with the query are discarded under budget pressure. We provide an operational definition of this failure, a reproducible deter

Best GPU Neoclouds 2026: CoreWeave, Nebius, Lambda, Crusoe, and Groq Ranked by Published Pricing and Contracted Power

The five largest GPU neoclouds now run on very different models. CoreWeave and Nebius report to the SEC; Lambda and Crusoe are private and heading toward IPOs; Groq rebuilt itself as an inference cloud after licensing its LPU technology to NVIDIA. This comparison checks each provider's live rate card, Q2 2026 financials, active and contracted gigawatts, anchor contracts, and SemiAnalysis ClusterMAX tier. Nebius posts the lowest H100 rate and the only published B300 price, Lambda has the cheapest

A Survey on Foundations and Frontiers of Multimodal Agentic Frameworks: Techniques and Applications

arXiv:2608.20379v1 Announce Type: new Abstract: Advances in large language models (LLMs) have fueled a wave of research into agency: the ability to reason, plan, and act. This effort has produced agentic frameworks that orchestrate perception, memory, and decision-making around powerful LLM backbones. With the advent of large multimodal models (LMMs), these systems can process and integrate diverse modalities, including images, audio, and video, thereby improving their real-world applicability.

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although these modes differ in reasoning budget, their delivered responses should satisfy the same user-facing standard. Correctness alone may not characterize this response quality; we therefore evaluate task accuracy and response-pattern failures as complementary outcomes. We study this gap through response-pattern alignment: whe

The Embedder's Dilemma: LLMs Are Better, but at What Cost?

Should you replace your text-embedding pipeline with a large language model? We answer this with a controlled, cost-aware comparison of ten LLMs across six families and 26 embedding models (118M to 14B parameters) on 37 tasks spanning classification, semantic textual similarity (STS), clustering, pair classification, and retrieval. In aggregate the two paradigms are effectively tied: the best LLM (Gemini 3.1 Pro, 77.6) and the best embedding model (77.2) differ by 0.4 points. Their strengths dif

Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent Harnesses

Modern LLM agents are often improved by modifying prompts, tools, or workflows manually, while the executable scaffold surrounding the model---the harness---is typically treated as a fixed artifact after deployment. This work studies an alternative where the harness is task-specific and continuously evolvable: each task family maintains its own harness, which is hot-swapped across iterations through a fixed task-injection seam and rewritten using environment feedback. We introduce Hierarchical S

I trained a 1.57B-parameter Dreamer 4 World Model from scratch for under $150

My first attempt didn't work. I built on Genie's architecture and the videos looked great, but the controls barely did anything. The effect of a keypress was basically zero. Genie learns its actions unsupervised into 8 codes, and that was too loose a grip for us. So I scrapped it and started again with Dreamer 4. The second attempt: Tokenizer at 40.41 PSNR (Genie's paper reports 35.7) FVD 32.19 end to end 144 frames before it falls apart 1.57B parameters, 9.6M frames, ~$150 Two important learnin

This is what Qwen 3.8 27b is capable of

Try it here: Model: Qwen 3.8 27b Q8_X_KL Unsloth Hardware: 3 x RTX3090 Harness: DeepSeek Harness Prompt: /goal I want you to create a **JavaScript + Node.js WebGL project** that renders a highly realistic real-time ocean in the browser. Use **JavaScript only, no TypeScript**. You may use WebGL2, GLSL, and Three.js. The ocean should include realistic waves, vertex displacement, Fresnel reflections, sun highlights, sky/environment reflection, foam/whitecaps, horizon treatment, atmospheric effects,

Hugging Face reportedly in talks to be acquired for $13B

Hugging Face has reportedly been fielding acquisition offers that would value the company at around $13B. But with the founders' feeling of responsibility to community, doubts arise as to whether a sale will happen.

OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs

Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continuously perceive environments and guide users to achieve specific goals. Unlike traditional passive video understanding, interactive assistants should actively combine visual states, user goals, and prior knowledge to provide effective help. Evaluating this is rather challenging, as the model's unpredictable response dynamically changes the user's subsequent actions, which static of

Representation Affects Retrieval: A Case Study of Skill Discovery and Routing in a Multimodal Agent Harness

arXiv:2608.20389v1 Announce Type: new Abstract: A production agent harness must discover and rank, from a growing library of skills, the one most appropriate for a user's task. At small scale this selection happens in context: the LLM planner chooses among skill representations exposed in its system prompt, without an explicit embedding-based retrieval step. We treat this in-context selection as the small-N counterpart to embedding-based skill retrieval at scale, and present a case study of how

Truth Lies Deep: Countering Semantic Camouflage via Latent Intent Verification

arXiv:2608.20378v1 Announce Type: new Abstract: Safety alignment in Large Language Models (LLMs) is often superficial, relying on refusal mechanisms that trigger only at the final stages of generation without erasing the foundational knowledge of harmful concepts acquired during pretraining. This study demonstrates that this architectural disconnect leaves models vulnerable to Semantic Camouflage -- adversarial attacks that wrap harmful intent in benign narrative contexts (e.g., creative writing

Scientific Data Analysis with LabPlot in Python: Signal Processing, Spectral Peak Fitting, Visualization, and Batch Automation

In this tutorial, we explore a LabPlot-inspired scientific data analysis workflow in Python while preserving the structure and terminology of LabPlot’s aspect tree, analysis kernels, plotting system, and project model. We build reusable components to import tabular data, compute descriptive statistics, smooth and differentiate signals, perform Fourier analysis and filtering, detect peaks, integrate curves, reduce […] The post Scientific Data Analysis with LabPlot in Python: Signal Processing, Sp

Every Coin Has Two Sides: On the Dual Nature of Generalization in On-Policy Distillation of Large Language Models

On-policy distillation (OPD) transfers teacher capabilities by supervising trajectories sampled from the student's own policy, yet its generalization behavior remains poorly understood, as most studies evaluate OPD on a single domain and on benchmarks close to the training data. We present a controlled study that varies one generalization factor at a time, from in-domain distribution shifts to cross-domain transfer and the multi-teacher setting. We find that OPD transfers a teacher's reasoning b

Qwen 3.8 27B, just wanted to say thanks to you guys

I commented on another Qwen 3.8 27B post that I was frustrated getting anything to work. You all gave some great comments. I nuked openwebui and straightened out my llama.cpp docker config. 1 hour of work and I have a model I can chat with, connected to my HomeAssistant server, which I have already updated dashboards with a short prompt and a screenshot (wtf vision built in?) Guess all I needed was the right push. I bought several GPUs in 2023 in impulse purchases for Folding@Home, but have alwa

Who would buy HuggingFace

Given OpenRouter.ai was snapped up by Stripe, who do we think would go after the "GitHib" of AI models? It is a big chunk of change they are looking ($13B). Apple may be a contender to give them a real chip in the AI race, given how they are focused on local AI execution.

De-Googled GrapheneOS is coming to Motorola’s foldables next year

GrapheneOS, an open source version of Android that prioritizes security and privacy, has detailed its plans for supporting Motorola smartphones. Official support is set to arrive next year, starting with traditional flagships, before rolling out to Motorola's foldable phones and perhaps cheaper models, eventually. In a Mastodon thread, the GrapheneOS Foundation announced that it will […]

Google Research Introduces ME-POIs: A Mobility-Informed Framework that Adds “How a Place Is Used” to Text-Based POI Embeddings

framework that folds aggregate human movement into text-based place embeddings. Language models describe what a place is; they miss how it is used. ME-POIs encodes each visit as a contextualized vector and aligns it with one learnable prototype per POI through contrastive learning, then transfers visit distributions from data-rich anchors to the long tail across three spatial scales. Across five map-enrichment tasks on Los Angeles and Houston mobility data, adding ME-POIs improved 34 of 35 model

EviRank: Structured Relevance Evidence for Multimodal Image Re-ranking

Real-world image search queries are multimodal and compositional: ``find this shirt in pink'' specifies an entity to retain, an attribute to modify, and context to ignore. Yet existing re-rankers either compress such multifaceted relevance into an opaque embedding or rely on free-form chain-of-thought that easily omits or hallucinates fine-grained constraints. Drawing on rubric- and checklist-based evaluation from NLP, we recast multimodal image re-ranking as a semantic constraint satisfaction p

DropstoneProduct Hunt1 minAI开发工具

The AI runtime that remembers, learns, and acts everywhere Discussion | Link

每天早晨,一份为你精选的科技日报