DawnSift
订阅日报
周二 · 科技日报 · 第 15 期

2026-07-21

— 开源模型冲击闭源护城河,安全护栏反成攻击者的掩护。

今日 TL;DR

中国开源模型 Kimi K3 与 Qwen 3.8 发布,性能逼近闭源前沿,引发美国对开源模型禁令的讨论。同时,AI 安全护栏被指阻碍防御者修复漏洞,而攻击者却能绕过。此外,罗马尼亚土地登记数据库遭黑客清空,WordPress 高危漏洞利用成本仅 25 美元,安全形势严峻。

非常可怕的是,作为防御者却被护栏阻挡,而你知道攻击者很可能正在绕过它们。

头条

1

中国开源模型 Kimi K3 与 Qwen 3.8 发布,冲击闭源模型格局多源事件 ×5

Moonshot AI 发布 2.8T 参数 MoE 模型 Kimi K3,阿里发布 Qwen 3.8,两者性能均接近 Anthropic Fable 5,且承诺在未来数周开源权重。此举将中美开源与闭源模型的性能差距缩短至 3-5 个月。 为什么重要:开源模型正在消解闭源 API 的技术护城河,迫使 OpenAI 等厂商考虑推动监管干预。这对依赖 API 的开发者意味着模型切换成本极低,供应商锁定难以为继。

社区普遍认为中国开源策略正在削弱美国闭源模型的竞争优势,但也有人质疑其安全性、政治偏见及实际采用率。

2

AI 安全护栏引争议:Kimi K3 修复 15 个高危漏洞,Codex 与 Fable 拒绝执行

安全研究员使用 Kimi K3 成功修复了 15 个 Codex 和 Fable 因“网络护栏”而拒绝处理的关键安全漏洞。Hugging Face 也证实遭遇类似情况,称作为防御者被护栏阻挡“非常可怕”。 为什么重要:模型安全护栏的过度保守正在阻碍合法的安全防御工作,而攻击者却能通过越狱绕过这些限制,形成攻防不对等的危险局面。

3

黑客清空罗马尼亚全国土地登记数据库,房地产市场停摆

一名黑客在勒索未遂后入侵罗马尼亚地籍局,清空了整个国家土地登记数据库。官方应用和网站已离线一周,公证处无法记录新交易,公民无法获取产权证明。 为什么重要:这是一起针对国家关键数字基础设施的毁灭性攻击,凸显了离线备份和纸质记录作为最后防线的关键作用,也为所有依赖集中式数据库的系统敲响警钟。

评论区普遍认为离线备份和纸质记录避免了灾难性后果,但也有人指出系统安全漏洞和数字化风险。

4

研究者用 GPT5.6 和 25 美元发现 WordPress 远程代码执行漏洞,黑市价值 50 万美元

安全研究员使用 GPT5.6 Sol Ultra 模型,仅花费 25 美元 API 费用就发现了一个 WordPress 远程代码执行漏洞,而漏洞经纪人对此类漏洞的出价高达 50 万美元。 为什么重要:LLM 大幅降低了漏洞挖掘的成本与门槛,使个人研究者能产出与国家级攻击团队媲美的成果,这既加速了防御也放大了威胁。

5

OpenCode 被指存在严重安全隐患,作者呼吁立即停用

一篇技术博文详细揭露了热门开源 AI 编码代理 OpenCode(GitHub 161k stars)的多项安全缺陷,包括令人担忧的权限处理和沙箱逃逸风险,作者结论是“所有使用者都应立即停止使用”。 为什么重要:随着 AI 编码代理在开发者工作流中深度集成,其自身的安全性成为软件供应链的新攻击面。盲目信任这些工具可能引入比其解决的问题更严重的风险。

每天早晨,一份为你精选的科技日报

网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。

已发布 15 期 · 每天筛过 150+ 条只留值得读的 30 条

AI 动态

Loop the Loopies!

Loopie 发布最强循环 Transformer,20B 参数 MoE 模型仅需 2B 活跃参数,在同等算力预算下显著优于标准 Transformer。

On-Policy Delta Distillation

NAVER 提出 On-Policy Delta Distillation,用教师模型与基模型的差异信号替代直接模仿,提升推理能力蒸馏效果。

AI is more likely than humans to form biases when hiring

普林斯顿与芝加哥大学研究发现,LLM 在模拟招聘中比人类更容易形成偏见,且能从经验中自行产生刻板印象。

开发与开源

Unsloth now supports AMD!

Unsloth 正式支持 AMD 硬件,覆盖 RX 9000/7000 系列和 Instinct MI350/MI300,支持本地推理、微调和 RL。

xHC: Expanded Hyper-Connections

xHC 扩展 Hyper-Connections 方法,突破 N=4 的限制,解决残差流扩展中的写回信息不足和混合成本瓶颈。

Reverse-engineering is cheap now

Simon Willison 指出编码代理大幅降低了逆向工程和自动化家庭设备的成本与风险,改变了 ROI 计算方式。

社区热议

Claude Fable produced a counterexample to the Jacobian Conjecture

Claude Fable 发现雅可比猜想反例引发数学界震动,评论区对 AI 发现的验证方式和数学价值存在分歧。

评论区普遍认为AI发现雅可比猜想反例是重大突破,但也有人质疑其验证方式及数学价值。

The EU is about to sell our most sensitive data to the US for visa-free travel

欧盟拟与美国签署增强边境安全协议,允许美方访问生物识别数据库以维持免签待遇,引发数据主权争议。

评论区主要认为这是数据共享而非出售,且美国入境本就需提供生物信息,但也有人认为这是欧盟向美国妥协的背叛行为。

GitHub Trending

Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

更多值得一看(内容池 55 条)
543 tok/s single-request Qwen3.6-35B-A3B on one RTX 5090 over a 65K-token decode

An example TL;DR I have open-sourced NInfer , a from-scratch C++/CUDA inference engine currently specialized for two exact Qwen3.6 checkpoints on a single RTX 5090. Both the engine and the converted model artifacts are publicly available: Github : The main result: Qwen3.6-35B-A3B sustained 542 tok/s while generating a full 65,536 token completion, on a single RTX 5090, single request. My goal was to find out how fast inference can get on a single GPU (in my case RTX 5090), with a fixed model and

AI’s most important protocol is getting a little bit easier to use

Under the new system, the protocol will take a looser, "stateless" approach to session IDs on the server side, similar to how most ordinary websites already work.

Safety and alignment in an era of long-horizon models

OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.

RoboTTT: Context Scaling for Robot Policies

Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturb

Understanding Reasoning from Pretraining to Post-Training

Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model? These questions are difficult to study in the standard LLM setting: pretraining corpora are vast and uncontrolled,

Recursive Harness Self-Improvement

Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both immediate agent performance and the quality of traces used for future model training. However, continually updating provider-built scaffolds is costly and labor-intensive. We therefore investigate whether optimizing user-constructed harnesses in a ta

I ran Ternary-Bonsai-27B (2-bit) and Bonsai-27B (1-bit) on Terminal-Bench 2.0, in 8GB VRAM

I asked myself where the Bonsai models actually land, so I ran them and compared to the results I already have for qwen-3.6-35b-a3b and qwen-3.5-9b on the same harness. Thought it might interest more people. Setup: little-coder harness via the harbor adapter, all 89 tasks of terminal-bench 2.0, single attempt (k=1), 40-turn cap, temp 0.2. RTX 5070 Laptop 8GB, i9-14900HX, 32GB RAM, CUDA 13.1. Runtime is PrismML's llama.cpp fork (stock llama.cpp can't load the 2-bit kernels). Results: Ternary-Bons

You can now train models on your own AMD hardware! (3GB VRAM)

Hey guys, we collaborated with AMD to enable you to train, run, and deploy LLMs across nearly all AMD hardware including Radeon, Instinct, Ryzen, and data center GPUs. It works on Windows , WSL, and Linux and we have optimized ROCm builds for both training and inference. If you don't know about Unsloth , we're a fully open-source local UI that enables you to do pretty much anything with local models (RAG, chat, train, coding etc)! For those who don't have AMD GPUs and only CPUs, we still support

China’s AI models have Trump’s AI world at war with itself

This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Over the weekend, several current and former advisors to President Donald Trump on AI publicly lobbed insults at the country’s leading AI companies. David Sacks, the president’s AI and crypto “czar” until…

RecGPT-V3 Technical Report

Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pioneered this paradigm on Taobao by centering user understanding, and RecGPT-V2 scaled it via coordinated multi-agent reasoning; both are deployed in production with consistent gains in user experience and commercial outcomes. However, operating RecGPT at scale reveals three challenges: (1) stateless behavior mo

AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery

arXiv:2607.15367v1 Announce Type: new Abstract: Desktop voice assistants are still dominated by cloud pipelines that ship raw audio off the machine and expose a fixed set of skills. We describe AnovaX, a small local-first assistant that runs entirely on the user's computer and treats the desktop itself as its action surface. A single Python process wires together a wake-word gate, a speech pipeline, an LLM planner (Gemini) that emits a JSON plan of tool calls, a whitelist-and-denylist safety lay

ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning

arXiv:2607.15660v1 Announce Type: new Abstract: While LLM agents demonstrate strong reasoning abilities in compact and well-defined scenarios, they struggle to maintain robustness and effectiveness when faced with large-scale, diverse, and dynamic real-world environments that demand seamless tool integration. To address this gap, we introduce ToolVerse, a comprehensive framework that scales up agentic RL environments and enables agents to perform complex long-horizon reasoning in Tool-Integrated

GraphDx: A Cost-Aware Knowledge-Enhanced Multi-Agent Framework for Sequential Diagnosis

arXiv:2607.15280v1 Announce Type: new Abstract: Sequential diagnosis requires balancing diagnostic accuracy against resource costs through iterative information gathering. Existing Large Language Model (LLM) approaches exhibit a critical knowledge-reasoning gap: despite encoding extensive medical knowledge, they struggle to reason systematically under cost constraints, often resorting to excessive testing. We propose GraphDx, a knowledge-enhanced framework with two core innovations. First, we de

DeepLoop: Depth Scaling for Looped Transformers

Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters. This reuse changes the residual-scaling problem: in an untied Transformer, each residual branch receives and applies its own parameter update, whereas in a looped Transformer one shared update aggregates gradients from repeated visits and is read back by those same visits in the next linearized forward pass. We formal

Running a 13M ASR conformer on a microcontroller

Hello everyone, I wanted to share a recent project of mine, which brings a 13.1 million parameter convolution transformer model to a Thanks to quantization, this model now fits into 14mb of flash memory and it now sits at 256kb of SRAM as well as 4mb of PSRAM to transcribe 8 seconds of audio. The speed is still painfully slow. It is lightning fast compared to my initial attempt however, which took 10 minutes of inference time to transcribe 5 seconds of audio. I also gave the whisper tiny model a

[Paper] xHC: Expanded Hyper-Connections - Scale Residual Streams Wider · Push Model Intelligence Further

Hyper-Connections (HC) expand the residual stream of Transformers into N parallel streams, providing a form of memory scaling beyond model width and depth. Manifold-Constrained HC (mHC) stabilizes this formulation at scale. The large gains from N{=}1 to N{=}4 suggest residual-stream expansion as a promising scaling axis. However, existing HC-family methods typically stop at N{=}4. Our experiments reveal why: scaling mHC beyond this point yields diminishing performance gains and rapidly increasin

With all the Kimi drama I feel like I want to download all the current best models in case there is a ridiculous knee jerk political move pulled

I haven't kept up since around February so I'm just not even sure... and there are quite a few options. I don't care about parameter size, from tiny to huge, what matters most is performance, I just want all the best safely locally stored, I'll worry about running them later. So, what do you consider some of the best of the best currently? Whether highly specialized, giant do everything well, or anywhere in between Edit: Also what you use any specific models for or the best you've found for any

SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction

arXiv:2607.15550v1 Announce Type: new Abstract: Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex tasks, yet they introduce critical safety risks where a single erroneous action can lead to irreversible consequences. Existing safety mechanisms are primarily reactive, lacking the ability to assess risks before execution. In this paper, we introduce SeerGuard, a consequence-aware safety framework designed to mitigate these risks through pr

InklingProduct Hunt1 minAI开源

Open weights 975B multimodal model built for fine-tuning Discussion | Link

My thoughts on qwen 3.8 so far with agentic coding.

So, im 50 and autistic. I learned C++ in college when I was young and didnt really ever use it that much. I also learned a bit of assembly from using cheat engine for game hacking (dont worry! single player or private servers!) which meant I got a little LUA experience. I've always disliked coding. Its such a "bang head against wall" type activity but I got into it because its pretty cool being able to make a computer do what you want. This entire attitude died a LONG time ago because coding was

From Black Box to Executable Logic: Explainable Reinforcement Learning through Prolog Expert Systems

arXiv:2607.15459v1 Announce Type: new Abstract: A trained deep reinforcement learning policy is a black box, and we ask whether it can be made explainable by rewriting it as an executable logic program that reproduces its behaviour and that a person can read, a logic engine can run, and an optimizer can edit. We present a three-stage post-hoc transformation that extracts a frozen proximal policy optimization teacher, induces an ordered rule list from its decisions in the manner of classical rela

Show HN: A canvas-based note taking and organizer app

So I've been working on this app for a very long time now. Started off by wanting an app that separates long form notes from short form notes visually and keeps everything organized on a canvas. The app features: -Sticky notes for quick, small notes and A4 notes for long form documents -A quick mode that uses local storage to quickly open the app and jot down a thought on the canvas -A PDF export option that exports your notes into 3 custom designed PDF styles -Visual hierarchy to replicate file

The Download: AI hiring biases, and weather data sabotage

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. AI is more likely than humans to form biases when hiring The next time you apply for a job, AI may screen your résumé before any human sees it. But there’s…

Quoting Sam Altman

We have been having extensive discussions around open source strategy. We will discuss it more at our next board meeting, but one thing we’d like to do soon is to create a language model with the approximate capability of GPT-3 that can run locally on consumer hardware and release that. We’d like to do it soon, before Stability or someone else does. In general, we think this helps discourage others from releasing similarly-powerful models, and makes it harder for new efforts to get funded. — Sam

CreedProduct Hunt1 min开发工具AI

Your personal context file for every agent. Discussion | Link

given the increasing likelihood of an open source AI ban, what are the alternative channels for downloading models?

the open ai exec in his "ai communism" post suggested a fraudulent FUD campaign and trump executice order against open source / chinese models. unfortunately that seems pretty likely to happen. i have always use huggingface for downloading models and they would be forced to comply. are there any established alternatives or mirrors that aren't us-based?

每天早晨,一份为你精选的科技日报