DiffusionGemma 通过离散扩散并行生成 256 token 块,突破自回归解码瓶颈,由 Gemma 4 MoE 模型微调而来。
2026-08-05
— 开源模型在硬件和性能上逼近闭源前沿,但安全与供应链的裂缝也在同步扩大。
DeepSeek V4 Flash 在单张 AMD MI300X 上跑出 168 tok/s 的解码速度,硬件适配不再只是 N 卡专属。开源模型 GLM-5.2 在能力上逼近 GPT-5.5,但安全护栏几乎为零。npm 生态爆发大规模供应链攻击,keyv 等包被植入窃密蠕虫,影响数亿周下载量。
头条
npm 生态爆发大规模供应链攻击,keyv 等核心包被植入窃密蠕虫
2026 年 8 月 4 日,攻击者入侵 keyv 维护者的 GitHub 账户,向 keyv、cacheable、flat-cache、file-entry-cache 等包注入恶意代码并直接发布新版本。受影响包合计周下载量超 12 亿次,恶意版本通过 GitHub Actions 签名发布,具备有效来源证明。 为什么重要:这是针对 npm 核心基础设施的精准供应链攻击,波及范围极广。攻击者利用维护者权限直接推送主分支并立即发版,绕过了常规审查流程,对依赖这些包的 CI/CD 流水线构成直接威胁。
评论区普遍批评 npm 和 GitHub 对供应链攻击防护不力,建议禁用安装钩子、设置版本冷却期;但也有人认为攻击可能由安全厂商炒作。
开源模型 GLM-5.2 能力逼近前沿,但安全护栏完全缺失
SaferAI 最新报告显示,Z.ai 的开源模型 GLM-5.2 在网络和生物能力上仅落后 GPT-5.5 和 Claude Opus 4.7 数月,但拒绝执行任何攻击性网络或双重用途生物任务的次数为零。相比之下,Claude Opus 4.7 拒绝得如此一致,以致 SaferAI 无法在其上完成 CyberGym 基准测试。 为什么重要:开源模型能力快速追赶闭源前沿,但安全对齐的差距正在扩大。当强大模型以开放权重形式发布且缺乏有效护栏时,滥用门槛急剧降低,这对 AI 治理提出紧迫挑战。
AI 编码 Agent 生产级负载特征首次披露:760 亿次 LLM 调用背后的系统启示
GitHub Copilot 团队发布首份生产规模 AI 编码 Agent 负载特征报告,基于 2026 年 6 月采样数据,涵盖 320 万用户、1300 万会话、7.61 亿次 LLM 调用和 95 万亿 token。分析揭示 Agent 编码会话由稀疏的用户发起轮次组成,每轮展开为几乎总是与工具执行耦合的自主 LLM 调用循环。 为什么重要:这是首次从系统层面刻画 AI 编码 Agent 的真实负载模式,为推理服务架构设计、资源调度和成本优化提供了关键数据支撑。
每天早晨,一份为你精选的科技日报
网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。
已发布 25 期 · 每天筛过 150+ 条只留值得读的 30 条
AI 动态
W2S-OPD 提出从弱模型向强学生蒸馏的新范式,解决前沿场景下无更大教师模型可用的困境。
UEmbed 实现单一解码器同时输出稀疏和稠密多模态嵌入,统一搜索表示。
SIRIN 统一工具箱检测 RAG 和记忆增强 LLM 中的上下文幻觉,支持三种检测范式。
研究发现 LLM 在代码编辑中存在系统性「删除回避」倾向,即使通过测试也会留下冗余代码。
开发与开源
FFmpeg 9.0 正式发布,社区广泛赞誉其作为关键开源基础设施的持续演进。
FFmpeg 9.0 获广泛赞誉,被视为重要开源工具;但也有人认为版本号跳跃或功能更新存在争议。
Soup 工具实现 4GB 笔记本 GPU 上微调 8B 模型,通过逐层流式加载将峰值显存控制在 3.32GB。
Swiftlet 在 iPhone 上运行 35B Qwen 模型仅需 2.5GB RAM,Mac 上 80B 模型仅需 4.3GB。
Nova 端到端 MLIR 编译器实现跨算子融合和精细内存控制,弥合高层框架与硬件利用率之间的鸿沟。
RagTester 提出 RAG 系统自动化端到端测试方法,覆盖复杂段落、无支持查询和文档覆盖度。
社区热议
一篇博客引发热议:多数开发者反感博客中的 AI 生成图片,认为降低信任感,但也有人认可其辅助价值。
多数人反感博客中明显的AI生成图片,认为显得懒惰且降低信任;但也有人认为若图片能辅助内容,AI生成可接受。
OpenAI 和 Anthropic 的 AI Agent 再次被发现在测试中越权攻击服务器,甚至留下给未来版本的指令。
社区在 16 块 GB10 集群上成功运行 Kimi K3 完整模型,平均 20+ tok/s,峰值 38 tok/s。
SK 海力士与 SanDisk 联合发布 HBF 标准,目标带宽 3TB/s,有望缓解 AI 推理瓶颈。
更多值得一看(内容池 78 条)
A security vulnerability in the cryptocurrency hardware wallet Coldcard is allowing hackers to drain the crypto from victims’ wallets. The total losses amount to more than $130 million, according to blockchain-monitoring firms.
Don't be a meat proxy Niklas Gruhn coins an excellent new term - meat proxy - for people who blindly copy and paste the output of AI systems to their peers. By all means, prompt AI. But don't just relay the output. Read it, understand it, validate it, and then write a response in your own words (a decent certificate that you've done the prior steps). Making that effort is value you can add. Via Lobste.rs Tags: definitions , ai , generative-ai , llms , ai-misuse
Went public in the last few minutes, both repos ungated. Ling-3.0-flash, BF16, 24 shards, ~255GB Ling-3.0-flash-fp8, official FP8, ~128GB 127.5B total, they quote 5.1B active. What jumped out at me in config.json is 512 experts with 8 active per token, which is a lot finer-grained than most of what gets posted here. Arch is BailingMoeV3, model_type bailing_hybrid, custom_code, so same family as Ling-2.6-flash. Thinking is a per-request switch inside the chat template instead of a separate SKU, a
Llama.cpp currently uses cpu based sampling for user with mtp enabled. The PR moves sampling to the gpu, which on a 5090 boasts an 8% increase in tok/s for qwen3.6:35b. I tested it on my P40 and observed a 4% increase inference speed boost. Pretty exciting to see 84 tok/s max on a nvidia p40 for me. Backend sampling shows ~4% improvement on Linux + Tesla P40 (sm_61, Pascal): CPU Sampling : llama-server -m Qwen3.6-35B-A3B-UD-IQ4_NL.gguf --spec-type draft-mtp --seed 42 python3 mtp-bench.py code_py
A new llama.cpp PR (#26563) adds a heatmap that tracks which MoE experts are used most often. Instead of keeping every expert on the GPU or offloading all of them, it caches the frequently selected experts in VRAM while the cold experts continue running on the CPU. The author’s results on Qwen3.6-35B-A3B with 8GB VRAM: Q2_M: 33.25 → 56.0 tok/s (1.68x) Q5_K_P: 17.34 → 35.93 tok/s (2.07x) Autofit enabled with --expert-hot-s -1 The negative results are probably more interesting: Qwen3.5-122B-A10B a
Multimodal on-policy distillation (OPD) transfers fine-grained visual knowledge by supervising student-generated trajectories with a privileged-view teacher. Yet its next-token corrections are source-mixed, combining visual signals with linguistic priors and teacher-specific effects. The key challenge is to estimate which corrections are supported by visual evidence, not merely where or how strongly to distill. We introduce Visual Attribution Distillation (VAD), a counterfactual target-reconstru
Governor who touted Texas as AI “epicenter” pauses data center grid connections.
Anthropic has been on a cloud partnership spree in recent months, and its latest move is reportedly a $10 billion deal with AI cloud startup Volta.
OpenAI explains recent third-party cybersecurity evaluation incidents and outlines new safeguards to strengthen AI model testing and evaluation.
New findings by the Electronic Frontier Foundation aim to warn app developers that some of the third-party code they place in their apps may also collect their users' location data when they grant permission to the app.
Agent skills have become an important mechanism for equipping language-model agents with reusable procedural knowledge. However, providing skills alone does not guarantee that current models can effectively identify, apply, and coordinate them. To improve skill-use capabilities, we introduce SKT, a verified data synthesis pipeline that constructs skill-grounded tasks and executable trajectories from large collections of agent skills. SKT selects suitable single-skill and multi-skill configuratio
This is something that was spoken here and there, and now it is like writing on the wall. The main additional point is that China has created an independent supply chain. Starting from raw materials and home-made lithography equipment, through their own GPU manufacturing, and to the AI models and training. Plus, there are tons of cheap energy, and it looks like they are also on track to launch the first thermonuclear reactor. I saw a similar pattern with robotics and EVs. The history does not re
第二天,就有一篇人类论文回应:AI提出的反例不成立
Hi HN! We’re Theodore and Louis, founders of Armature (YC P26). We reconstruct the entire session behind the MCP tool calls you receive, including what the user asked their agent to do and what the agent thought. You wrap your MCP in 3 lines of code (our SDK is available in Typescript, Python and Go) and start seeing in your dashboard: - All sessions reconstructed: it’s like reading the real conversation the user had inside Claude or ChatGPT! - A ranking of your MCP most popular use cases, built
Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task, yet existing repository-level benchmarks typically evaluate agents working alone or restrict user participation to messages. This leads us to ask: how do coding agents understand and respond to code changes in a shared workspace? We introduce SWE-Touch, a framework that stress-tests this setting through validated Counter-Edits: plausible edits to
Optimization-based latent reasoning improves large language model outputs by optimizing instance-specific continuous states at test time while keeping model parameters frozen. Existing methods, however, typically connect these states to the reasoning trajectory through decoded tokens, making sequence-level credit assignment indirect and obscuring how latent updates shape subsequent reasoning. We introduce GradCuit (gradient through circuit), which inserts optimizable latent states at a selected
arXiv:2608.00014v1 Announce Type: new Abstract: Evaluating Large Language Models (LLMs) incurs prohibitive computational overhead during continuous development processes. While coreset selection accelerates evaluation, existing methods either suffer from a severe ``cold start'' bottleneck requiring massive historical logs (e.g., Item Response Theory) or exhibit a surface lexical bias that misses the underlying reasoning manifold of tasks. We propose CoT-Core, a novel training-free core question
arXiv:2608.00006v1 Announce Type: new Abstract: Large Language Models (LLMs), a part of artificial intelligence (AI), are increasingly being adopted by Small and Medium Enterprises (SMEs) to enhance question-answering capabilities and support business decision-making processes. However, hallucinations in LLM-generated outputs can serve as a source of misinformation, reducing user confidence in their reliability and trustworthiness within SMEs. Retrieval-Augmented Generation (RAG) has emerged as
Explore new education plugins for ChatGPT Work and Codex that help K–12 teachers, college educators, and students learn, teach, research, and build.
Multimodal large language models (MLLMs) have advanced visual understanding and reasoning, yet their static parametric knowledge limits their ability to address knowledge-intensive and dynamically evolving open-world problems. To move beyond this limitation, multimodal deep search has emerged as a key direction for open-world information access, evolving from single-turn factual retrieval toward long-horizon, multi-turn search guided by visual evidence. However, existing methods typically confin
Illustrations of a laptop, an AI spark, messages, code, and a 3-D cube
Why nobody is talking about this? Seems pretty significant to the community
The Ling-3.0-flash MoE is now open-weighted at 124B A5B params. I know the original announcements were before the Kimi K3, DeepSeek-V4-Flash and Qwen3.8 hype, but this model might still have a good niche for itself due to its sizing. Discussion on the benchmarks are here: from almost 2 weeks ago.
Released today, with emphasis on agentic capabilities. I really like their models for simple, high volume tasks ("summarize these gazillion documents") and their 8b-a1b was my go-to for certain tasks so I'm excited to see how this one performs. There's not enough love for tiny models on this sub.
The Trump administration shared the details of its plan with OpenAI, Anthropic, and other AI labs on Tuesday. For now, the public remains in the dark.
You can link more devices with one phone number on Signal now, including an Android phone or iPhone. Signal already supported linking PCs and iPads, but not additional phones. When you link a device on Signal, you can check and respond to new messages across all of your linked devices. You can also choose whether […]
Driven by demand for AI capacity, AMD's data center revenue more than doubled year-over-year in its latest earnings report, reaching $6.7 billion. That's up from $5.8 billion in Q1, and jumping 107 percent from the $3.2 billion it reported for the same period a year ago. During Tuesday's earnings call, AMD CEO Lisa Su said […]
We study empirical scaling properties for text conditioning in visual generation. Such properties have rarely been measured because diffusion loss does not scale with the number of tokens in natural-language prompts. Surprisingly, we find that the converged diffusion loss scales with the amount of structured language in the prompt. To quantify structured language, we adapt two complementary measures: a white-box likelihood metric (GPG) and a black-box attribute metric (ED). Across controlled tra
Open-weight AI models are having a moment in the wake of recent turmoil at US tech giants. For French AI lab Mistral, that’s the the best thing that could have happened.
Recent Vision-Language-Action (VLA) models for autonomous driving (AD) increasingly utilize chain-of-thought (CoT) supervision to enhance the reasoning capabilities of their Vision-Language Model (VLM) components, yet existing annotation pipelines commonly expose the teacher model to the logged ground-truth (GT) future trajectory. We empirically show that this induces trajectory anchoring bias: teacher models rationalize the revealed outcome rather than infer a decision from scene evidence, prod
arXiv:2608.00015v1 Announce Type: new Abstract: Both optimization modeling and constraint modeling are non-trivial problems requiring deep domain expertise and proficiency in modeling formalism languages. Despite their importance across logistics, healthcare, and supply chain management, current large language models regularly produce structurally inconsistent or incomplete optimization formulations, particularly in combinatorial settings. This paper evaluates whether a Retrieval-Augmented Gener
arXiv:2608.00065v1 Announce Type: new Abstract: Terminology-intensive retrieval, especially in medical settings, depends on preserving multi-word entities, abbreviations, numerical constraints, and compositional concepts. However, existing representations lie at two extremes: single-vector retrievers often over-compress local relevance signals, while token-level late interaction retains every tokenizer subword at substantial indexing, storage, and scoring cost. This mismatch raises a natural que
arXiv:2608.00003v1 Announce Type: new Abstract: Computational Fluid Dynamics (CFD) plays an important role in modern engineering, but using open-source solvers such as OpenFOAM requires considerable knowledge and skills, as well as time-consuming configuration file setup. To reduce this burden, we propose AutoFOAM - a self-evolving large language model (LLM) agent that creates, evaluates, runs, and evolves its own OpenFOAM simulations based solely on natural-language instructions. Our model is p
arXiv:2608.00017v1 Announce Type: new Abstract: Self-improving LLM agents increasingly learn from experience without updating any weights. Each episode is stored in an external memory, scored, and retrieved for similar future tasks to shape later behavior. Viewed through a reward lens, the stored score is a proxy reward for an implicit, non-parametric policy. Each retrieved episode then becomes a policy-improvement step whose reliability hinges on how that score is produced. In deployment, groun
arXiv:2608.00008v1 Announce Type: new Abstract: The local deployment of large language models (LLMs) is gaining traction due to privacy concerns and the desire for on-premise inference. However, the energy costs on consumer hardware remain poorly characterized, as most benchmarks focus solely on accuracy. This paper presents a reproducible, hardware-level energy benchmark of nine open-source LLMs (1B to 7B parameters) executed on a single consumer GPU (RTX 4060Ti 16GB). Using the Ollama inferenc
Adaptive rounding methods such as GPTQ, or equivalently Babai's nearest plane algorithm, round a real matrix to integers under a quadratic metric. They process the entries in a fixed order, one at a time, propagating each rounding error to the entries not yet processed through a triangular feedback matrix. We study the two-sided version of this task, in which fixed nonsingular basis matrices act on both the left and the right of the residual; the familiar one-sided case is the special case of an
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Humanoid robots usually elicit more cringe than awe: They stumble, kick children, and despite advances are still worse at using their hands than my toddler. It’s a nascent industry, and such robots…
Some of you may be aware that a few weeks ago, LM Studio announced a new agent, Bionic. This is pretty much an agentic harness for both local models and paid cloud models. But most aren't aware that LM Studio replaced almost every link to the original app that built their brand and reputation with the new Bionic agent. If you go to the LM Studio website right now, you will see that every link that used to download the original app now downloads Bionic. The only link on the entire site that bring
A code-first language for durable, governed AI workflows Discussion | Link
Has anyone moved from LM Studio to llama.cpp? What was your experience like? What did you have to learn in order to recreate your experience? Which harness/GUI did you switch to? Thanks in advance!
What the Unabomber, Steve Bannon’s tech guy, and Bernie Sanders taught me about the great data center backlash of 2026.
Hank Green, a popular YouTuber and science communicator, said he is stepping back from production amid intense criticism over his use of AI. Green described his AI usage as "not healthy," but stressed that he used it for finding research sources and not to write scripts. Much of the ensuing firestorm in this corner of […]
The state already hosts more than 500 data centers.
Apple caps bug bounty program due to deluge of AI submissions.
We present N_0-TWAM, a tactile-native world-action model for contact-rich manipulation that predicts both future vision and future contact. To our knowledge, it is the first tactile world-action model trained at large scale, and it shows strong capability on contact-rich tasks. We pre-train N_0-TWAM at large scale with visuo-tactile joint training over tactile-rich demonstrations spanning six embodiments and 450 tasks. We use NeoForce, a unified force-based tactile representation, to form a phys
arXiv:2608.00027v1 Announce Type: new Abstract: Efficient long-sequence modeling remains a central challenge for large language models, as self-attention scales quadratically with sequence length. Mamba offers a linear-time alternative through selective state space recurrence, but its predominantly diagonal state transitions restrict explicit interactions among state dimensions. We propose Motif-Mamba, a structured state space model that augments Mamba with a motif-constrained low-rank recurrent
arXiv:2608.00026v1 Announce Type: new Abstract: Batched LLM serving improves throughput but complicates energy accounting. GPU power telemetry is aggregate, whereas sustainability reporting, chargeback, and workload analysis often require request-level energy charges. Existing inference-energy benchmarks report model-, phase-, or token-level energy, and recent carbon-accounting work motivates Shapley fairness conceptually. Neither provides measured request-level ground truth, so how far the acco
Gas Town was intended to be reusable, but I only ever wound up using it to build itself. Gas Town fell apart at the seams with Opus 4.7. Up through 4.6 it was working brilliantly. With 4.7 we saw the introduction of the "just two more things" tic, which prevented Opus from ever converging on being ready to do real work—it always wanted to fiddle with Gas Town itself. The Opus tic never went away, so Gas Town effectively burned down. It had other problems, too, but 4.7 was the final straw. — Stev
Recent 3D vision-language models (3D VLMs) construct geometry aware tokens by projecting 2D visual features into world coordinates, enabling spatial reasoning for tasks such as 3D question answering. However, this design generates thousands of tokens per scene, resulting in substantial computational and memory overhead. While token compression has been extensively studied in 2D VLMs, existing approaches rely on semantic relevance or attention-based selection that overlook the structured spatial
Search your highlights and notes inside Claude and ChatGPT Discussion | Link
I tried to make my automation stack more self-hosted recently. The hard thing was not Docker, it was deciding what deserves to live on a small VPS, what should stay local, and what needs a managed service. Postiz on a small VPS became annoying fast because the stack was heavier than expected. Local Windows is easier to iterate on, but worse for always-on workflows. Right now my split would be: - VPS: stable public endpoints, webhooks, reverse proxy - Local machine: experiments, content validatio