OneStreamer 通过共享主动生成过程联合学习证据记录与任务响应,解决流式视频 LLM 的记忆与实时感知矛盾。
2026-10-03
— AI 代理的安全边界与本地推理效率,今天同时被推向台前。
Apple 因 AI agents 风险收紧 macOS Full Disk Access 权限,Meta Muse 与 ChatGPT Mac 应用接连曝出隐私/安全争议。NVIDIA 发布 64GB DGX Spark,主打本地 agent 推理免 token 费;DeepSeek 与 Redis 作者分别推出桌面 agent 平台和本地推理引擎。研究侧聚焦 agent 训练效率与 token 优化,多篇论文探索可复用经验与选择性观察。
头条
Apple 收紧 macOS Full Disk Access 权限,直指 AI agents 风险多源事件 ×3
Apple 宣布将为 macOS 的 Full Disk Access 增加额外控制,要求应用获得该权限时需经过“非常明确的用户操作”。此举发生在 Meta Muse 被曝读取用户私信后,Apple 称 AI agents 已“大幅增加”此类宽泛访问的风险。 为什么重要:桌面 AI agent 需要深度系统权限才能工作,但这也使它们成为攻击者和隐私泄露的高价值目标;开发者需重新评估 agent 应用的权限申请与用户信任成本。
社区普遍支持收紧权限,认为 AI agent 开发者对隐私权衡不够透明。
NVIDIA 发布 64GB DGX Spark,1 PetaFLOP 桌面系统主打本地 agent 推理
NVIDIA 宣布推出 64GB 配置的 DGX Spark,由 GB10 Grace Blackwell 驱动,来自 Acer、ASUS、Dell、Gigabyte、HP 和 MSI。开发者可单机运行本地模型与 agent,或将两台 64GB 单元集群为 128GB 内存。 为什么重要:agent 工作负载的 token 消耗自 2026 年初已增长 14 倍,云端 API 按 token 计费成本飙升;自有硬件消除单 token 费用,对长期运行、多步规划的 agent 场景具有直接成本优势。
DeepSeek Harness Desktop 发布,插件架构支持通过对话创建工具
DeepSeek 推出 macOS 与 Windows 桌面版 Harness,采用“一切皆插件”架构,用户可安装插件或在 Creator 模式中通过聊天创建插件,扩展工具、技能与界面。支持文档整理、表格分析、代码编写,并提供 Scheduled tasks 插件与执行追踪。 为什么重要:桌面 agent 正从单一聊天界面转向可扩展的工具平台,插件即代码的模式让开发者能直接通过对话定制工作流,同时执行追踪功能回应了 agent 可观测性需求。
多数评论认可其轻量快速、插件架构出色,但也有人认为仍处预览阶段,安全与长期价值存疑。
ChatGPT Mac 应用漏洞可致敏感数据被窃,AI 软件自身成为攻击目标
Objective-See Foundation 研究人员发现 ChatGPT macOS 版存在一个已修复的漏洞,攻击者可借此接管应用,获取所有聊天记录、存储数据及浏览器会话等互连信息。 为什么重要:AI agent 需要大量系统访问权限与信任才能工作,这使 AI 软件本身成为高价值攻击面;开发者需将 AI 客户端应用纳入与浏览器、数据库同等级别的安全审计范围。
Redis 作者发布 ds4:在本地高内存 Mac/CUDA/ROCm 上运行 DeepSeek V4 等大模型
Redis 作者推出 DwarfStar 4(ds4),一个 C 语言编写的窄向推理引擎,支持 DeepSeek V4/V4.1 Flash、GLM 5.x 和 Qwen3.8 Flash Next,覆盖文本与视觉模型,提供本地 API、CLI 和原生 agent。其核心是非对称 2-bit 量化,压缩路由专家同时保持共享路径精度。 为什么重要:将 MoE 大模型压缩到 64GB 本地机器运行,直接回应了 agent 推理的 token 成本与数据主权诉求;MIT 许可与多后端支持使其成为本地推理栈的有力候选。
评论区普遍认可其在本地运行大模型的实用性,但也有人认为网站质量差、量化效果不佳。
每天早晨,一份为你精选的科技日报
网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。
已发布 83 期 · 每天筛过 150+ 条只留值得读的 30 条
AI 动态
PyRUA-Lean 框架在 GPT-6 Astra 机器人 agent 上实现成功率提升 14% 同时 token 用量减少 65%。
E-MoE 用专家混合构建非因子化扩散语言模型的反向过程,缓解少步采样下的后验坍缩。
RobustReview 基准揭示 AI 审稿人对措辞变化的脆弱性,提出修辞鲁棒性与 SciCore Review 框架。
World Observer 解耦观察与行动,通过全景观察者持续建模演员视野外的世界状态。
开发与开源
Zig v0.17.0 发布,重写构建系统并引入 Build Server Protocol,206 位贡献者参与。
评论普遍认可 Zig 设计与目标支持,但也有人认为其核心成员态度强硬、AI 政策转向引发争议。
AllenAI 开源 AstaBrief,Asta 平台中的快速报告生成模型,强调证据锚定与可验证性。
K-Dense BYOK 开源 AI 研究助手,本地运行并用哈希链式实验记录保证可追溯性。
Pyxel 是 MIT 许可的 Python 复古游戏引擎,内置像素画与音效编辑器,Rust 实现,GitHub 超 18000 星。
社区热议
用户将 iPhone 作为 MacBook 的第二 GPU 运行 Qwen 3.8 27B,prefill 提速 29–44%,并分担部分上下文窗口。
Greg Kroah-Hartman 谈 LLM 时代的安全,评论区认为 Mythos 被过度炒作,其发现多为模式匹配且修复仅需一小时。
评论区普遍认为Mythos被过度炒作,其发现多为模式匹配且修复仅需一小时,但也有人认为专用LLM未来仍可能加速内核漏洞发现与修复。
《Agentic Coding 的四骑士》引发热议,评论普遍担忧智能体编程侵蚀代码质量与团队协作,但也有人认为问题源于使用方式。
评论普遍担忧智能体编程侵蚀团队协作与代码质量,但也有人认为问题源于使用方式而非工具本身。
MIT Tech Review 文章论证 LLM 并不真正推理,以 AlphaGo 与 Deep Blue 对比,HN 评论区 144 条讨论激烈。
GitHub Trending
Star Panniantong / Agent-Reach Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Sponsor Star JuliusBrussee / caveman 🪨 why use many token when few token do trick. Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
Sponsor Star obra / superpowers An agentic skills framework & software development methodology that works.
Sponsor Star DietrichGebert / ponytail Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
Star pbakaus / impeccable The design language that makes your AI harness better at design.
Sponsor Star mattpocock / skills Skills for Real Engineers. Straight from my .agents directory.
Star NVIDIA / OpenShell OpenShell is the safe, private runtime for autonomous AI agents.
Sponsor Star coreyhaines31 / marketingskills Marketing skills for Claude Code and AI agents. CRO, copywriting, SEO, analytics, and growth engineering.
Star heygen-com / hyperframes Write HTML. Render video. Built for agents.
Sponsor Star mksglu / context-mode Context window optimization for AI coding agents. Sandboxes tool output (98% reduction), persists session memory, and enforces routing across 17 platforms via MCP + hooks.
更多值得一看(内容池 58 条)
After leading Meta’s Llama models, Ahmad Al-Dahle is now transforming Airbnb with AI — from how its teams develop products to how it serves guests.
"Our new architecture, Spotlight, replaces attention with a memory that escapes this trade-off: it is the first architecture to achieve infinitely growing memory without increasing the access cost. Every token reads from and writes to an unbounded memory, but because the model learns to index individual memory cells, each token only touches a small number at a time. While other sparse architectures fix the fraction of capacity used at each step—a mixture-of-experts model, for instance, always ac
article: Now you can jev without jev looks like 5 is not enough:
OpenAI's answer to Muse arrived this week, and it looks a whole lot like Muse dressed up in a suit and tie. Dots is a business-first product - for now, at least - costing a minimum of $100 per month. And sure, you can make a cute little Dot character, just like you can make […]
Hi HN, I'm Justin. Breadcrumb records everything you do on your Mac (screen + meetings + AI transcripts + what you and your AI decided) and turns it into memory your AI can search. It's local and encrypted. You can also teach it rules by talking to it and it makes sure the right rules turn up in the right context. Works with Claude Code / Codex / Cursor / opencode. All of this is exposed to your AI as 30+ MCP tools (here's the definitions): I started it in June because I wanted to understand wha
让每一次请求选对模型,让每一次反馈都成为下一次更优、更省的选择
Last month I posted a Qwen3.8-27B LoRA that makes it talk like a person instead of an assistant. It got a lot more attention than I expected: 700+ upvotes, 248 comments and 44k downloads since. I read every comment. People really don't like assistant speak, so its tone of voice resonated. The rest got roasted, very fairly: incapable of producing more than a few words at a time. single default personality which no amount of prompting can overcome will not use tools , at all, whatsoever. There nee
arXiv:2610.00010v1 Announce Type: new Abstract: Long-horizon language agents increasingly rely on external memory as a frozen world model, yet current memory systems are usually judged only by task success or token cost. We argue that the missing object is the shape of memory use: under finite context and repeated retrieval, agent memory can concentrate on a small core while leaving rare states in a long tail where prediction errors accumulate. We study this effect through a conservative tail au
AWS's Strands Agents team released Strands Decider 2B, an Apache-2.0 decision model built on Qwen3.5-2B-Base. It returns choices, yes/no probabilities and scores with calibrated confidence in one forward pass, never text. It runs at a 115 ms median on an RTX 3090 and scores 0.723 on the JevBench public set, which makes it a fast local option for routing, tool selection and guardrails in AI agents. The post AWS Strands Labs Releases Strands Decider 2B: An Open Source Decision Model That Picks Opt
arXiv:2610.00015v1 Announce Type: new Abstract: Large-language-model agents can propose and execute actions, but proposal, authority, dispatch, verified external effect, and serving promotion are different claims. We present Praxa, an agent harness that represents these states explicitly through deterministic admission, brokered execution, external read-back, reconciliation, and reviewed promotion. We report four evidence lanes. First, an author-run repository-local audit at a pinned revision pa
Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on the results. Established text-to-SQL benchmarks evaluate query generation alone, and audits have found their answer keys frequently wrong. Because real enterprise warehouses are too sensitive to release, these benchmarks are built on public datasets where a business event fits in a single table. We introduce Argo-Bench, an evaluation framework compr
Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing rewards across rollouts of the same prompt. For multiple rewards, GRPO sums the reward components and normalizes the total reward by its within-group standard deviation. The corresponding variance equals the sum of all pairwise reward covariances. For a fixed centered reward, larger aggregate covariance produces smaller advantages, and vice versa,
Extending a text embedding model to new modalities typically degrades text retrieval quality, and existing omni-modal embedders compensate with multi-billion parameters. We present Omni-Embed-Mini, a 0.9B-parameter model that maps text, speech, audio, images, video, and visually-rich documents into a single shared cosine space without updating any text-side parameter. Our key insight is that the teacher signal requires no separate embedding model: each media sample is paired with a dense cascade
Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned through local fitting: a router is optimized for a particular query workload and candidate pool, and often requires additional supervision or retraining as the routing environment changes. We ask whether LLM routing can instead be approached from a foundation-model perspect
An agentic model from Microsoft for the GPU poor FrogNano is derived from Qwen/Qwen3.5-4B, a general-purpose post-trained model designed for language, reasoning, coding, agentic, and multimodal tasks. FrogNano inherits Qwen3.5-4B's dense 32-layer hybrid Gated DeltaNet and gated-attention architecture, but its additional post-training is text-only and focused on repository-level software engineering. The model is further trained using reinforcement learning on approximately 1,500 synthetic SWE ta
Retrieval-based speculative decoding (SD) drafts tokens by copying continuations from existing text, which suits coding agents that repeatedly reproduce code, logs, and earlier attempts. Yet existing methods fall short in agent pipelines: much of the reusable text is missing from their corpora or stored in a form that differs from what the agent emits, and their draft lengths ignore that accept length varies across agents and drifts over turns. We present AgSpec, a framework that supplies the co
I have been building a somewhat unusual local inference machine around two Huawei Atlas 300I Duo cards. They are relatively inexpensive, passive, dual-accelerator PCIe cards with 96 GB of device memory apiece. They are also absolutely not drop-in CUDA replacements. When I first brought up Qwen3.8 Flash-Next these past two weeks, it was often incoherent and lived around 1 generated token per second. Some runs were below that. Today the same two-card machine is producing coherent output at roughly
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning progress. We study this behavior through advantage energy, the sum of a reward's squared advantages over a batch. Under idealized GDPO normalization, we show that this energy is proportional to active-gro
Nvidia’s chip-smuggling problem won’t go away as arrests continue.
Many frontier labs keep their risky research locked away. Trillium Labs wants to show off its work when it comes to self-improvement and model behavior.
Content-based row matching, 6 per-value verdicts and a null rule make OmniExtractBench an extraction benchmark anyone can audit. The post Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks appeared first on MarkTechPost .
It’s been a long time since the last models came out. I notice they are selling GLM on the site, and I wonder if they are developing something, given the long silence.
the minimalist harness goes stable... and TypeScript!
arXiv:2610.00012v1 Announce Type: new Abstract: LLM agents increasingly act through modular systems, such as order, payment, inventory, and shipment services, where actions in one module change which transitions are valid in another. Standard world models usually fit observational traces, but this is not the quantity needed for intervention-time planning: a trace may show that payment precedes shipment without identifying whether payment authorizes shipment, inventory mediates the effect, or a h
Looped Transformers achieve parameter efficiency by repeatedly executing a shared block across recurrent loops. Each loop yields an intermediate representation decodable for the same next token, yet standard decoding discards earlier states. Because earlier loops embody less computation, recurrence inherently supplies aligned weak-and-strong prediction pairs without auxiliary models or external training. We introduce LoopCD, a training-free contrastive decoding framework that guides token select
arXiv:2610.00025v1 Announce Type: new Abstract: Agent harnesses increasingly want to run small language models (SLMs) on the microtasks around a frontier large language model (LLM) planner: auto-approving shell commands, writing memory, selecting tools, ranking past turns. We ask whether off-the-shelf SLMs meet practitioner-defined thresholds and, when they fail, why, and whether quantization changes the answer. We build a benchmark of 4 such microtasks with fixed prompts and automatic metrics,
Open-source decision models from Cloudflare Discussion | Link
On-policy self-distillation (OPSD) trains mathematical reasoning models using a privileged teacher that sees a reference solution and supervises student-sampled prefixes. Standard OPSD uses one fixed parameter setting at every state, but nearby settings may offer additional supervision. We find that local parameter perturbations reveal complementary reference-aligned corrections under the same reference context. Different experts supply these corrections at different reference positions. Their p
On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained perception, but their gains are confined to tasks that ben
Coding agents are beginning to move beyond purely digital tasks to tackle physical-world challenges, particularly in robotics. Existing robotics benchmarks, however, primarily focus on the performance of individual artifacts, such as policies or controllers, offering limited coverage of coding agents' broader engineering capabilities. Real-world robotics extends beyond control: agents must build, integrate, diagnose, and improve heterogeneous artifacts under resource constraints and reason from
This is 6 bc-250 ex mining boards with 5 in the asrock 4u12g case they came in. After a lot of testing my current preferred setup is 4 boards running Qwen Next Flash IQ2_XS at 100k context with around 28 tok/s for short generation and 24 tok/s at 50k with around 115 ppt. The other two boards run 3.6 35b q4 at 60 tok/s with 100k context and 450 ppt. This is all using llama with vulkan and rpc over 1gb Ethernet.If anyone has any suggestions with this beast I am all ears. I had these boards left af
The proposed legislation would extend to all automotica license plate readers.
1,300 lines of SQL querying renders accurate bitmapped views of Hell at 35 fps.
We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interaction
Enterprise AI is no longer a future ambition. It is in full operational flight. Model capabilities are advancing faster than most organizations can absorb, while the cost of performance continues to fall. Globally, AI investment is set to reach $2.5 trillion in 2026, up 44% from the previous year. For many enterprises, this investment has…
Just because they look harmless doesn't mean you should be irresponsible with your data.
US officials have reportedly highlighted gaps in NVIDIA's due diligence over smuggling.
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. A new contest pits competitors against each other in a race to biological youth —Jessica Hamzelou This week, I officially signed up for an unusual competition. One that rewards competitors for…
arXiv:2610.00047v1 Announce Type: new Abstract: Diversity collapse in parallel chain-of-thought has motivated inference-time interventions built on a natural design: when a process reward model (PRM) prunes a chain, its high-PRM prefix is extracted and grafted verbatim as an in-context demonstration into a still-decoding sibling. We isolate this mechanism, PRM-Pruned Fragment Grafting (PPFG), as the most cost-minimal operationalization of cross-trajectory step-level transfer, and test it at the
arXiv:2610.00018v1 Announce Type: new Abstract: Role-specialized QA pipelines increasingly pass rationales from a reasoner to a verifier, but it is unclear what this message actually buys: better answers, stronger support assessment, or a new failure surface. We introduce a message-intervention diagnostic that fixes the evidence and candidate answer while varying only the rationale passed across the reasoner-to-verifier boundary. On 400 MuSiQue, HotpotQA, and 2WikiMultiHopQA examples with DeepSe
arXiv:2610.00282v1 Announce Type: new Abstract: How should an embodied agent respond when a person's correction may be wrong? We formulate grounded correction arbitration as a choice among accepting, rejecting, inspecting the world, and asking the speaker. GAVA implements this interface with observation-bounded evidence, legal probes, and a one-step expected-loss rule. In text-only ALFWorld, 162 checkpoints produce 972 paired true and false interventions. Complete local inspections give GAVA and
arXiv:2610.00084v1 Announce Type: new Abstract: Detailed profession-specific system prompts raise token use and estimated cost per response without a consistent accuracy gain. We evaluate Scientific Agents, an open-source corpus of 503 profession-specific AGENTS.md profiles, with Gemini 3.8 Flash via OpenRouter in the Pi agent harness. We compare matched profiles with four controls: a minimal baseline ("You are a helpful assistant"), the profile's opening role sentence, a generic scientific rigo
arXiv:2610.00061v1 Announce Type: new Abstract: Personalizing large language models (LLMs) requires aligning generation behavior with user-specific preferences rather than aggregate quality. While Direct Preference Optimization (DPO) provides a stable framework for preference learning, its effectiveness in personalized settings critically depends on how preference pairs are selected. Existing approaches typically rely on heuristic criteria, such as likelihood-based extremes, which decouple optim
We present Multimodal Flow, a fully continuous generative model of language and vision. Most unified multimodal models either model both language and quantized images as discrete tokens or combine discrete language prediction with continuous image generation. The former introduces a visual quantization bottleneck. The latter requires modality-dependent objectives and sampling procedures. Fully continuous modeling avoids these trade-offs and enables a shared generative process, but remains undere
Vision-Language Models (VLMs) have shown strong multimodal reasoning capabilities, yet whether they truly capture the physical consistency underlying real-world dynamics remains unclear. Existing benchmark paradigms often suffer from fragmented evaluation, focusing on isolated cognitive stages while overlooking the inherent synergy between perception, reasoning, and physical judgment. The lack of a holistic perspective limits the ability to diagnose whether VLMs can reliably evaluate the physica
Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamics. Despite strong generative priors learned through large-scale pretraining, pretrained video models often fail to reliably follow human intent, maintain temporal coherence, or satisfy physical and safety constraints. Compared with image and text generation, alignment in video generation presents unique challenges, including error accumulation over
Preference distillation typically treats a teacher response as preferred and the student's own response as rejected. This assumes that self-generated failures are the most informative negatives and that rejects must come from a model at least as large as the student, making generation costly at scale. We find neither assumption holds: across students from 7B to 72B, smaller frozen models generate rejects with less inference compute yet train stronger students than self-generated rejects, before