EnvHarness 通过可编程插件层动态重塑静态环境,针对 agent 弱点进行强化学习协同进化。
🤖EnvHarness and EnvRigger dynamically reshape static environments via programmable plugins to target agent weaknesses and improve reinforcement learning co-evolution.
— 今天的主线:Agent 的“外壳”比模型本身更受关注,从 harness 到环境生成都在抢戏。
Nvidia 研究显示,通过精心设计的 harness 和 supervisor 组件,Claude Opus 5 在 ARC-AGI-3 上达到 100% 得分,凸显 agent 框架的重要性。DeepSeek 发布支持视觉输入的 deepseek-v4-flash-vision-exp 模型,开发者可通过 OpenAI 兼容 API 调用。Anthropic 将 Claude Mythos 5 引入 Claude Security,企业团队可在不直接访问模型的情况下进行漏洞扫描。
Nvidia 发布研究,通过自定义 harness 优化内存管理并加入 supervisor 组件,让 Claude Opus 5 在交互式推理基准 ARC-AGI-3 上取得 100% 得分,即使底层模型本身并非最强。 为什么重要:对软件工程师而言,这意味着在构建 AI agent 时,工具编排、内存管理和监督机制的设计可能比单纯升级模型更能决定任务成败,值得重新审视 agent 架构的投入重点。
HN 讨论热度一般,但 TechCrunch 报道指出该结果对 OpenAI 构成压力,因为 ARC-AGI-3 是 OpenAI 尤为在意的基准。
Anthropic 于 2026 年 8 月 21 日起,将此前仅限受信任防御者的 Mythos 级模型 Claude Mythos 5 部署到 Claude Security 扫描中,面向 Claude Enterprise 客户公开测试,无需额外模型附加组件。扫描连接 GitHub 仓库,追踪跨文件数据流,返回带 CWE 分类、置信度和严重性评级及建议补丁的结果。 为什么重要:安全团队可以在不直接接触模型的情况下使用前沿模型进行漏洞扫描,且产品设计上限制了同一模型被用于编写 exploit 的可能性,这对企业安全工具链是一个值得关注的模式。
OpenAI 推出 Apple Messages 应用插件,允许 ChatGPT 在 Mac 上搜索消息、起草并发送回复,目前仅限 ChatGPT Work 和 Codex 用户使用,且为 opt-in 机制。 为什么重要:这标志着 ChatGPT 开始深入本地通信数据,对开发者而言既是生产力工具的扩展,也引发对隐私边界的关注——用户需主动授权才能让模型读取设备上的消息历史。
Proliferate(YC S25)发布开源自托管 AI IDE,可在同一工作区并行运行 Claude Code、Codex、OpenCode、Cursor、Grok 等编码 agent,每个任务获得独立的 git worktree 分支、终端和对话状态,支持子 agent 委派和定时工作流。 为什么重要:对需要同时管理多个编码 agent 的工程师来说,这种统一 harness 能降低 agent 编排的复杂度,worktree 隔离机制也提供了更干净的并行开发环境。
Show HN 上获得 35 分和 14 条评论,社区对多 agent 并行工作区的实用性表现出兴趣。
网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。
已发布 44 期 · 每天筛过 150+ 条只留值得读的 30 条
EnvHarness 通过可编程插件层动态重塑静态环境,针对 agent 弱点进行强化学习协同进化。
🤖EnvHarness and EnvRigger dynamically reshape static environments via programmable plugins to target agent weaknesses and improve reinforcement learning co-evolution.
FACET 框架在共享修复环境中保持源意图与可执行状态,为终端 agent 训练生成高质量可执行任务。
🤖FACET constructs executable terminal tasks by preserving source intent and grounding instructions, solutions, and verifiers in a shared repaired environment to enable scalable agent training.
SWE-bench Science 基准覆盖 20 个科学领域 119 个任务,揭示编码 agent 修复科学软件时的失败机制。
🤖SWE-bench Science benchmarks coding agents on scientific software repair, revealing failure mechanisms and mixed effects of scientific guidance.
MemTrapBench 发现检索到的记忆可诱导 LLM 推理错误和信念扭曲,并提出推理时策略避免认知陷阱。
🤖Retrieved memories can induce reasoning errors and belief distortions in large language models, and an inference-time strategy helps avoid these cognitive traps while maintaining benchmark performance.
Claudette 是一个 Claude Code skill,用 Gemini CLI 将 Claude 的 BuzzFeed 式回复翻译成正常英文。
llm 0.32.1 修复因 OpenAI Python 库弃用 httpx 导致的新安装故障,0.33 将切换到 httpx2。
llm-openrouter 0.7 兼容 LLM 0.32,改用 OpenRouter 的 Responses API,并新增 Shell、WebFetch、WebSearch 三个服务端工具。
Seed 是一个极简自修改 agent harness,只提供一个 exec 工具,让 agent 在 self/ 目录中逐步生长出工具、记忆和技能。
Bun 1.4 发布首个 Rust 重写后的稳定版,新增 Bun.WebView、Bun.Image 等 API,Node.js 兼容性大幅提升。
Encore 在 Apple Silicon 上重建 Linux MicroVM 栈,让 Mac 开发者能在本地运行 Firecracker 微虚拟机。
Kagi 新增自动移除付费墙链接的设置,评论区普遍赞赏其优于谷歌,但也有人希望保留优质付费内容或提供白名单。
评论区普遍赞赏Kagi移除付费墙链接功能,认为实用且优于谷歌;但也有人认为应保留优质付费内容或需白名单选项。
Anna's Archive 呼吁志愿者在 AI 公司销毁实体书前扫描稀有书籍,引发对训练数据与文化保存的讨论。
安全研究员披露通过接管 e164.arpa 域名意外记录数十万通军事基地电话的经过,引发对 ENUM 基础设施安全的关注。
OpenRouter 上线匿名 stealth 模型 Ox Alpha,评论区普遍猜测为中国 GLM 系列模型,认可能力但质疑匿名做法。
评论区普遍猜测Ox Alpha为GLM系列中国模型,认可其能力但质疑匿名做法,也有人认为可能是西方模型。
大厂工程师建议将自托管经历写入简历,认为其涵盖部署、监控、安全等真实工程问题的缩影。
基于官方 DeepSeek Harness 打造的 Electron 桌面端,深度适配 macOS 和 Windows,提供最佳的,开箱即用的体验。
29 editorial diagram types for Claude Code. Self-contained HTML + SVG. No shadows, no Mermaid-slop.
The DeepSeek adapter adds the multimodal visual understanding model DeepSeek-V4-Flash-Vision-Exp. It also supports configuring native image requests. Commands such as /goal and /plan can accept text and image input, and the @ menu can reference files and sessions; MCP/ACP also supports persistent image attachments, and PTC Mode supports forwarding nested images.
离开谷歌的原因之一:小团队可以极致聚焦!
Three ~300M drafters bring speculative decoding to LFM2.5, delivering up to 3.18x faster decoding with identical greedy output. The post Liquid AI Releases LFM2.5-DSpark Draft Models That Deliver Up to 3.18x Faster Decoding Without Changing Model Outputs appeared first on MarkTechPost .
I tried this model yesterday, and it felt to me like the best one I've tried for a local model for interactive use; the responses and reasoning are very fast, and it actually performs agentic tasks well. The speed is phenomenal. I am running this on Ninfer for Windows
A quick feedback after a really major test: nearly 20 hours of non-stop goal-oriented work with Qwen3.8-27B Q6, running across an RTX 3090 and an RTX 3060. It maintained a speed of around 60–63 tokens/s throughout the session.
Here's a brand new Blackwell-native, prefill-optimized 4-bit quant that runs 50% faster on compatible hardware than a Q4 quant of the same memory footprint. And it runs 4-7% faster than other NVFP4 quants as benchmarked on RTX 5090 32GB. Quant Benchmark Speed NVFP4 pp2048 6250 t/s unsloth NVFP4 pp2048 6010 t/s Q4_0 pp2048 4130 t/s Q6_K pp2048 3210 t/s This GGUF also includes a quantized MTP draft head for a good measure. Check it out for all details and specifically recommended settings for 15%
Run Claude Code, Codex or Opencode in cloud with your repo Discussion | Link
Artificial Analysis just benchmarked them and the scores are crazy good, proving the earlier success wasn't only enabled by overthinking.
Nvidia continues to pour money into data center development — just as AI data centers bring lots of money into Nvidia.
Google DeepMind partners with game studios to prototype breakthrough AI gameplay.
When the biotech company Insilico Medicine used its computer models to propose a promising drug for pulmonary fibrosis, it enthusiastically claimed in a press release that the molecule had been “discovered by” its generative AI platform. Insilico leads a pack of companies using AI to rapidly come up with drug ideas humans might never think…
Yes, we’re confused too.
arXiv:2608.18111v1 Announce Type: new Abstract: Foundation models such as GPT and Claude now solve olympiad-level mathematics with remarkable proficiency, so much so that geometry problem solving has become a standard proxy for their mathematical reasoning. Yet solving a geometry problem and drawing the figure it depends on are not the same skill: progress often hinges on a faithful diagram with the right auxiliary constructions and incidences, and it is unclear that a model which reasons its wa
ChatGPT search now uses the site:operator at scale Promptwatch is part of the emerging "GEO" space, for Generative Engine Optimization - the chatbot version of SEO, where companies offer tools and consulting to help your site increase its presence in replies to prompts inside tools like ChatGPT. The Promptwatch product uses automation to track responses to prompts across end-user chat products like ChatGPT, Claude, and Gemini. They publish aggregate reports on this as part of their own content m
Action-conditioned video world models require low-latency causal generation and reliable responses to game-native controls. Although causal distillation enables one- or few-step video synthesis, extending it to interactive world models remains challenging, as discrete keyboard states and continuous mouse motion must remain aligned with temporally compressed latent chunks during causal training and autoregressive rollout. We introduce ForgeWM, a progressive framework that transforms a bidirection
S1-mini is a 462 MB open-weights normalizer that sits after ASR, removing fillers and resolving self-corrections locally. The post Meet S1-mini: Superwhisper’s 462 MB Open-Weights Text Normalizer That Turns Raw ASR Transcripts Into Clean Written Text appeared first on MarkTechPost .
So usually I avoid Q3 quants because I have had bad experiences with it, models were usually too degraded, so the smallest I normally do is Q4, since I only have rtx 4060 ti 16gb. But since there hasn't been a 35b-3ab released yet, I had to try it. I don't use LLMs in agentic workflows, just on Textgen since I'm not a coder so this is not the primary use case of LLMs for me - but sometimes I really need some coding capabilities or help. I'm very impressed how it one shot multiple serious coding
LinkedIn says its AI slop button is working.
My Pro subscription expired today, they killed my access at 1pm local time. I'm now using Qwen3.8-27b w/ 5090m 24gb vram and pi to do everything i was doing in claudecode. The only downside is claudecode let me code without using my gpu, meaning I have to plan things now. Last night I had ChatGPT write up a prompt for a fancy aurora predictor for Canadians. I fed it to local pi and claude sonnet 5. They took about the same time, pi's app looked better, but claude's had better science. I asked th
Identity-preserving image generation becomes increasingly unreliable when a scene must contain many specified people. Beyond retaining each identity, the model must bind every reference to a distinct person and location, while training-time identity losses must establish correspondence among several noisy predicted faces. We introduce WithEveryone, a unified framework for generating group images up to ten reference identities. WithEveryone injects each selected identity as an addressed token, pr
I have been using omarchy on my tower since nearly a year now, shortly after it was released first. I really love the experience I am having with it but I still use my macbook for daly work, so I wanted to recreate a similar experience on it. Thats why I created omacosy, a setup for tiling windows, custom menu bar, some themes from omarchy, focus follows mouse, focus rings around windwos, some mac flavors with trackpad events and a custom mission control overview for your workspaces. I used Aero
Your personal AGI on your PC. Built to finish real work. Discussion | Link
arXiv:2608.18080v1 Announce Type: new Abstract: We present a review on the applications of large language models (LLMs) in health, e.g., social media analysis, clinical conversational agents, therapy support tools, prompt engineering, multimodal learning, and ethical considerations. We integrate findings from interdisciplinary studies utilizing diverse data sources such as social media posts, electronic medical records, and multimodal inputs to enable early detection of depression, suicide risk