DawnSift
订阅日报
周一 · 科技日报 · 第 57 期

2026-09-07

— 今天的主线:开源模型在本地跑出花样,OpenAI 则在研究加速上自曝家底。

今日 TL;DR

OpenAI 宣布达成“自动化研究实习生”目标,并公开内部编码 agent 加速研究的数据;GPT-6 Astra 在机器人操作上表现惊艳但成本与稳定性仍存疑。本地 LLM 社区围绕 Qwen 3.8 27B 的 abliterated 变体展开大规模对比评测,同时出现多个本地 agent 与推理优化实践。Nitter 在收到 X Corp 停止侵权函后经法律咨询决定继续运营,引发开源社区对平台垄断的讨论。

If you continue to add floors and rooms to a building forever, it will collapse. Software faces no such constraint. The code can always get worse. — Zach Kehs

头条

1

OpenAI 达成“自动化研究实习生”目标,公开内部 agent 加速研究数据

OpenAI 宣布已实现“自动化研究实习生”——一个能在人类指导下完成需熟练研究员数天工作的系统,并称正朝着 2028 年 3 月前打造“自动化 AI 研究员”的目标迈进。同时,OpenAI 发布内部数据,展示 coding agents 如何提升实验速度、任务复杂度与研究加速。 为什么重要:这标志着前沿实验室开始用 agent 自我加速研究,对软件工程师而言,意味着 AI 辅助研发的边界正从代码补全扩展到完整研究任务,值得关注其工程实践与可靠性。

Hacker News 评论区认可 agent 在研究加速上的潜力,但也有人对“自动化研究员”的可信度与安全性持保留态度。

2

GPT-6 Astra 机器人操作评测:碗任务 19/20 成功,拼图任务仅 2/20

GPT-6 Astra 在 YAM 机械臂上执行“拾取红块放入碗中”任务时,20 次试验成功 19 次,远超 Claude Fable 5.1 的 8/20,单次成本约 $0.94;但在拼图任务中仅成功 2/20,与 Fable 5.1 持平,最终插入步骤停滞。 为什么重要:这表明前沿模型在简单操作上已接近实用,但精细操作仍是瓶颈;对关注具身智能与 agent 的工程师,这是模型能力与成本权衡的典型样本。

多数人认可 Astra 在编程和游戏构建上的表现,但也有人认为其在机器人控制上成本高、效果差,且评测方法有局限。

3

8 个 Qwen 3.8 27B abliterated 变体对比:167 GPU 小时,11 天评测

社区对 Hugging Face 上 8 个 Qwen 3.8 27B 的 abliterated(去审查)变体进行对比,耗时 11 天、167 GPU 小时,通过权重比较与 KL 散度等方法验证其声称的去审查效果。 为什么重要:本地 LLM 用户对“uncensored”模型需求旺盛,但质量参差不齐;该评测为选择可靠变体提供了量化依据,也反映了开源模型微调生态的活跃度。

社区对评测结果兴趣浓厚,认为部分变体名不副实,但也有人强调 abliteration 会损害模型整体能力。

4

Nitter 与 XCancel 在收到 X Corp 停止侵权函后经法律咨询决定继续运营

X Corp 于 2026 年 8 月 24 日向 Nitter 项目发出停止侵权函,要求永久下架实例与仓库;经法律咨询后,Nitter 项目宣布将继续运营,更多细节将稍后公布。 为什么重要:Nitter 是注重隐私与性能的 Twitter 前端替代方案,其存续关系到开源社区对抗平台封闭与追踪的实践,也考验 AGPL 等许可证在现实法律压力下的韧性。

评论区普遍支持 Nitter 恢复服务,认为 X 垄断不公,但也有人认为法律风险仍存,需谨慎。

5

Asahi Linux 官方支持 M3 系列 Mac,GPU 与 DCP 仍待突破

Asahi Linux 宣布 M3 系列 SoC 支持已合并入安装器,摄像头、麦克风、USB 3 10 Gb/s、硬件视频解码(含 AV1)、WiFi、蓝牙等几乎全部可用;主要例外仍是完整 DCP 支持与 GPU。 为什么重要:这是逆向工程在苹果封闭硬件上的重大进展,为希望在 M 系列 Mac 上运行 Linux 的开发者提供了更完整的选择,但 GPU 缺失意味着 3D 加速与功耗效率仍受限。

评论区普遍赞赏 Asahi Linux 团队的逆向工程成就,但也有人认为支持不足、目标用户模糊,且担忧苹果法律风险。

每天早晨,一份为你精选的科技日报

网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。

已发布 58 期 · 每天筛过 150+ 条只留值得读的 30 条

AI 动态

开发与开源

社区热议

How I feel about AI

一篇关于 AI 的复杂感受引发热议:对神经网络涌现推理的惊讶、对超级智能风险的恐惧、对 AI 公司摧毁开放网络的厌恶并存。

评论区主要围绕AI的恐惧与乐观展开,多数人担忧其社会冲击,但也有人认为AI将带来进步。

GitHub Trending

affaan-m/ECC★ 251280

Sponsor Star affaan-m / ECC The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

Sponsor Star mattpocock / skills Skills for Real Engineers. Straight from my .agents directory.

Star cathrynlavery / diagram-design 38 editorial diagram types for Claude Code, Codex, and Pi. Self-contained HTML + SVG. No shadows. No Mermaid slop.

openai/skills★ 25608

Star openai / skills Skills Catalog for Codex

Sponsor Star anomalyco / opencode The open source coding agent.

Star blader / humanizer Agent skill that removes signs of AI-generated writing from text

Sponsor Star llvm / llvm-project The LLVM Project is a collection of modular and reusable compiler and toolchain technologies.

Sponsor Star DietrichGebert / ponytail Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.

ruvnet/ruflo★ 70973

Star ruvnet / ruflo 🌊 The original agent meta-harness. Deploy intelligent multi-player swarms, coordinate autonomous workflows, and build conversational AI systems. Features adaptive memory, self-learning intelligence, RAG integration, and native Claude Code / Codex / Hermes and many more Integrated

更多值得一看(内容池 6 条)
Can we have a weekly "wins" thread?

Feels like AI has been especially taxing on the mental states of developers and the outlook of this industry in general (understandably so). Would anyone else be interested in a weekly thread to share your "wins" or other good things as a brief reprieve?

Reproducibility seems to be headed towards irrelevance in ML research. Is it too late? [D]

I feel that reproducibility is now a lost cause in machine learning research for three reasons: Many research is moving towards the physical AI territory, where you need expensive hardwares or even entire laboratories with high-speed cameras, in order to perform an experiment. You truly have no idea if the experiment can be reproduced and have to trust the demo. But demos are not perfectly reliable. Plus people are incentivized to only show the part of the demo that works. The entire system can

每天早晨,一份为你精选的科技日报