PersonaDose 通过校准激活转向实现分级人格特征控制,在 Llama-3.1-8B、Qwen3-8B、Gemma-3-4B 上验证有效。
2026-10-04
— 主权模型与安全警报齐飞,今天 AI 圈既在秀肌肉也在敲警钟。
Aleph Alpha 发布 78B 开源 MoE 模型 Kolibri,主打欧洲主权与 1M 上下文;OpenAI 安全员工辞职并公开批评公司文化,引发行业安全讨论;一篇论文发现预训练 Transformer 推理深度严重不足,但一个 rank-8 LoRA 即可修复;Simon Willison 呼吁所有按量计费服务必须提供默认硬性预算上限,以防 AI agent 失控烧钱。
头条
Aleph Alpha 发布 Kolibri:78B 开源 MoE 模型,主打欧洲主权与 1M 上下文多源事件 ×3
Aleph Alpha 在德国统一日发布 Kolibri,一个英德双语 Mixture-of-Experts Transformer,总参数 78B、每 token 激活 3.46B,支持最长 1M token 上下文,权重以 Apache 2.0 协议在 Hugging Face 开放下载,并在德国与芬兰的基础设施上从零训练。为什么重要:这是欧洲在主权 AI 领域的一次实质性开源投入,技术报告长达 189 页,对关注模型训练管线、MoE 架构和合规场景的工程师有直接参考价值。
评论区肯定其开放性和技术报告透明,但也有人认为性能不及 Qwen3.8 27B,且“主权”说法存疑。
OpenAI 安全员工辞职并公开批评公司文化“已破碎”
曾负责撰写 OpenAI 主要模型发布安全报告的 David Robinson 本周辞职,并在《The Atlantic》发表文章称公司“culture is broken”,直言这是比单个产品更深的行业问题。为什么重要:安全团队核心成员的公开出走与警告,对依赖 OpenAI 模型构建应用的开发者而言,是评估模型治理与长期稳定性的重要信号。
评论区普遍认为此类警告虽已成“cliché”,但不代表可以忽视其内容。
论文发现 Transformer 推理深度严重不足,一个 rank-8 LoRA 即可修复
一篇新论文《Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It》指出,13 个预训练基础模型在上下文中可靠跟随引用仅 1.4-3.6 行,但在一个早期层加入任务训练的 rank-8 LoRA 后,Qwen3-8B 在 24 行链上的准确率从 15.5% 提升至 99%,更长训练的 LoRA 可达 50 行。为什么重要:该发现为推理优化和 agent 长程任务提供了极低成本的微调路径,且所有模型权重保持冻结,对资源受限的工程团队极具吸引力。
Simon Willison 呼吁:所有按量计费服务必须提供默认硬性预算上限
Simon Willison 发文指出,随着 coding agents 和个人 agent 大幅降低启动付费代码的门槛,所有 pay-by-usage 服务和 API 都需要提供“超过 $X/月即切断并返回错误”的硬性默认上限,软性警告邮件远远不够。为什么重要:对正在集成 agent 工作流的工程师来说,这是防止自动化系统意外烧穿云账单的关键工程实践,直接关系到生产环境的成本控制。
评论区讨论热烈,多数人认同硬上限的必要性,但也有人担心默认值设置过高或过低都会带来问题。
每天早晨,一份为你精选的科技日报
网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。
已发布 84 期 · 每天筛过 150+ 条只留值得读的 30 条
AI 动态
NEEDLE 提出无需训练的 LLM 后门移除方法,通过权重正交化抑制后门行为且不损害正常性能。
Hugging Face 发布多 harness RL 训练指南,基于 TRL 和 Harbor 框架解决开源模型在多种编码 harness 中的训练问题。
GPT-6 Astra 的 3D 建模能力引发关注,但专业 3D 生成公司 Meshy 的 ARR 不到两年从 100 万美元增至 1 亿美元,显示垂直模型仍有壁垒。
开发与开源
FTL 是一个面向云的新型操作系统,用户态 OS 以库形式构建,兼容 Linux 二进制,隔离性优于传统 monolithic kernel。
评论普遍认为FTL思路有趣、作者可信,但也有人认为其定位更像gVisor或微内核,且缺乏与Firecracker等对比。
社区讨论“过拟合推理引擎”的兴起:Strata、ninfer、DwarfStar 等专门针对少数模型和特定硬件优化的运行时正在涌现。
Anyworld 是一个自托管多人文字 RPG,本地 LLM 通过 llama.cpp 担任 Dungeon Master,也支持 OpenAI 等云 API。
开发者搭建 WoW 私服并构建浏览器客户端与 MCP agent harness,让 LLM 通过 WebSocket 控制角色游玩。
社区热议
联邦法官裁定 Flock 车牌搜索构成“无差别大规模监控”,但该裁决不构成有约束力的先例,社区对公共道路执法合法性存在分歧。
评论区普遍认为Flock属无差别大规模监控,应受严格监管,但也有人认为其在公共道路执法中合法有效。
Apple 修改 macOS 全盘访问权限以遏制 AI agent 滥用,Meta 声称 FDA 不足以阻止 Muse 读取消息,双方各执一词。
Meta 的 AI agent Muse 会为用户的亲友建立详细档案,研究人员已提取其内部操作指令,隐私争议持续发酵。
ICE 被曝使用 Palantir 数据库为抗议者建立档案,律师称此举侵犯第一修正案权利。
GitHub Trending
Sponsor Star DietrichGebert / ponytail Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
Star pbakaus / impeccable The design language that makes your AI harness better at design.
Sponsor Star affaan-m / ECC The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
Star Effect-TS / effect Build production-ready applications in TypeScript
Sponsor Star JuliusBrussee / caveman 🪨 why use many token when few token do trick. Viral skill + proxy for coding agents that cuts 65% of tokens by talking like a caveman.
Star Panniantong / Agent-Reach Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Sponsor Star thedotmack / claude-mem Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Star cloudflare / cloudflare-os Agent workspace built on Cloudflare Workers for creating documents, building apps, and running agents with your company’s context and systems.
Star addyosmani / agent-skills Production-grade engineering skills for AI coding agents.
更多值得一看(内容池 23 条)
Qwen Flash Next now uses less VRAM
IQ3_XXS weights are just under 80GB and my slowww DDR4+7900XTX is stabilizing around 45-70/s (sometimes higher while coding depending on mtp). Looking online I'm seeing similar results for users with 12GB and 16GB cards, and significantly faster numbers for owners of DDR5. (In comparison, Llama CPP with tuning was maxing out around 22.5t/s on the same rig. Quality seems reliably superior (I wouldn't recommend the Q2 weights though)) Seriously. Ask to set it up for your specs. If 27B doesnt fit w
EDIT: Same with the website. EDIT 2: New commits added on clarifying what you can and can't do. "Apache 2.0 in 2029" -> "Apache 2.0" now changed on the website. Also, u/jotkaPL (creator of Dockhand) wrote something along the lines of: "No worries, it will convert to Apache 2.0 in 2029" but later deleted the comment here . Original post: The most important lines have been changed: Previously New Change Date: January 1, 2029. Change License: Apache License, Version 2.0 Change Date: Four years from
I've noticed a trend with most new models with regards to their writing style. They are creating a new style, and this seems common among them. It's very information-dense. Here is an example from GLM 5.3 Flash. I'm gonna be honest here and say that my prompt was kinda silly; my prompt was 'Why wouldn't you just name your Chinese restaurant 'Chinese Food' instead of 'Ming Dynasty' or 'Szechuan Garden' or whatever?' the idea being that someone searching for 'Chinese food' on Google Maps would put
I love Qwen3.8 27B so much that I decided to show my gratitude to the Alibaba ecosystem by building a dedicated RAG/search node using a cheap Mini PC from AliExpress. Turns out, my ecosystem loyalty got rewarded with an absolute masterpiece of fraud: Promised: Intel N150 + DDR4/DDR5 Delivered: Core i3-7020U (2018 Kaby Lake, 2C/4T) + DDR3 1600MHz The Scam: The seller literally hardcoded New_N150 into the BIOS release string ( HSHW_M6_DDR3_EC_Intel_Com_New_N150_K001 ). So now my Qwen3.8 RAG stack
Just published this post about how we’re going to need default hard budget caps on pretty much everything simonwillison.net/2026/Oct/3/d...
The company's former safety lead said frontier AI model releases should have "layers of redundancy and careful, time-consuming planning."
Amazon praised for ending NDAs but slammed for downplaying data center pollution.
A video showing the September AI updates
Video generation models are increasingly being explored as world simulators for embodied planning and learning. To do so effectively, these models must not only generate visually appealing frames, but also predict how environments dynamically evolve when executing goal-directed actions. While evaluating these capabilities is crucial, existing benchmarks focus mainly on single short actions or step-by-step instructions. This leaves multi-step physical reasoning underexplored, especially in egocen
Meta's open-source kit for building your own AI gadgets Discussion | Link
CookTrace is a self-hosted recipe manager, pantry and shopping list, an alternative to Mealie, Tandoor and Paprika. AGPL-3.0, single Docker container, native Android app and a Wear OS app, no telemetry and no cloud sync. Your recipes are a SQLite file on your own machine. Part of the TraceApps family: NutriTrace (nutrition), CookTrace (recipes / pantry / shopping), LiftTrace (strength / lifting), NoteTrace (notes / tasks / reminders). New here? It keeps your recipes and cooks them with you. Impo
I recently finished The Principles of Diffusion Models , and honestly I think it’s exceptional. The authors strike a really good balance between mathematical rigor and intuition, with dedicated appendices for anyone who wants to go deeper into the math. It’s aimed at researchers, graduate students, and practitioners with basic deep learning knowledge, so you don’t need to already specialize in diffusion models (in my case, a strong background in Information and Probability Theory and a solid und