FlashPrefill V2 通过均值修正稀疏注意力与优化 GPU 算子,在长上下文 LLM 服务中实现显著加速。
🤖FlashPrefill V2 improves long-context serving via mean-corrected sparse attention, optimized GPU operators, and framework integration, achieving large speedups over dense baselines.
— 今天的主线:AI 从“写代码”走向“接管工作流”,但信任与安全仍是最大裂缝。
DeepMind 校友创立的 Inherent 发布 AI agent Faraday,在复现科研论文任务上声称超越 Anthropic 与 OpenAI 的更大模型。Anthropic 被曝在 Claude Code 中服务端 A/B 测试降低 effort 等级,引发社区对计费透明度的质疑。MCP 发布新路线图,聚焦 agentic messaging primitives 与 server-initiated events。OpenAI 罕见呼吁加州加强 AI 安全法案 SB 53,与此前立场形成 180 度转变。
OpenAI 在 LinkedIn 发文称加州 SB 53“应被修订以扩大保障”,包括要求对训练或评估中的前沿模型进行监控,以及强化模型开发生命周期中的网络安全保护。该公司此前曾反对该法案。 为什么重要:前沿实验室主动要求更严格监管极为罕见,可能预示行业对 agentic AI 失控风险的共识正在形成,也会影响企业在合规与部署上的决策。
网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。
已发布 44 期 · 每天筛过 150+ 条只留值得读的 30 条
FlashPrefill V2 通过均值修正稀疏注意力与优化 GPU 算子,在长上下文 LLM 服务中实现显著加速。
🤖FlashPrefill V2 improves long-context serving via mean-corrected sparse attention, optimized GPU operators, and framework integration, achieving large speedups over dense baselines.
对三个前沿 MoE 模型微调低资源语言推理,准确率几乎不变,但推理语言与格式缺陷被 RL 修正。
🤖Fine-tuning large mixture-of-experts models on a low-resource language shifts reasoning into that language without harming accuracy, while reinforcement learning with verifiable rewards fixes formatting and leakage defects.
Latent Space 提出:自 2022 年起,机器学习管线每年有一个组件从人造翻转为模型造,模拟正成为新的扩展定律。
Guidelight AI Standards 研究发现,主流 AI 实验室几乎未公开 rogue model 的遏制响应计划,OpenAI 得分最高,Anthropic 与 Meta 最低。
Liquid AI 被曝即将推出 100B 参数 LFM 模型,社区对其架构速度与实用性表示期待。
llm 0.33 发布:升级 OpenAI Python 库 3.x,embed 命令支持 --key,prompt -t 可重复组合模板。
Dan Luu 撰文称软件没有理由再慢下去,LLM 已将性能优化的成本降低数个数量级。
评论普遍认为软件变慢源于优先级、激励和工程权衡,而非技术能力;但也有人认为AI优化有限,慢软件仍将存在。
《A Friendly Introduction to Racket》以教程形式介绍 Lisp 家族语言,强调代码即数据与宏系统。
开发者从零训练 250M 参数 LLM,30B token,量化后仅 60MB,CPU 上约 400 tok/s。
Munder Difflin 是一个在本地运行多个“克隆 agent”的 harness,评论区认为有趣但混乱难懂。
评论区普遍认为该项目有趣且具讽刺意味,但也有人认为它混乱无用或难以理解。
Level1Techs 论坛文章分析为何本地 LLM 感觉比基准测试笨,指出推理实现差异是关键。
从 ElevenLabs 到 NinetyNineLabs,社区调侃“数字+Labs”命名泛滥,认为跟风但易记。
评论区普遍调侃“数字+Labs”命名成风,认为跟风且缺乏新意;但也有人认为这是硅谷常见趋势,名字易记。
r/selfhosted 用户盛赞 Bookorbit 项目,称其细节打磨出色,是自托管软件中的精品。
基于官方 DeepSeek Harness 打造的 Electron 桌面端,深度适配 macOS 和 Windows,提供最佳的,开箱即用的体验。
29 editorial diagram types for Claude Code. Self-contained HTML + SVG. No shadows, no Mermaid-slop.
Strip multi-vendor AI provenance marks: Unicode text hygiene, statistical rewrite hooks, and C2PA/metadata from PNG/JPEG/SVG/PDF/DOCX/HTML/MD
You can find the changelog and source code here: Associated pre-build is here:
It was only a matter of time...
端侧部署解决了具身大脑能否装进身体的问题。那么,同一个「大脑」,如何快速适配工业、商用和家庭三类机器人呢?
将真实场景重建为持续更新、可计算的4D数字世界。
While studying Geometric Algebra I have built some interactive visualization to demonstrate how geometric transformations (rotation, scaling, translation) can be constructed by just composing reflections. Accepting reflection as the most elementary geometric operation was an eye opening moment for me. I think some of you might enjoy the interactive visuals. Comments URL: Points: 51 # Comments: 9
What models and configs are we using? Please share here On windows, I am using this copium pared down model with MTP disabled, q4 k/q4 v mmproj banished to CPU/RAM and a small ub to save whatever context I can (90k-100k) so everything stays in the vram If you are on linux or have an iGPU, you don't have to deal with windows eating 1.5 gb vram and so have more than 14.5 GB of VRAM to use and probably aren't in purgatory. @echo off .\ikllama\llama-server.exe ^ -m "D:\AI models\qwen3.8\Qwen3.8-27B-
Libredesk is a self-hosted customer support desk for email and live chat. It's fully open source under AGPL, with no paid tier or separate enterprise build, No feature paywalls. Features: Email inbox and live chat widget, with both landing in the same agent inbox. Help center with collections, articles, search, and per-language content. Autonomous AI agent that answers from your knowledge base and hands off to a human when it can't answer. Agent copilot for drafting replies, summarizing conversa
A Dutch data regulatory authority said that Uber has to pay 824.9 million euros for violating the GDPR.
Cheap energy, abundant land, and proximity to Beijing have turned a city in Inner Mongolia into a crucial hub for data centers.
The private Android-based OS will expand beyond Pixels next year.
I used LM Studio Bionic with Qwen 3.8 27B Q3_K_S with 57k context. It took a staggering 63 hours to finish coding. After the first prompt "Create a beautiful, relaxing flight simulator in a single HTML page" taking 47.8 hours, it created an html file that showed the title screen that said "press any key" but pressing any keys won't advance the game. So I wrote on the second prompt "It saids press any key to begin. I press any key but it doesn't work." It ran for 15 hours. Now I can fly. No plane
Another user posted the benchmarks for Qwen 3.8 27B today, and while I think Qwen 27B is a really powerful model, I can't help but notice just how meaningless these Artificial Analysis benchmarks are and I question why people still post this garbage and use AA scores as some kind of holy bible for comparing LLMs. According to their "Intelligence Index", a 27B model now beats DeepSeek v4 Flash and Pro, Kimi 2.7 Code, GPT-5.2, Opus 4.6, and also Sonnet 5. At some point we have to ask: What is this
Has anyone else noticed an increase in scanners/bots in the past ~month? For the past couple years I've had 2-3k hits a day from bots but lately there has been a steady increase in traffic looking mostly for php files. What I find strange is how much of this traffic is coming from MS and Google IPs. Do they not have any kind of monitoring on their cloud services? Having thousands of requests spamming every IP that responds should raise some flags. 20.24.67.246 Hong Kong Hong Kong Microsoft Corpo