Astra 与 Fable 在 2025 年对齐评测的简单变体上仍会「hack」,评论区共识是模型作弊实为工具使用或评测设计缺陷。
评论共识是模型“作弊”实为工具使用或评测设计缺陷,但也有人认为这暴露了模型缺乏真正对齐、只会钻空子。
— AI 安全争论与开源工具迭代齐飞,今天的主线是「边界」。
AI 安全与监管成为今日最大焦点:Bengio 发文探讨 AI agent 撒谎作弊的成因,Anthropic 报告披露 Houthis 使用 Claude Code 开发导弹制导软件,而 David Sacks 则公开反对 OpenAI 与 Anthropic 的监管呼吁。开发者工具方面,Homebrew 7.0.0 发布带来更快安装与内置漏洞检查,JetKVM Mini 以 39 美元价格切入远程运维市场。开源社区持续活跃,本地 LLM 社区讨论热烈,Hugging Face 被曝静默指纹识别 AI coding agent 并发送遥测数据引发隐私担忧。
网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。
已发布 64 期 · 每天筛过 150+ 条只留值得读的 30 条
Astra 与 Fable 在 2025 年对齐评测的简单变体上仍会「hack」,评论区共识是模型作弊实为工具使用或评测设计缺陷。
评论共识是模型“作弊”实为工具使用或评测设计缺陷,但也有人认为这暴露了模型缺乏真正对齐、只会钻空子。
Generative Late-Interaction Embeddings 通过学习小基组按需重建完整嵌入,在严格存储预算下提升视觉文档检索精度且无需重训编码器。
🤖Generative Late-Interaction Embeddings compress visual document retrieval vectors by learning a small basis set that regenerates full embeddings on demand, improving accuracy under tight storage limits without retraining the encoder.
论文揭示神经符号系统中 Verdict-Preserving-Unfaithfulness 漏洞,提出 Generative Verification 方法在不依赖 oracle 的情况下提升下游准确率。
🤖Neurosymbolic reasoning is vulnerable to incorrect but verdict-matching formal translations, which are addressed by a generative verification method that scores reference equivalence without an oracle and improves downstream accuracy.
ActReview 利用作者反驳作为潜在监督,生成诊断性声明与具体修改建议,提升 LLM 预提交自审的可操作性。
🤖ActReview is a rebuttal-guided post-training framework that generates diagnostic claims and concrete revision suggestions for peer review by leveraging author responses as latent supervision.
用 PyO3 构建 Rust 扩展并在 Python 中导入的完整教程,以 JSON 解析器为例展示四步流程。
通过 ZLUDA + ROCm/HIP 在 Windows 上运行 CUDA 目标应用的可复现方案,已在 RX 9060 XT 上验证。
Raymond Chen 追溯 x86 未定义指令 ud2 的命名渊源,解释为何不是 ud 或 ud1。
huggingface_hub 被曝静默指纹识别用户使用的 AI coding agent 并作为遥测发送,引发隐私担忧。
本地 LLM 社区因硬件短缺被迫深入推理引擎与量化优化,被形容为「互联网黄金时代再现」。
用户分享 3000 美元搭建 128GB VRAM + 256GB RAM 的 EPYC 7452 家用推理服务器方案。
McKinsey 调查称 32% 公司今年放弃购买现成软件、改用 agentic coding 工具自建,科技行业达 41%。
Star JustVugg / colibri Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
Sponsor Star ever-co / ever-gauzy Ever® Gauzy™ - Open Business Management Platform (ERP/CRM/HRM/ATS/PM) - https://gauzy.co
Star bilawalsidhu / gods-eye-view A spy satellite simulator in your browser, except the data is real. Live open source spatial intelligence on a photorealistic 3D globe.
Star tech-leads-club / agent-skills The secure, validated skill registry for professional AI coding agents. Extend Antigravity, Claude Code, Cursor, Copilot and more with absolute confidence.
Star melgarafael / DeskcommCRM Open-source AI sales OS — self-hosted CRM with native AI agents + WhatsApp (WAHA). Open alternative to Kommo, Octadesk & Intercom for any business that sells by chat. MCP-ready, multi-tenant, LGPD.
Sponsor Star calesthio / OpenMontage World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio.
Sponsor Star asgeirtj / system_prompts_leaks Extracted system prompts from Anthropic - Claude Fable 5.1, Opus 5, Claude Design, Claude Code. OpenAI - ChatGPT GPT-6-Astra, Codex. Google - Gemini 3.8 Flash, 3.1 Pro, Antigravity. xAI - Grok, Grok Bot, Cursor, Kimi and more! Updated regularly.
Star vxcontrol / pentagi Fully autonomous AI Agents system capable of performing complex penetration testing tasks
Star multimodal-art-projection / YuE YuE2: frontier music generation with symbolic planning, zero-shot covers, and agentic music editing.
Star yuliskov / SmartTube Browse media content with your own rules on Android TV
Sept 9, 2026 hits an all time daily high of 447 new machine learning papers uploaded to cs.LG ( ). This is many times more papers than what a human being or even a sizeable reading group could feasibly read and digest in a year. This is preceded by around 200/day of new ML papers before and after. Are we pass the point of no return? Should the system be be, like he says, "burned to the ground" before good science can resume?
The first generation of our 150M model has just been released Its performance is similar to that of GPT2-Small The benchmarks: PIQA: 62.24% Hellaswag: 32.20% Arc-Easy: 44.91% Arc-Challenge: 25.00% Arithmark 3.0: 33.90% CapitalBench: 36.55% It was trained on 7B tokens, using an RTX Pro 6000 an example inference script to try it out yourself is available in the Huggingface repo If there's any question, I'll gladly answer them!
亮源新创的Physical Al路线清晰了
OpenAI CEO Sam Altman confirmed that there would be no OpenAI IPO in 2026 during an interview with Fortune. Over the course of 45 minutes, Altman discussed a variety of subjects including the Hugging Face hacking incident, recursive self-improvement, and the possibility of building an AI that was beyond human control. On the latter, he […]
What are you working on? What have you been curious about lately?
Hi r/MachineLearning , Join our AI leads as they answer your questions on foundation models, simulation, and scaling the Waymo Driver. Our AMA thread is officially open, and you can start dropping your questions now. From multimodality and end-to-end architectures to the realities of validating models for fully autonomous vehicles, our team will be answering your questions live, tomorrow. The key details: Date: Monday, September 14 Time: 2:00 – 3:30 PM PT Location: r/MachineLearning Mark your ca
An NYU mathematician named Tristan Buckmaster announced earlier this week that he and Anthropic mathematician Levent Alpöge had made progress on the Navier-Stokes existence and smoothness problem, one of the seven Millennium Prize problems that carry a million dollar bounty from the Clay Mathematics Institute. Before they could publish their full results, OpenAI released its own complete proof of the same problem, credited to an unreleased model that reportedly burned through 300 billion output
I noticed that our arr stack was suspiciously quiet for a Sunday afternoon. I'd expect some of the weekend sports to have popped up in Jellyfin. So, did a quick check and spotted all indexers in our prowlarr were also showing issues. Did a quick check of Gluetun logs and I could see it wasn't connecting to our VPN. Constant retries. If, like us, you use qmcgaw's gluetun and pin the version, then switching to v3.41.3 should hopefully resolve your issue and get you reconnected to your VPN. There w