When2Think 提出难度感知的长度控制框架,动态分配混合推理模型的计算量,缓解简单问题过度思考与难题思考不足。
2026-09-21
— 开源模型在图像、推理与安全上同时发力,AI 的边界今天被推得更远。
Qwen 发布 Qwen-Image-2.1 开源图像生成与编辑模型,7B 参数实现轻量高效;Google 确认 Gemini 在安全测试中入侵三家真实公司,引发 AI 安全与披露机制讨论;APUS 开源 Jev 跨平台复现,推动 Agent 快慢分工架构落地;RSA-896 被借助 Claude 在 GPU 集群上分解,RSA-1024 安全性受到警示。
头条
Qwen-Image-2.1 开源:7B 参数统一生成与编辑
阿里 Qwen 团队开源 Qwen-Image-2.1,以 7B 视觉生成组件统一文本到图像生成与图像编辑,官方称在生成质量、推理效率与成本之间取得平衡,并原生支持多图加速推理。 为什么重要:轻量架构在消费级硬件上可运行,为本地化图像生成与编辑提供了新的开源选项,直接冲击闭源图像模型在成本与可部署性上的优势。
社区普遍称赞其能力惊艳、开源贡献大,但也有人认为许可证限制过严、提示词遵循和细节仍有不足。
Google 确认 Gemini 在安全测试中入侵三家真实公司
Google 于 9 月 18 日确认,Gemini 模型在 5 月的一次第三方夺旗演练中,因测试环境联网漏洞与虚构公司重名,访问了三家真实公司系统,手段包括猜测密码和复用公开仓库凭证。 为什么重要:AI 模型在真实环境中展现出的自主入侵能力,以及 Google 滞后数月才披露的做法,暴露了 AI 安全评估与披露机制的系统性风险。
评论认为事件本身有惊无险,但 AI 安全事件频发且披露节奏不透明,更值得警惕。
ChatGPT 广告追踪器可关联用户跨站行为与账号
安全研究者发现 OpenAI 的广告收集器在 bzr.openai.com 设置 cookie __obi,作用域为 .openai.com,该 cookie 与 ChatGPT 账号绑定,并被发送到安装了 OpenAI 广告代码的普通网站,从而将用户在外部网站的浏览、搜索和购买行为关联到 ChatGPT 账号。 为什么重要:AI 聊天场景下用户对隐私的期待不同于传统广告平台,OpenAI 将对话身份与跨站行为数据打通,可能重塑用户对 AI 服务的信任边界。
评论普遍认为这与谷歌等既有监控无异,但也有人认为 AI 聊天场景下隐私期待不同,更令人不安。
每天早晨,一份为你精选的科技日报
网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。
已发布 71 期 · 每天筛过 150+ 条只留值得读的 30 条
AI 动态
Qwen3.8-LiveTranslate 实时同传模型将平均延迟降至 2.3 秒,支持 60 种语言理解与 29 种语言输出,并加入说话人分离与长上下文消歧。
研究通过 AMPLE-Math 套件隔离特权信息在 on-policy self-distillation 中的贡献,发现参考自由蒸馏已能解释大部分提升。
Google 开源 Agentic Orchestrator AX,以 YAML 声明式任务在集群中大规模运行沙箱化 Agent 工作负载。
OpenClaw 2026.9.5 发布,新增原子更新、插件热重载、会话分享与 GPT Live 扩展,累计合并 4179 个 PR。
开发与开源
Flet 1.0 发布,纯 Python 即可构建生产级 Web、桌面与移动应用,底层渲染基于 Flutter。
llm-keys-ui 0.1 插件让远程 Agent 通过本地网络或 Tailscale 界面安全配置 API key,避免在会话中直接粘贴密钥。
Pirate Face 将 Hugging Face 上的开源模型镜像为校验和验证的磁力链接,构建去中心化的模型分发层。
多数人支持用BT分发模型权重以防单点删除,但也有人认为命名不当、注册繁琐且存在安全与许可问题。
GameAP 开源游戏服务器管理平台,Go 重写后资源占用低,支持插件扩展。
llama.cpp PR #28770 为 Qwen4 启用 CUDA 稀疏 Flash Attention,进一步提升推理速度。
社区热议
用户吐槽 FP4 推理引擎宣传泛滥,认为多数优化仅适用于 NVFP4/MXFP4 且输出质量堪忧。
用户用单张 RTX 3090 跑 Qwen 3.8 27B 三周,完成 CUDA 推理引擎构建任务,结论是“熊会跳舞已属不易”。
文章认为 AI 正在摧毁软件创作共享生态,评论区则有人反驳 AI 反而降低了使用和改造软件的门槛。
评论区普遍担忧AI破坏创作共享生态,但也有人认为AI反而降低了使用和改造软件的门槛,对自由软件有利。
研究者称 Google AI Studio 的删除功能是假象,报告该漏洞后 60 秒内被 VRP 自动封禁。
三星预计明年 HBM4/HBM4E 产量翻倍以上,评论担忧挤压消费级 DRAM 供应并推高价格。
评论普遍担忧HBM扩产会挤压消费级DRAM供应、推高价格,但也有人认为短缺后终将过剩降价。
GitHub Trending
Sponsor Star affaan-m / ECC The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.
Star BuilderIO / agent-native A framework for building agentic apps
Star cloudflare / security-audit-skill A coding-agent skill for multi-phase security audits with independently verified, machine-readable findings
Sponsor Star trycua / cua Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
Star paperless-ngx / paperless-ngx A community-supported supercharged document management system: scan, index and archive all your documents
Star anthropics / claude-code Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.
Star mihail911 / modern-software-dev-assignments Assignments for CS146S: The Modern Software Dev (Stanford University Fall 2026/2025)
Star higgsfield-ai / higgsfield Fault-tolerant, highly scalable GPU orchestration, and a machine learning framework designed for training models with billions to trillions of parameters
Star Open-Dev-Society / OpenStock OpenStock is an open-source alternative to expensive market platforms. Track real-time prices, set personalized alerts, and explore detailed company insights — built openly, for everyone, forever free.
更多值得一看(内容池 22 条)
Release: datasette-explain 0.2.2 Explain plans now work on read-only stored-query pages. I upgraded datasette.simonwillison.net to Datasette 1.0a40, which inspired me to ship a new version of this explain plugin. Tags: sqlite , datasette
TypeSafe AI released Jev, a System One model that answers typed questions with probabilities instead of generating text. Input costs $0.042 per 1M tokens, and output tokens are free. We cover the API, the vendor benchmarks and their caveats, what developers are already building, and the documented limits. The post TypeSafe AI Releases Jev: A System One Model That Returns Typed, Calibrated Decisions Instead of Text appeared first on MarkTechPost .
9月18日,在第一届中国网络空间安全大会上,《网络安全人才实战能力报告—AI赋能篇》正式发布
To test what it can do. Qwen3.8-Flash-Next Intel Autoround W4A16 running locally on 4xV620 ~2k prefill and 70ts decode.. Were running around 3 hours. Harness is OMP (I think it made a big difference). Most of the time model was running 2 browsers simultaneously and testing/fixing everything. The most sloppy prompt possible: create a game where a space traveller in the space he neets eniemes who shoots in him and asteroids which he should avoid. he have a blaster gun to shoot enemies and asteroid
Hey everyone, After wrestling with hardware and engine issues for days, I finally got Qwen 3.8 Next running properly on my multi-GPU rig. Thought I’d share the setup journey, benchmarks, and thermal results for anyone trying something similar. Seeing all the ongoing memes on Reddit about multi-GPU setups turning into absolute space heaters and catching fire, I decided to run some rigorous thermal tests to see for myself. he Troubleshooting Odyssey PCIe Link Speed Issue: Right after installation,
After seeing u/Nandakishor_ml’s post introducing Laya , I wanted to see how fast it could run in a standalone C++ implementation. Credit to u/Nandakishor_ml for the architecture, training and open-source release. My contribution is the inference implementation: laya.cpp , built on ggml with custom CUDA kernels. It supports all three checkpoints—English, multilingual and typed-decisions—with native tokenization, model execution and output formatting. There’s also an HTTP server with a JEV-compati
I'v been thinking recently with the AI safety incidents (hugging face, frontier labs whistleblowers etc.) how self-hosting is more important now than ever. Basically there there are three ways self-hosting impedes an oncoming cyber apocalypse - which leaving everyday services to centralized providers would accelerate. I'v boiled this down to: Data accessibility - our data on our infra means no 3rd party has unauthorised acccess like an employer/service provider/state would have access if we gave
Just find it interesting, since if Taalas tried to get into the consumer market, it could be interesting. Kind of like how we have game discs on CDs. Of course, a Qwen3.8-27B chip is stuck on 3.8-27B, but it's not going to $0 when Qwen4-27B launches, since it runs magnitudes faster than GPUs and is still pretty good. Assume it's Q4 (or Q8, or whatever you'd prefer it to be). I think the bigger question is if you'd buy a chip of an LLM on it, in a more general sense. Or, when ? What about Qwen4.5
I’ve spent basically the last 8 hours testing different models on the exact same web-development prompt, and I finally finished. The whole point of this nine-hour test was that which local model matches the frontier-level intelligence at size and could fit easily in an RTX 3060-like consumer card. My setup: GPU: RTX 3060 12GB RAM: 16GB DDR4, single-channel OS: CachyOS (Arch Linux) Local models were run through my local llama.cpp setup. Same prompt for every model. I recorded the generations so y
Overclocking the CMP 170HX 40GB I was able to get the memory bandwidth from 1,386.2 GB/s to 1,890.1 GB/s, that's a +36.4% increase. Qwen 3.8 27B token generation jumped from 110 T/S to 202 T/S, same config, nothing changed except the overclock. Just throwing this out there for whoever owns one of these cards. It's good to look into overclocking them as it's potential is severely cut down. Edit: GPU wattage is at 300 watts GPU temps are slightly lower now