Google 发布 Gemini Omni 1.1 Flash,场景扩展可读取 10 秒前文,支持首尾帧控制与 4K 超分。
— 腾讯开源 770B 巨模,OpenAI 与 Cursor 决裂,AI 圈今天火药味十足。
腾讯发布并开源 Hy4 Preview,770B 总参数、49B 活跃参数、1M token 上下文,社区已将其压缩至 200GB GGUF 且性能保持约 98%。OpenAI 宣布将于 2026 年 11 月 12 日切断 Cursor 对其模型的访问,原因是 SpaceXAI 收购 Cursor 后 OpenAI 不信任其会遵守服务条款。Sony Music 与 Warner Chappell 联合起诉 Anthropic,指控其大规模侵权训练 Claude,索赔金额可能高达数十亿美元。vLLM v0.28.0 发布,重点优化 Kimi-K3 与 DeepSeek V4 推理性能。
头条
腾讯开源 Hy4 Preview:770B 参数、1M 上下文,社区已压缩至 200GB多源事件 ×3
腾讯发布并开源 Hy4 Preview,总参数 770B、活跃参数 49B,上下文窗口超过 1M token,Hugging Face 上权重体积达 1.56TB。社区已将其压缩为约 200GB 的 GGUF 格式,并声称保持约 98% 性能。 为什么重要:这是目前开源模型中参数规模最大的之一,1M 上下文与 MoE 架构对长文档处理、代码生成等生产力任务有直接价值;社区量化方案则大幅降低了本地部署门槛。
评论区普遍认可 Hy4 的性能与性价比,但质疑其推理速度、开源定义及官方图表呈现方式,也有人认为其编码能力有限。
OpenAI 宣布 11 月 12 日切断 Cursor 模型访问,SpaceXAI 收购成导火索
OpenAI 宣布将于 2026 年 11 月 12 日停止向 Cursor 提供模型访问,原因是 Cursor 已被 SpaceXAI 收购,OpenAI 称基于与 Elon Musk 旗下公司违反合同的经验,无法确信 SpaceX 会遵守其服务条款。 为什么重要:Cursor 是当前最流行的 AI 编程工具之一,模型断供将直接影响大量开发者的日常工作流,也标志着 AI 基础设施层的竞争从技术延伸到商业与政治层面。
Sony Music 与 Warner Chappell 联合起诉 Anthropic,索赔或达数十亿美元多源事件 ×3
Sony Music Publishing 与 Warner Chappell 在加州北区联邦法院对 Anthropic 提起诉讼,指控其通过非法 torrent、爬取和下载受版权保护的作品来训练 Claude,要求每部侵权作品最高 15 万美元赔偿,总额可能达数十亿美元。 为什么重要:这是针对 AI 模型训练数据版权问题的最新且规模最大的诉讼之一,判决结果可能为整个行业的数据使用合规边界设定先例。
vLLM v0.28.0 发布:Kimi-K3 与 DeepSeek V4 推理性能大幅优化
vLLM v0.28.0 发布,584 个 commits 来自 270 位贡献者。Kimi-K3 获得 Decode Context Parallel 支持、融合 FlashKDA 内核、自适应投机 token 预算(DSpark TTFT 提升约 60%)等优化;DeepSeek V4 的 sparse MLA 已端到端支持 plain decode、MTP 与 DSpark 投机解码。 为什么重要:vLLM 是生产环境最主流的 LLM 推理引擎之一,对前沿 MoE 模型的持续优化直接决定了企业在有限 GPU 资源下的服务吞吐与延迟。
每天早晨,一份为你精选的科技日报
网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。
已发布 58 期 · 每天筛过 150+ 条只留值得读的 30 条
AI 动态
Terminal Bench 4.0 发布,GLM-5.3 与 Fable 5 处于同一水平(考虑误差范围)。
Google 论文提出 SKILL.state,通过追踪状态而非历史将长会话 agent token 用量削减 94%。
实验揭示本地部署大模型不如官方版的根因:推理软件栈微小差异(如注意力后端切换)即可改变输出 token。
Exo labs 声称通过 m5u Mac Studio 集群实现 4.8 TB/s 内存带宽,并强调延迟而非带宽才是 RDMA 集群关键。
开发与开源
TurboKV:异步嵌入式 Rust KV 存储,支持原子批处理、有序范围扫描与后台压缩。
StemDeck:免费开源本地 AI 音轨分离工具,可将音频拆分为最多 6 个 stems。
Firecrawl 开源 OCR It,20ms 将 PDF 转为 Markdown,比 Docling 快近 300 倍。
Samsung 在 Hot Chips 2026 展示 Processing-in-Memory 方案,在 LPDDR5X 芯片内集成 MAC 单元。
评论普遍认可PIM技术潜力,但质疑其实际应用、能效与软件适配,认为缺乏杀手级应用;但也有人认为这是未来方向。
Debian 投票允许在开发、维护与文档中“负责任地使用生成式 AI”。
评论区多数支持Debian允许负责任使用生成式AI,认为开发者应对代码负责;但也有人认为“负责任”定义模糊,担忧质量下降。
社区热议
用户分享在 16GB RTX 4070 Ti SUPER 上以 50 tok/s 运行 Qwen 3.8 27B、100k 上下文的配置方案。
GrapheneOS 团队称 Pixel 11 不再支持硬件内存标记(MTE),Google 为省钱砍掉了关键安全特性。
评论区普遍对Pixel 11取消MTE表示失望,认为其性价比低、升级有限,但也有人认为可转向其他品牌。
开发者反思 LLM 让自己失去“手艺感”:prompt-评估-调整的循环取代了真正的工程与设计。
评论认为 GDPR 本身合理但执行差,cookie 横幅是恶意合规,大企业受益、小企业受损。
评论区普遍认为GDPR本身合理,但执行差、cookie横幅是恶意合规,且大企业受益、小企业受损;但也有人认为其沦为数据交易合法化工具。
GitHub Trending
Agent skill for beautiful, verifiable architecture, workflow, sequence, data-flow, and lifecycle diagrams—self-contained HTML with motion and crisp export.
A spy satellite simulator in your browser, except the data is real. Live open source spatial intelligence on a photorealistic 3D globe.
Prompt as Code | GPT-Image2 工业级提示词引擎与模板库,470+ 个案例逆向工程,20+ 套工业级模板,并提炼出Skills,持续更新中
FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.
Beautiful, Modern & Opinionated Linux
更多值得一看(内容池 20 条)
Tencent’s 770B open model for long-horizon work Discussion | Link
You can beat SOTA Time Series Anomaly Detection methods with a 100 year old algorithm Time Series Anomaly Detection (TSAD) seems to be one of the hottest topics in NeurIPS, SIGKDD, VLDB etc. Many (perhaps most) papers evaluate on Paparrizos’ TSB-AD-M benchmark… However, I tested these benchmark datasets and found that in most cases I could beat the SOTA TSAD methods with a 100-year-old algorithm, simple Statistical Process Control (SPC). In the attached example, SPC gets perfect results. If we c
Have anyone tried this yet? Looks promising, seems too good to be true with no performance loss.
The other day I was trying out a distillation of DS4 Pro, and it came with MTP. It was slow as hell on my hardware, barely 2-3 t/s, BUT the speed got bumps every once in a while, and I noticed it was in moments like: United States of America First law of thermodynamics The enshittification of the internet Basically, every time a very predictable phrase came up, it was instantly written. A fun thing to see. But it also has me wondering - would MTP work together with n-grams? Since n-grams are Mar
Turn complex docs, tables & images into AI-ready data Discussion | Link
The new generation of data center systems is increasing efficiency with smarter traffic control instead of just more processor cycles.
Coding正在变成Al世界的数字执行力
In mid-August, Ramp published spending data collected from 70,000 U.S. companies: Fable 5 ,the most powerful and expensive model in Anthropic’s lineup, accounts for just 11% of what those businesses spend on the company’s tools. The remaining 79% is worth its weight in gold. With the new releases from Qwen and GLM, we are likely close to Opus 4.8, and certainly ahead of Sonnet and the other LLMs shown at the top of the image. The "anti-open-source crusade" therefore comes as no surprise: it is a
Folks! We're just 50 PRs away from more faster inference . Hopefully by end of year. Experts!, please chip in there. List of Open/Ongoing PRs(and also Discussions) related to CPU/RAM/Disk/Hybrid: [Discussion] RFC: MoE expert cache, VRAM caching of hot CPU-resident experts with hybrid hit/miss execution #24528 AVX2: Speed up large batch size prompt processing of IQ models #27402 llama: add Maple 20B-A1B ternary MoE architecture (CPU)- #27000 ggml-cpu: tiled mul_mat for k-quants- #27851 ggml-cpu:
Came across the subreddit /singularity the other days and many commented that they had 20+ years of experience and hadn’t written one single line of code since 2025. I feel living in a parallel universe, because at my work, no matter how we incorporate AI into our workflow (we literally tried every way people recommended), AI rarely produces the exact or the same quality of codes without decent amount of human developers’ intervention. And at the end of the day, the amount of time we spend revie
In this tutorial, we build an ensemble weather forecasting workflow with NVIDIA Earth2Studio. We install the required Earth2Studio components while preserving Colab’s existing CUDA-enabled PyTorch environment, load the FCN prognostic model, and retrieve atmospheric initial conditions from GFS. We then implement a custom wind-power diagnostic that converts 10-meter wind components into turbine capacity factors, along […] The post Building Custom Batched Ensemble Weather Forecasting with NVIDIA Ea
Investors, founders, and operators from across Europe arrived for the annual Nordic TechBBQ conference to talk about how humans can have agency over AI.
Google's Pixel 11 phone uses a Tensor G6 processor with a powerful TPU. How is it different from a GPU, and what does that mean in real-world use?
Testing 100 companies found privacy requests often led to confusion and dead ends.
Pollen Robotics, the Bordeaux robotics team at Hugging Face, opened pre-orders for Microduck — a 25 cm bipedal robot where every movement is a neural policy trained in MuJoCo and exported to ONNX. At $399, it puts the full sim-to-real loop on a desk: 15 motors, camera, LiDAR, two IMUs, and an Apache-2.0 training stack you can retrain yourself. The post Hugging Face Unveils Microduck: A $399 Open-Source 25 cm Biped You Train with Reinforcement Learning appeared first on MarkTechPost .
I've spent the last week or two pondering whether I am in the wrong, or the AI tools are really that uncontrollable, or maybe I am just losing my mind over it. Perhaps someone has similar experiences or found a way to actually do something about it. Thing is, I cannot keep up with all of those new "the" AI tools you are supposed to use to succeed that pop up every other week. Normally I just stick to the simplest things like a CLI agent for my daily work, because I didn't like the UX of AI IDE p