DawnSift
订阅日报
周日 · 科技日报 · 第 49 期

2026-08-30

— 腾讯开源 770B 巨模,OpenAI 与 Cursor 决裂,AI 圈今天火药味十足。

今日 TL;DR

腾讯发布并开源 Hy4 Preview,770B 总参数、49B 活跃参数、1M token 上下文,社区已将其压缩至 200GB GGUF 且性能保持约 98%。OpenAI 宣布将于 2026 年 11 月 12 日切断 Cursor 对其模型的访问,原因是 SpaceXAI 收购 Cursor 后 OpenAI 不信任其会遵守服务条款。Sony Music 与 Warner Chappell 联合起诉 Anthropic,指控其大规模侵权训练 Claude,索赔金额可能高达数十亿美元。vLLM v0.28.0 发布,重点优化 Kimi-K3 与 DeepSeek V4 推理性能。

OpenAI 在其博客中表示,做出这一决定是因为基于“与 Elon Musk 旗下公司违反合同的经验”,无法确信 SpaceX 会在其服务条款范围内使用其技术。

头条

1

腾讯开源 Hy4 Preview:770B 参数、1M 上下文,社区已压缩至 200GB多源事件 ×3

腾讯发布并开源 Hy4 Preview,总参数 770B、活跃参数 49B,上下文窗口超过 1M token,Hugging Face 上权重体积达 1.56TB。社区已将其压缩为约 200GB 的 GGUF 格式,并声称保持约 98% 性能。 为什么重要:这是目前开源模型中参数规模最大的之一,1M 上下文与 MoE 架构对长文档处理、代码生成等生产力任务有直接价值;社区量化方案则大幅降低了本地部署门槛。

评论区普遍认可 Hy4 的性能与性价比,但质疑其推理速度、开源定义及官方图表呈现方式,也有人认为其编码能力有限。

2

OpenAI 宣布 11 月 12 日切断 Cursor 模型访问,SpaceXAI 收购成导火索

OpenAI 宣布将于 2026 年 11 月 12 日停止向 Cursor 提供模型访问,原因是 Cursor 已被 SpaceXAI 收购,OpenAI 称基于与 Elon Musk 旗下公司违反合同的经验,无法确信 SpaceX 会遵守其服务条款。 为什么重要:Cursor 是当前最流行的 AI 编程工具之一,模型断供将直接影响大量开发者的日常工作流,也标志着 AI 基础设施层的竞争从技术延伸到商业与政治层面。

3

Sony Music 与 Warner Chappell 联合起诉 Anthropic,索赔或达数十亿美元多源事件 ×3

Sony Music Publishing 与 Warner Chappell 在加州北区联邦法院对 Anthropic 提起诉讼,指控其通过非法 torrent、爬取和下载受版权保护的作品来训练 Claude,要求每部侵权作品最高 15 万美元赔偿,总额可能达数十亿美元。 为什么重要:这是针对 AI 模型训练数据版权问题的最新且规模最大的诉讼之一,判决结果可能为整个行业的数据使用合规边界设定先例。

4

vLLM v0.28.0 发布:Kimi-K3 与 DeepSeek V4 推理性能大幅优化

vLLM v0.28.0 发布,584 个 commits 来自 270 位贡献者。Kimi-K3 获得 Decode Context Parallel 支持、融合 FlashKDA 内核、自适应投机 token 预算(DSpark TTFT 提升约 60%)等优化;DeepSeek V4 的 sparse MLA 已端到端支持 plain decode、MTP 与 DSpark 投机解码。 为什么重要:vLLM 是生产环境最主流的 LLM 推理引擎之一,对前沿 MoE 模型的持续优化直接决定了企业在有限 GPU 资源下的服务吞吐与延迟。

5

Anthropic 研究:Claude 以 4 美元/小时训练 Claude,跑赢 150 美元/小时人类研究员

Anthropic 发布研究,基于 Claude Opus 4.8 构建自动化对齐研究员系统 AAR,让 Claude 自主查论文、提方案、造数据、训练模型,在 10 类 AI 安全问题上全部找到改进方案,部分任务表现优于 28 名人类安全研究员。 为什么重要:这展示了 AI 自我改进的可行路径,对 AI 安全研究范式和自动化 ML 工作流都有深远影响,也意味着模型能力提升可能进入加速循环。

每天早晨,一份为你精选的科技日报

网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。

已发布 58 期 · 每天筛过 150+ 条只留值得读的 30 条

AI 动态

开发与开源

Samsung's Processing-in-Memory (PIM)

Samsung 在 Hot Chips 2026 展示 Processing-in-Memory 方案,在 LPDDR5X 芯片内集成 MAC 单元。

评论普遍认可PIM技术潜力,但质疑其实际应用、能效与软件适配,认为缺乏杀手级应用;但也有人认为这是未来方向。

Debian votes to allow "responsible use of generative AI"

Debian 投票允许在开发、维护与文档中“负责任地使用生成式 AI”。

评论区多数支持Debian允许负责任使用生成式AI,认为开发者应对代码负责;但也有人认为“负责任”定义模糊,担忧质量下降。

社区热议

You Know GDPR Is Good Based on Who Hates It

评论认为 GDPR 本身合理但执行差,cookie 横幅是恶意合规,大企业受益、小企业受损。

评论区普遍认为GDPR本身合理,但执行差、cookie横幅是恶意合规,且大企业受益、小企业受损;但也有人认为其沦为数据交易合法化工具。

GitHub Trending

tt-a1i/archifyHTML★ 99

Agent skill for beautiful, verifiable architecture, workflow, sequence, data-flow, and lifecycle diagrams—self-contained HTML with motion and crisp export.

A spy satellite simulator in your browser, except the data is real. Live open source spatial intelligence on a photorealistic 3D globe.

Prompt as Code | GPT-Image2 工业级提示词引擎与模板库,470+ 个案例逆向工程,20+ 套工业级模板,并提炼出Skills,持续更新中

FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.

更多值得一看(内容池 20 条)
Hy4 预览Product Hunt1 minAI开源
Hy4 preview

Tencent’s 770B open model for long-horizon work Discussion | Link

You can beat SOTA Time Series Anomaly Detection methods with a 100 year old algorithm [R]

You can beat SOTA Time Series Anomaly Detection methods with a 100 year old algorithm Time Series Anomaly Detection (TSAD) seems to be one of the hottest topics in NeurIPS, SIGKDD, VLDB etc. Many (perhaps most) papers evaluate on Paparrizos’ TSB-AD-M benchmark… However, I tested these benchmark datasets and found that in most cases I could beat the SOTA TSAD methods with a 100-year-old algorithm, simple Statistical Process Control (SPC). In the attached example, SPC gets perfect results. If we c

If your t/s is low enough, you can see speculative decoding with your own eyes

The other day I was trying out a distillation of DS4 Pro, and it came with MTP. It was slow as hell on my hardware, barely 2-3 t/s, BUT the speed got bumps every once in a while, and I noticed it was in moments like: United States of America First law of thermodynamics The enshittification of the internet Basically, every time a very predictable phrase came up, it was instantly written. A fun thing to see. But it also has me wondering - would MTP work together with n-grams? Since n-grams are Mar

How important is it for Chinese LLMs to reach the Opus 4.8 level?

In mid-August, Ramp published spending data collected from 70,000 U.S. companies: Fable 5 ,the most powerful and expensive model in Anthropic’s lineup, accounts for just 11% of what those businesses spend on the company’s tools. The remaining 79% is worth its weight in gold. With the new releases from Qwen and GLM, we are likely close to Opus 4.8, and certainly ahead of Sonnet and the other LLMs shown at the top of the image. The "anti-open-source crusade" therefore comes as no surprise: it is a

llama.cpp Open PRs list - CPU/RAM/Disk/Hybrid Related - Better for CPU-only & Hybrid inference

Folks! We're just 50 PRs away from more faster inference . Hopefully by end of year. Experts!, please chip in there. List of Open/Ongoing PRs(and also Discussions) related to CPU/RAM/Disk/Hybrid: [Discussion] RFC: MoE expert cache, VRAM caching of hot CPU-resident experts with hybrid hit/miss execution #24528 AVX2: Speed up large batch size prompt processing of IQ models #27402 llama: add Maple 20B-A1B ternary MoE architecture (CPU)- #27000 ggml-cpu: tiled mul_mat for k-quants- #27851 ggml-cpu:

Why does it feel like some people live in a parallel universe?

Came across the subreddit /singularity the other days and many commented that they had 20+ years of experience and hadn’t written one single line of code since 2025. I feel living in a parallel universe, because at my work, no matter how we incorporate AI into our workflow (we literally tried every way people recommended), AI rarely produces the exact or the same quality of codes without decent amount of human developers’ intervention. And at the end of the day, the amount of time we spend revie

Building Custom Batched Ensemble Weather Forecasting with NVIDIA Earth2Studio

In this tutorial, we build an ensemble weather forecasting workflow with NVIDIA Earth2Studio. We install the required Earth2Studio components while preserving Colab’s existing CUDA-enabled PyTorch environment, load the FCN prognostic model, and retrieve atmospheric initial conditions from GFS. We then implement a custom wind-power diagnostic that converts 10-meter wind components into turbine capacity factors, along […] The post Building Custom Batched Ensemble Weather Forecasting with NVIDIA Ea

What's the difference between TPU vs. GPU?

Google's Pixel 11 phone uses a Tensor G6 processor with a powerful TPU. How is it different from a GPU, and what does that mean in real-world use?

Hugging Face Unveils Microduck: A $399 Open-Source 25 cm Biped You Train with Reinforcement Learning

Pollen Robotics, the Bordeaux robotics team at Hugging Face, opened pre-orders for Microduck — a 25 cm bipedal robot where every movement is a neural policy trained in MuJoCo and exported to ONNX. At $399, it puts the full sim-to-real loop on a desk: 15 motors, camera, LiDAR, two IMUs, and an Apache-2.0 training stack you can retrain yourself. The post Hugging Face Unveils Microduck: A $399 Open-Source 25 cm Biped You Train with Reinforcement Learning appeared first on MarkTechPost .

Have been using genAI for a few years now and it still feels like a slot machine

I've spent the last week or two pondering whether I am in the wrong, or the AI tools are really that uncontrollable, or maybe I am just losing my mind over it. Perhaps someone has similar experiences or found a way to actually do something about it. Thing is, I cannot keep up with all of those new "the" AI tools you are supposed to use to succeed that pop up every other week. Normally I just stick to the simplest things like a CLI agent for my daily work, because I didn't like the UX of AI IDE p

每天早晨,一份为你精选的科技日报