DawnSift
订阅日报
周五 · 科技日报 · 第 4 期

2026-07-10

— 模型井喷,监管收紧,本地推理破门槛

今日 TL;DR

OpenAI 发布 GPT-5.6 三型号及 ChatGPT Work,并宣布成为微软 365 Copilot 首选模型;Meta 推出 Muse Spark 1.1,定价仅 $1.25/百万输入 token;GLM-5.2 (744B MoE) 在 25GB RAM 消费级机器上成功本地运行;NYT 指控 OpenAI 在版权诉讼中隐藏证据,要求法院制裁;欧盟通过 Chat Control 1.0,允许无嫌疑批隐私扫描。

头条

1

GPT-5.6 正式发布:三型号 + ChatGPT Work + 微软 Copilot 集成多源事件 ×7

OpenAI 推出 GPT-5.6 家族,包含旗舰 Sol、平衡 Terra 和经济 Luna,其中 Sol 在编码、网络安全等基准上超越前代与竞品。同时发布 ChatGPT Work agent,基于 Codex 技术,可跨应用完成复杂任务,并宣布 GPT-5.6 成为微软 365 Copilot 首选模型。 为什么重要:软件工程师可立即低成本试用新模型(Sol $5/$30 每百万 token),ChatGPT Work 将 agent 能力扩展至非开发者,微软集成意味着大规模企业部署即将落地。

社区普遍认可成本效率提升,但部分开发者认为编码能力未达预期,对营销宣传存疑。

2

Meta 发布 Muse Spark 1.1:多模态 agent 模型,定价极具竞争力多源事件 ×4

Meta 推出 Muse Spark 1.1,面向 agentic 任务的多模态推理模型,在工具调用、计算机使用和编码能力上有显著提升。开放 Meta Model API 预览,定价 $1.25/百万输入 token,支持 1M token 上下文窗口。 为什么重要:以极低价格进军 AI 编码市场,可能推动整体降价;1M 上下文适合复杂 agent 工作流,对开发者的 agent 架构选择产生直接影响。

评论区认可性价比,但对其基准测试公正性与可用性存在分歧。

3

GLM-5.2 (744B MoE) 在 25GB 内存消费级机器上成功本地运行多源事件 ×3

开发者 JustVugg 发布 colibrì 项目,以纯 C、零依赖的方式在约 25GB RAM 的消费级机器上流式运行 GLM-5.2(744B 参数 MoE 模型),仅 9.9GB 驻留内存,其余专家按需从磁盘加载。 为什么重要:大幅降低本地运行超大模型的门槛,无需高端 GPU 即可使用 744B 级别模型,为边缘计算和隐私敏感场景提供可能性。

高度赞赏工程实现,但也有人质疑实际推理速度与应用价值。

4

NYT 指控 OpenAI 在版权诉讼中隐藏证据,要求法院制裁

纽约时报等新闻机构在版权诉讼中提交制裁动议,称 OpenAI 多年谎称无法搜索训练数据以隐瞒侵权证据,其隐私工程师在作证中暴露矛盾;OpenAI 被指隐藏了数十亿条聊天日志。 为什么重要:若法院制裁,OpenAI 可能需交出关键数据,直接影响 AI 训练合理使用的法律边界,对整个行业的大规模爬取行为具有深远影响。

社区观点两极,有人支持版权方,也有人认为这是对 AI 发展的阻碍。

5

欧盟通过 Chat Control 1.0:允许无嫌疑前提的私聊批量扫描

欧洲议会以特殊多数通过 Chat Control 1.0,允许对私人通信进行无嫌疑的批量扫描;加密通信被豁免但实际不扫描,措施有效至 2028 年。 为什么重要:对依赖端到端加密的应用和开源项目构成监管压力,开发者需重新评估在欧洲市场的隐私合规策略。

开发者社区强烈批评,认为侵犯隐私、破坏民主程序;但也有人指出该法案尚未最终生效。

每天早晨,一份为你精选的科技日报

网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。

已发布 44 期 · 每天筛过 150+ 条只留值得读的 30 条

AI 动态

开发与开源

My thoughts on the Bun Rust rewrite

Bun 创始人谈 Rust 重写争议,Zig 作者 Andrew Kelley 发文批评工程水准,社区认为是人身攻击 vs 必要批评。

评论区主要认为该文章是对Jarred的人身攻击而非技术讨论,但也有人认为这种直率批评是必要的。

Why developers are ditching GitHub for Codeberg and self-hosting alternatives

因微软收购后审查、AI 数据滥用等问题,部分开发者转向 Codeberg、自建 Gitea/Forgejo 替代 GitHub。

用户因微软收购后GitHub的审查、AI训练数据滥用、服务不稳定等问题,转向自托管Gitea/Forgejo或Codeberg等替代方案,但也有人认为GitHub的便利性仍难以替代。

社区热议

GitHub Trending

AI-powered job application framework built on Claude Code. Fork it, fill in your profile, and let Claude evaluate jobs, tailor CVs, write cover letters, and prepare you for interviews.

Give Claude the ability to watch any video. /watch downloads, extracts frames, transcribes, hands it all to Claude.

World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio.

LLM 驱动的多市场股票智能分析系统:多源行情、实时新闻、决策看板与自动推送,支持零成本定时运行。 LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost-free scheduled runs.

microsoft/flint-chartTypeScript★ 3

🪄 Flint is a visualization language that lets AI agents reliably create expressive, good-looking charts from simple, human-editable chart specs.

更多值得一看(内容池 50 条)
Building a real-time AI tutor for 5-year-olds

Hey HN! We've spent the good part of this past year building an AI tutor that teaches kids ages 4-9 reading, math, ESL and more. Getting an AI tutor to effectively teach a child turns out to be a really hard technical challenge, this took getting the underlying architecture right. Our tutor steers the UX in real-time and makes complex decisions on the fly. Doing both at conversation speed required us to replace the standard tool-use loop. We built our own tutor harness that utilizes a streaming

NVIDIA Puzzle-75B-A9B NVFP4 at 132 t/s on 3×3090 — Why is this size category a desert otherwise?

TLDR: 75B-total / 9B-active MoE is the perfect shape for multi-24GB rigs, and almost nobody ships it. Qwen 27B is a great model and punches way above its weight-class, it is a frequent fallback for me. Nemotron-3-Puzzle-75B-A9B, NVFP4, vLLM 0.22.1 (the new Marlin fallbacks run FP4 on Ampere), pipeline-parallel across 3×3090 capped at 200W each. The 4th card runs a speech sidecar untouched - 3 seats × 256K ctx, fp8 KV — hybrid Mamba keeps the cache tiny - 132 t/s decode across 3 streams (~65 sing

From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents

arXiv:2607.07321v1 Announce Type: new Abstract: Tool utilization enables Large Language Model (LLM) agents to interact with the real world and resolve complex tasks. However, existing agent frameworks predominantly rely on static toolsets composed of granular atomic actions (e.g., basic file I/O or single-turn search), which forces agents to reinvent low-level logic for every recurring workflow, leading to increased reasoning overhead and failure rates. In this study, we propose that agents can

The untuned 27B beat the tuned 75B as an agent

I have to admit, a lot of people we're 100% correct to make the suggestion to try this model. I am sorry I ever doubted. The 27B passed every agentic task on a neutral system prompt in 6-9 tool calls. The 75B needed a hand-tuned profile to pass at all and used 2x the turns. For agents, fewer turns beat faster tokens. The two contenders - Nemotron Puzzle-75B-A9B NVFP4, vLLM, PP=2000 across 3 cards, ~65 t/s decode. I made a post about this model. I still think its good for throughput on chatbots a

OpenMed 1.8: Apache-2.0 clinical de-identification that runs fully local, now on Android, iOS, and in the browser. 400+ open issues if you want in on 1.9

Maintainer here. OpenMed is an Apache-2.0 toolkit for clinical NLP with one hard rule: patient data never leaves your hardware. No cloud calls, no API keys, works in airplane mode. What shipped in 1.8 this week: OpenMedKit for Android (Kotlin, ONNX Runtime Mobile + ML Kit OCR): read a document, strip every name/MRN/date, entirely on the phone. iOS/Swift and React Native bridges landed too. Browser runtime : de-identification via Transformers.js / ONNX Runtime Web with wasm + WebGPU backends. Ful

Deepseek V4 Flash on a single RTX 6000 Pro - vLLM-Moet

Wow... Using this customized vllm provided as a docker, I'm able to run DS V4 Flash on a single RTX 6000 Pro (apparently it also works on a single 5090 - check his readme, but I haven't tried). Apparently this also works with GLM 5.2 (though you need at least two 6000 pros, which is still amazing). Setting 130K context, I needed around 150 GB of RAM to get past the safetensor sharding, but once it is fully loaded in VRAM I am able to fit it all in the GPU (If you have less than this much RAM, cr

Postgres is enough for more than we admit

I came across this on Hacker News and felt like I needed to share it with the dev community. The main point is simple: a lot of teams reach for extra databases, queues, search engines, caches, and services before they actually need them. This page lays out where Postgres is usually enough, and where you may actually need something else: Postgres is not perfect for everything, but it is good enough for a surprising amount of real-world work. The more I build and maintain systems, the more I appre

If You Already Pay for an LLM Service, Running Local Embeddings and Rerankers Feels More Useful Than Running Local LLMs

This post was originally written in Korean, then polished and translated into English using ChatGPT. I do run llama.cpp locally on a Tesla P40, but as someone who already pays for ChatGPT Pro, I was gradually losing the practical reason to keep running local LLMs like Qwen 3.6 27B or Gemma 4 31B. If I need access to OpenAI models through an API-like workflow, I can usually just use Codex OAuth instead. But then I realized that embedding models and reranker models are not something I can access t

Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

arXiv:2607.07097v1 Announce Type: new Abstract: Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single "pipeline effect." We argue that this aggregate is difficult to interpret because it conflates three mechanisms: harmful intent may be reframed as plausible operational work, the planner may refuse or transform the request, and the executor may act under delegation prompts implying prior approval. To sep

Pangolin 1.20: Resource Launcher & Global Command Palette

Hello everyone! Pangolin 1.20 is focused on how people find and reach their resources in the UI. Here's what's new: Pangolin is an open-source, identity-based remote access platform that lets you securely expose your infrastructure to your team. It supports browser based remote access and a remote access VPN in one platform with strong authentication controls. GitHub: Resource Launcher The landing page non-admins see when they sign in now is vastly more capable. Resources are grouped by site or

Meta are apparently working on an open source variant of Muse Spark.

No real details or timescales yet, but this article has confirmation from Alexandr Wang that Meta are working on an open source variant of Muse Spark. One to keep an eye on.

MOSS-Transcribe-Diarize 0.9B is an end-to-end audio understanding model for long-form multi-speaker transcription, diarization, timestamps, and acoustic event awareness. Given an audio or video file, the model generates a compact speaker-aware transcript in one pass, including timestamps and anonymous speaker labels such as [S01] , [S02] , and beyond. Introduction MOSS-Transcribe-Diarize 0.9B turns real-world long-form audio into structured, speaker-aware transcripts in one pass. Instead of stit

Benchmarking Coding Agents on Databricks’ Multi-Million Line Codebase

The main conclusions from analysis were: The Pareto frontier for coding tasks (i.e. best quality for a given cost) includes models from OpenAI, Anthropic, and open source. This means today, only a mix of tools can provide frontier performance. Open models, and GLM 5.2 in particular, are now able to handle even the highest level of task difficulty. The token price of a model is a poor indicator of actual costs incurred on end-to-end tasks. Larger models can be far more token efficient and have lo

Microsoft’s patch Tuesdays are about to get bigger

Windows 11 updates could soon include fixes for more security issues at once. Microsoft said in a blog post on Thursday that it's now using AI to "identify potential issues earlier," which means "customers will see a higher volume of security updates included in each security release." Hackers, even amateurs, have increasingly been using AI […]

llm 0.31.1

Release: llm 0.31.1 Fix for a bug with OpenAI Chat Completion endpoints where a tool call with empty arguments could result in a JSON error from some providers. #1521 This bug came up when I was testing llm-meta-ai . Tags: llm

Launch HN: Context.dev (YC S26) – API to get structured data from any website

Hi Hacker News, I’m Yahia. I built Context.dev ( ) to make it really easy to integrate web data into your products and agents. Here’s a demo video: Since it’s an API, here are the docs: . You can send us a URL and get back clean Markdown, rendered HTML, screenshots, extracted images, etc.. You can also send us a domain and get company or brand context: name, description, logos, colors, fonts, social links, screenshots, style information, and related metadata. For more custom use cases, you can s

Reasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

arXiv:2607.07229v1 Announce Type: new Abstract: Prior work has shown that chain-of-thought (CoT) reasoning is often unfaithful: a model's stated reasoning does not reliably reflect the process that produced its output. Detecting unfaithfulness, though, requires controlled experimental interventions, which cannot be applied to evaluation transcripts after the fact. We turn instead to a more tractable question that has received less attention: whether the stated reasoning is logically consistent w

6x MI50's (96gb) vs 6 P40's (144gb) running MiniMax M2.7 REAP 139B Q3_K_L

Hey Guys, As promised here are the results from running MiniMax M2.7 REAP 139B Q3_K_L on llama-bench on 6x MI50's. Memory Load: Hardware: Asus X99-E-WS ( Modded BIOS to support a large number GPU's ) Intel(R) Xeon(R) CPU E5-2680 v4 @ 2.40GHz 128GB DDR4 RAM SSD 6x MI50's 96GB VRAM (Gen3 x8,x8,x8,x8,x8,x8) Results GPU Setup Model Test Result 6x MI50 / Pro VII 16GB MiniMax M2.7 REAP 139B Q3_K_L pp512 139.27 t/s 6x MI50 / Pro VII 16GB MiniMax M2.7 REAP 139B Q3_K_L tg128 24.87 t/s 6x MI50 / Pro VII 1

GLM 5.2 generated most of this playable 3D game in the first iteration

I made this simple 3D Geometry Wars-style game using my coding agent, Jarvis Code, with GLM 5.2. You can play it here: I was honestly surprised by the result. Most of the game came together in the first iteration, and I only needed about four small follow-up tweaks afterward. I'd be curious to hear what you think about the gameplay, the code quality, or GLM 5.2's coding ability.

Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics

arXiv:2607.06820v1 Announce Type: new Abstract: Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS) in agentic LLM workflows underexplored. We propose a ReAct-style agentic setup that combines LLM reasoning with verifiable feedback from SageMath, together with Context7 for the up-to-date documentation. We evaluate this agentic setup across frontier models for solving research-level mathematical pr

引用OpenAISimon Willison1 minAI产品
Quoting OpenAI

[...] Work on web and mobile runs in the cloud. Work in the desktop app can also use local files and desktop apps with your permission. At launch, cloud Work conversations do not appear in desktop Work; desktop Work threads and local files remain on that computer. — OpenAI , trying (unsuccessfully) to clarify ChatGPT Work Tags: openai , chatgpt , ai

OpenAI is shutting down Atlas, but its AI browser ambitions are still growing

OpenAI is sunsetting its AI-powered browser after less than a year. But it's moving some agentic browsing features to its desktop app and a Chrome extension.

The ChatGPT browser is already dead

OpenAI is already shutting down ChatGPT Atlas, its browser that could do tasks for you on your behalf, less than a year after launching it. Atlas was announced in October, but as part of its wave of news about ChatGPT Work today, the company confirmed that it will be "sunsetting" Atlas and is targeting an […]

MispherProduct Hunt1 min产品AI

Dictate, rewrite, translate, and an agent in a single device Discussion | Link

Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks

arXiv:2607.07189v1 Announce Type: new Abstract: Vision-language models (VLMs) and agentic AI have shown strong performance on semantic visual tasks, but it remains unclear whether they can handle the physics and inverse problems that underlie computational imaging. We present ImagingBench, a benchmark of 20 computational imaging tasks spanning five categories: ray and wave optics, image signal processing, inverse reconstruction, computational sensing, and calibration. ImagingBench evaluates thre

Repurposed a $30 Dell XPS 13 with a dead battery into a 6W homelab server. Power efficient, small, silent. Roast me!

I finally decided to set up a dedicated home lab server a while ago. I priced out built-from-scratch x86 ITX configurations and low-power NAS builds, and they were easily running $500–$800 just for CPU/RAM/chassis. I hesitated. I wanted something cheap, extremely low-power, and compact. So I did what any frugal self-hoster does: I looked at old hardware I already owned. I had a 2017 Dell XPS 13 laptop (i5, 8GB RAM) sitting in my drawer. I had tried selling it locally, but since the battery was c

每天早晨,一份为你精选的科技日报