OpenAI 启动 GPT-5.5/5.6 生物漏洞赏金计划,通用 jailbreak 奖励提高至 $50,000。
2026-07-10
— 模型井喷,监管收紧,本地推理破门槛
OpenAI 发布 GPT-5.6 三型号及 ChatGPT Work,并宣布成为微软 365 Copilot 首选模型;Meta 推出 Muse Spark 1.1,定价仅 $1.25/百万输入 token;GLM-5.2 (744B MoE) 在 25GB RAM 消费级机器上成功本地运行;NYT 指控 OpenAI 在版权诉讼中隐藏证据,要求法院制裁;欧盟通过 Chat Control 1.0,允许无嫌疑批隐私扫描。
头条
GPT-5.6 正式发布:三型号 + ChatGPT Work + 微软 Copilot 集成多源事件 ×7
OpenAI 推出 GPT-5.6 家族,包含旗舰 Sol、平衡 Terra 和经济 Luna,其中 Sol 在编码、网络安全等基准上超越前代与竞品。同时发布 ChatGPT Work agent,基于 Codex 技术,可跨应用完成复杂任务,并宣布 GPT-5.6 成为微软 365 Copilot 首选模型。 为什么重要:软件工程师可立即低成本试用新模型(Sol $5/$30 每百万 token),ChatGPT Work 将 agent 能力扩展至非开发者,微软集成意味着大规模企业部署即将落地。
社区普遍认可成本效率提升,但部分开发者认为编码能力未达预期,对营销宣传存疑。
Meta 发布 Muse Spark 1.1:多模态 agent 模型,定价极具竞争力多源事件 ×4
Meta 推出 Muse Spark 1.1,面向 agentic 任务的多模态推理模型,在工具调用、计算机使用和编码能力上有显著提升。开放 Meta Model API 预览,定价 $1.25/百万输入 token,支持 1M token 上下文窗口。 为什么重要:以极低价格进军 AI 编码市场,可能推动整体降价;1M 上下文适合复杂 agent 工作流,对开发者的 agent 架构选择产生直接影响。
评论区认可性价比,但对其基准测试公正性与可用性存在分歧。
NYT 指控 OpenAI 在版权诉讼中隐藏证据,要求法院制裁
纽约时报等新闻机构在版权诉讼中提交制裁动议,称 OpenAI 多年谎称无法搜索训练数据以隐瞒侵权证据,其隐私工程师在作证中暴露矛盾;OpenAI 被指隐藏了数十亿条聊天日志。 为什么重要:若法院制裁,OpenAI 可能需交出关键数据,直接影响 AI 训练合理使用的法律边界,对整个行业的大规模爬取行为具有深远影响。
社区观点两极,有人支持版权方,也有人认为这是对 AI 发展的阻碍。
每天早晨,一份为你精选的科技日报
网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。
已发布 44 期 · 每天筛过 150+ 条只留值得读的 30 条
AI 动态
Anthropic 开发 Jacobian 透镜,发现 Claude Opus 4.6 内部隐藏的「J-space」概念空间,揭示模型预推理过程。
AgentLens 提出基于轨迹评分的代码 agent 评估基准,结合形式化验证与 LLM 审查,提供可解释的评分。
新论文提出「Harners Effect」:编排设计是控制企业 agent 系统 token 开销的关键杠杆,实验横跨六大模型。
AI agent 初创公司 Lyzr 用自己的 agent SivaClaw 完成了 $1 亿美元 B 轮融资,全程处理投资者问答与文档。
开发与开源
Bun 创始人谈 Rust 重写争议,Zig 作者 Andrew Kelley 发文批评工程水准,社区认为是人身攻击 vs 必要批评。
评论区主要认为该文章是对Jarred的人身攻击而非技术讨论,但也有人认为这种直率批评是必要的。
因微软收购后审查、AI 数据滥用等问题,部分开发者转向 Codeberg、自建 Gitea/Forgejo 替代 GitHub。
用户因微软收购后GitHub的审查、AI训练数据滥用、服务不稳定等问题,转向自托管Gitea/Forgejo或Codeberg等替代方案,但也有人认为GitHub的便利性仍难以替代。
Mitchell Hashimoto 访谈:聊 Ghostty 终端、Zig 语言和开源创业经历,提到 15 年 CLI 经验是偶然探索。
社区热议
Reddit 资深开发者讨论:雇佣初级工程师的价值——倒逼团队保持文档、规范与代码质量。
GLM-5.2 在 VAT 审计任务中达到接近人类会计师的准确度,且成本仅为人工的 1%。
GitHub Trending
AI-powered job application framework built on Claude Code. Fork it, fill in your profile, and let Claude evaluate jobs, tailor CVs, write cover letters, and prepare you for interviews.
Mellea is a library for writing generative programs.
Give Claude the ability to watch any video. /watch downloads, extracts frames, transcribes, hands it all to Claude.
agent multiplexer that lives in your terminal.
Clone any website with one command using AI coding agents
World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio.
LLM 驱动的多市场股票智能分析系统:多源行情、实时新闻、决策看板与自动推送,支持零成本定时运行。 LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost-free scheduled runs.
Use Codex from Claude Code to review code or delegate tasks.
🪄 Flint is a visualization language that lets AI agents reliably create expressive, good-looking charts from simple, human-editable chart specs.
Instant, Concurrent, Secure & Lightweight Sandbox for AI Agents.
更多值得一看(内容池 50 条)
Rebranded Codex promises independent workflows that can run "for hours if needed."
Hey HN! We've spent the good part of this past year building an AI tutor that teaches kids ages 4-9 reading, math, ESL and more. Getting an AI tutor to effectively teach a child turns out to be a really hard technical challenge, this took getting the underlying architecture right. Our tutor steers the UX in real-time and makes complex decisions on the fly. Doing both at conversation speed required us to replace the standard tool-use loop. We built our own tutor harness that utilizes a streaming
TLDR: 75B-total / 9B-active MoE is the perfect shape for multi-24GB rigs, and almost nobody ships it. Qwen 27B is a great model and punches way above its weight-class, it is a frequent fallback for me. Nemotron-3-Puzzle-75B-A9B, NVFP4, vLLM 0.22.1 (the new Marlin fallbacks run FP4 on Ampere), pipeline-parallel across 3×3090 capped at 200W each. The 4th card runs a speech sidecar untouched - 3 seats × 256K ctx, fp8 KV — hybrid Mamba keeps the cache tiny - 132 t/s decode across 3 streams (~65 sing
arXiv:2607.07321v1 Announce Type: new Abstract: Tool utilization enables Large Language Model (LLM) agents to interact with the real world and resolve complex tasks. However, existing agent frameworks predominantly rely on static toolsets composed of granular atomic actions (e.g., basic file I/O or single-turn search), which forces agents to reinvent low-level logic for every recurring workflow, leading to increased reasoning overhead and failure rates. In this study, we propose that agents can
I have to admit, a lot of people we're 100% correct to make the suggestion to try this model. I am sorry I ever doubted. The 27B passed every agentic task on a neutral system prompt in 6-9 tool calls. The 75B needed a hand-tuned profile to pass at all and used 2x the turns. For agents, fewer turns beat faster tokens. The two contenders - Nemotron Puzzle-75B-A9B NVFP4, vLLM, PP=2000 across 3 cards, ~65 t/s decode. I made a post about this model. I still think its good for throughput on chatbots a
Maintainer here. OpenMed is an Apache-2.0 toolkit for clinical NLP with one hard rule: patient data never leaves your hardware. No cloud calls, no API keys, works in airplane mode. What shipped in 1.8 this week: OpenMedKit for Android (Kotlin, ONNX Runtime Mobile + ML Kit OCR): read a document, strip every name/MRN/date, entirely on the phone. iOS/Swift and React Native bridges landed too. Browser runtime : de-identification via Transformers.js / ONNX Runtime Web with wasm + WebGPU backends. Ful
Wow... Using this customized vllm provided as a docker, I'm able to run DS V4 Flash on a single RTX 6000 Pro (apparently it also works on a single 5090 - check his readme, but I haven't tried). Apparently this also works with GLM 5.2 (though you need at least two 6000 pros, which is still amazing). Setting 130K context, I needed around 150 GB of RAM to get past the safetensor sharding, but once it is fully loaded in VRAM I am able to fit it all in the GPU (If you have less than this much RAM, cr
I came across this on Hacker News and felt like I needed to share it with the dev community. The main point is simple: a lot of teams reach for extra databases, queues, search engines, caches, and services before they actually need them. This page lays out where Postgres is usually enough, and where you may actually need something else: Postgres is not perfect for everything, but it is good enough for a surprising amount of real-world work. The more I build and maintain systems, the more I appre
This post was originally written in Korean, then polished and translated into English using ChatGPT. I do run llama.cpp locally on a Tesla P40, but as someone who already pays for ChatGPT Pro, I was gradually losing the practical reason to keep running local LLMs like Qwen 3.6 27B or Gemma 4 31B. If I need access to OpenAI models through an API-like workflow, I can usually just use Codex OAuth instead. But then I realized that embedding models and reranker models are not something I can access t
arXiv:2607.07097v1 Announce Type: new Abstract: Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single "pipeline effect." We argue that this aggregate is difficult to interpret because it conflates three mechanisms: harmful intent may be reframed as plausible operational work, the planner may refuse or transform the request, and the executor may act under delegation prompts implying prior approval. To sep
“token efficiency needs to drop to as much as 20% over the next 12 months, and 90% by the following year”
Hello everyone! Pangolin 1.20 is focused on how people find and reach their resources in the UI. Here's what's new: Pangolin is an open-source, identity-based remote access platform that lets you securely expose your infrastructure to your team. It supports browser based remote access and a remote access VPN in one platform with strong authentication controls. GitHub: Resource Launcher The landing page non-admins see when they sign in now is vastly more capable. Resources are grouped by site or
No real details or timescales yet, but this article has confirmation from Alexandr Wang that Meta are working on an open source variant of Muse Spark. One to keep an eye on.
MOSS-Transcribe-Diarize 0.9B is an end-to-end audio understanding model for long-form multi-speaker transcription, diarization, timestamps, and acoustic event awareness. Given an audio or video file, the model generates a compact speaker-aware transcript in one pass, including timestamps and anonymous speaker labels such as [S01] , [S02] , and beyond. Introduction MOSS-Transcribe-Diarize 0.9B turns real-world long-form audio into structured, speaker-aware transcripts in one pass. Instead of stit
The main conclusions from analysis were: The Pareto frontier for coding tasks (i.e. best quality for a given cost) includes models from OpenAI, Anthropic, and open source. This means today, only a mix of tools can provide frontier performance. Open models, and GLM 5.2 in particular, are now able to handle even the highest level of task difficulty. The token price of a model is a poor indicator of actual costs incurred on end-to-end tasks. Larger models can be far more token efficient and have lo
The feud between NightmareEclipse and Microsoft shows no signs of resolving soon.
Broadcom accuses Allstate of dodging VMware audits.
The agent-native way to ship software Discussion | Link
Windows 11 updates could soon include fixes for more security issues at once. Microsoft said in a blog post on Thursday that it's now using AI to "identify potential issues earlier," which means "customers will see a higher volume of security updates included in each security release." Hackers, even amateurs, have increasingly been using AI […]
Release: llm 0.31.1 Fix for a bug with OpenAI Chat Completion endpoints where a tool call with empty arguments could result in a JSON error from some providers. #1521 This bug came up when I was testing llm-meta-ai . Tags: llm
Hi Hacker News, I’m Yahia. I built Context.dev ( ) to make it really easy to integrate web data into your products and agents. Here’s a demo video: Since it’s an API, here are the docs: . You can send us a URL and get back clean Markdown, rendered HTML, screenshots, extracted images, etc.. You can also send us a domain and get company or brand context: name, description, logos, colors, fonts, social links, screenshots, style information, and related metadata. For more custom use cases, you can s
arXiv:2607.07229v1 Announce Type: new Abstract: Prior work has shown that chain-of-thought (CoT) reasoning is often unfaithful: a model's stated reasoning does not reliably reflect the process that produced its output. Detecting unfaithfulness, though, requires controlled experimental interventions, which cannot be applied to evaluation transcripts after the fact. We turn instead to a more tractable question that has received less attention: whether the stated reasoning is logically consistent w
Hey Guys, As promised here are the results from running MiniMax M2.7 REAP 139B Q3_K_L on llama-bench on 6x MI50's. Memory Load: Hardware: Asus X99-E-WS ( Modded BIOS to support a large number GPU's ) Intel(R) Xeon(R) CPU E5-2680 v4 @ 2.40GHz 128GB DDR4 RAM SSD 6x MI50's 96GB VRAM (Gen3 x8,x8,x8,x8,x8,x8) Results GPU Setup Model Test Result 6x MI50 / Pro VII 16GB MiniMax M2.7 REAP 139B Q3_K_L pp512 139.27 t/s 6x MI50 / Pro VII 16GB MiniMax M2.7 REAP 139B Q3_K_L tg128 24.87 t/s 6x MI50 / Pro VII 1
I made this simple 3D Geometry Wars-style game using my coding agent, Jarvis Code, with GLM 5.2. You can play it here: I was honestly surprised by the result. Most of the game came together in the first iteration, and I only needed about four small follow-up tweaks afterward. I'd be curious to hear what you think about the gameplay, the code quality, or GLM 5.2's coding ability.
arXiv:2607.06820v1 Announce Type: new Abstract: Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS) in agentic LLM workflows underexplored. We propose a ReAct-style agentic setup that combines LLM reasoning with verifiable feedback from SageMath, together with Context7 for the up-to-date documentation. We evaluate this agentic setup across frontier models for solving research-level mathematical pr
[...] Work on web and mobile runs in the cloud. Work in the desktop app can also use local files and desktop apps with your permission. At launch, cloud Work conversations do not appear in desktop Work; desktop Work threads and local files remain on that computer. — OpenAI , trying (unsuccessfully) to clarify ChatGPT Work Tags: openai , chatgpt , ai
OpenAI is sunsetting its AI-powered browser after less than a year. But it's moving some agentic browsing features to its desktop app and a Chrome extension.
OpenAI is already shutting down ChatGPT Atlas, its browser that could do tasks for you on your behalf, less than a year after launching it. Atlas was announced in October, but as part of its wave of news about ChatGPT Work today, the company confirmed that it will be "sunsetting" Atlas and is targeting an […]
Dictate, rewrite, translate, and an agent in a single device Discussion | Link
arXiv:2607.07189v1 Announce Type: new Abstract: Vision-language models (VLMs) and agentic AI have shown strong performance on semantic visual tasks, but it remains unclear whether they can handle the physics and inverse problems that underlie computational imaging. We present ImagingBench, a benchmark of 20 computational imaging tasks spanning five categories: ray and wave optics, image signal processing, inverse reconstruction, computational sensing, and calibration. ImagingBench evaluates thre
AI cheating leads to "a failed society," professor says.
I finally decided to set up a dedicated home lab server a while ago. I priced out built-from-scratch x86 ITX configurations and low-power NAS builds, and they were easily running $500–$800 just for CPU/RAM/chassis. I hesitated. I wanted something cheap, extremely low-power, and compact. So I did what any frugal self-hoster does: I looked at old hardware I already owned. I had a 2017 Dell XPS 13 laptop (i5, 8GB RAM) sitting in my drawer. I had tried selling it locally, but since the battery was c