DawnSift
订阅日报
周三 · 科技日报 · 第 73 期

2026-09-23

— 大模型价格战开打,安全与军事 AI 的阴影同日浮现。

今日 TL;DR

OpenAI 发布 GPT-6 Sol 和 Luna,价格较前代减半;Anthropic 发布 Claude Opus 5.5,性能对标 Fable 5.1 且成本降 40%。两家公司同日发布,标志前沿模型竞争进入性价比阶段。安全方面,ShinyHunters 声称入侵 FBI 并窃取全部员工数据;五角大楼承认过度依赖 AI 导致伊朗学校误炸。

GPT-6 Sol and Luna are half the price of their GPT-5.6 equivalents

头条

1

OpenAI 发布 GPT-6 Sol 和 Luna,价格减半多源事件 ×5

OpenAI 推出 GPT-6 Sol 和 GPT-6 Luna,作为 GPT-6 Astra 的轻量级补充,主打成本效率。Sol 面向复杂任务如编码,Luna 面向高容量文书工作;两者输入/输出价格均为 GPT-5.6 同档模型的一半。为什么重要:对于以 API 构建应用的开发者,这意味着在保持接近前沿性能的同时,推理成本大幅下降,尤其利好 agent 和批量处理场景。

评论区普遍认可降价与高性价比,但也有人认为性能提升有限、发布时机疑似针对竞品。

2

Anthropic 发布 Claude Opus 5.5,性能对标 Fable 5.1 且成本降 40%多源事件 ×3

Anthropic 发布 Claude Opus 5.5,宣称在多数任务上达到 Claude Fable 5.1 水平,运行成本比 Opus 5 低 40%,输出 token 定价从 25 美元/百万降至 20 美元/百万。该模型是 Anthropic 自呼吁放缓前沿竞赛以来的首个发布,经 Frontier Design 和 METR 外部评估。为什么重要:旗舰级模型降价并提升推理效率,直接降低复杂编码与知识工作负载的 API 开销,同时强化了安全测试叙事。

多数人认可 Opus 5.5 更自然、更便宜且性能强,但也有人认为版本号跳跃、定价仍高,且对基准测试持怀疑态度。

3

ShinyHunters 声称入侵 FBI,窃取全部员工及申请人数据

黑客组织 ShinyHunters 声称入侵多个 FBI 相关服务,窃取所有 FBI 员工和申请人的数据,包括姓名、家庭住址、电话号码及配偶信息。404 Media 已验证部分数据与公共记录匹配。为什么重要:若属实,这是美国执法机构史上最严重的数据泄露之一,可能引发大规模反情报风险,并再次暴露政府系统的安全短板。

评论区普遍认为此次入侵暴露了美政府网络安全严重失职,但也有人认为数据真实性存疑,需等实际泄露内容才能确认。

4

五角大楼承认过度依赖 AI 导致伊朗学校误炸

五角大楼调查发现,情报缺陷、过时图像和对 AI 的过度依赖共同导致 2026 年 2 月 28 日导弹袭击伊朗 Minab 一所学校,造成 123 名儿童死亡。为什么重要:这是 AI 参与军事杀伤链导致重大平民伤亡的典型案例,对 AI 在国防系统中的自动化决策边界提出了尖锐的问责问题。

评论普遍认为 AI 只是替罪羊,真正责任在决定使用 AI 的人和军方,必须追究人类责任;但也有人认为技术提供方 Palantir 同样难辞其咎。

5

GPT-6 Astra 破解 2005 年以来未解的恩尼格玛消息

2026 年 9 月 15 日,Carter Leffer 使用 GPT-6 Astra 破解了 1941 年 7 月 10 日德国陆军恩尼格玛消息 MVUEH,该消息自 2005 年以来一直无法破解。为什么重要:展示了 LLM 在密码分析中的辅助能力,但社区对其真实性存疑,认为可能涉及训练数据泄露或炒作。

评论普遍质疑 GPT-6 破解恩尼格玛消息的真实价值,认为可能是炒作或训练数据泄露,但也有人认为这展示了人机协作的潜力。

每天早晨,一份为你精选的科技日报

网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。

已发布 73 期 · 每天筛过 150+ 条只留值得读的 30 条

AI 动态

开发与开源

llm 0.36Simon Willison1 minAI开发工具

llm 0.36 新增 gpt-6-sol 和 gpt-6-luna 模型,并支持声明不支持对话的模型。

llm-typesafe 0.1a0 插件支持 TypeSafe AI 的 Jev 决策模型,可输出结构化 yes/no 或 choice 答案。

社区热议

Qwen 4 Announced at Apsara Conference

阿里在 Apsara 大会正式宣布 Qwen 4,LocalLLaMA 社区关注其开源权重与本地部署潜力。

GitHub Trending

Star dream-num / univer The Office Harness for AI Agents — Spreadsheets, Docs, Slides, Canvas, Relational Tables, and PDF in one runtime.

google/ax★ 7589

Star google / ax Google's open agentic orchestration runtime

Star mvt-project / mvt MVT (Mobile Verification Toolkit) helps with conducting forensics of mobile devices in order to find signs of a potential compromise.

Star superdesigndev / treg OpenRouter for agent tools. Join community here: https://discord.gg/6mQYYfFMAn

更多值得一看(内容池 62 条)
onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction

We present onPanda, an interactive tool for efficiently annotating LLM alignment data and agent trajectories. onPanda adopts token-level correction as its core interaction: while reading a model response, the annotator locates the first inappropriate token and either picks a substitute from the model's candidate tokens or types the correct text via free-form editing. The system then truncates everything after that position and continues generation from the corrected prefix, repeating this locate

Designer-RSI: Evolving Procedural Memory from User Traffic for Agentic Graphic Design

Professional graphic design is a long-horizon agentic task in which structured, editable artifacts emerge from many interdependent actions, yet outcomes admit no reliable programmatic oracle. We introduce a continual adaptation framework in which a frozen frontier model operates professional design software through more than 230 tools, while an external procedural memory of natural-language skills accumulates and refines reusable design procedures from experience. The memory widens by acquiring

Ngram and world knowledge - why are we just building a coding model?

This post is written by a human and I'd appreciate it if you treated it as such. Thanks. So, I've been noticing a pretty clear interest in developing as good a coding and agentic tool-calling model as possible, especially at smaller sizes, sub-50 gigs. However, I'm finding that at least for my use of AI, if I really want to move away from big providers, I am going to require a model that has better world knowledge than the current offerings. Qwen 3.8 27B is a truly fantastic model for tons and t

One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents

Repository-level software engineering (SWE) comprises heterogeneous task categories, whose progress under pooled agentic reinforcement learning can be uneven: gains in some categories coincide with regressions in others, while aggregate resolution obscures these changes. Motivated by this category see-saw, we develop a category-aware expert-training and policy-integration framework. Executable task construction and SWE Labeler, an evidence-grounded multi-axis labeling system, organize the traini

Jev-Mem: System-One-Controlled Agentic Memory for Efficient AI Agents

Agentic memory is becoming essential for long-horizon AI agents, yet many existing systems rely on autoregressive LLMs to control how memories are organized, retrieved, and used, placing expensive generation on the critical path of memory operations. We introduce \method, a new agentic memory architecture inspired by System-One/System-Two cognition. System One captures fast, lightweight decision-making, whereas System Two performs slower, deliberative reasoning. Jev-Mem brings this division of l

Harness-Zero: Harness Distillation via Agent-as-Harness

Agent harnesses, the external systems that mediate model-environment interaction, can substantially improve agent performance, but their gains remain tied to the harness at deployment. Because the best harness varies across domains, instances, and models, a general-purpose agent must either settle for a suboptimal shared harness or route among an ever-growing set of specialized ones. We therefore study agent harness distillation: using a domain- or instance-optimized harness as training-time gui

1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation

Sparse on-policy distillation (OPD) allocates teacher supervision to a small subset of tokens in student-generated trajectories. However, useful teacher guidance can yield a noisy update when its gradient is estimated from a sampled next token. We study this estimation problem at a fixed prefix in information geometry and propose an information-efficiency ratio (IER) based on a signal-to-noise decomposition. IER characterizes relative gradient estimation error under an optimal scalar baseline. A

GraphSkillEvo: Evolutionary Optimization of Graph-Structured Agent Skills

Skills can improve the performance of Large Language Model (LLM) agents by providing task-specific procedural guidance, while skill optimization further improves their effectiveness through iterative refinement. However, existing skill optimization methods typically represent skills as unstructured natural-language instructions, creating two key challenges: 1) Unstructured skills often lack explicit workflow-level guidance and contain substantial redundancy, making them difficult for LLMs to exe

ACLArena: Agent Continue Learning in Multi-stage Post-training

Building general-purpose agents for industrial deployment requires integrating multiple capabilities, each typically acquired at a distinct stage of training. Yet there is currently no well-established recipe for Agent Continual Learning (ACL), with little understanding of the trade-offs among existing integration paradigms. To address this gap, we introduce ACLArena, a framework for comprehensively studying, analyzing, and evaluating ACL. We first build a sequential training pipeline and conduc

AntLing open sourced the Ming-Image-0.1-Design family

AntLing open sourced the Ming-Image-0.1-Design family: • Ming-Image-0.1-Design, 6B • Ming-Image-0.1-Design-Layer, 6B • Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard

OpenAI wants to consult elite mathematicians about how to not fumble again

After turning a string of spectacular mathematical results into a reputational crisis, OpenAI is consulting human mathematicians to help it figure out a less disastrous path forward. On Monday, the company announced a new independent panel of mathematicians tasked with advising it and other AI companies on their interactions with mathematical research and the wider […]

Grounded Action Model: 3D Grounding as a Foundation for Robotics

Manipulation policies must know which objects matter and where they are, yet the pretrained backbones that current robot foundation models build on, from language in vision-language-action models (VLAs) to video generation in world-action models (WAMs), do not directly require this metric grounding, leaving it to be learned implicitly from robot demonstrations. We propose Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding. GAM can be conditioned usin

BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence

Business intelligence (BI) is a cornerstone of enterprise decision-making and is widely used by enterprise users in software such as Power BI and Tableau. In traditional BI workflows, users need to prepare data by (1) identifying relevant tables, (2) performing data transformations, and (3) building join relationships, before they can (4) answer their business questions. These steps can be complex and time-consuming, making BI challenging. Given the strong capabilities of large language models (

Realtime-Venus: A full-duplex interaction system with asynchronous delegation

Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requires grounding the conversation in evolving visual context. We present Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction. Each model serves as a complete co

New 6B image model coming, AntLing just open sourced the Ming-Image-0.1-Design family

• Ming-Image-0.1-Design, 6B • Ming-Image-0.1-Design-Layer, 6B • Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard.

Qwen image 2.1 (Fast FP8) generates premium quality images

Don't know how they did it, but for under 10GB model, the results are astonishing. I am running it on Unsloth Studio. They just released the update, so if you are not seeing the option, I recommend updating your Unsloth Studio. Cheers!

Rabbit’s new AI agent doesn’t need an R1 to run

Rabbit, the company behind the underwhelming R1 device, is rolling out a standalone AI agent that you don't need its hardware to use, as reported earlier by Wired. The startup says its new OS3 "agentic operating system" runs in the cloud but operates locally across Windows, Mac, and Linux devices. According to Rabbit, you can […]

Microsoft killed FoxPro in 2007. Anyway, here's FoxPro revived

Visual FoxPro stopped at version 9 in 2007. A surprising amount of it is still running, in 32 bits, because rewriting a 20-year-old business app is how you lose the business. A customer wanted to keep milking their app for the foreseeable future, so here it is: the same language on a new runtime (Rust, compiled to wasm, checked against the real vfp9.exe), tables no longer stopped at 2 GB, the old 32-bit .fll add-ins still loading, and lambdas, JSON and an HTTP server bolted on for good measure.

The Download: why AI’s latest breakthroughs and fears may be more hype than reality

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. Don’t be fooled by this summer of AI hype —Timnit Gebru, executive director of the Distributed AI Research Institute (DAIR), and Emily M. Bender, professor of linguistics at the University of…

HuRo: Robotizing Human Videos for Scalable VLA Pretraining

Human video datasets offer an abundant and diverse source of interaction data that can complement expensive real-robot data. To bridge the human-to-robot embodiment gap, existing approaches either robotize videos in task-matched settings or address observation and action alignment separately at scale. In this work, we systematically examine whether robotized human videos can serve as an effective and scalable source of supervision for VLA pretraining. To this end, we develop a robotization pipel

Priorities and principles for effective third party assessments

OpenAI outlines priorities and principles for rigorous, secure, and independent third-party AI safety assessments of frontier models and safeguards.

Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations

Existing approaches to persona simulation with Large Language Models (LLMs) mostly rely on shallow character descriptions that fail to sustain coherent character behavior across extended interactions. We introduce Deep Persona, a psychologically grounded, three-layered architecture that organizes personas into hierarchical levels of observable expression, latent beliefs, and core motivational drives, for constructing highly convincing role-playing agents. Governed by the principles of scripted d

Basically the title.We did not get a new moe model with qwen 3.8 and Alibaba did not announce any small moe models on apsara.I know we might get an announcement later but ngl I kinda lost hope

Gricea: An Open Science Platform for Conversational AI Research

We need studies on conversational AI (CAI) at scale to understand human behavior and shape CAI design. However, fragmented reporting of systems and study configurations hinders replication, extension, and knowledge accumulation. We present Gricea, an open-science platform representing studies as configurable, deployable research artifacts that researchers can run, inspect, share, and reuse. Informed by a formative analysis of prior CAI research, Gricea couples study procedures, participant-facin

Self hosted Bitwarden vs Vaultwarden

I have been successfully hosting Vaultwarden for the last year or so and have had no major issues to write home about. It serves myself and my mum, but I'm intending to expand that to other family. What I'm interested in is understanding if I might be better of just using the official Bitwarden for self hosting instead of Vaultwarden. I'm concerned that if the Vaultwarden maintainer is let go from Bitwarden or changes priorities or whatever, then I'm at risk. One of the initial reasons for choos

每天早晨,一份为你精选的科技日报