DawnSift
S’abonner
jeu · Quotidien tech · Numéro 60

2026-09-10

— OpenAI burned $40 million in 88 hours to solve a millennium problem, while an Anthropic researcher resigned that day warning humanity.

TL;DR du jour

OpenAI announced that GPT-6 Astra found a Navier-Stokes singularity in 88 hours using about 10,000 agents and 130B tokens, vying for the second Millennium Prize, but was accused of academic scooping. Anthropic pre-training researcher Jacob Coxon resigned and publicly warned that the AI race is out of control, sparking millions of discussions. DeepSeek announced that V4.1 Flash fully surpasses V4 Pro and plans to softly retire the Pro version. Google open-sourced the Mantis security skills toolkit, enabling coding agents to handle the full vulnerability lifecycle.

À la une

1

OpenAI finds Navier-Stokes singularity in 88 hours with Astra-next, vying for second Millennium Prize

OpenAI reported using Astra-next to find a singularity in the Navier-Stokes equations within 88 hours, using about 10,000 agents and 130B tokens (costing over $40 million), making it a strong contender for the second Millennium Prize. Why it matters: This marks AI leaping from a supporting tool to a primary driver in mathematical proof-level research tasks, but The Verge reported that OpenAI rushed ahead after learning of other researchers' progress, raising concerns about scooping and norm-breaking in academia.

Latent Space believes OpenAI's achievement is real but the process is controversial; The Verge emphasizes that academia worries this scooping behavior could cool the entire field.

2

Anthropic pre-training researcher Jacob Coxon resigns, warning the AI race is putting humanity at risk

Jacob Coxon announced his resignation from Anthropic on X, stating that three years of pre-training research at OpenAI and Anthropic convinced him both companies are irresponsibly racing toward self-improving superintelligence. Why it matters: Coxon's post garnered over 100 million views, and in a WIRED interview he revealed colleagues often describe the next year or two as 'endgame' and 'crunch time,' providing a rare insider perspective on AI safety discussions.

HN commenters generally believe the AI race is out of control and hard to stop, but some think the resigner's words and actions are inconsistent or just for show.

3

DeepSeek announces V4.1 Flash fully surpasses V4 Pro, plans soft retirement of Pro version

DeepSeek plans to officially release V4.1 Flash around September 10, 2026, claiming it fully surpasses V4 Pro on all key metrics including performance, cost, speed, and task completion time. Why it matters: This means DeepSeek is replacing its flagship Pro with a cheaper Flash model, while Reddit users discovered Pro requests are being automatically routed to Flash, a soft retirement strategy that could change API users' cost expectations.

HN commenters generally look forward to V4.1 Flash's cost-performance improvement, but some are concerned about Pro requests being auto-routed to Flash and price adjustments.

4

Google open-sources Mantis: enabling coding agents to handle the full vulnerability discovery, reproduction, and patching process

Google open-sourced Mantis, a stack-agnostic security review skills toolkit that lets AI coding agents execute the full vulnerability lifecycle: scanning code, filtering false positives, reproducing bugs in a sandbox, writing patches, re-attacking patched code, and scoring residual risk. Why it matters: Mantis is released under Apache 2.0 and works with agent frameworks like Gemini CLI, Antigravity CLI, or Google ADK, offering developers a modular way to embed security review into existing agent workflows.

5

GPT-6 Astra officially released: OpenAI calls it the strongest model for work scenarios

OpenAI officially released GPT-6 Astra, available in ChatGPT Work, Codex, and API, claiming state-of-the-art performance in computer use, browsing, professional work, software engineering, cybersecurity, and scientific tasks. Why it matters: Astra uses a looped transformer architecture and reportedly hides its reasoning chain; Sebastian Raschka's analysis notes this architecture, while not new, reduces the observability of the reasoning process, directly impacting developers who rely on CoT transparency.

HN commenters note looped transformers are not a new mechanism but worry they reduce reasoning observability; others find Astra's actual experience underwhelming and its capabilities inconsistent.

Chaque matin, un digest tech fait pour vous

Le web montre la vue d’ensemble ; les abonnés reçoivent la leur — sélection IA selon vos intérêts, votre RSS privé intégré, avec les avis de la communauté, livrée chaque matin. Gratuit à vie.

60 numéros publiés · 150+ infos filtrées à 30 chaque jour

Actu IA

NeoHorse-1 explores recursive self-improvement (RSI) through agentic post-training and intelligent routing, converting user interaction records into training samples that preserve reasoning and tool-calling context.

🤖NeoHorse-1 uses agentic post-training with intelligent routing, structured feedback loops, and curriculum-based distillation to improve model capabilities across agent benchmarks.

Tencent Hunyuan open-sources the AuK speech foundation model, using 3.03 billion instruction-audio instances to unify speech generation and editing, combining a multimodal LLM, joint VAE, and hybrid rectified-flow Transformer.

🤖AuK is an open-source foundational model that unifies speech generation and editing via natural-language instructions and audio context, using a multimodal language model, joint VAE, hybrid rectified-flow Transformer, and efficient distillation for fast inference.

Gander is an end-to-end model unifying full perception, real-time interaction, and agent capabilities, using a Cerebellum-Brain architecture and chunk-level token stream for full-duplex streaming interaction.

🤖Gander is an end-to-end framework that integrates continuous multi-modal streaming, real-time full-duplex interaction, and agentic reasoning through a Cerebellum-Brain architecture and a chunk-level token stream design.

Miles v0.1 open-sources a production-grade post-training system based on the SGLang rollout engine, supporting Megatron-LM and PyTorch FSDP dual backends and three weight synchronization transports.

🤖Miles is an open-source, production-ready system for large-scale reinforcement learning and post-training that supports diverse backends, weight synchronization, LoRA, distillation, and diffusion models.

BeaconKV uses compact beacon queries to predict which KV pairs will be revisited in long reasoning chains, compressing cache without losing accuracy and alleviating memory bottlenecks in LRM inference.

🤖BeaconKV improves memory efficiency for long reasoning traces by using compact beacon queries to predict which past key-value pairs will be revisited, reducing cache size without sacrificing accuracy.

Dev & open source

Échos de la communauté

Satirical site opusfived.dev precisely hits the pain point of Claude overcomplicating and adding unrequested actions, though some find it inconsistent with their own experience.

Most believe the satire precisely hits the pain point of Claude overcomplicating and adding unrequested actions, but some find it inconsistent with their own experience.

The author demonstrates how to run malicious ads on Google Ads; commenters widely criticize Google's lax review and ineffective appeal mechanisms.

Commenters widely criticize Google's lax ad review and ineffective appeal mechanisms, but some argue malicious advertisers should not receive feedback information.

GitHub Trending

Star ayghri / i-have-adhd A skill to stop your coding agent from burying the answer. ADHD-friendly output.

Sponsor Star obra / superpowers An agentic skills framework & software development methodology that works.

Star pascalorg / editor Open-source 3D architectural editor with a local CLI, MCP tools, and practical workflows for humans and AI agents.

Star cathrynlavery / diagram-design 38 editorial diagram types for Claude Code, Codex, and Pi. Self-contained HTML + SVG. No shadows. No Mermaid slop.

Star freestylefly / awesome-gpt-image-2 Prompt as Code | GPT-Image2 工业级提示词引擎与模板库,530+ 个案例逆向工程,20+ 套工业级模板,并提炼出Skills,持续更新中

Aussi à voir(58 de plus)

Speculative decoding accelerates rollout generation, which dominates the cost of reinforcement learning (RL) post-training. Online co-training can further increase the draft's accuracy, yielding greater speedups. However, scaling this approach to co-training on large models with long contexts poses two obstacles: (1) branch attention is unsupported by standard causal context-parallel (CP) implementations, and (2) target features span across pipeline-parallel (PP) stages. We address both with an

I have several years of software experience, following best practices, design patterns, KISS, DRY, BDD, OOP, etc. My code is actually very readable since I worked on many open source projects and heard praise overall, and always had the time for quality control and refactoring. The truth is I haven't worked at a normal company since AI hit so I'm a bit detached currently from the industry. For the last 1 year I've been using the Pro subscription for 20$ on Codex and Claude on my projects, but I

I have been a dwarfstar fan for awhile and I really liked glm 5.3 flash but needed it to be materially faster to feel good using it. In the screenshot you can see the outcome of using the model with a claude code harness at ~200k depth, with many tool calls and averaging over 38tps output. Yes, I put 60tps in the headline and you will get that if you ask it to write SQL. Main ds4 was single-stream serially decoding GLM-5.3-Flash at about 59 percent of the M3 Ultra's measured memory bandwidth. We

Linear attention is increasingly used in frontier language models for efficient long-context inference and constant-memory decoding. Its fixed-size recurrent memory, however, requires an online decision at each token: what to write and how strongly to overwrite existing associations before knowing which information future queries will require. Delta-rule models learn this strength from the current token embedding but do not track confidence in the memory estimate, preventing each write from adap

Visual generation is evolving from generative models used through a single invocation into agentic control processes that can plan, select tools, inspect intermediate synthesized outputs, revise failures, and reuse prior experience. In most existing systems, the controller is an LLM or VLM, while visual generation models serve as tools or executors. However, existing work lacks a consistent criterion for determining when a generation system becomes agentic. Planning depth, tool use, multi-role c

arXiv:2609.05441v1 Announce Type: new Abstract: Long-term memory for LLM agents is evaluated today by conversational recall benchmarks (LoCoMo, LongMemEval), which measure question answering over dialogue history, not whether remembered facts change what a tool-using agent does. We present MERIT (Memory Evaluation for Realistic Instrumented Tasks), a benchmark and harness that measures the marginal utility of memory for task-executing agents under explicit cost accounting. MERIT provides episodi

Today, Meta has introduced Muse, a personal AI agent that takes actions rather than just answering questions. Muse can send emails, book travel, negotiate bills, and pursue long term goals. It keeps working after you close the app and returns only when it needs approval. The bigger story for AI devs is architectural. Each user […] The post Meta Introduces Muse, a Personal AI Agent That Runs on Its Own Dedicated Secure Cloud Computer appeared first on MarkTechPost .

Causal inference is the practice of estimating the effect of a treatment or intervention from data. It traditionally requires a bespoke pipeline for every new problem: first proposing a causal mechanism, selecting a compatible estimator, and finally training it. Meanwhile, across diverse settings and modalities, much of machine learning has shifted to the paradigm of foundation models: networks pretrained once at scale and applied to new tasks without fine-tuning. Causal foundation models (CFMs)

arXiv:2609.05511v1 Announce Type: new Abstract: Web agents need to navigate visually rich, long-horizon interfaces that change across sites, yet most previous agents still learn each task in isolation and discard the procedural knowledge they accumulate. Recent skill-augmented frameworks take an important first step, but they treat the skill library as a flat or two-tier prompt-side cache and offer no principled mechanism for compressing redundancy or composing skills recursively. We introduce \

We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, February to March 2026) and the DXAP live alpha fleet (500 to 599 user-created agents all-history, 91 to 117 concurrently active, trading Hyperliquid perpetuals, June to August 2026). The record spans roughly six months, 7.5M sing

A few posts tagged with "new model" present models that are finetunes. My opinion : I'd rather have the "new model" tag reserved for new "major" releases, like a new Qwen model, Deepseek V4 -> Deepseek V4.1, etc., that involved a new pretrain or intensive post-training (in opposition to a small finetune). Otherwise, maybe prepend "[Finetune]" to the title to indicate that the new model is "less of a big news", a use a "new finetune" tag, to differentiate between the two kinds of new models. I re

I have a successful app I built myself with a solid user base. I had been working on a new version via Claude on and off for about six months In work we use LLMs exclusively. Nobody writes code anymore. It's all hands off and we have a high level understanding of how things work but no more than that. When it comes to my side project, I was adding features at breakneck speed with Claude but I realized I have no clue how the new code works or what it changes or breaks. I spent many years of my li

Contract gig, mid-size fintech. Their security team decided any control panel counts as unnecessary attack surface. So no cPanel, no Plesk, nothing with a web UI on prod. Servers get configured by hand over SSH, using a shared root account for the whole team. I run BeAdmin on my own boxes at home. But that's a different world here. Two people editing the same nginx.conf in one week wiped out three days of one guy's changes, and nobody noticed until a client site went down. How common is this? Wh

Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than explicit textual chains of thought. However, visual information being present in a latent state does not imply that the model actually relies on that state when producing its answer, especially when alternative image-conditioned paths remain available. We introduce Causal Visual Recurrent Reasoning (CVRR), which preserves pretrained visual competence while making recurrent computation the re

As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuning methods (e.g., RLHF, DPO) for behavioral control. However, existing work typically validates steering on isolated behaviors, leaving it unclear whether steering vectors encode coherent semantic structure or merely exploit behavior-specific shortcuts. We investigate whether the latent geometry of LLM steering vectors

arXiv:2609.05513v1 Announce Type: new Abstract: Web agents have achieved significant success in automating complex internet tasks but deploying them in real-world environments requires continuous online adaptation. Given that deploying powerful proprietary models remains commercially cost-prohibitive, practitioners must rely on lightweight local models that evolve post-deployment via online teaching from a stronger teacher. However, standard interactive feedback imposes prohibitive costs. We sho

Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produces thousands of redundant visual tokens whose cost grows with every view. Existing visual token pruners fall into two families, each limited in the 3D multi-view setting. Learned importance methods rank tokens by attention or encoder features; because redundancy here is fundamentally spatial, they keep near-duplicate to

arXiv:2609.05439v1 Announce Type: new Abstract: Current evaluation methods for large language models are coarse-grained and decoupled from generation, producing generic explanations that fail to provide actionable feedback for model improvement. We propose CriticGen, a fine-grained, generation-aware evaluation framework that turns evaluation into actionable control for answer improvement. CriticGen first generates sample-specific evaluation dimensions and scoring criteria under high-level catego

arXiv:2609.05505v1 Announce Type: new Abstract: Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review stages in isolation. We introduce SciLitBench, a multi-stage benchmark spanning title and abstract screening, full-text screening, and schema-guided data extraction, with 42,981 retrieved records, 1,012 full texts, and annotations for 888 included papers. Across 22 open-weight LLMs from s

I wrote recently about how the collection of good, fruitful open problems is now being mined in a non-renewable fashion, leading to the potential scenario of these problems becoming scarce. [...] We have now seen that even the rumor of someone working on a problem can trigger a massive amount of AI-powered effort to flatten it before the original research project has time to reach its full potential. The incentives may now be pointing in the direction of no longer sharing any promising research

I can close an eye on using LLMs for research only. These days I feel I am the only one that hasn’t changed their way of writing software at all. I’d honestly switch careers rather than manage agents. But I’m currently out of work (contract ended) and wondering if there is still a sane place to work, or is it truly time to pivot to another career.

Autoregressive (AR) video diffusion models have shown great potential in real-time video generation. Recent methods distill pretrained bidirectional video diffusion models into causal AR students through Distribution Matching Distillation (DMD), but the generated videos often suffer from over-saturation and over-smoothing issues, resulting in limited visual quality and realism. The key contributing factor is the mode-seeking behavior of the reverse KL objective in DMD, which can cause the studen

OpenAI’s latest mathematical milestone has quickly become mired in controversy. Today, the company announced that its agents have solved one of the Millennium Prize Problems, some of the most important open problems in mathematics. Under normal circumstances, that solution would be a huge feather in OpenAI’s cap. But the announcement has been overshadowed by accusations…

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. What OpenAI’s latest controversy tells us about the future of math OpenAI says its agents have solved one of the most important open problems in mathematics. Under normal circumstances, that would…

arXiv:2609.05448v1 Announce Type: new Abstract: Structured post-training pruning of transformers requires selecting complete functional units whose suppression causes limited degradation. We formulate structured-unit selection for language and vision transformers as a damage-aware multi-armed bandit problem under a fixed candidate-evaluation budget. Attention heads and MLP channel groups are temporarily masked on calibration batches. Paired damage is the masked loss minus the base loss on the sa

Vision-Language-Action (VLA) models have shown promising progress in language-conditioned robotic manipulation. However, existing datasets and benchmarks mainly evaluate task completion under predefined settings, offering limited insight into model reasoning under increasing spatial and procedural complexity. We introduce RoboSPA (Robot Spatial-Procedural Assessment), a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in VLA models. RoboSPA focuses on two

arXiv:2609.05527v1 Announce Type: new Abstract: Wherever a coding agent works under engineer supervision, or a clinical model assists a radiologist, the deployment question is whether to keep the human-AI workflow or replace it with the human alone or the agent alone. The human-AI workflow is worth keeping only if it beats both of those alternatives. Yet once it is deployed, neither alternative outcome is observed: recovering one means replaying the task under that alternative, and every replay

Hi r/machinelearning . Nice to meet you! My name is Chris Piech and I'm a professor at Stanford University in the AI lab. I built a class called Probability for AI: pai.stanford.edu. It starts Oct 9th and applications are due end of Sept. Its (hopefully) cool for a few reasons: The plan is to have one volunteer teacher for every 10 students! Apps have been open for a week and over 1,000+ folks have applied to teach. So we might actually be able to make this pretty big. I have built a lot of fun

Chaque matin, un digest tech fait pour vous