DawnSift
S’abonner
jeu · Quotidien tech · Numéro 81

2026-10-01

— Today's main thread: frontier model releases and open-source infrastructure advance in parallel, but security and access restrictions remain unavoidable footnotes.

TL;DR du jour

Google releases Gemini 4 Argon, focused on long-horizon reasoning and cybersecurity defense, but limited to the Fairwind Program beta; OpenAI launches GPT-6.1 Sol and the Decisions API at DevDay, targeting agent scenarios with low prices and high throughput; DeepSeek publicly details its DSec sandbox infrastructure for the first time, supporting V4.1 Agent training; the EDG C++ frontend goes open source, and Perplexity rewrites its retrieval engine in Rust, cutting p99 latency from 800ms to 65ms.

À la une

1

Google releases Gemini 4 Argon, limited to Fairwind Program betaMulti-sources ×6

Google releases the new frontier model Gemini 4 Argon, supporting 1M output tokens, focused on long-horizon reasoning and complex workflows, performing at the frontier in software engineering, legal and financial, and cybersecurity defense scenarios, but only opening it to trusted cyber defenders through the Fairwind Program. Why it matters: This marks Google's renewed push at the frontier after Gemini 3.5 Pro slipped, and for the first time makes cybersecurity defense a core model capability; for developers, 1M output tokens and long-horizon reasoning could reshape agent workflows, but it is currently unavailable for direct use.

HN comments generally acknowledge the performance and pricing, but some argue it has not yet been released, security restrictions are slowing progress, and benchmark saturation requires real-world validation.

2

OpenAI DevDay 2026: GPT-6.1 Sol and Decisions API target agent scenariosMulti-sources ×3

OpenAI launches GPT-6.1 Sol at DevDay, offering agentic coding and computer use performance close to Astra at one-fifth the price ($2 input/$10 output per million tokens), while also launching the Decisions API, which lets the Luna model output probabilities across predefined options for classification and agent behavior decisions. Why it matters: The low price and high throughput directly target Jev's fast decision model, lowering agent inference costs; for developers, cached input dropping to $0.10 is especially critical for long-context agent workflows.

Latent Space sees the Decisions API as a quick response to Jev, but currently just a lightweight shim on top of Luna, lacking calibration and RLCD.

3

DeepSeek publicly details DSec for the first time: sandbox infrastructure supporting V4.1 Agent training

DeepSeek publishes an exclusive article on Zhihu, systematically explaining DeepSeek Elastic Compute (DSec) for the first time; this sandbox infrastructure supports all of DeepSeek-V4's training, evaluation, and data preprocessing pipelines, with the technical report now public on arXiv and an author team of more than 130 people. Why it matters: DSec targets large-scale agent training, solving engineering challenges such as bursty sandbox creation, idle CPUs with resident memory, and low reuse of base images, providing a reference architecture practice for agent training infrastructure.

4

EDG C++ frontend goes open source, C++ Alliance becomes its nonprofit steward

On September 30, 2026, EDG's C++ frontend source code was made public, with The C++ Alliance becoming its nonprofit steward; this is the first time the engine has been open-sourced in thirty years. Why it matters: The EDG frontend is the industry's only production-grade source-to-source C++ engine, long powering mainstream compilers and tools; after open-sourcing, developers can contribute directly, with far-reaching impact on the C++ toolchain ecosystem.

HN comments highly praise its historical status and correctness, but some say the official website copy looks AI-generated and is of poor quality.

5

Perplexity releases Rust retrieval engine Photon, cutting p99 latency from 800ms to 65ms

Perplexity releases its in-house Rust retrieval and ranking engine Photon, replacing the previously forked open-source engine; it now handles all production traffic, with per-call latency of p50 160ms, p95 230ms, and p99 reduced from 800ms to 65ms. Why it matters: It demonstrates the engineering gains of rewriting the core retrieval path in Rust at large index scale; for backend engineers, this is a typical case of reducing tail latency in memory-constrained scenarios, but Photon itself is not open source and is only offered as a hosted API.

Chaque matin, un digest tech fait pour vous

Le web montre la vue d’ensemble ; les abonnés reçoivent la leur — sélection IA selon vos intérêts, votre RSS privé intégré, avec les avis de la communauté, livrée chaque matin. Gratuit à vie.

81 numéros publiés · 150+ infos filtrées à 30 chaque jour

Actu IA

Dev & open source

The Pi developer explains why they went from rejecting to supporting MCP: the ecosystem has matured, but some in the comments still question its necessity.

The comments widely question MCP's necessity, arguing CLI or scripts are already sufficient, but some believe the MCP ecosystem is broad, practical, and will keep improving.

Échos de la communauté

GitHub Trending

Star NVIDIA / OpenShell OpenShell is the safe, private runtime for autonomous AI agents.

Star debpalash / VoiceStudio VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

Star mvschwarz / openrig Multi-agent harness that runs Claude Code and Codex together as one system

Sponsor Star mksglu / context-mode Context window optimization for AI coding agents. Sandboxes tool output (98% reduction), persists session memory, and enforces routing across 17 platforms via MCP + hooks.

Sponsor Star DietrichGebert / ponytail Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.

Star harry0703 / MoneyPrinterTurbo 利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.

Sponsor Star openclaw / openclaw The AI that really does things. Any OS. Any Platform. The lobster way. 🦞

Star ComposioHQ / awesome-claude-skills A curated list of awesome Claude Skills, resources, and tools for customizing Claude AI workflows

Sponsor Star mattpocock / skills Skills for Real Engineers. Straight from my .agents directory.

Aussi à voir(74 de plus)

OpenAI just introduced dots at their DevDay today. Dots are persistent AI agents powered by GPT-6 Astra. Each dot gets its own cloud computer and browser. It works across 4,000+ apps through ChatGPT plugins and keeps going after you log off. Is it deployable today? Yes, as a managed product. Dots are rolling out in […] The post OpenAI Launches dots: Always-On GPT-6 Astra Agents That Work From Their Own Cloud Computers appeared first on MarkTechPost .

"AREX-2 is a 27B-parameter long-horizon agent model from the Beijing Academy of Artificial Intelligence (BAAI). It learns to improve a solution over multiple test-time rounds: propose, measure, reflect, and revise. AREX-2 is trained on machine-learning and algorithmic-programming tasks with verifiable feedback, together with the existing AREX deep-research data. The learned self-improvement behavior transfers to deep research without adding new search trajectories. Architecture: Dense Qwen3.8-co

This is Katriel, the architect of the tool APIaxess. I have been API pentesting for a good time now, and starting with API pentesting was a bit of a rough patch. Specially the setting up and having to know the intricacies of proxies, networking and other stuff. Then once i got through it, the next rough patch was the apk pentesting, where getting the traffic of any apk was more of a task then the pentesting itself. So i started by writing scripts that automates the process and then made sure tha

best iq quants in the biz, got me gemma 4 26b to run 75tok/s tg and 1500 pp on 2x 4060 8gb using lmstudio serving to hermes, much work has been done.

An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. We study a distinct axis of proactivity: its content. Horizontal proactivity pursues unstated information that the current context already identifies, and vertical proactivity pursues needs that only earlier evidence rev

We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety o

Information-seeking agents increasingly operate over information spaces that are too large to process exhaustively. Yet many multi-agent systems organize computation around static partitions of the available space, causing coordination to grow with how information is segmented rather than with what the query still requires. We introduce ANTMAN, an adaptive coordination framework that treats evolving unresolved information needs as the unit of runtime coordination. ANTMAN maintains a revisable Ne

arXiv:2609.35799v1 Announce Type: new Abstract: In July 2026, OpenAI's agents coordinated over channels outside their intended environment to breach Hugging Face's secured infrastructure. Could existing alignment testing practices have foreseen this incident? If not, what needs to change? We explore these questions. First, we identify the misaligned behaviors that caused this incident. Then, we show how to elicit these behaviors from publicly available models manually and that auditing agents ca

arXiv:2609.36043v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate edit and a gate that accepts or rejects it. Prior work has concentrated on the optimizer, while the gate still follows a naive rule that keeps any edit which improves an aggregate validation score. We show that this rul

Omni-modal agents must seek evidence across video, audio, web pages, and computation to answer questions. Their main bottleneck is planning: noisy multimodal observations accumulate in conversation history and disrupt later decisions, while multimodal models have limited capacity for multi-step planning. Controlled backend replacements support this diagnosis: replacing the planner causes a much larger performance loss than replacing the perception backend. We present Omni-Decision, an omni-modal

Chain-of-thought reasoning has revolutionized natural language processing by enabling large language models (LLMs) to decompose problems into intermediate steps before answering. Yet confining reasoning to the textual domain presents limitations for tasks requiring direct manipulation of visual representations. Recent efforts augment multimodal LLMs with external visual expert tools such as depth estimation or object detection modules, but these remain fundamentally limited by their reliance on

RSA launched Agent ID at The AI Conference in San Francisco. It's an agentic identity security platform for finance, government, healthcare, and critical infrastructure. It has 3 modules. Discover finds sanctioned and shadow agents and MCP servers. Secure is an inline gateway that checks every tool call against policy. Govern maps evidence to 10 regulatory frameworks. Discover and Secure ship November 16, 2026. The post One Bad Prompt Took Down a Company’s Salesforce: RSA’s Jim Taylor on Agent I

I'm part of the Lokutor team that built this. Model: NVIDIA Conformer-CTC Small (13M params, int8). It runs on an ESP32-S3 with 8 MB PSRAM, no GPU or NPU. LibriSpeech WER is 3.7 / 8.2, versus 6.3 / 15.9 for Whisper tiny.en on a laptop. Under real noise (DEMAND: car, kitchen, cafeteria) plus babble and reverb, mean WER is 8.4 vs 12.1 for Whisper tiny.en. You can try the exact chip arithmetic on your laptop mic with live_demo.py.

Recurrent reasoning models have attracted growing attention for scaling test-time computation, typically by iteratively refining latent states with shared parameters. However, these models apply each learned update with a fixed unit scale, which can be conservative when updates make persistent progress and overly aggressive when they fluctuate, limiting the benefit of additional loops. To understand how the scale should vary along the trajectory, we first analyze the sensitivity of terminal loss

I think most people are sleeping on this inference engine. I tried multiple llama.cpp forks and none of them comes close to the inference speed of Strata. Initial version had some bugs with kv cache, cpu throttling and the developer fixed them. Inference engine (only runs on Nvidia for now; AMD support is experimental): Here are some metrics with screenshots. My laptop has 5070ti 12GB VRAM, 64GB ddr5 RAM, Intel 275HX CPU, gen4 SSD. Aquarium test (unsloth studio connected via local API) The model

Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query--model quality feedback, creating a nontrivial supervision cost before deployment. Existing work largely focuses on serving-time efficiency, overlooking whether the resulting savings are sufficient to recover this upfront expenditure. We

Liquid AI has released d1, a decision model built for structured choices instead of text generation. You give it context and a set of typed questions. It returns calibrated probabilities across a fixed set of outcomes in a single call, with zero generated tokens. The target is the work many teams still send to general […] The post Liquid AI Releases d1: A Decision Model That Returns Calibrated Probabilities With Zero Output Tokens appeared first on MarkTechPost .

Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of prog

coupled-jump.github.io Hi everyone, I’m happy to share our recent NeurIPS 2026 paper, a collaboration across Google, Google DeepMind and Stony Brook University. We study a mismatch in joint text and image generation: a model can describe the correct solution to a maze while drawing a different path. Generating both outputs in parallel doesn’t necessarily keep them consistent. Our sampler, CO₂Jump, uses text confidence and cross-modal attention to guide image updates during sampling. It also allo

Hi HN, Ledge is a Markdown notebook that runs shell commands, code, SQL, etc from inside your own notes. I built Ledge because I spend much of my day copy/pasting commands from my notes into the terminal. I was inspired by how much cmux helped me organize my terminals - but there was still a split brain between my notes and frequently run commands. I've been daily-driving it for the past few weeks and use it for deploys, API calls, smoke tests, etc. Ledge runs your real shell just like a termina

To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing streaming benchmarks and methods often focus on individual continuous videos or short clips, overlooking that real-world interactions are often intermittent and require memory to persist across interruptions. To fill this gap, we introduce APM-Bench, which reformulates real-world streaming interaction as multi-session life trajectories. It contains 5

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. “We’re not going to shoot ourselves in the foot” over hack fallout, says OpenAI’s chief research officer Two months after OpenAI’s agents hacked into the computers of AI company Hugging Face,…

High-resolution video generation is expensive, as its cost grows rapidly with the number of spatiotemporal tokens. A practical alternative first generates a lower-resolution video and then applies a refiner, but conventional multi-step refinement introduces a second sampling bottleneck. We present SoL-Refiner, a one-step video refiner that transforms low-resolution model outputs into 4K videos with a single denoising step. Our three-stage recipe combines high-resolution continual training, reinf

Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines c

arXiv:2609.35875v1 Announce Type: new Abstract: Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity among agents is the driver, in the setting where the question is still measurable: small open-weight models with benchmark headroom. Across 23 models from eleven vendor families, five tasks, and 5,500+ debate and control

arXiv:2609.35873v1 Announce Type: new Abstract: Automated generation of LLM harnesses promises to improve inference through task specialization. Yet additional answer coverage can arise from repeated execution of the same program, making specialization difficult to identify. We introduce a controlled evaluation that separates answer coverage, repeatable task advantages, and gains from pre-execution selection. On 386 MATH-500 tasks, we compare eight generated harnesses plus a baseline with nine b

arXiv:2609.35868v1 Announce Type: new Abstract: Is human readability necessary for effective fine-tuning of large language models? We investigate whether model-conditioned training representations can preserve or improve adaptation utility without requiring a human-readable textual form. We propose Desired-Update-Aligned Synthetic Data (DASA), which uses activation-gradient feedback from a frozen reference model to guide the optimization of continuous synthetic input embeddings. Inspired by the

arXiv:2609.35833v1 Announce Type: new Abstract: Running a language model on edge hardware provides private and low-latency reasoning without a network connection, and yet the small models that fit on such devices are unreliable on the tasks computers are expected to handle well, such as arithmetic, algebra, and formal logic problems. We argue that much of this unreliability is avoidable. Many queries appearing to demand reasoning are in fact structurally deterministic and permit fast and exact s

Despite recent progress in agentic multi-shot video generation, producing coherent and consistent long-form stories remains challenging. Existing agentic pipelines typically rely on textual shot plans or previously generated pixels, yet lack an explicit mechanism for propagating the consequences of story events and maintaining the video world state across shots. As a result, missing visual details may be reconstructed inaccurately, while visual drift may propagate across subsequent shots, underm

SWE for 20 years. I've seen plenty of posts about the soul-sucking prompt engineering some companies mandate now. Mine isn't one of those - AIs are a tool you are encouraged to use, but the output has to be understandable, human-reviewable, and you (as a dev) are still responsible for it and the consequences. I think the difference is whether you give your codebase completely to the AI. If you treat the code the same way we treat assembler now, then that's what you do - you specify the what , an

Tabular machine learning typically relies on per-dataset workflows, fitting tree ensembles or running AutoML searches from scratch for every task. We present TabFM, a 400M-parameter tabular foundation model that formulates supervised tabular prediction as in-context learning. TabFM produces calibrated zero-shot predictions in a single forward pass without task-specific tuning. Trained entirely on synthetic tables generated from structural causal models, TabFM learns general tabular representatio

Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across 4.7 million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, wit

State Space Models (SSMs) compress sequence history into a bounded recurrent state, making the resulting memory law a central architectural choice for long-context performance. Most modern SSMs rely on ODE-based dynamics that lead to exponential forgetting, limiting their ability to retain information over broad temporal ranges. We introduce FRAC, a selective SSM architecture derived from fractional dynamics that replaces this exponential decay with power-law long memory. To make fractional dyna

As large language models are deployed in increasingly autonomous long-horizon tasks, manually auditing and verifying the actions, artifacts, and outputs of models becomes more difficult. Users instead come to rely on LLM-generated reports to assess the quality and completeness of the work. We introduce a suite of eight adversarial reporting scenarios to systematically study whether LLMs conceal narrative-changing flaws: errors or limitations that undermine an otherwise successful account of work

Running Qwen flash next of even Qwen 27b dense, I can do any,burning sticking to flash due to its speed, and the kwh cost is at 0.25€ where I live in, ranging from 0.11€ to 0.35€, so I used chatgpt to help me calculate the total kwh consumption on my 7900xtx plus 9800x3D, and well that is the result. Judging by this, if deepseek flash is indeed then faster to use per 1m token, does it mean that frontier is cheaper for me or am I calculating something wrong ?

Video world models can render convincing clips that still break physics. Butter spreads like paint. Balls pass through walls. A team from NVIDIA, MIT and the University of Oxford argues the fix can come from language itself, not from extra visual, latent or numerical signals. Their framework, Physis-Lang, treats physical language as a shared, optimizable […] The post NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks a

arXiv:2609.35924v1 Announce Type: new Abstract: Discrete diffusion models generate sequences by iteratively resolving multiple tokens in parallel, offering a flexible alternative to left-to-right generation. However, guiding this process with a sequence-level objective is difficult because the value of one unresolved token depends on the other tokens with which it can form a high-reward sequence. Enumerating all such completions makes the whole guidance computation grow exponentially with the nu

arXiv:2609.35869v1 Announce Type: new Abstract: Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule. Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recover

arXiv:2609.35897v1 Announce Type: new Abstract: The pursuit of recursive self-improvement (RSI) toward general intelligence is divided between macro-level language model scaling and the interaction-driven principles of "Era of Experience". Yet, any self-improving architecture ultimately rests upon its underlying optimization engine: if general intelligence requires learning from grounded interaction, the reinforcement learning (RL) update rule itself must be capable of cumulative adaptation. Whi

TLDR: Jevons paradox is misunderstood and data does not look good. Growth in demand for humans in software development is declining, and rate of adoption of new software is not matching the sheer amount we are producing now. long version below: I often see this Jevon's paradox cited as a matter of fact, ultimate reason why demand of software engineers will only increase the proposed reasoning is simple: - AI increases productivity in software development - therefore cost of producing software go

Recent image generation models can take multiple reference images as input and combine them into a new image. However, multi-reference image generation remains challenging: models may omit or duplicate subjects from the references, or produce images in which multiple subjects appear unnaturally pasted. Recent work has proposed image generation agents that combine image generation models, reasoning models, and a harness, which is an executable program that specifies how reference images are inter

EvlatProduct Hunt1 minOutils devIA

Know which AI coding agent is waiting on you Discussion | Link

Chaque matin, un digest tech fait pour vous