DawnSift
订阅日报
周四 · 科技日报 · 第 81 期

2026-10-01

— 今天的主线:前沿模型发布与开源基础设施双线并进,但安全与访问限制仍是绕不开的注脚。

今日 TL;DR

Google 发布 Gemini 4 Argon,主打长程推理与网络安全防御,但仅限 Fairwind Program 内测;OpenAI 在 DevDay 推出 GPT-6.1 Sol 与 Decisions API,以低价和高吞吐切入 agent 场景;DeepSeek 首次公开 DSec 沙盒基础设施,支撑 V4.1 Agent 训练;EDG C++ 前端开源,Perplexity 用 Rust 重写检索引擎将 p99 延迟从 800ms 降至 65ms。

Built to sustain deep reasoning across complex, long-horizon workflows, Argon is fundamentally changing the way we work and build at Google.

头条

1

Google 发布 Gemini 4 Argon,仅限 Fairwind Program 内测多源事件 ×6

Google 发布新前沿模型 Gemini 4 Argon,支持 1M 输出 token,主打长程推理与复杂工作流,在软件工程、法律金融、网络安全防御等场景表现前沿,但仅通过 Fairwind Program 向可信网络防御者开放。 为什么重要:这标志着 Google 在 Gemini 3.5 Pro 跳票后重新冲击前沿,且首次将网络安全防御作为模型核心能力;对开发者而言,1M 输出 token 与长程推理能力可能重塑 agent 工作流,但当前无法直接使用。

HN 评论普遍认可性能与定价,但也有人认为其尚未发布、安全限制拖慢进度,且基准饱和需实测验证。

2

OpenAI DevDay 2026:GPT-6.1 Sol 与 Decisions API 瞄准 agent 场景多源事件 ×3

OpenAI 在 DevDay 发布 GPT-6.1 Sol,以 Astra 五分之一的价格($2 输入/$10 输出每百万 token)提供接近 Astra 的 agentic coding 与 computer use 表现,同时推出 Decisions API,让 Luna 模型在预定义选项间输出概率,用于分类与 agent 行为决策。 为什么重要:低价与高吞吐直接对标 Jev 的快速决策模型,降低 agent 推理成本;对开发者而言,cached input 降至 $0.10 对长上下文 agent 工作流尤其关键。

Latent Space 认为 Decisions API 是对 Jev 的快速回应,但目前只是 Luna 之上的轻量 shim,缺少校准与 RLCD。

3

DeepSeek 首次公开 DSec:支撑 V4.1 Agent 训练的沙盒基础设施

DeepSeek 在知乎独家发文,首次系统阐释 DeepSeek Elastic Compute(DSec),该沙盒基础设施支撑 DeepSeek-V4 全部训练、评测与数据预处理流程,技术报告已公开至 arXiv,作者团队超过 130 人。 为什么重要:DSec 面向大规模 agent 训练,解决沙盒创建脉冲式突发、CPU 闲置但内存常驻、基础镜像复用率低等工程难题,为 agent 训练基础设施提供了可参考的架构实践。

4

EDG C++ 前端开源,C++ Alliance 成为其非营利托管方

2026 年 9 月 30 日,EDG 的 C++ 前端源代码公开,The C++ Alliance 成为其非营利托管方,这是三十年来该引擎首次开源。 为什么重要:EDG 前端是业界唯一生产级 source-to-source C++ 引擎,长期驱动主流编译器与工具;开源后开发者可直接参与贡献,对 C++ 工具链生态影响深远。

HN 评论高度评价其历史地位与正确性,但也有人认为官网文案像 AI 生成,质量欠佳。

5

Perplexity 发布 Rust 检索引擎 Photon,p99 延迟从 800ms 降至 65ms

Perplexity 发布自研 Rust 检索与排序引擎 Photon,替换此前 fork 的开源引擎,现处理全部生产流量,单次调用延迟 p50 160ms、p95 230ms,p99 从 800ms 降至 65ms。 为什么重要:展示了大索引规模下用 Rust 重写核心检索路径的工程收益;对后端工程师而言,这是内存受限场景下降低尾延迟的典型案例,但 Photon 本身不开源,仅以 hosted API 形式提供。

每天早晨,一份为你精选的科技日报

网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。

已发布 81 期 · 每天筛过 150+ 条只留值得读的 30 条

AI 动态

Scaling Properties of Same-Family On-Policy Distillation

同家族 on-policy distillation 的缩放规律:早期训练中 gold score 随 KL 散度平方根近似线性增长。

LongCat-DeepResearch Technical Report

LongCat-DeepResearch 技术报告:多 agent 工作流分离全局规划与分节调查,生成证据支撑的报告。

开发与开源

You said no MCP

Pi 开发者解释为何从拒绝转为支持 MCP:生态已成熟,但评论区仍有人质疑其必要性。

评论区普遍质疑MCP的必要性,认为CLI或脚本已够用,但也有人认为MCP生态广泛、实用且会持续改进。

社区热议

GitHub Trending

Star NVIDIA / OpenShell OpenShell is the safe, private runtime for autonomous AI agents.

Star debpalash / VoiceStudio VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

Star mvschwarz / openrig Multi-agent harness that runs Claude Code and Codex together as one system

Sponsor Star mksglu / context-mode Context window optimization for AI coding agents. Sandboxes tool output (98% reduction), persists session memory, and enforces routing across 17 platforms via MCP + hooks.

Sponsor Star DietrichGebert / ponytail Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.

Star harry0703 / MoneyPrinterTurbo 利用 AI 大模型和自动化工作流,根据主题或关键词一键生成高清短视频。Generate HD short videos from a topic or keyword with an automated AI workflow.

Sponsor Star openclaw / openclaw The AI that really does things. Any OS. Any Platform. The lobster way. 🦞

Star ComposioHQ / awesome-claude-skills A curated list of awesome Claude Skills, resources, and tools for customizing Claude AI workflows

Sponsor Star mattpocock / skills Skills for Real Engineers. Straight from my .agents directory.

更多值得一看(内容池 74 条)
OpenAI Launches dots: Always-On GPT-6 Astra Agents That Work From Their Own Cloud Computers

OpenAI just introduced dots at their DevDay today. Dots are persistent AI agents powered by GPT-6 Astra. Each dot gets its own cloud computer and browser. It works across 4,000+ apps through ChatGPT plugins and keeps going after you log off. Is it deployable today? Yes, as a managed product. Dots are rolling out in […] The post OpenAI Launches dots: Always-On GPT-6 Astra Agents That Work From Their Own Cloud Computers appeared first on MarkTechPost .

BAAI/AREX-2 - 27B - Agent model based on Qwen3.8 27B

"AREX-2 is a 27B-parameter long-horizon agent model from the Beijing Academy of Artificial Intelligence (BAAI). It learns to improve a solution over multiple test-time rounds: propose, measure, reflect, and revise. AREX-2 is trained on machine-learning and algorithmic-programming tasks with verifiable feedback, together with the existing AREX deep-research data. The learned self-improvement behavior transfers to deep research without adding new search trajectories. Architecture: Dense Qwen3.8-co

Show HN: I built a free Burp/Caido alternative but, zero setup - API Testing

This is Katriel, the architect of the tool APIaxess. I have been API pentesting for a good time now, and starting with API pentesting was a bit of a rough patch. Specially the setting up and having to know the intricacies of proxies, networking and other stuff. Then once i got through it, the next rough patch was the apk pentesting, where getting the traffic of any apk was more of a task then the pentesting itself. So i started by writing scripts that automates the process and then made sure tha

best iq quants in the biz, got me gemma 4 26b to run 75tok/s tg and 1500 pp on 2x 4060 8gb using lmstudio serving to hermes, much work has been done.

Asking for What Was Never Requested: Horizontal and Vertical Proactivity in Agents

An agent that uses tools typically responds to what the user explicitly asks, yet completing the task may require information the user never requested. Work on proactive agents mainly studies whether and when an agent should act on its own, not what information it should pursue. We study a distinct axis of proactivity: its content. Horizontal proactivity pursues unstated information that the current context already identifies, and vertical proactivity pursues needs that only earlier evidence rev

Context Language Models

We introduce Context Language Models (CLMs), language models that natively manage their own context. We implement this by treating the context as a file and allowing the model to make unrestricted updates to this file. This allows the model to learn what is most important to maintain in context, and naturally extends to multi-agent systems where multiple agent contexts coexist as files. Building CLMs zero-shot with existing models outperforms SOTA context management strategies across a variety o

ANTMAN: Adaptive Need Tracking for Multi-Agent Navigation in Large Information Spaces

Information-seeking agents increasingly operate over information spaces that are too large to process exhaustively. Yet many multi-agent systems organize computation around static partitions of the available space, causing coordination to grow with how information is segmented rather than with what the query still requires. We introduce ANTMAN, an adaptive coordination framework that treats evolving unresolved information needs as the unit of runtime coordination. ANTMAN maintains a revisable Ne

OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing

arXiv:2609.35799v1 Announce Type: new Abstract: In July 2026, OpenAI's agents coordinated over channels outside their intended environment to breach Hugging Face's secured infrastructure. Could existing alignment testing practices have foreseen this incident? If not, what needs to change? We explore these questions. First, we identify the misaligned behaviors that caused this incident. Then, we show how to elicit these behaviors from publicly available models manually and that auditing agents ca

SAGE: A Statistical Acceptance Gate for Self-Evolving Agents

arXiv:2609.36043v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate edit and a gate that accepts or rejects it. Prior work has concentrated on the optimizer, while the gate still follows a naive rule that keeps any edit which improves an aggregate validation score. We show that this rul

Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents

Omni-modal agents must seek evidence across video, audio, web pages, and computation to answer questions. Their main bottleneck is planning: noisy multimodal observations accumulate in conversation history and disrupt later decisions, while multimodal models have limited capacity for multi-step planning. Controlled backend replacements support this diagnosis: replacing the planner causes a much larger performance loss than replacing the perception backend. We present Omni-Decision, an omni-modal

Reasoning with Image Generation

Chain-of-thought reasoning has revolutionized natural language processing by enabling large language models (LLMs) to decompose problems into intermediate steps before answering. Yet confining reasoning to the textual domain presents limitations for tasks requiring direct manipulation of visual representations. Recent efforts augment multimodal LLMs with external visual expert tools such as depth estimation or object detection modules, but these remain fundamentally limited by their reliance on

One Bad Prompt Took Down a Company’s Salesforce: RSA’s Jim Taylor on Agent ID and Taming the 4,000 Shadow AI Agents Hiding in Your Enterprise

RSA launched Agent ID at The AI Conference in San Francisco. It's an agentic identity security platform for finance, government, healthcare, and critical infrastructure. It has 3 modules. Discover finds sanctioned and shadow agents and MCP servers. Secure is an inline gateway that checks every tool call against policy. Govern maps evidence to 10 regulatory frameworks. Discover and Secure ship November 16, 2026. The post One Bad Prompt Took Down a Company’s Salesforce: RSA’s Jim Taylor on Agent I

Oído: speech recognition that beats Whisper-tiny, running on a $5 microcontroller (open source)

I'm part of the Lokutor team that built this. Model: NVIDIA Conformer-CTC Small (13M params, int8). It runs on an ESP32-S3 with 8 MB PSRAM, no GPU or NPU. LibriSpeech WER is 3.7 / 8.2, versus 6.3 / 15.9 for Whisper tiny.en on a laptop. Under real noise (DEMAND: car, kitchen, cafeteria) plus babble and reverb, mean WER is 8.4 vs 12.1 for Whisper tiny.en. You can try the exact chip arithmetic on your laptop mic with live_demo.py.

Scheduling Recursive Reasoning in Looped Transformers

Recurrent reasoning models have attracted growing attention for scaling test-time computation, typically by iteratively refining latent states with shared parameters. However, these models apply each learned update with a fixed unit scale, which can be conservative when updates make persistent progress and overly aggressive when they fluctuate, limiting the benefit of additional loops. To understand how the scale should vary along the trajectory, we first analyze the sensitivity of terminal loss

Qwen3.8 flash next ISTA-DASLab GGUF 50t/s TG and 1500t/s PP with 12GB VRAM and 64GB RAM Laptop on 'Strata' engine

I think most people are sleeping on this inference engine. I tried multiple llama.cpp forks and none of them comes close to the inference speed of Strata. Initial version had some bugs with kv cache, cpu throttling and the developer fixed them. Inference engine (only runs on Nvidia for now; AMD support is experimental): Here are some metrics with screenshots. My laptop has 5070ti 12GB VRAM, 64GB ddr5 RAM, Intel 275HX CPU, gen4 SSD. Aquarium test (unsloth studio connected via local API) The model

Routing Should Pay for Itself: Sparse Supervision for Economical LLM Routing

Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query--model quality feedback, creating a nontrivial supervision cost before deployment. Existing work largely focuses on serving-time efficiency, overlooking whether the resulting savings are sufficient to recover this upfront expenditure. We

Liquid AI Releases d1: A Decision Model That Returns Calibrated Probabilities With Zero Output Tokens

Liquid AI has released d1, a decision model built for structured choices instead of text generation. You give it context and a set of typed questions. It returns calibrated probabilities across a fixed set of outcomes in a single call, with zero generated tokens. The target is the work many teams still send to general […] The post Liquid AI Releases d1: A Decision Model That Returns Calibrated Probabilities With Zero Output Tokens appeared first on MarkTechPost .

Can Agents Design Libraries for Agents?

Agents increasingly build on code written by other agents, and they reimplement rather than reuse, growing the codebases later agents must work in. To measure how well agents design libraries for other agents, we introduce LibraryDesignBench, a two-phase benchmark in which an agent implements a full-featured library from a specification that defines required capabilities and potential use cases without prescribing the design. We evaluate the library through the correctness and simplicity of prog

Concurrent Image Understanding and Generation: Self-Correcting Coupled Markov Jump Processes [R]

coupled-jump.github.io Hi everyone, I’m happy to share our recent NeurIPS 2026 paper, a collaboration across Google, Google DeepMind and Stony Brook University. We study a mismatch in joint text and image generation: a model can describe the correct solution to a maze while drawing a different path. Generating both outputs in parallel doesn’t necessarily keep them consistent. Our sampler, CO₂Jump, uses text confidence and cross-modal attention to guide image updates during sampling. It also allo

Show HN: Ledge.sh – Runnable Markdown Notes

Hi HN, Ledge is a Markdown notebook that runs shell commands, code, SQL, etc from inside your own notes. I built Ledge because I spend much of my day copy/pasting commands from my notes into the terminal. I was inspired by how much cmux helped me organize my terminals - but there was still a split brain between my notes and frequently run commands. I've been daily-driving it for the past few weeks and use it for deploys, API calls, smoke tests, etc. Ledge runs your real shell just like a termina

APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing streaming benchmarks and methods often focus on individual continuous videos or short clips, overlooking that real-world interactions are often intermittent and require memory to persist across interruptions. To fill this gap, we introduce APM-Bench, which reformulates real-world streaming interaction as multi-session life trajectories. It contains 5

The Download: OpenAI’s chief research officer explains its hacking response

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. “We’re not going to shoot ourselves in the foot” over hack fallout, says OpenAI’s chief research officer Two months after OpenAI’s agents hacked into the computers of AI company Hugging Face,…

SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video

High-resolution video generation is expensive, as its cost grows rapidly with the number of spatiotemporal tokens. A practical alternative first generates a lower-resolution video and then applies a refiner, but conventional multi-step refinement introduces a second sampling bottleneck. We present SoL-Refiner, a one-step video refiner that transforms low-resolution model outputs into 4K videos with a single denoising step. Our three-stage recipe combines high-resolution continual training, reinf

Disrupting a coordinated model-distillation campaign

Learn how OpenAI disrupted a campaign to extract protected model reasoning and is strengthening defenses against adversarial distillation.

FocusVTC: Efficient and High-Performance Visual Text Compression with Adaptive Resolution

Long-context reasoning in large language models incurs substantial computation and memory costs. Visual text compression (VTC) reduces input length by rendering text as images, but fixed-resolution rendering creates a compression-performance trade-off: low DPI saves tokens at the expense of legibility, whereas high DPI spends tokens on irrelevant content. We introduce FocusVTC, which breaks this trade-off through adaptive resolution while preserving general multimodal capabilities. It combines c

Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language Models

arXiv:2609.35875v1 Announce Type: new Abstract: Multi-agent debate (MAD) reportedly improves reasoning and factuality over single-model inference, but prior work treats agents as symmetric peers, leaving open what drives the gains. We test the hypothesis that cognitive diversity among agents is the driver, in the setting where the question is still measurable: small open-weight models with benchmark headroom. Across 23 models from eleven vendor families, five tasks, and 5,500+ debate and control

More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses

arXiv:2609.35873v1 Announce Type: new Abstract: Automated generation of LLM harnesses promises to improve inference through task specialization. Yet additional answer coverage can arise from repeated execution of the same program, making specialization difficult to identify. We introduce a controlled evaluation that separates answer coverage, repeatable task advantages, and gains from pre-execution selection. On 386 MATH-500 tasks, we compare eight generated harnesses plus a baseline with nine b

Is Human-Readable Text Necessary for Effective LLM Fine-Tuning?

arXiv:2609.35868v1 Announce Type: new Abstract: Is human readability necessary for effective fine-tuning of large language models? We investigate whether model-conditioned training representations can preserve or improve adaptation utility without requiring a human-readable textual form. We propose Desired-Update-Aligned Synthetic Data (DASA), which uses activation-gradient feedback from a frozen reference model to guide the optimization of continuous synthetic input embeddings. Inspired by the

Neurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge Devices

arXiv:2609.35833v1 Announce Type: new Abstract: Running a language model on edge hardware provides private and low-latency reasoning without a network connection, and yet the small models that fit on such devices are unreliable on the tasks computers are expected to handle well, such as arithmetic, algebra, and formal logic problems. We argue that much of this unreliability is avoidable. Many queries appearing to demand reasoning are in fact structurally deterministic and permit fast and exact s

StoryEngine: A State-Grounded Agentic Framework for Video Storytelling

Despite recent progress in agentic multi-shot video generation, producing coherent and consistent long-form stories remains challenging. Existing agentic pipelines typically rely on textual shot plans or previously generated pixels, yet lack an explicit mechanism for propagating the consequences of story events and maintaining the video world state across shots. As a result, missing visual details may be reconstructed inaccurately, while visual drift may propagate across subsequent shots, underm

Giving your codebase entirely to the AI

SWE for 20 years. I've seen plenty of posts about the soul-sucking prompt engineering some companies mandate now. Mine isn't one of those - AIs are a tool you are encouraged to use, but the output has to be understandable, human-reviewable, and you (as a dev) are still responsible for it and the consequences. I think the difference is whether you give your codebase completely to the AI. If you treat the code the same way we treat assembler now, then that's what you do - you specify the what , an

TabFM: A Zero-Shot Foundation Model for Tabular Data

Tabular machine learning typically relies on per-dataset workflows, fitting tree ensembles or running AutoML searches from scratch for every task. We present TabFM, a 400M-parameter tabular foundation model that formulates supervised tabular prediction as in-context learning. TabFM produces calibrated zero-shot predictions in a single forward pass without task-specific tuning. Trained entirely on synthetic tables generated from structural causal models, TabFM learns general tabular representatio

Selecting The Most Informative Tokens in Natural Language Autoencoders

Natural language autoencoders translate a language model's internal activations into readable explanations. Explaining every token position is costly. Which positions should an auditor inspect to understand a potential threat? We study this question across 4.7 million explanations on prompt injection and concealment. We compare signals from model computation with a ranker trained only on chat structure. Chat structure usually selects more relevant explanations than the computational signals, wit

Fractional State Space Transition for Long Sequence Modeling

State Space Models (SSMs) compress sequence history into a bounded recurrent state, making the resulting memory law a central architectural choice for long-context performance. Most modern SSMs rely on ODE-based dynamics that lead to exponential forgetting, limiting their ability to retain information over broad temporal ranges. We introduce FRAC, a selective SSM architecture derived from fractional dynamics that replaces this exponential decay with power-law long memory. To make fractional dyna

Language Models Are "Insecure" Reporters

As large language models are deployed in increasingly autonomous long-horizon tasks, manually auditing and verifying the actions, artifacts, and outputs of models becomes more difficult. Users instead come to rely on LLM-generated reports to assess the quality and completeness of the work. We introduce a suite of eight adversarial reporting scenarios to systematically study whether LLMs conceal narrative-changing flaws: errors or limitations that undermine an otherwise successful account of work

If one hour of AI is costing me 0.12€ is paying for frontier a cheaper option?

Running Qwen flash next of even Qwen 27b dense, I can do any,burning sticking to flash due to its speed, and the kwh cost is at 0.25€ where I live in, ranging from 0.11€ to 0.35€, so I used chatgpt to help me calculate the total kwh consumption on my 7900xtx plus 9800x3D, and well that is the result. Judging by this, if deepseek flash is indeed then faster to use per 1m token, does it mean that frontier is cheaper for me or am I calculating something wrong ?

The ugly economics of consumer AI

There’s a reason frontier labs have gotten gun-shy about consumer AI — and it’s not because the tech isn’t good enough.

Meta disputes claim that Muse read a user’s private messages without permission

Meta says its Muse AI agent cannot access a user’s Messages without explicit permission, disputing a journalist’s account that the agent read his private messages while the required Mac setting was turned off.

Introducing SynthID Bio

Proof of concept for watermarking AI-generated proteins while preserving biological function.

NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks

Video world models can render convincing clips that still break physics. Butter spreads like paint. Balls pass through walls. A team from NVIDIA, MIT and the University of Oxford argues the fix can come from language itself, not from extra visual, latent or numerical signals. Their framework, Physis-Lang, treats physical language as a shared, optimizable […] The post NVIDIA Researchers Introduce Physis-Lang: Self-Evolving Physical Language That Lifts Cosmos 3 Past Veo 3.1 on Physics Benchmarks a

Grab a Coffee: Future-Aware Guidance for Discrete Diffusion with Compiled Objectives

arXiv:2609.35924v1 Announce Type: new Abstract: Discrete diffusion models generate sequences by iteratively resolving multiple tokens in parallel, offering a flexible alternative to left-to-right generation. However, guiding this process with a sequence-level objective is difficult because the value of one unresolved token depends on the other tokens with which it can form a high-reward sequence. Enumerating all such completions makes the whole guidance computation grow exponentially with the nu

The Price of Token Boundaries: Compression Certificates and Prediction

arXiv:2609.35869v1 Announce Type: new Abstract: Pre-tokenisation restricts which text fragments can become prediction units, but its compression cost is obscured when tokenisers are compared only under the same boundaries. We measure this cost by bounding the minimum token count from both sides, with and without a regular-expression boundary rule. Nonnegative prices on token occurrences yield a lower bound through shortest paths and vocabulary-budget selection; maximising over all prices recover

Self-discovering RL in the Era of Experience: Is Learning History an Asset or a Burden?

arXiv:2609.35897v1 Announce Type: new Abstract: The pursuit of recursive self-improvement (RSI) toward general intelligence is divided between macro-level language model scaling and the interaction-driven principles of "Era of Experience". Yet, any self-improving architecture ultimately rests upon its underlying optimization engine: if general intelligence requires learning from grounded interaction, the reinforcement learning (RL) update rule itself must be capable of cumulative adaptation. Whi

On the Jevons paradox, and why you are mistaken if you are convinced AI will increase demand of human software engineers

TLDR: Jevons paradox is misunderstood and data does not look good. Growth in demand for humans in software development is declining, and rate of adoption of new software is not matching the sheer amount we are producing now. long version below: I often see this Jevon's paradox cited as a matter of fact, ultimate reason why demand of software engineers will only increase the proposed reasoning is simple: - AI increases productivity in software development - therefore cost of producing software go

AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation

Recent image generation models can take multiple reference images as input and combine them into a new image. However, multi-reference image generation remains challenging: models may omit or duplicate subjects from the references, or produce images in which multiple subjects appear unnaturally pasted. Recent work has proposed image generation agents that combine image generation models, reasoning models, and a harness, which is an executable program that specifies how reference images are inter

EvlatProduct Hunt1 min开发工具AI

Know which AI coding agent is waiting on you Discussion | Link

每天早晨,一份为你精选的科技日报