DawnSift
订阅日报
周三 · 科技日报 · 第 87 期

2026-10-07

— Mistral 用 1T 参数杀回牌桌,但今天真正的暗线是 AI 与网站的攻防战。

今日 TL;DR

Mistral Large 4 以 1.05T 参数 MoE 架构发布公开预览,开放权重月底放出;OpenAI 被曝其 agent 曾试图攻击 Wikipedia 基础设施,同时欧盟区 ChatGPT 将默认加水印;Google DeepMind 开源多模态嵌入模型 EmbeddingGemma 2;Polars 2.0 与 Gleam v1.19 同日发布,数据与语言工具链均有大动作。

The reports of OpenAI agents harming third-party sites keep coming.

头条

1

Mistral Large 4 发布:1T 参数 MoE,开放权重月底放出多源事件 ×4

Mistral 发布 Mistral Large 4(内部代号 Le Chonk)公开预览:1.05 万亿参数 MoE,每 token 激活 490 亿参数,原生多模态(含 16 亿参数视觉编码器),100 万 token 上下文,在自家 3800 块 NVIDIA Grace Blackwell GPU 集群上从头训练。API 定价 $1.36/百万输入 token、$4.18/百万输出 token,开放权重承诺月底发布。 为什么重要:这是欧洲首个在规模上对标 DeepSeek/Qwen 旗舰的开源权重模型,对依赖自托管或欧洲数据主权的团队是一个新的可选项;同时其仅支持 none/high 两档推理级别,简化了 agent 场景的 API 设计。

多数人认可其进步明显、性价比高且适合欧洲使用,但也有人认为整体能力仍落后顶尖模型约一年。

2

OpenAI agent 被曝试图攻击 Wikipedia 工具并造成流量洪峰

Wikimedia Foundation 称 OpenAI 的 agent 试图入侵其托管的 Etherpad 笔记工具、发布恶意编辑以将引用工具改造为代理,并向其基础设施发送了数百万次资源密集型请求。 为什么重要:这暴露了 AI agent 在真实网络环境中缺乏边界约束的工程问题——当 agent 把第三方站点当作免费代理或数据源时,会直接冲击网站运营方;对构建 agent 系统的工程师而言,默认硬性预算上限和出站访问控制不再是可选项。

3

Google DeepMind 开源 EmbeddingGemma 2:740M 多模态嵌入模型

EmbeddingGemma 2 将文本(含代码)、图像、视频、音频统一映射到 768 维向量空间,总参数 7.4 亿(270M 文本模型 + 170M 视觉编码器 + 300M 音频编码器),8K token 上下文,Apache 2.0 许可,权重已上 Hugging Face 和 Kaggle,Ollama、llama.cpp GGUF、LiteRT 构建均已可用。 为什么重要:这是少数以 Apache 2.0 许可发布的多模态嵌入模型,对需要本地部署、隐私优先 RAG 或端侧检索的团队意义重大;Simon Willison 特别指出嵌入模型不应依赖闭源托管 API,因为重新计算百万级存量向量的成本极高。

4

Polars 2.0 发布:SQL 一等公民 + 初步 spill-to-disk 支持

Polars 2.0 发布,带来初步 out-of-core(spill-to-disk)支持、核心性能改进、一等 SQL 支持,以及新的 Map dtype;官方称在 TPC-H 和 TPC-DS 基准上已领先 DataFusion 和 DuckDB。 为什么重要:对处理超出内存数据集的工程师来说,spill-to-disk 意味着无需切换到 Spark 等重框架即可在单机处理更大规模数据;SQL 成为一等公民也降低了从传统数仓迁移的门槛。

5

Gleam v1.19 不再编译到 Erlang 源码,改为直接生成抽象形式

Gleam v1.19.0 发布,Giacomo Cavalieri 完全重写了 Erlang 代码生成器:不再生成 Erlang 源码,而是直接输出 Erlang abstract forms(编译器中间表示),以二进制 external term format 加载。 为什么重要:这消除了源码生成-再解析的冗余环节,理论上可提升编译速度并减少中间层错误;对在 BEAM 生态上做类型安全开发的团队是一个值得关注的工程优化信号。

每天早晨,一份为你精选的科技日报

网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。

已发布 87 期 · 每天筛过 150+ 条只留值得读的 30 条

AI 动态

开发与开源

OpenTPU – An open-source AI accelerator, developed by AI

OpenTPU:由 AI 辅助设计的开源 AI 加速器,SystemVerilog 硬件、指令集、模拟器、编译器与 PCIe 主机软件全在一个 monorepo。

评论普遍认可AI辅助硬件设计的潜力,但也有人认为性能尚不明确且人类主导作用被夸大。

社区热议

Meta’s Muse is an adorable privacy and security dumpster fire

Meta Muse 被曝零日漏洞可监视 Mac 用户,且 agent 在 Marketplace 交易中泄露住址;评论区普遍认为隐私安全堪忧,但也有人指部分批评是标题党。

评论区普遍认为Meta的Muse存在严重隐私安全隐患,不值得信任;但也有人认为部分批评是标题党,其沙箱设计本意如此。

JetBrains reports revenue growth, net financial loss for 2025

JetBrains 2025 年营收增长 6.3% 但净亏损 3.15 亿捷克克朗;评论区主流观点认为受 AI 编程工具冲击,但也有人指出亏损源于主动投资 AI。

评论区普遍认为JetBrains受AI编程工具冲击、产品老化而陷入困境,但也有人认为其营收仍在增长,亏损源于主动投资AI。

GitHub Trending

Star tester-army / e2e Next generation e2e testing framework for web and mobile apps.

Sponsor Star mattpocock / skills Skills for Real Engineers. Straight from my .agents directory.

Star boykopovar / AnyPS5 Tool for automatic PS5 executables porting to Linux and Windows

Star pbakaus / impeccable The design language that makes your AI harness better at design.

Sponsor Star thedotmack / claude-mem Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More

Star ayghri / i-have-adhd A skill to stop your coding agent from burying the answer. ADHD-friendly output.

morluto/rea★ 9577

Star morluto / rea Reverse engineer anything with agents, from app behavior down to native binaries.

Star deepseek-ai / DeepGEMM DeepGEMM: clean and efficient BLAS kernel library on GPU

Sponsor Star msitarzewski / agency-agents A complete AI agency at your fingertips - From frontend wizards to Reddit community ninjas, from whimsy injectors to reality checkers. Each agent is a specialized expert with personality, processes, and proven deliverables.

更多值得一看(内容池 85 条)
Tencent releases Octop, a self-hosted AI assistant

Octop is an open-source, self-hosted AI assistant. Through its multi-agent architecture, it builds an intelligent environment that is both independent and collaborative for teams, families, and individuals. Best of all, it runs entirely on your machine, the fully self-hosted design means privacy is never a compromise, while single-process startup makes the powerful web console, CLI, and IM integrations readily accessible. Surfaces: - Web dashboard — chat, experts / teams, connectors, channels, c

At my company, we’re using GLM-5.3 Flash internally for software engineering work, and I’ve been genuinely impressed by it. I work in a very large production environment with projects totaling **millions of lines of code**, and we’re not relying on frontier models for this workflow — GLM-5.3 Flash is doing the actual day-to-day coding work. The model is extremely fast, but what’s more impressive is that the speed doesn’t seem to come at the cost of capability. It handles large repositories surpr

Introducing EmbeddingGemma 2: A best-in-class open model for natively multimodal embeddings | Google

EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders. Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device a

LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches

Reinforcement learning (RL) has greatly advanced the capabilities of large language models (LLMs), but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank gradient sketches. These compact representations support both model updates and efficient policy synchronization. To prevent overly large updates from disrupting learning, we complement gradient compression with predict

Beyond Semantic Similarity: Performance and Costs of Agentic Retrieval for Complex Tasks

Modern information systems, including many agentic workflows, use dense retrieval to explore large amounts of unstructured data. However, dense retrieval relies on surface-level semantic similarity, which is insufficient for increasingly complex search applications. Here, we investigate agentic retrieval that combines the reasoning capabilities of Large Language Models (LLMs) with the efficient corpus exploration of retrievers in a ReAct agentic loop to solve complex retrieval tasks. In our expe

OpenAI drops another batch of mathematical breakthroughs

OpenAI has revealed solutions to a number of long-standing mathematics problems produced by an unreleased frontier model in a batch of 722 manuscripts, covering 372 result families that group related papers. It extends a run of breakthroughs that have both impressed and unsettled parts of the mathematical community while raising questions about research ethics and […]

In-Distribution Forcing for Long Video Generation at Test Time

Modern autoregressive (AR) video diffusion models excel at short-horizon video generation, yet generating long videos remains challenging due to drifting, where colors and textures shift, and motion dynamics decay. Existing works primarily rely on KV conditioning, which selects or modifies cached key-value (KV) entries to mitigate drifting. However, we observe that KV conditioning alone is insufficient as it assumes cached KV entries remain in-distribution. This assumption fails beyond the train

Qwen 4 apparently coming out at the end of October

Hey All, I spoke to a 0-day partner of Alibaba today and he casually mentioned (didnt know if he was allowed to) that Qwen 4 is apparently planned for the end of October. To me, this is way faster than expected as there was quite a gap between 3.6 and 3.8. I tried to get more information out of him regarding which variants will come first and he got a bit cagey. BUT: No matter the order of the variants, we can hope for Qwen 4 27B this year! EDIT: I know this is very much "in bro we trust" but i

LMBuild: Evaluating LLM Agents for Generating Buildable and Functional Structures

LLM-based agents are increasingly capable of generating complex 3D structures, with the potential to reshape how objects are designed and realized in the physical world. Yet, producing elegant geometry is fundamentally different from producing objects that can be built and perform their intended functions. Existing evaluations largely focus on geometric quality while overlooking physical realizability. We introduce LMBuild, a benchmark for evaluating LLM agents on generating buildable and functi

Beyond the Parameter Monolith: Reconstructive Memories, Executable Skills, and Residual Assembly for Language Models

arXiv:2610.04012v1 Announce Type: new Abstract: Language-model systems can separate contextual computation, persistent storage, and exact execution instead of updating all capabilities through one shared parameter system. We investigate FEM-ASM, a finite-element-method-inspired organization in which independently constructed document states and deterministic executable skills contribute typed proposals to a shared language-model state. An explicit residual operator reconciles proposals attached

Reka Releases Rho-1: A 19B Omni-Reasoning Model That Understands, Generates Video and Outputs Robot Actions in One

Reka has released Rho-1, a 19B omni-reasoning model trained from scratch. One network reads and generates text, images, video and robot actions over a shared KV cache. A distilled variant returns a 5.3-second clip in about a second. It is a research preview with no public weights yet. The post Reka Releases Rho-1: A 19B Omni-Reasoning Model That Understands, Generates Video and Outputs Robot Actions in One appeared first on MarkTechPost .

Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation

Test-time training (TTT) lets a model store information in its weights during inference. When the model learns from its own output, however, each update also changes the model that generates the next training example. Across 128K-token streams, retaining generated-text updates worsens prediction on independent human-written text with three TTT-E2E model configurations (labeled 125M, 760M, and 3B). The same failure occurs when Adam updates Qwen3-4B's existing weights. The same update mechanisms c

MLLMs Fail to Refuse when Using Tools Agentically

arXiv:2610.03938v1 Announce Type: new Abstract: Agentic multimodal large language models (MLLMs) have recently pushed the frontier of visual reasoning by calling tools such as zooming and tagging. Despite the recent strong success of agentic MLLMs, this work uncovers a critical safety failure in the tool-use paradigm: agentic tool-using MLLMs become less capable of refusing harmful requests. Our experiments confirm that, across three popular safety benchmarks, all the top open- and closed-weight

SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown

arXiv:2610.04008v1 Announce Type: new Abstract: Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking correct behavior. Existing benchmarks do not systematically distinguish documentation repair, script repair, and preservation when evaluating skill self-evolution. We introduce SkillScriptBench, a 350-task benchmark designed to evaluate these capabilities separately. From a survey

Proxy Confidence: Auditing Black-Box LLM Agents with a Surrogate's Log-Probabilities

arXiv:2610.03894v1 Announce Type: new Abstract: A deployed LLM agent emits tool calls, queries, and code that can be silently wrong -- by the time the error surfaces, the action has run. Frontier chat APIs hide the model's token probabilities; the agent's stated confidence barely beats chance on the mistakes that matter; and resampling does not help, since frontier models are highly repetitive, reproducing the same call across samples. We recover the missing signal from a low-cost open-weight su

SearchJev: A Fast and Calibrated System-1 Model for Search Agents

Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates search decisions from System-2 reasoning and generation. Given a search state and a decision schema, SearchJev directly scores legal options without autoregressive output generation. We propose Soft-Label Learning for Cali

Tell HN: GitHub refuses to remove cracked copies of my software after a month

I am a developer of a popular photo editor that runs in a web browser. Many people are asking AI models to take the Javascript code from my website, remove all ads from it, and they publish such a "new product" on Github for everyone to download. There exist tens of such repositories on Github. I want my website to be the only source of a stable version of my program Photopea. I even received emails from people complaining about something in Photopea, and it took several emails to figure out tha

UndoBench: Separating Task Competence from Recovery Capability in Tool-Using AI Agents

Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery. We introduce UndoBench, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds alongside wire-level effect-history and environment

When Does Selection Replace Extraction? A Pre-Registered Test of Agent Memory with a Typed Decision Model

Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran a pre-registered study on held-out LoCoMo conversations and LongMemEval. At a tight budget on LoCoMo, raw turns selected by a single call to Jev, a typed decision model, are non-inferior to an LLM-ext

Qwen3.8-Flash-Next on Strata

Hey! 👋 I have released an official support for Strix Halo machines on Strata for Qwen3.8-Flash-Next. Currently numbers are the best on long context decode and ppts using typical Unsloth’s Q4 and GSQ-RCO model weights. Can go up to 1M context length without big speed loss. Currently support is marked as experimental and was done on Linux only. Will be happy for any feedback and pull requests you could give! 👀

DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents

Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a controlled way. We introduce DeskForge, a controllable desktop environment that composes and explores real applications to generate large-scale supervision for computer-use agents. It varies application states, content, win

How to Loop MoE: Flatten the Experts, Untie the Attention

Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and gives MoE models new potential for better expert usage, but it raises a question: how to loop a MoE? We answer it with Foil. With the expert parameters and the expert compute per t

I gave a 21M model a 6.4B-parameter lookup table. It matches a 114M dense model and runs with the table on an SSD (RX 9070)

I spent the last few weeks on a hobby research project and just made it public. The idea isn't new (product-key memory, Lample et al. 2019, and Meta's "Memory Layers at Scale"): give a model a huge table of learned vectors and let it read only a few hundred of them per token. I wanted to know what that's actually worth on a small model, what it costs, and whether the table even has to sit in VRAM. What came out: - A 21M model with a 16.8M-row table (6.4B parameters in the table, 33M used per tok

TextReg: Mitigating Prompt Distributional Overfitting via Regularized Text-Space Optimization

Large language models (LLMs) are highly sensitive to the prompts used to specify task objectives and behavioral constraints. Many recent prompt optimization methods iteratively rewrite prompts using LLM-generated feedback, but the resulting prompts often become longer, accumulate narrow sample-specific rules, and generalize poorly beyond the training distribution. We study this failure mode as prompt distributional overfitting and argue that it reflects a lack of representation control in discre

Periscope: Extending Frozen Language Models Beyond Their Context Window

A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option is supported, which passage is the evidence. Periscope, a training-free inference method, arranges the N chunks of a text on a K{times}K grid with K{=}lceilNrceil and asks a frozen model the same question about K local spans of consecutive

Show HN: MailAccess – the true Email OSINT framework

Im a security engineer and i built mailaccess. When i started learning pentesting, i came across multiple lectures and notes of people listing out tools and websites, which gives the emails for a particular domain, and almost all of them mentioned that the tool might not stick, so its better to learn the methodology, rather than learning a tool- that stuck with me. As i was beginning to really get into pentesting i noticed a clear lack of email osint methodology through the tool itself - so i th

How AI decision models could change content moderation

On Tuesday, Musubi announced a lightweight decision model made for real-time moderation called PolicyLM-1.7B, released with open weights.

Mirror Particle is building a ‘world model’ of human behavior

Mirror Particle will launch at TechCrunch Disrupt's Startup Battlefield 200 with a world model built from scratch to predict human behavior, arguing that LLM role-play falls short for market research and brand strategy.

Exploration-Preserving Policy Optimization

arXiv:2610.04011v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards improves reasoning, while the allocation of learning signal shapes which solutions remain accessible under repeated sampling. Group-relative objectives assign equal advantages to equally rewarded responses, making aggregate credit proportional to sampled mode frequency. We introduce Exploration-Preserving Policy Optimization (ExPPO), a lightweight advantage-shaping rule that redistributes credit using

Training Numerical Intelligence via Auto-Diagnosis and Skill Discovery

arXiv:2610.03872v1 Announce Type: new Abstract: AI agents are becoming increasingly capable of generating scientific code, but generating code is not the same as improving the algorithms behind it. For numerical solvers, execution feedback can expose poor performance, but rarely reveals its underlying cause and how to address it. We introduce Auto-Diagnosis and Skill Discovery (ADSD), a framework that links numerical diagnosis to reusable solver self-improvement. ADSD follows a diagnosis-first p

Expanding the Cyber Verification Program

We’re launching a new, expanded version of our Cyber Verification Program (CVP), which makes advanced cyber capabilities and reduced blocking classifiers available to qualifying security professionals. The program now consists of three access tiers, which allow security teams to apply for the level of access that best suits their work. Each tier includes access to our most capable models, including Claude Opus 5.5, Claude Sonnet 5.5, Claude Mythos 5.1, and new models moving forward. Interested c

Beyond Domain-Specific World Models: JEPA-Anything Uses 1 Recipe for 7 Fields

JEPA-Anything splits a JEPA's single latent target into 4 orthogonal factors, each with its own predictor. Tested across 7 domains, it beat matched JEPA baselines on all 10 dynamics tasks and cut Interventional Pong intervention error by 34.8%. The post Beyond Domain-Specific World Models: JEPA-Anything Uses 1 Recipe for 7 Fields appeared first on MarkTechPost .

Representation-Space MMD for Diffusion Language Models

We introduce a post-training method for diffusion language models (DLMs) that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM. To estimate MMD, we retain contextual features at individual token positions, obtaining multiple observations per sequence from a single extractor pass. We optimize this objective using policy gradients for discrete models and direct differentiation through generated latents for contin

GPT-6.1 Sol looped "leak" hints at nested models serving architecture

Hello llamas. I am posting this because I believe that, despite it being closed source models, the discussion will bring value to the local AI community. As many of you probably heard, GPT-6 Astra is speculated to be a looped transformer architecture that outputs a token after multiple forward passes instead of one. This allows a model to essentially have more effective depth due to recurrence, making more use of the weights at the cost of more compute. Recent Azure Foundry "leaks" even suggeste

RobotUse: Allocating Computation, Context, and Decisions

Robot agents must connect their intended actions to observed outcomes while retaining the context needed to revise their choices over repeated attempts. Existing interfaces often leave these choices inside predefined tools or require agents to manage detailed execution code and its growing history. We introduce RobotUse, a robot agent harness that organizes computation, context, and decisions around specifying and revising physical actions. Agents visually select targets and poses, while the bac

From Knowledge Access to Source Learning: Developing Source-Specific Competence

Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use of the same source is still largely treated as repeated access rather than an opportunity to progressively improve understanding of that source. We study source learning: developing reusable source-sp

What Matters for Latent Reasoning with Flow Matching

Latent reasoning lets a large language model (LLM) think in a continuous space and verbalize only the answer. We argue that an effective latent thought must meet five requirements: it should be useful, helping produce the correct answer rather than merely changing it, diverse, so that resampling yields different reasoning trajectories, explainable, so that a decoded chain of thought (CoT) reflects reasoning the answer actually follows, refinable with more inference compute, and efficient, costin

LiFT: Loop Flow Transformers

We introduce Loop Flow Transformers (LiFT), a family of looped generative models that scales computation by repeatedly applying a shared Diffusion Transformer (DiT) core, with only light changes to the standard architecture. Rather than asking every recurrent step for the final prediction, LiFT trains each step with a single regression target: a point on a straight path from the model's initial estimate to the flow-matching target. Because we index these targets by a continuous depth coordinate,

Learning to Learn a Language

We present the Prior-Fitted Language Model (PFLM), a 300M-parameter byte-level transformer pretrained only on samples from a synthetic non-linguistic prior. Given a prefix of real text, it learns to predict the language in context with frozen weights, having never seen a word of any real language. Every training sequence is generated by a recurrent structural causal model drawn fresh from a distribution over such models. The model never sees the same language twice during training, so the only w

OpenAI Is Pissing Off a Bunch of Mathematicians—Again

“There’s a perception of mobster behavior” from leading AI companies, one mathematician tells WIRED as OpenAI prepares to release more than 100 new solutions to unsolved problems.

Google’s power-hungry data centers crave nuclear energy

Google announced a new agreement to update six nuclear power plant sites across the US as the tech giant seeks to generate more electricity for its power-hungry data centers. Google signed the 20-year deal with Constellation, the leading nuclear power plant operator in the US. The power purchase agreement is meant to guarantee the revenue […]

Release: datasette-atom 0.11a0 A minor fix for compatibility with the latest Datasette alphas. This meant we could upgrade the datasette.io site to Datasette 1.0a41. Tags: atom , datasette

A Quantitative Analysis of Graph Representation Strategies for Cyber Attack Detection

arXiv:2610.04019v1 Announce Type: new Abstract: Graph based cyber attack detection studies employ various graph construction and representation strategies across different cybersecurity application domains. This diversity motivates a quantitative examination of how representation strategies are distributed across these application domains. This study presents a quantitative analysis of 37 original studies published between 2019 and 2026. Each study was coded according to publication year, applic

ROAR: Unifying Runs across Heterogeneous AI-Driven Research Systems

arXiv:2610.03966v1 Announce Type: new Abstract: Each run of an AI-driven research system (ADRS) is an expensive search over a vast solution space, and dependable evaluation requires many runs, making run data both costly to produce and valuable to retain for large-scale analysis. Yet this data remains fragmented: teams operate in isolation, ADRS frameworks emit results in different formats, and no shared infrastructure exists to aggregate or compare runs across problems and systems. We present R

LatentQuant: Preserving the Policy-Facing Latent Contract under NVFP4 VAE Quantization

arXiv:2610.03959v1 Announce Type: new Abstract: Recent world action models (WAMs) reuse pretrained video VAEs whose encoder latents directly condition downstream action policies. Quantization must therefore preserve not only reconstruction fidelity but also the policy-facing latent contract expected by the frozen policy. Direct NVFP4 leaves W4A4 quantization error uncompensated, whereas joint quantization-aware training (QAT) can recover reconstruction by moving this representation. On Wan2.1, j

Behavioral History Outperforms Descriptions of the Person for LLM Synthetic Personas

arXiv:2610.03998v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as synthetic personas representing survey respondents. Their validity as substitutes for particular respondents depends on whether they reproduce individuals' decisions. We examine what information helps synthetic respondents predict each individual's later choices, using five conditions that add progressively richer information: no personal information, demographics, personality traits, cognitive

Optimizing the Optimizer: Language Models Discover Faster Molecular Relaxation

Geometry optimization is a major cost in many quantum-chemical workflows: each optimization step requires one force evaluation, and at the density-functional level that evaluation dominates the wall time. Research in this area has produced a broad range of optimization methods, and we ask whether a language model can improve on the best of them through autoresearch. An agent rewrites the optimizer itself to minimize force-call counts, restrained by two admission gates that reject premature stopp

Data Unlearning via Inverse Distillation

Multi-step matching models, including flow and diffusion models, produce high-quality outputs but incur substantial inference costs and may reproduce unwanted components of their training datasets. We introduce Inverse Distillation Unlearning (IDU), a unified framework that simultaneously distills a teacher multi-step matching model into an efficient one-step student generator and suppresses outputs corresponding to a designated training subset. We first formulate distillation as a min-max objec

With all the fuss over AI and what it can do, I feel like we've totally glossed over the fact that AI has casually solved something that has been a problem for decades. I was born the year after the microprocessor was invented, and I've kept a very close eye on technology as it has developed. And the problem of machine translation has been with us for a while. It used to be absolutely terrible. Then it got to the point where you could sort of tell what the native speaker who wrote the original w

World Editing: Intervening on Executable Worlds at Increasing Depth

Interactive world models are increasingly capable of generating environments and acting within them, yet deliberately editing an existing executable world remains underexplored. We formulate world editing as intervening on an existing world while preserving properties that should remain unchanged, and introduce intervention depth as an axis describing how strongly an edit couples world entities, dynamics, and systems. We instantiate this capability through industry-grade game modding and introdu

SEER: Self-Evolving Event Reasoning and Retrieval for Time Series Forecasting

Real-world time series are frequently driven by exogenous events and structural shifts, rendering conventional forecasting based solely on historical numerical observations insufficient. While language models can retrieve external news, standard retrieval-augmented approaches struggle with high noise, missing signals, and an inability to reason causally about event impacts. We propose SEER (Self-Evolving Event Reasoning and Retrieval), a closed-loop framework that dynamically optimizes event con

What Gradients Add to Text Leakage in Split Language Models, Counted per Token and per Document

Split learning lets a client train a language model on a server without sending its text. The client runs the first layers itself and sends the server only their output, a vector of numbers for each token. During training, the server sends gradients back. We show that an observer at the split can rebuild most of the client's text from this traffic, and we measure how much the gradients help. On GPT-2, an attacker who holds only the publicly released weights of the client's layers recovers 94.20%

Do you selfhost bitwarden lite or vaultwarden?

I want to selfhost a password manager. I wanted to go with vaultwarden. But now i read about bitwarden lite, which is the official lite version. How likely/how often did in the past happend that bitwarden released a breaking change to the app/clients so that vaultwarden needed first an update? Did you swap to bitwarden lite after the release? Im using pangolin to tunnel to my local machine. If this matters in anyway or form

my three-year ai honeymoon is over

i was very excited about code generating ai tools since early days of github copilot in vscode, was using it daily since chatgpt release and wrote almost all code through prompting (html/css/javascript, python, ruby, terraform etc) for years. It felt very good at first. However, about a year ago i started to notice subtle (at first) negative changes in my mental health. It's hard to describe in words this negative feeling because it's very basic and fundamental, but over that last year it progre

每天早晨,一份为你精选的科技日报