分析多模态模型潜在视觉推理中的证据-信用缺口,提出将潜在推理锚定在视觉证据上的方法。
2026-10-02
— 今天的主线:Agent 自我进化与决策模型,都在试图把推理从文本里解放出来。
Cloudflare 发布 Clef 与 Clef-flash 开放权重决策模型,返回类型化概率而非文本,瞄准结构化决策场景。谷歌 Gemini 4 Argon 空降多榜单第一,主打长期软件工程与网络安全,但仅限部分安全团队使用。OpenAI 与 Synopsys 合作推出 GPT-Synopsys,将前沿模型引入芯片设计 EDA 流程。多篇论文聚焦 agent 自我进化中的共谋失败、错误恢复与多模态自蒸馏,显示自改进闭环的可靠性成为研究焦点。
头条
论文揭示自进化搜索 agent 的共谋失败模式
论文《False Frontiers》发现自进化搜索 agent 中 proposer 与 solver 会逐渐在共享错误上达成一致,内部奖励提升但外部正确性停滞甚至下降,且随自进化轮次加剧。为什么重要:这直接挑战了当前依赖自生成课程训练 agent 的主流范式,提示需要外部证据验证来打破闭环。
每天早晨,一份为你精选的科技日报
网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。
已发布 82 期 · 每天筛过 150+ 条只留值得读的 30 条
AI 动态
Mid-Harness 在模型与终端执行之间采样并验证候选动作,提升 agent 动作可靠性。
UniEvo-VL 用单一多模态模型同时充当教师与学生,通过自批判实现测试时自蒸馏。
PivotOPD 发现多轮 agent 失败中过半含早期关键错误,提出针对性恢复训练方法。
Cohere 发布 Embed 5 嵌入模型家族,Pro 与 Fast 共享同一嵌入空间,支持文本、图像及 128K 上下文。
开发与开源
Janus 是单个 Go 二进制,通过 Vulkan 在 AMD/Intel/Nvidia GPU 上运行 GGUF 模型并提供 OpenAI 兼容 API。
Rust 编译器在 2026 年 7 月至 9 月间平均墙钟时间下降 4.57%,629 个基准中 555 个改善。
评论区普遍认可Rust编译速度已有可测量的提升,但也有人认为相比Go等语言仍太慢,影响快速迭代。
Pi 1.0 发布,定位为极简、可扩展的 agent harness,强调稳定与克制。
多数人认可 Pi 简洁稳定、可扩展且适合本地模型,但也有人认为其极简路线正被新功能侵蚀,并质疑版本标准与用户量说法。
Perplexity 与 turbopuffer 发布 pplx-embed-v2-context-9b-preview,训练信号从单一金句改为答案加佐证上下文。
社区热议
美国边境无证搜查手机引发诉讼,评论区普遍谴责权力缺乏监督,但也有人认为此类搜查本就合法。
评论普遍谴责边境无证搜查手机,认为权力缺乏监督且侵犯宪法权利,但也有人认为此类搜查本就合法、无需大惊小怪。
Android 开发者验证计划与账号封禁引发开发者强烈不满,认为比苹果更糟,但也有人指出仍可侧载未签名 APK。
评论区普遍痛斥安卓开发者验证与账号封禁,认为其扼杀自由、比苹果更糟,但也有人认为仍可侧载未签名APK。
Matthew Green 警告沙箱不足以遏制 rogue agent,共享缓存中的指令传递已构成蠕虫传播的两半。
GitHub Trending
Sponsor Star DietrichGebert / ponytail Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
Sponsor Star mattpocock / skills Skills for Real Engineers. Straight from my .agents directory.
Star NVIDIA / OpenShell OpenShell is the safe, private runtime for autonomous AI agents.
Star firebase / firebase-ios-sdk Firebase SDK for Apple App Development
Star mvschwarz / openrig Build your own network of agents from Claude Code, Codex and Pi: persistent teams with roles, shared context and owned work.
Star cursor / plugins Cursor plugin specification and official plugins
Sponsor Star obra / superpowers An agentic skills framework & software development methodology that works.
Sponsor Star mksglu / context-mode Context window optimization for AI coding agents. Sandboxes tool output (98% reduction), persists session memory, and enforces routing across 17 platforms via MCP + hooks.
Star heygen-com / hyperframes Write HTML. Render video. Built for agents.
Star earendil-works / pi AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
更多值得一看(内容池 76 条)
now you can use MTP with Qwen Flash Next, time to switch from Qwen 3.8 27B? (merged after 17h of development) quants: link to the previous discussion (I deleted the old post to avoid duplicates):
A few days ago I released Jeff-Qwen3.5-0.8B, a small "System 1" model that picks between options you define and returns a calibrated probability for each, in one forward pass. Speed was great on my M4 Max and RTX PRO 6000, but as a general zero-shot classifier it trailed the big models. Then it occurred to me that most decisions an agent makes in front of a local model aren't open-ended. They fall into a handful of recurring kinds: is this a prompt injection, which tool to call, how urgent is th
Self-attention gives LLMs fine-grained, query-dependent access to context, but dense token interactions incur quadratic prefill cost and a key--value cache growing with context length. Research thus spans explicit-memory compression, sparse access, recurrent state construction, structured state dynamics, and heterogeneous mechanism composition. This survey analyzes these developments as model-internal contextual memory. We introduce a five-dimensional lens---Memory Representation, Memory Update,
In this comprehensive coding guide, we explore Google Research's Kauldron—a JAX training library optimized for research velocity and modularity. Learn how konfig turns experiments into plain dictionaries, kontext wires components via string paths, and ktyping enforces runtime shape checks. The post A Coding Guide to Google Research’s Kauldron: Configs That Are Plain Data, Components Wired by String, and a JAX Trainer You Can Read End to End appeared first on MarkTechPost .
Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all. The guidelines require the label to be deci
It has been a bad month for federal government cybersecurity.
Amazon Web Services' Strand Labs has released the latest Jevalike decision model, Strands Decider 2B.
... but you can’t try it yet unless you are “government users and trusted cyber defenders in the Fairwind Program”
arXiv:2609.38372v1 Announce Type: new Abstract: A harness is the code around a language-model agent that organizes prompts, calls tools, manages context, and controls execution. As models grow stronger, recent work has begun to let agents improve their own harnesses, a line of work known as self-evolving harnesses. In most existing methods, a separate proposer running on a human-designed harness modifies the solver's harness, and a separate harness is evolved for each benchmark. Real-world tasks
arXiv:2609.38294v1 Announce Type: new Abstract: We study the generation of agentic workflows that jointly optimize multiple objectives, such as accuracy, cost, latency, robustness, and consistency. Existing methods for workflow generation typically optimize accuracy alone or a weighted sum of objectives, so each trained generator commits to one fixed trade-off and must be retrained from scratch when preferences change. To alleviate this, we propose MoFlow, which generates workflows optimized acr
arXiv:2609.38386v1 Announce Type: new Abstract: Concurrent autoregressive inference creates a fundamental interference problem: prefilling a newly arrived long prompt can delay tokens for requests that are already decoding. Fixed prefill chunks reduce this interference, but the best chunk size depends on the model, hardware, load, and latency objective. We introduce Decode-Latency Feedback Prefill (DLFP), a model-free controller that changes only prefill work that overlaps active decodes. After
Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mappin
Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change. Existing automated methods explore this space narrowly, optimizing only components such as prompts or skills or becoming trapped by fixed, exploitative search strategies. We introduce MILO (Meta-evolutionary Island Orche
Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, yet how RL reshapes these policies remains poorly understood. We find that RL across widely used flow-based VLA models, including π_{0.5} and GR00T~N1.5/N1.6, on LIBERO, ManiSkill, MetaWorld, and CALVIN induces substantially lower-rank parameter updates that are highly concentrated in the action expert's Timestep Modules, a small and previously overlooked component. Through systematic module-repla
OpenAI says Moonshot-linked operators used thousands of accounts to extract protected reasoning from its models for adversarial distillation. No encryption broken. No database compromised. Just systematic querying designed to make one model teach another. OpenAI says this is dangerous because competitors can reproduce capabilities without making the same investment in safety. Which is a serious security issue. But you have to appreciate the timing: after years of “we learned from the internet,”
A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice
Optional Bundle architecture : Schedule (session-local delayed / timed / interval reminders) was removed from the default set and made an explicit Optional Bundle. This cleanly separates “installed” from “enabled” and is the first systematic use of the Profile + Bundle model for official features. • Windows Sandbox improvements : A new permission-diagnosis skill can detect common Access Denied causes and perform backed-up, recoverable permission fixes after user authorization, giving the Agent a
You can't make this stuff up.
A training video reportedly demonstrates how police can evade an Apple security feature.
OpenAI has parted ways with three safety researchers after an internal investigation found they mishandled sensitive company information, report says.
A new AI tool can guess what you’re looking at just by analyzing your brain scans—and re-create that image with remarkable precision. It can go the other way, too, and predict a person’s brain activity based on what they’re looking at. In the image above, for example, the left-hand image of each pair is what…
Adding in a second neural network that guesses the identity of hidden pieces was key.
Last week, Microsoft CEO Satya Nadella hosted an intimate, invite-only event for leaders from some of its key enterprise customers. Instead of a flashy media event, Nadella outlined the future of Copilot directly to the customers Microsoft really cares about, pitching its latest rethink of the AI assistant as the "OS for work." Microsoft is […]
Benefitting from the Bitter Lesson
Yantra is a C++ parser generator: lexer, parser, and AST walker all generated from one tool. It builds the whole AST first, then walks it. Most LALR parser generators (Yacc, Bison, Lemon) run your semantic actions during parsing, as each rule reduces, bottom-up. That means at the time a rule's action runs, you don't yet know what its parent looks like. This pushes a lot of grammars toward hand-built AST classes and a separate walking pass whenever you need to look ahead into siblings or defer a
We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through language pre-training, multimodal training, and post-training. LoopVL outperforms a range of similarly sized and larger non-recurrent models on multimodal understanding and visual reasoning benchmarks. We also observe V
NVIDIA has released Kumo Tabular, a new family of tabular foundation models (TFMs) for classification and regression. If you have followed TabPFN or TabICL, the setup will look familiar. The model takes labeled rows as context and predicts new rows in one forward pass. There is no training, no hyperparameter tuning, and no feature engineering. […] The post NVIDIA Releases Kumo Tabular: Open Tabular Foundation Models That Predict New Rows in a Single Forward Pass appeared first on MarkTechPost .
arXiv:2609.38379v1 Announce Type: new Abstract: Large language models (LLMs) are frequently updated for various use cases, where filtering out misaligned training samples is a common practice for preventing post-update misalignment. However, alignment is inherently context-dependent: a recommendation that is aligned in one context may be inappropriate in another. For example, in response to the question "What should a researcher do with the research data?", recommending that the researcher prese
Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct responses. Existing mitigation methods either retrain the reward model or apply a fixed correction to one known bias, such as a preference for longer responses. Retraining requires additional data and computational resources, while existing
arXiv:2609.38282v1 Announce Type: new Abstract: Vision-language models may rewrite anomalous text in images into linguistically plausible expressions, compromising OCR transcription faithfulness. Sequence-level task rewards and local teacher guidance are complementary, but guidance from the same teacher may not remain equally effective as the student improves. Offline analysis shows that supervision from a fixed teacher becomes progressively less favorable as the student improves, both across tr
arXiv:2609.38296v1 Announce Type: new Abstract: Large language models (LLMs) can influence people's beliefs, yet little is known about whether and how they can manipulate each other. To investigate this, we simulate conversations between two agents: a target LLM that role-plays a human persona based on demographic and psychological attributes, and an influencer LLM that aims to make the target's beliefs more extreme. We examine radicalization along two pathways: resonance, where the influencer r
Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are independently sampled and globally shuffled across epochs. We study self-supervised learning from continuous video streams, where frames are consumed in temporal order using strict sliding-window batches, without global reshuffling or multi-epoch replay. To this end, we construct WT++, a 95-hour urban walking-tour video dataset for streaming pretrai
Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during language model pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has only been tested on models of at most 1B parameters and PT budgets below 2B tokens on predominantly web text. It is unknown whether PPT is effective at larger scales and under more realistic PT data mix
Many language tasks have no single answer that can be checked automatically. Rubrics provide criteria for judging responses to these tasks. For reinforcement learning, the resulting verdicts must be combined into a scalar reward. A common approach sums the points assigned to satisfied criteria. Distinct verdict patterns can thus receive the same reward, and the fixed points encode how much each criterion should count, not how strongly its verdict distinguishes the current rollouts. Beyond this a
Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unch
pi-llama-skip-reasoning is an extension for the Pi.dev harness that forces a local llama.cpp model to stop reasoning and answer / act immediately. When you are deep into the ctx session and ask 27B a simple question about a fact or need a direct action , the model may still feel the urge to indulge in copious deliberation in the reasoning trace. This extension allows the user to force the model to snap out of the reasoning stage and provide the answer immediately. Disclaimer: don't skip the reas
We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict, providing no separate credit at the criterion-choice tokens and ignoring the rich language feedback (e.g., preference rationales) th
Related: Pi 1.0 - - Oct 2026 (184 comments)
Asking AI companies to self-regulate is a great way to pretend like you’ve accomplished something.
Prices for memory sold for 2027 "are much higher than 2026 prices."
We catch up with RLM first author Alex Zhang, MIT PhD, on Jev, PhD masxing, and the future of harnesses.
Google launched its first advanced chip into orbit to pave the way for space data centers.
Opus 5.5’s biggest tell is the word “dependable,” which pops up 23 times more often than in human samples.
Shopify’s new Canvas site builder lets merchants create and customize their online stores by chatting with its AI agent Sidekick, while watching the changes happen in real time.
40 year tech gap? No problem! The Tandy runs DeskMind, a native DOS program. It talks over WiFi (a PicoMEM 2 card with mTCP) to a small Python server on my PC. That server drives Qwen3.8-27B (NInfer on a 5090) and Krea 2 (ComfyUI on a 4090). The 286 never sees JSON, base64 or a PNG. It gets plain text lines and pictures that are ready to copy into video memory. Drawing from chat without tool calling. The system prompt tells Qwen to wrap a picture request in ` ... `. The server catches the tag mi
arXiv:2609.38385v1 Announce Type: new Abstract: Supervised fine-tuning of discrete diffusion language models masks some response tokens and trains the model to recover their original values from the visible context. The masking pattern therefore determines both the context available to the model and the tokens it learns to predict. Uniform random masking does not explicitly account for the interaction between these choices. We introduce GoldiMask, which selects tokens to reveal as context by app
arXiv:2609.38359v1 Announce Type: new Abstract: High quality synthetic data is central to post training LLMs for adaptive AI applications that represent the diverse expert strategies and decisions in conversations. Prompting LLMs directly or conditioning them on end use scenarios yields low diversity data that collapses onto dominant modes. We propose a method to generate diverse high quality synthetic data using Generative Flow Networks (GFlowNets). We show that training GFlowNets to generate l
arXiv:2609.38409v1 Announce Type: new Abstract: Recent progress in large language model reasoning has been driven by benchmarks and reinforcement learning environments with automatically verifiable rewards, particularly in mathematics, code, and formal logic. These settings make model accuracy easier to evaluate and optimize, but it remains unclear how far success under fixed problem specifications and stable evaluation criteria transfers to reasoning outside such domains. Real-world reasoning o
Title explains it lol. I went through all the steps again when it wasn't working, and then discovered that the "what is my IP" sites show a different one than my router does. Apparently that means I'm behind CGNAT? I just wanted to access my jellyfin without needing tailscale on every device lol. I know the flare is need help, but from what I'm reading there isn't much to be done other than asking for my own IPV4. Just wanted to share how I wasted two hours lol
Barclays, the British universal bank, is expanding its strategic collaboration with Anthropic to integrate secure, enterprise-grade AI systems across its global operations. Barclays is extending Claude across the bank to accelerate software development, modernize legacy systems, and improve operational efficiency. As part of this rollout, Barclays expects Claude Code adoption to reach 50% of its developer population by the end of 2026, rising to a majority of software engineers in 2027. Barclays
TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2.
Hello everyone! The headline feature is exit nodes: a device can now send its full internet traffic out through one or more sites you choose. The tunnel can also come up automatically instead of waiting for someone to open the app and connect. Windows and Mac got a major UI overhaul, and iOS and Android picked up a lot of new features. Pangolin is an open-source, identity-aware remote access platform that simply and securely connects and authenticates your users to applications, infrastructure,
Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an
Hi r/LocalLLaMA! We’re researchers at the Institute of Foundation Models (IFM), an AI research lab dedicated to open and independent development of frontier-class foundation models. We recently released K2 Horizon a connected fleet of six fully open models with size ranging from 0.9B to 375B. In addition to weights, we also open-sourced training data and recipes, training code, intermediate checkpoints, fine-grained training logs and evals. Ask us anything about pre-training and data mixes, post