Analyzes the evidence-credit gap in latent visual reasoning of multimodal models and proposes a method to anchor latent reasoning in visual evidence.
2026-10-02
— Today's main thread: agent self-evolution and decision models are both trying to free reasoning from text.
Cloudflare releases Clef and Clef-flash open-weight decision models that return typed probabilities instead of text, targeting structured decision scenarios. Google's Gemini 4 Argon lands at the top of multiple leaderboards, focusing on long-horizon software engineering and cybersecurity, but is limited to select security teams. OpenAI partners with Synopsys to launch GPT-Synopsys, bringing frontier models into the chip design EDA workflow. Several papers focus on collusion failures, error recovery, and multimodal self-distillation in agent self-evolution, showing that the reliability of self-improvement loops has become a research focus.
トップニュース
Cloudflare releases Clef and Clef-flash open-weight decision models
Cloudflare releases Clef (27B) and Clef-flash (9B), two open-weight decision models that return typed probabilities instead of free text, are compatible with the Jev API, and support yes/no, choice, and numeric questions, with median latency of 209.3 ms and 38.8 ms respectively on Workers AI. Why it matters: decision models offer a more deterministic and cheaper alternative to LLMs for steps in workflows that require structured judgment, and the Apache 2.0 license supports self-hosting.
Most people acknowledge Cloudflare's entry and are optimistic about competition, but some argue it is slow, expensive, and not truly open source.
Google releases Gemini 4 Argon, topping multiple leaderboards but limited to security teams
Google releases Gemini 4 Argon, scoring 77.9% on DeepSWE v1.1, above Claude Opus 5.5's 74.2% and GPT-6 Astra's 74.1%; it reaches 68% on the CWE-bench v1 cybersecurity vulnerability remediation test, tying for first with GPT-6 Astra. Why it matters: the model focuses on long-horizon software engineering and cybersecurity defense, with a minimum cost of $1.99 per individual task, but it is currently only available to vetted cybersecurity teams.
OpenAI and Synopsys partner to launch GPT-Synopsys chip design model
OpenAI and Synopsys sign a multi-year agreement to jointly develop GPT-Synopsys, a specialized model for chip design; OpenAI will be licensed to use Synopsys's EDA tools for model development, and the two sides will share a revenue framework. Why it matters: this is the first time a frontier large model has entered the core semiconductor EDA workflow through deep collaboration, and it could change the path toward intelligence in chip design tools.
Paper reveals collusion failure mode in self-evolving search agents
The paper "False Frontiers" finds that in self-evolving search agents, the proposer and solver gradually converge on shared errors, with internal rewards rising while external correctness stagnates or even declines, worsening with each round of self-evolution. Why it matters: this directly challenges the current mainstream paradigm of training agents with self-generated curricula, suggesting that external evidence verification is needed to break the loop.
turbopuffer announces v3 architecture, moving beyond pure vector database positioning
turbopuffer releases its v3 storage architecture, rewriting the layout, writes, compaction, and query paths for documents and indexes, comprehensively speeding up text, regex, and vector search, and laying the foundation for pushing more SQL queries down into turbopuffer. Why it matters: this marks the evolution of vector search infrastructure toward a general-purpose query engine, with direct implications for RAG and hybrid retrieval architecture choices.
毎朝、あなた仕様のテックダイジェストを
ウェブは全体像、購読者にはあなた専用を——興味に合わせた AI 精選、プライベート RSS の統合、コミュニティの見解付きで毎朝配信。ずっと無料。
82 号配信 · 毎日150件超から読む価値ある30件に厳選
AI動向
Mid-Harness samples and verifies candidate actions between the model and terminal execution, improving agent action reliability.
UniEvo-VL uses a single multimodal model as both teacher and student, achieving test-time self-distillation through self-critique.
PivotOPD finds that more than half of multi-turn agent failures contain early critical errors and proposes targeted recovery training methods.
Cohere releases the Embed 5 embedding model family, with Pro and Fast sharing the same embedding space and supporting text, images, and 128K context.
開発とOSS
Janus is a single Go binary that runs GGUF models on AMD/Intel/Nvidia GPUs via Vulkan and provides an OpenAI-compatible API.
The Rust compiler's average wall-clock time fell 4.57% between July and September 2026, with 555 of 629 benchmarks improving.
Commenters generally agree that Rust compile speed has improved measurably, but some still consider it too slow compared with languages like Go, affecting rapid iteration.
Pi 1.0 is released, positioned as a minimalist, extensible agent harness that emphasizes stability and restraint.
Most people acknowledge that Pi is simple, stable, extensible, and suitable for local models, but some argue its minimalist path is being eroded by new features and question its versioning standards and user-count claims.
Perplexity and turbopuffer release pplx-embed-v2-context-9b-preview, changing the training signal from a single golden sentence to an answer plus supporting context.
コミュニティの話題
Warrantless phone searches at the U.S. border spark litigation; commenters broadly condemn the lack of oversight over power, though some argue such searches are already legal.
Comments broadly condemn warrantless phone searches at the border, arguing that power lacks oversight and constitutional rights are violated, though some argue such searches are already legal and not a big deal.
Android developer verification and account bans spark strong developer backlash, with many saying it is worse than Apple, though some note that unsigned APKs can still be sideloaded.
The comment section broadly condemns Android developer verification and account bans, arguing that they stifle freedom and are worse than Apple, though some note that unsigned APKs can still be sideloaded.
Matthew Green warns that sandboxes are insufficient to contain rogue agents, and that instruction passing through shared caches already constitutes both halves of worm propagation.
GitHub Trending
Sponsor Star DietrichGebert / ponytail Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
Sponsor Star mattpocock / skills Skills for Real Engineers. Straight from my .agents directory.
Star NVIDIA / OpenShell OpenShell is the safe, private runtime for autonomous AI agents.
Star firebase / firebase-ios-sdk Firebase SDK for Apple App Development
Star mvschwarz / openrig Build your own network of agents from Claude Code, Codex and Pi: persistent teams with roles, shared context and owned work.
Star cursor / plugins Cursor plugin specification and official plugins
Sponsor Star obra / superpowers An agentic skills framework & software development methodology that works.
Sponsor Star mksglu / context-mode Context window optimization for AI coding agents. Sandboxes tool output (98% reduction), persists session memory, and enforces routing across 17 platforms via MCP + hooks.
Star heygen-com / hyperframes Write HTML. Render video. Built for agents.
Star earendil-works / pi AI agent toolkit: unified LLM API, agent loop, TUI, coding agent CLI
その他の注目(あと76件)
now you can use MTP with Qwen Flash Next, time to switch from Qwen 3.8 27B? (merged after 17h of development) quants: link to the previous discussion (I deleted the old post to avoid duplicates):
A few days ago I released Jeff-Qwen3.5-0.8B, a small "System 1" model that picks between options you define and returns a calibrated probability for each, in one forward pass. Speed was great on my M4 Max and RTX PRO 6000, but as a general zero-shot classifier it trailed the big models. Then it occurred to me that most decisions an agent makes in front of a local model aren't open-ended. They fall into a handful of recurring kinds: is this a prompt injection, which tool to call, how urgent is th
Self-attention gives LLMs fine-grained, query-dependent access to context, but dense token interactions incur quadratic prefill cost and a key--value cache growing with context length. Research thus spans explicit-memory compression, sparse access, recurrent state construction, structured state dynamics, and heterogeneous mechanism composition. This survey analyzes these developments as model-internal contextual memory. We introduce a five-dimensional lens---Memory Representation, Memory Update,
In this comprehensive coding guide, we explore Google Research's Kauldron—a JAX training library optimized for research velocity and modularity. Learn how konfig turns experiments into plain dictionaries, kontext wires components via string paths, and ktyping enforces runtime shape checks. The post A Coding Guide to Google Research’s Kauldron: Configs That Are Plain Data, Components Wired by String, and a JAX Trainer You Can Read End to End appeared first on MarkTechPost .
Vision-language models (VLMs) are increasingly used in place of human annotators, making it important that substitutability tests reflect the model rather than incidental evaluation conditions. We introduce MIST, the Misleading-Image Stress Test: 200 English sentences, each built around a phrase readable either figuratively or literally and shown with an aligned image depicting its reading, a misleading image depicting the opposite, or no image at all. The guidelines require the label to be deci
It has been a bad month for federal government cybersecurity.
Amazon Web Services' Strand Labs has released the latest Jevalike decision model, Strands Decider 2B.
... but you can’t try it yet unless you are “government users and trusted cyber defenders in the Fairwind Program”
arXiv:2609.38372v1 Announce Type: new Abstract: A harness is the code around a language-model agent that organizes prompts, calls tools, manages context, and controls execution. As models grow stronger, recent work has begun to let agents improve their own harnesses, a line of work known as self-evolving harnesses. In most existing methods, a separate proposer running on a human-designed harness modifies the solver's harness, and a separate harness is evolved for each benchmark. Real-world tasks
arXiv:2609.38294v1 Announce Type: new Abstract: We study the generation of agentic workflows that jointly optimize multiple objectives, such as accuracy, cost, latency, robustness, and consistency. Existing methods for workflow generation typically optimize accuracy alone or a weighted sum of objectives, so each trained generator commits to one fixed trade-off and must be retrained from scratch when preferences change. To alleviate this, we propose MoFlow, which generates workflows optimized acr
arXiv:2609.38386v1 Announce Type: new Abstract: Concurrent autoregressive inference creates a fundamental interference problem: prefilling a newly arrived long prompt can delay tokens for requests that are already decoding. Fixed prefill chunks reduce this interference, but the best chunk size depends on the model, hardware, load, and latency objective. We introduce Decode-Latency Feedback Prefill (DLFP), a model-free controller that changes only prefill work that overlaps active decodes. After
Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacity at fixed active compute. Yet existing scaling laws model recurrence or sparsity in isolation. In this work, we introduce Loop Scaling Laws, the first scaling law to jointly model recurrence and sparsity alongside model size and data. At its core is a bounded, sparsity-conditional recurrence mappin
Modern agentic systems combine an AI model with a harness that controls execution and environmental interactions. Harness design strongly affects long-horizon performance, yet its combinatorial search space demands substantial human effort that must be repeated as models change. Existing automated methods explore this space narrowly, optimizing only components such as prompts or skills or becoming trapped by fixed, exploitative search strategies. We introduce MILO (Meta-evolutionary Island Orche
Reinforcement learning (RL) is increasingly used to post-train vision-language-action (VLA) models, yet how RL reshapes these policies remains poorly understood. We find that RL across widely used flow-based VLA models, including π_{0.5} and GR00T~N1.5/N1.6, on LIBERO, ManiSkill, MetaWorld, and CALVIN induces substantially lower-rank parameter updates that are highly concentrated in the action expert's Timestep Modules, a small and previously overlooked component. Through systematic module-repla
OpenAI says Moonshot-linked operators used thousands of accounts to extract protected reasoning from its models for adversarial distillation. No encryption broken. No database compromised. Just systematic querying designed to make one model teach another. OpenAI says this is dangerous because competitors can reproduce capabilities without making the same investment in safety. Which is a serious security issue. But you have to appreciate the timing: after years of “we learned from the internet,”
A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice
Optional Bundle architecture : Schedule (session-local delayed / timed / interval reminders) was removed from the default set and made an explicit Optional Bundle. This cleanly separates “installed” from “enabled” and is the first systematic use of the Profile + Bundle model for official features. • Windows Sandbox improvements : A new permission-diagnosis skill can detect common Access Denied causes and perform backed-up, recoverable permission fixes after user authorization, giving the Agent a
You can't make this stuff up.
A training video reportedly demonstrates how police can evade an Apple security feature.
OpenAI has parted ways with three safety researchers after an internal investigation found they mishandled sensitive company information, report says.
A new AI tool can guess what you’re looking at just by analyzing your brain scans—and re-create that image with remarkable precision. It can go the other way, too, and predict a person’s brain activity based on what they’re looking at. In the image above, for example, the left-hand image of each pair is what…
Adding in a second neural network that guesses the identity of hidden pieces was key.
Last week, Microsoft CEO Satya Nadella hosted an intimate, invite-only event for leaders from some of its key enterprise customers. Instead of a flashy media event, Nadella outlined the future of Copilot directly to the customers Microsoft really cares about, pitching its latest rethink of the AI assistant as the "OS for work." Microsoft is […]
Benefitting from the Bitter Lesson
Yantra is a C++ parser generator: lexer, parser, and AST walker all generated from one tool. It builds the whole AST first, then walks it. Most LALR parser generators (Yacc, Bison, Lemon) run your semantic actions during parsing, as each rule reduces, bottom-up. That means at the time a rule's action runs, you don't yet know what its parent looks like. This pushes a lot of grammars toward hand-built AST classes and a separate walking pass whenever you need to look ahead into siblings or defer a
We introduce LoopVL to study whether Loop Transformers can be effectively extended to vision- language models. LoopVL combines Module-Loop and Model-Loop computation to iteratively update a unified vision-language state through shared modules. We train LoopVL from scratch through language pre-training, multimodal training, and post-training. LoopVL outperforms a range of similarly sized and larger non-recurrent models on multimodal understanding and visual reasoning benchmarks. We also observe V
NVIDIA has released Kumo Tabular, a new family of tabular foundation models (TFMs) for classification and regression. If you have followed TabPFN or TabICL, the setup will look familiar. The model takes labeled rows as context and predicts new rows in one forward pass. There is no training, no hyperparameter tuning, and no feature engineering. […] The post NVIDIA Releases Kumo Tabular: Open Tabular Foundation Models That Predict New Rows in a Single Forward Pass appeared first on MarkTechPost .
arXiv:2609.38379v1 Announce Type: new Abstract: Large language models (LLMs) are frequently updated for various use cases, where filtering out misaligned training samples is a common practice for preventing post-update misalignment. However, alignment is inherently context-dependent: a recommendation that is aligned in one context may be inappropriate in another. For example, in response to the question "What should a researcher do with the research data?", recommending that the researcher prese
Reward models score responses from large language models (LLMs) and guide LLM training toward human preferences. However, reward models can favor superficial attributes such as length or confidence, leading LLMs to produce higher-scoring but not more correct responses. Existing mitigation methods either retrain the reward model or apply a fixed correction to one known bias, such as a preference for longer responses. Retraining requires additional data and computational resources, while existing
arXiv:2609.38282v1 Announce Type: new Abstract: Vision-language models may rewrite anomalous text in images into linguistically plausible expressions, compromising OCR transcription faithfulness. Sequence-level task rewards and local teacher guidance are complementary, but guidance from the same teacher may not remain equally effective as the student improves. Offline analysis shows that supervision from a fixed teacher becomes progressively less favorable as the student improves, both across tr
arXiv:2609.38296v1 Announce Type: new Abstract: Large language models (LLMs) can influence people's beliefs, yet little is known about whether and how they can manipulate each other. To investigate this, we simulate conversations between two agents: a target LLM that role-plays a human persona based on demographic and psychological attributes, and an influencer LLM that aims to make the target's beliefs more extreme. We examine radicalization along two pathways: resonance, where the influencer r
Self-supervised learning draws inspiration from infant visual development, yet standard training pipelines bear little resemblance to it: images are independently sampled and globally shuffled across epochs. We study self-supervised learning from continuous video streams, where frames are consumed in temporal order using strict sliding-window batches, without global reshuffling or multi-epoch replay. To this end, we construct WT++, a 95-hour urban walking-tour video dataset for streaming pretrai
Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during language model pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has only been tested on models of at most 1B parameters and PT budgets below 2B tokens on predominantly web text. It is unknown whether PPT is effective at larger scales and under more realistic PT data mix
Many language tasks have no single answer that can be checked automatically. Rubrics provide criteria for judging responses to these tasks. For reinforcement learning, the resulting verdicts must be combined into a scalar reward. A common approach sums the points assigned to satisfied criteria. Distinct verdict patterns can thus receive the same reward, and the fixed points encode how much each criterion should count, not how strongly its verdict distinguishes the current rollouts. Beyond this a
Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unch
pi-llama-skip-reasoning is an extension for the Pi.dev harness that forces a local llama.cpp model to stop reasoning and answer / act immediately. When you are deep into the ctx session and ask 27B a simple question about a fact or need a direct action , the model may still feel the urge to indulge in copious deliberation in the reasoning trace. This extension allows the user to force the model to snap out of the reasoning stage and provide the answer immediately. Disclaimer: don't skip the reas
We study training LLM judges from natural language feedback, especially for subjective tasks where the verdict depends strongly on which evaluation criteria the judge invokes and how it weighs them. The dominant approach, outcome-supervised RL (e.g., GRPO), credits every token in the rollout with a single scalar determined only by the accuracy of the final verdict, providing no separate credit at the criterion-choice tokens and ignoring the rich language feedback (e.g., preference rationales) th
Related: Pi 1.0 - - Oct 2026 (184 comments)
Asking AI companies to self-regulate is a great way to pretend like you’ve accomplished something.
Prices for memory sold for 2027 "are much higher than 2026 prices."
We catch up with RLM first author Alex Zhang, MIT PhD, on Jev, PhD masxing, and the future of harnesses.
Google launched its first advanced chip into orbit to pave the way for space data centers.
Opus 5.5’s biggest tell is the word “dependable,” which pops up 23 times more often than in human samples.
Shopify’s new Canvas site builder lets merchants create and customize their online stores by chatting with its AI agent Sidekick, while watching the changes happen in real time.
40 year tech gap? No problem! The Tandy runs DeskMind, a native DOS program. It talks over WiFi (a PicoMEM 2 card with mTCP) to a small Python server on my PC. That server drives Qwen3.8-27B (NInfer on a 5090) and Krea 2 (ComfyUI on a 4090). The 286 never sees JSON, base64 or a PNG. It gets plain text lines and pictures that are ready to copy into video memory. Drawing from chat without tool calling. The system prompt tells Qwen to wrap a picture request in ` ... `. The server catches the tag mi
arXiv:2609.38385v1 Announce Type: new Abstract: Supervised fine-tuning of discrete diffusion language models masks some response tokens and trains the model to recover their original values from the visible context. The masking pattern therefore determines both the context available to the model and the tokens it learns to predict. Uniform random masking does not explicitly account for the interaction between these choices. We introduce GoldiMask, which selects tokens to reveal as context by app
arXiv:2609.38359v1 Announce Type: new Abstract: High quality synthetic data is central to post training LLMs for adaptive AI applications that represent the diverse expert strategies and decisions in conversations. Prompting LLMs directly or conditioning them on end use scenarios yields low diversity data that collapses onto dominant modes. We propose a method to generate diverse high quality synthetic data using Generative Flow Networks (GFlowNets). We show that training GFlowNets to generate l
arXiv:2609.38409v1 Announce Type: new Abstract: Recent progress in large language model reasoning has been driven by benchmarks and reinforcement learning environments with automatically verifiable rewards, particularly in mathematics, code, and formal logic. These settings make model accuracy easier to evaluate and optimize, but it remains unclear how far success under fixed problem specifications and stable evaluation criteria transfers to reasoning outside such domains. Real-world reasoning o
Title explains it lol. I went through all the steps again when it wasn't working, and then discovered that the "what is my IP" sites show a different one than my router does. Apparently that means I'm behind CGNAT? I just wanted to access my jellyfin without needing tailscale on every device lol. I know the flare is need help, but from what I'm reading there isn't much to be done other than asking for my own IPV4. Just wanted to share how I wasted two hours lol
Barclays, the British universal bank, is expanding its strategic collaboration with Anthropic to integrate secure, enterprise-grade AI systems across its global operations. Barclays is extending Claude across the bank to accelerate software development, modernize legacy systems, and improve operational efficiency. As part of this rollout, Barclays expects Claude Code adoption to reach 50% of its developer population by the end of 2026, rising to a majority of software engineers in 2027. Barclays
TTS systems with autoregressive semantic modeling have demonstrated strong zero-shot voice cloning performance and rich expressive variation, but their sequential decoding incurs substantial latency. Non-autoregressive alternatives offer much faster generation, yet often rely on more restrictive reference conditioning, such as requiring transcripts of the reference speech during inference. We present Tacit-TTS, an efficient transcript-free zero-shot voice cloning system distilled from IndexTTS2.
Hello everyone! The headline feature is exit nodes: a device can now send its full internet traffic out through one or more sites you choose. The tunnel can also come up automatically instead of waiting for someone to open the app and connect. Windows and Mac got a major UI overhaul, and iOS and Android picked up a lot of new features. Pangolin is an open-source, identity-aware remote access platform that simply and securely connects and authenticates your users to applications, infrastructure,
Unified Multimodal Models (UMMs) often rely on separate visual representations for understanding and generation, increasing visual context length and complicating integration with established vision-language pretraining pipelines. Recent advances in pixel-space modeling offer an encoder-free alternative, but extending this paradigm from images to videos is non-trivial: video understanding and generation adopt different temporal representations, leaving the design of a unified visual interface an
Hi r/LocalLLaMA! We’re researchers at the Institute of Foundation Models (IFM), an AI research lab dedicated to open and independent development of frontier-class foundation models. We recently released K2 Horizon a connected fleet of six fully open models with size ranging from 0.9B to 375B. In addition to weights, we also open-sourced training data and recipes, training code, intermediate checkpoints, fine-grained training logs and evals. Ask us anything about pre-training and data mixes, post