Cross-tokenizer online policy distillation study: strict 1:1 alignment already covers most student-generated tokens, and expanding alignment coverage may not improve learning.
2026-10-08
— The model price war and agent reliability both hit the accelerator today.
Anthropic releases Claude Haiku 5.5, entering high-throughput scenarios at $0.10/M input tokens; OpenAI opens GPT-6 and Intelligent UI to all users, with interactive visualization becoming the new default. Liquid AI open-sources the d1 decision model series, focused on real-time decisions with zero output tokens. Meanwhile, multiple studies focus on agent reliability: from tool-call evidence chains to cross-tokenizer distillation, and weight synchronization optimization for trillion-parameter RL.
헤드라인
Anthropic releases Claude Haiku 5.5: 1M context, $0.10/M input tokens다중 소스 ×3
Anthropic launches Claude Haiku 5.5, positioned as its cheapest and fastest small model yet, supporting 1M token context and up to 128K output, with input pricing at $0.10/M tokens, about 90% lower than Haiku 4.5 for prompts within 100K tokens. At the same time, Sonnet 5.5's cache reads price is halved. Why it matters: For high-throughput, cost-sensitive engineering tasks (summarization, compression, classification, sub-agent calls), this price point directly changes the economics of small-model selection; the 1M context and adjustable effort parameter also make it better suited as a subagent in coding workflows.
Most people acknowledge the performance improvements and the Max subscription API quota, but some worry about price increases after 100K tokens and cybersecurity restrictions.
OpenAI opens GPT-6 and Intelligent UI to all users
OpenAI rolls out GPT-6 to 1.2 billion weekly ChatGPT users and introduces Intelligent UI capabilities: the model can combine text, charts, forms, clickable buttons, and other interactive elements to answer questions. Why it matters: This marks a paradigm shift in LLM output from plain text to interactive interfaces, with direct implications for frontend generation, data visualization, instant tool building, and similar scenarios; but excessive visualization may also affect text exportability and auditability.
The comment section generally agrees that interactive UI is a natural evolution, but some believe it may over-visualize, reduce text exportability, and raise concerns about accuracy and dependency risks.
Liquid AI open-sources d1 decision models: real-time decisions with zero output tokens다중 소스 ×3
Liquid AI releases the Open d1 series: d1-3B supports text and images, and d1-omni-600M supports text+image or text+audio. Neither generates text; instead, they return calibrated typed answers in a single forward pass, with zero output tokens. d1-3B scores 48.57 on Decision Index 0.2.1, surpassing all 4B and 9B models. Why it matters: Decision models bypass token generation and directly output structured results, potentially more efficient than traditional LLMs on edge devices (16ms on Jetson AGX Thor) and in real-time control scenarios; first-day llama.cpp support also lowers the barrier to local deployment.
Research focuses on agent reliability: from evidence chains to tool-call failure modes다중 소스 ×3
Several papers simultaneously address agent reliability: CheckerBench uses 300 tasks from 297 CVEs to evaluate whether long-horizon agents can synthesize static analysis checkers; From Evidence to Action studies the breakpoints in the evidence-to-action chain for tool-calling agents; UNREAL proposes using a single model to unify retrieval and long-context evidence selection. Why it matters: The key bottleneck for agents moving from demo to production is precisely reliability—correct results do not equal evidence-backed actions. These works provide reproducible benchmarks and methods for evaluating and fixing agent failure modes.
Chrome 155 officially supports JPEG XL decoding
Starting with version 155, Chrome supports decoding the JPEG XL (.jxl) image format, offering 30-50% better compression than JPEG, lossless compression, built-in HDR support, and lossless JPEG transcoding. The implementation uses Rust to ensure memory safety. Why it matters: For web developers, JPEG XL is an important complement to AVIF in high-fidelity and lossless scenarios, especially suitable for photographic images and progressive decoding needs; the Rust implementation also reaffirms browser investment in memory-safe languages.
The comment section generally welcomes Chrome's renewed support for JPEG XL, seeing it as important progress, but some believe AVIF is better at low bitrates and that ecosystem support remains incomplete.
매일 아침, 당신을 위한 테크 다이제스트
웹은 전체 그림을, 구독자에게는 당신만의 것을 — 관심사 맞춤 AI 큐레이션, 개인 RSS 통합, 커뮤니티 반응과 함께 매일 아침 배달. 영원히 무료.
88호 발행 · 매일 150개+ 중 읽을 가치 있는 30개로 선별
AI 소식
TRACE proposes an FP4 quantization framework for MoE language model RL training, directly reducing the quantization gap between training and rollout paths.
NeMo-DCR enables bit-exact incremental compressed weight synchronization for trillion-parameter agentic RL, greatly compressing cross-region transfer of 1T checkpoints from 87.5 minutes.
Mistral releases the 1-trillion-parameter open-source model Large 4 (Le Chonk), focused on coding and cyber defense, with a preview already available.
Nous Research closes a $90 million Series B at a $1.5 billion valuation; its open-source Hermes Agent has been cloned over 24 million times.
개발·오픈소스
Docker releases the docker-agent CLI plugin, using declarative YAML configuration to build multi-agent collaboration without writing code.
Commenters generally question Docker Agent's vague positioning and unclear connection to Docker, seeing it as a bandwagon agent framework, but some believe it has potential value for secure and reproducible containerized development workflows.
Meta open-sources Rebalancer: a C++ assignment solver that handles about 40 million shard/server/traffic placement problems per day.
Artcraft uses AI reverse engineering to release 7 open-source Adobe alternatives, replicating the interfaces and features of Photoshop, Illustrator, and others.
커뮤니티 화제
Meta and Microsoft reduce internal use of Claude AI and shift to in-house coding tools; commenters believe the main reason is cost control and internal model self-use, not a decline in Claude quality.
Commenters generally believe this move is driven by cost control and promoting internal model self-use, rather than questioning AI's value; but some think it is more about data governance and interface adjustments, and does not mean Claude's quality is poor.
Three engineers used GPT-6 Astra to drive a Toyota Corolla through a drive-thru pickup, and only GPT successfully completed the driving task.
OpenAI Dots hands-on: always-on agents still have reliability issues such as misremembering names and failing captchas.
GitHub Trending
Star morluto / rea Reverse engineer anything with agents, from app behavior down to native binaries.
Sponsor Star mattpocock / skills Skills for Real Engineers. Straight from my .agents directory.
Star boykopovar / AnyPS5 Tool for automatic PS5 executables porting to Linux and Windows
Star ayghri / i-have-adhd A skill to stop your coding agent from burying the answer. ADHD-friendly output.
Sponsor Star cathrynlavery / diagram-design Editorial diagram design for Claude Code, Codex, GitHub Copilot, Factory Droid, and Pi. 42 diagram types. Self-contained HTML + SVG. No shadows. No Mermaid slop.
Star addyosmani / agent-skills Production-grade engineering skills for AI coding agents.
Star EpicGames / raddebugger A native, user-mode, multi-process, graphical debugger.
Sponsor Star thedotmack / claude-mem Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Star manaflow-ai / cmux Open source Ghostty-based macOS terminal with vertical tabs and notifications for AI coding agents. Built for multitasking, organization, and programmability.
Sponsor Star trycua / cua Scale computer-use 2.0 with open-source drivers, cross-OS fleets, and benchmarks for training, evaluation, and data generation.
더 볼만한 소식(49건 더)
Potentially big speedup for MoE models that don’t fully fit in VRAM. Are you GPU Poor? Show your speedups ;)
LLM agents often lack the operational knowledge to act reliably in new environments, as they must discover specific tool behaviors or environment conventions on their own. Without memory of past attempts, they repeat the same mistakes across tasks, leading to more task failures and longer trajectories. To address this, agentic systems typically rely on human-written guidelines or on procedural memory built from training tasks and an oracle verifier, both of which require prior knowledge of the e
Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a t
EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in TypeScript, and made a small photo gallery where search runs entirely on your GPU in the browser. ruNNtime also supports plenty of other vision-like models, and you can play with them in the interactive docs source:
now you can use GLM 5 Flash MTP locally
TP-Link can't sell latest routers in US, still needs exemption from FCC ban.
Enterprises adopting retrieval-augmented generation (RAG) face a recurring operational decision: promote, revise, or block a system version. The evidence is incomplete and the metrics come from fallible LLM judges. We report on AGO AI Quality Gate (AGO), an evidence-first quality-gate framework deployed in industrial RAG assessment engagements. AGO integrates four key components: a four-state decision model that treats missing data and judge errors as explicit outcomes; layered scoring combining
On-policy distillation (OPD) has emerged as a widely used paradigm for post-training large language models, reducing the train--test mismatch of conventional distillation by supervising the student on its own generated trajectories. However, existing OPD objectives remain largely token-local and outcome-agnostic, optimizing teacher--student agreement at each prefix despite reasoning quality being determined at the trajectory level. Reinforcement learning with verifiable rewards (RLVR), particula
Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughput benefit. We present SlimWise, a serving framework that tailors the expert pool to each inference phase. SlimWise performs prefill with the full model and decode with a pruned mo
On-policy reinforcement learning has become a central paradigm for improving the reasoning abilities of large language models. However, its effectiveness is often limited by reward sparsity: when a model fails to discover correct trajectories for difficult problems, the optimization process receives little useful signal and may stagnate. Existing approaches mitigate this issue by incorporating off-policy demonstrations, expert traces, or model-generated solutions, but they typically require the
Self-evolving reasoning models learn from their own generated questions, yet repeated self-training can lead to performance collapse. In this paper, we investigate why performance deteriorates over successive rounds and how to sustain self-evolution. Our analysis identifies two recurring quality problems in self-generated questions: invalid questions and repeated variants of the same mathematical questions. First, invalid questions become more prevalent across rounds, and answer-consistency filt
Low-precision execution can substantially accelerate reinforcement learning (RL) for large language models, but discrepancies between learner and sampler execution can destabilize policy optimization. In this paper, we characterize the interaction between mismatch and the policy-gradient direction, distinguishing locally amplifying from contracting update contributions that mismatch magnitude alone cannot identify. In native NVFP4 runs, we observe an early imbalance between the two amplifying re
Developers can build LLM agents by adapting third-party models through benign post-training. We study a supply-chain threat in which an attacker supplies a model with a backdoor: hidden behavior that produces malicious outputs when a particular input pattern appears. Focusing on software-engineering agents, we ask whether such backdoors survive the developer's supervised fine-tuning (SFT) and subsequent task-level reinforcement learning (RL). We observe that benign SFT substantially reduces atta
Article URL: Comments URL: Points: 50 # Comments: 19
The last mile toward enterprise AGI is a company that runs itself. Training and adapting such agents require longitudinal enterprise data, which remain scarce, costly to acquire, and often restricted by privacy constraints. Historical archives are also frequently incomplete and record only what actually happened. They cannot show the outcomes of alternative decisions. We introduce MiniCorp, an office simulator for studying how agents can collectively run a company while generating enterprise dat
A coding agent CLI designed around small local models first Discussion | Link
Explore a comprehensive coding guide to Laya, the open-source zero-shot decision engine. Learn how to implement typed decisions, fit custom temperatures, and build reliable abstention gates using real-world CLINC150 banking data. The post A Developer’s Guide to Laya: Zero-Shot Decisions and Calibration appeared first on MarkTechPost .
Recently Large Language Models (LLMs) and LLM-based agents increasingly need to incorporate knowledge acquired after pretraining, e.g., domain facts, user preferences, documents, and interaction experience. In-context learning (ICL) and ICL-based agent harness remain flexible, but they consume context capacity and incur repeated discretized encoding cost that grows with context length. In-parameter memory offers a complementary substrate: reusable memory information is represented in model param
Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, but this compression can make selective access to sparse and distant information difficult. Existing chunk-based extensions increase memory capacity, yet learned chunk-mixing coefficients may remain fixed with respect to input content and therefore cannot adapt historical access to each query. We introduce Hybrid Linear Attention (HLA), a query-dependent chunk-level attention mec
Language model agents are deployed with a harness, the software around the model that manages its context, tools, and feedback. When such an agent is distilled into a smaller one, the harness stays in place, so the student mainly needs the teacher-specific abilities that the harness cannot provide, such as acting correctly on harness information. Standard distillation, however, imitates the teacher's full outputs and treats the harness as part of the input. We propose Harness-Aware Distillation
Many useful language-model tasks cannot be evaluated by exact outcome verification. Rubric-based reinforcement learning (RL) addresses this issue by scoring open-ended responses against explicit criteria. However, because the reward is assigned after the complete response, the training signal does not directly identify which individual decisions contributed to the final score. We propose a two-stage training framework that uses rubrics first as privileged teacher context for dense token-level su
The US government’s Tradewinds initiative has made it easier to throw millions of dollars at “nontraditional” defense contractors, including OpenAI, Anthropic, and Google.
OpenAI is launching a new user interface that will bring interactive visuals to ChatGPT.
At today's Windows and Surface event, Microsoft showed off an upgrade to its Copilot AI system that will give it access to local files on your PC and the ability to take actions across the OS. It's part of an idea Microsoft is calling "Hybrid Intelligence," where apps and tools rely on a mix of […]
Hey Pocket-ID Dev-Team, Hey stonith404 , I just wanted to say thank you for your work. After a couple of months of intensive use, I have to say: Pocket ID is the service that makes using my self-hosted apps so much easier and more comfortable. I came across Pocket ID while searching for an encrypted file transfer service, and I've been following its development ever since. In the beginning, I didn't have much trust that a young developer could build secure and reliable software. Back then, I did
OpenAI is updating ChatGPT for all users with an “Intelligent UI” that’s more visual—generating interactive elements as part of the chatbot’s outputs.
Microsoft's Copilot search redesign arrives this fall.
The new SynthID website can now identify AI content from Google, OpenAI, and more.
Vision-language-action (VLA) policies often fail when a robot's executed motion deviates from their commanded action. Such execution errors arise from the robot's mechanics and operating conditions, such as wear and payload changes. We propose self-compensating VLA, a deployment-time adaptation method that enables a VLA policy to pre-compensate for the robot's execution errors when generating commands. Without task rewards or labels, it updates the policy online using the residual between the ac
I’ve been pushing TinyStories-style models downward in size, and this is the smallest one so far: MacroStories — 19,969 parameters, 81 KB FP32 For scale: → ~50× smaller than the 1M TinyStories model → ~3,000× smaller than AlexNet → 32-dim hidden state → 378-token vocabulary → one decoder block, recurrently applied 4 times with shared weights It’s obviously not a general-purpose LM, but within its constrained story distribution it can maintain a 100–300 word narrative with a goal, problem, releva
We introduce Adaptive LeWorldModel (ALeWM), a world model based on a joint-embedding predictive architecture (JEPA) that learns to concentrate predictive information in compact prefixes of a wide latent representation. To encourage this ordering, ALeWM learns a sequence-conditioned distribution over prefix lengths and trains the predictor to estimate the full next embedding from a sampled input prefix. As standard anti-collapse objectives encourage variation across latent coordinates and do not
Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks like. However, these signals are often not grounded in individual student learning outcomes, where effective teaching strategies can vary substantially across learners. To address t
I set up borg via borgmatic like a year ago. 3-2-1 strategy. Confirmed it was backing up and did a quick extract test. That was it. That was a year ago. Well I set a vm in Proxmox and because I was still learning I didn’t set up some directories correctly. Later I installed Immich but apparently installed it under a directory owned by Nextcloud. I never updated Nextcloud because it was local and then I decided to make it available remotely via a reverse proxy and all that fun stuff. So I wanted
I thought this was a very interesting article, of relevance to the readers here.
A safe action is not necessarily a viable one. A frozen vision-language-action (VLA) policy can favor a locally admissible move that leaves no policy-supported route to safe task completion. We call this the feasibility-likelihood gap: likelihood ranks the next move, while feasibility depends on the futures it leaves open. To bring those futures into the decision, we derive the exact next-block marginal of the history-conditioned policy-environment trajectory law restricted to safe task completi