How bad do you think models like Qwen3.8-27B or GLM-5.3-Flash would be with H-Neurons disabled?
TL;DR: this paper proposes a method to fix hallucination rates to very low levels or zero by disabling neurons which contribute to hallucination. This discovery has been out for a while now, but it hasn't been that popular, since it kind of lobotomises parts of the LLM. I honestly don't care too much about talking to AI, but instead care about it producing working and good code. I wonder what percentage models would get on e.g. DeepSWE if we found their H-Neurons and disabled them?
A CVE Dispute
Debian won’t ban AI code from its Linux distribution
Debian voted to allow developers to use AI tools in their contributions to the Linux distribution's "development, maintenance, [and] documentation." The new policy on AI acknowledges that "responsible" use of AI can improve developers' productivity, and goes on to say, "generative AI is neither exempt from nor subject to special rules beyond the standards already […]
What are your hopes for the new Mistral?
Mistral is to be release a new model this summer, they still are working on it. What are your hopes?
Think twice before installing this device promising free movies
In exchange for free stuff, devices make home connections part of a proxy network.
Hackers claim millions of patient records stolen during data breach at healthcare giant McKesson
The company, which distributes medicines and medical devices to hospitals and healthcare practices across the U.S., said it was hacked and expects intermittent service degradation.
SlopTV: an infinite livestream of AI slop generated from youtube chat comments, Minimax H3 on 2x5090
SlopTV: a YouTube live stream where the chat writes the programming. You type "capybara dj underwater rave", an LLM inflates it into a 400-word structured video prompt, one of my 5090s renders 15 seconds of it with MiniMax H3, and it airs on the same stream you typed into. Then people comment on that clip, and the ouroboros keeps eating. Inspired by infiniteslop from @levelsio, but running fully locally. Numbers: H3 open weights, 66GB on disk, the int8 pruned diffusion model (19.5GB) and the nvf
Fast Weight Attention for Continual Learning
Recurrent fast-weight memories and selective state-space models compress an expanding context into a fixed-size recurrent state, making the state transition an online learning rule. We study this rule under read-after-write autoregressive semantics. For the prefix-prediction objective considered here, the local fast-memory example revealed at step t is the prefix-aligned pair (x_t,y_t)=(ϕ(k_{t-1}),v_t). The common same-step association (ϕ(k_t),v_t) remains causal, but optimizes a different inter
WM-R1: Training GUI Agents to Reason and leverage World Models with Reinforcement Learning
arXiv:2608.27508v1 Announce Type: new Abstract: GUI agents trained with reinforcement learning (RL) have showcased strong environment learning capabilities on mobile platforms. However, RL typically demands extensive real-environment interactions, leading to high resource costs and instability, especially in GUI scenarios. To address these, we propose WM-R1, the first reinforcement learning framework that trains mobile GUI agents with world models instead of real environments. Specifically, worl
Run macOS Software on Linux
自主研究系统的上限,不只取决于模型有多聪明,也取决于系统能否分辨什么是新证据,什么只是一次偶然的高分。
Blind Men and the Elephant: Probing the Epistemic Myopia of LLMs under Long-Tail Divergent Knowledge
Factual question answering (QA) typically assumes a single canonical answer, obscuring whether large language models (LLMs) retain divergent accounts of long-tail facts. To address this gap, we introduce ElephantBench, a closed-book knowledge probe comprising 1,094 questions generated through an auditable graph-based pipeline. The pipeline retrieves related documents from a low-exposure web corpus, identifies naturally occurring disagreements, and converts them into multi-account QA records. Eac
Paint What You See: Benchmarking Dexterous Visual Tool Use in Multimodal Agents
Evaluation is shifting from static QA toward agentic settings where models act through external tools. We identify a critical yet underexplored capability within this space - dexterous visual tool use: fine-grained, closed-loop parameterized visual action in which models infer tool parameters from visual evidence, and those parameters directly govern the final result. Existing benchmarks cover web navigation, GUI operation, and software engineering, but rarely target this coupling between visual
Meta Security Researcher's AI Agent Accidentally Deleted Her Emails
AI-written code is still your code
Please, I beg you, we need to stop using Stored Procedures (from applications)
Got this for cheap, what do i do with it other than AI?
Got this 32GB VRAM card (basically 2x WX7100 in one card) for 100 USD. Gigasteal. I want to be running local LLMs for the most part but i was wondering what else can i do with it? Do you have any other uses
uv: Deduplicate all files in the wheel cache
Apple adds more allegations to its trade secrets lawsuit against OpenAI
It claims a laptop proves that a former employee accessed and used proprietary files for the AI company.
Internet centralization and the original sin of NAT
Harvard Law dropout raises $6M for Blue Voice to build a ‘Harvey for police officers’
Blue Voice is trained on department-specific laws, local ordinances, protocols, and guidelines that general-purpose AI tools can't access on the public internet.
Transfer files over an Ethernet patch cable
OpenClaw 2.0, Accidentally
Import AI 471: Why Hugging Face worries me; space mining; FIve Eyes on AI
Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now Import AI reader giveaway! Upcoming event: Fiction and the Future with Robin SloanI’ll be chatting with my chum Robin Sloan on the evening of Monday September […]
First time running local models
Sad that I only have 12gb of vram but this ik_llama is so fast
范式与华为达成重磅算力战略合作,成为首批拥抱国产最高端算力底座的AI企业
Malleable software = solid bases and custom code
Nemotron 3.5 Content Safety Moderator: A Compact Multimodal, Multilingual, and Reasoning Enabled Content Safety Moderator
arXiv:2608.27548v1 Announce Type: new Abstract: Safety moderation for deployed AI applications is moving beyond text-only prompts: systems increasingly need to judge images, documents, screenshots, and generated responses under policies that vary across domains. Existing guardrails usually cover only part of this setting, making it difficult to combine broad coverage, custom policy control, and low compute cost. We present Nemotron 3.5 Content Safety Moderator, also referred to as Nemotron 3.5 C
Hypothesize, Evaluate, Refine: A Scientific Agent for PDE Discovery with Unknown Spatial Coefficient Fields
arXiv:2608.27475v1 Announce Type: new Abstract: Discovering PDEs in heterogeneous media requires jointly identifying the governing operator and the unknown spatial fields that parameterize it. These tasks are coupled: changing field placement changes the differential law, while a sufficiently flexible field can conceal structural error on a single trajectory. We present Hypothesize, Evaluate, Refine for PDE Discovery (HER-PDE), a scientific-agent framework that discovers compositional PDE struct
Rating the Raters: Rasch Measurement Theory for LLM Evaluation
arXiv:2608.27463v1 Announce Type: new Abstract: LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by raters. Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our understanding of what is b
OpenAI买几万台Mac搞强化训练!英伟达的活被苹果抢了
什么样的AI业务,英伟达GPU和谷歌TPU搞不定,非得用Mac
C++26: Standard Library Hardening Experiments
Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding
Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the target entity throughout that interval. Existing multimodal large language models typically serialize dense localization trajectories autoregressively, causing decoding latency to grow with tube length and allowing localization errors to propagate across time. We introduce Parallel Tube Decoding (PTD), a generative formulation that decomposes grounding into a temporal block followed b
Claude Code reduces it's weekly limit by 17% – compared to today
35 years later, Torvalds' hobby project remains developed worldwide.
I think the military commissary's freezers were hacked
ChatGPT and Reddit now face EU's toughest online safety rules
Explosive growth comes with a new regulatory burden in the European Union.
vote for the Qwen 3.8
Remember to vote and comments guys ;)
Act with Intent: Distilling Behavior Intent for Vision-Language-Action Models
Vision-Language-Action (VLA) models can turn multimodal context into robot actions, but their action decoders are still trained largely by behavior cloning. This supervises which motor command was demonstrated while leaving implicit the local objective served by the behavior under the instruction. Future-based supervision enriches action learning with frames, latent observations, trajectories, or motion representations, but these signals capture particular realizations of what may happen rather
LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation
Autoregressive video diffusion enables scalable long-video generation by producing chunks from a bounded recent context. While recency-based caching preserves local continuity, it evicts historical cues needed when subjects, objects, scenes, or attributes reappear. Existing memory mechanisms expose models to nonlocal history, but access alone does not ensure effective use. Our analysis reveals that video DiT layers exhibit distinct preferences for current, recent, and distant context, suggesting
Polimill builds Japan's next-generation public AI infrastructure
Polimill uses OpenAI GPT models and Codex to help municipalities search and use administrative knowledge while accelerating development.
The safest job from AI may be writing
LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and Orientation
arXiv:2608.27472v1 Announce Type: new Abstract: Bayesian network structure learning (BNSL) from observational data struggles with orientation identifiability, while large language models (LLMs) offer broad but often unreliable causal knowledge. We propose combining these complementary sources through a novel representation, termed Probabilistic Dependency Graphs (PDGs). In a PDG, each edge is associated with a distribution over directed, undirected, and absent states, enabling fusion via weighte