Loopie 发布最强循环 Transformer,20B 参数 MoE 模型仅需 2B 活跃参数,在同等算力预算下显著优于标准 Transformer。
2026-07-21
— 开源模型冲击闭源护城河,安全护栏反成攻击者的掩护。
中国开源模型 Kimi K3 与 Qwen 3.8 发布,性能逼近闭源前沿,引发美国对开源模型禁令的讨论。同时,AI 安全护栏被指阻碍防御者修复漏洞,而攻击者却能绕过。此外,罗马尼亚土地登记数据库遭黑客清空,WordPress 高危漏洞利用成本仅 25 美元,安全形势严峻。
头条
中国开源模型 Kimi K3 与 Qwen 3.8 发布,冲击闭源模型格局多源事件 ×5
Moonshot AI 发布 2.8T 参数 MoE 模型 Kimi K3,阿里发布 Qwen 3.8,两者性能均接近 Anthropic Fable 5,且承诺在未来数周开源权重。此举将中美开源与闭源模型的性能差距缩短至 3-5 个月。 为什么重要:开源模型正在消解闭源 API 的技术护城河,迫使 OpenAI 等厂商考虑推动监管干预。这对依赖 API 的开发者意味着模型切换成本极低,供应商锁定难以为继。
社区普遍认为中国开源策略正在削弱美国闭源模型的竞争优势,但也有人质疑其安全性、政治偏见及实际采用率。
每天早晨,一份为你精选的科技日报
网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。
已发布 15 期 · 每天筛过 150+ 条只留值得读的 30 条
AI 动态
RAGU 开源模块化 GraphRAG 引擎,通过两阶段类型化提取和紧凑领域适配 LLM 解决知识图谱噪声问题。
NAVER 提出 On-Policy Delta Distillation,用教师模型与基模型的差异信号替代直接模仿,提升推理能力蒸馏效果。
研究发现多智能体数学推理中,评审者精度高并不保证批评被采纳,广播式同行讨论在难题上优于层级式管线。
普林斯顿与芝加哥大学研究发现,LLM 在模拟招聘中比人类更容易形成偏见,且能从经验中自行产生刻板印象。
开发与开源
微软开源 Resource2Skill,从教程视频和仓库等多模态资源中蒸馏可执行 Agent 技能,构建层级化 Skill Wiki。
Unsloth 正式支持 AMD 硬件,覆盖 RX 9000/7000 系列和 Instinct MI350/MI300,支持本地推理、微调和 RL。
xHC 扩展 Hyper-Connections 方法,突破 N=4 的限制,解决残差流扩展中的写回信息不足和混合成本瓶颈。
研究分析 102 万 PR 发现,从纯人工审查转向 LLM 和 AI Agent 参与审查,对审查质量和效率的影响仍缺乏实证。
Simon Willison 指出编码代理大幅降低了逆向工程和自动化家庭设备的成本与风险,改变了 ROI 计算方式。
社区热议
Claude Fable 发现雅可比猜想反例引发数学界震动,评论区对 AI 发现的验证方式和数学价值存在分歧。
评论区普遍认为AI发现雅可比猜想反例是重大突破,但也有人质疑其验证方式及数学价值。
消息称特朗普政府部分部门正重新推动对外国开源模型实施事实性禁令,社区担忧中国模型势头引发政策反弹。
研究通过四组嵌套 Codex 代理消融实验,揭示可执行世界模型和验证机制对 ARC-AGI-3 性能的各自贡献。
欧盟拟与美国签署增强边境安全协议,允许美方访问生物识别数据库以维持免签待遇,引发数据主权争议。
评论区主要认为这是数据共享而非出售,且美国入境本就需提供生物信息,但也有人认为这是欧盟向美国妥协的背叛行为。
Causal-Audit 提出可审计的因果推理框架,通过目标感知因果链构建解决 LLM 隐式推理不透明问题。
GitHub Trending
Codex Dream Skin
Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦
Skills for Design Engineers.
"Vibe-Trading: Your Personal Trading Agent"
更多值得一看(内容池 55 条)
An example TL;DR I have open-sourced NInfer , a from-scratch C++/CUDA inference engine currently specialized for two exact Qwen3.6 checkpoints on a single RTX 5090. Both the engine and the converted model artifacts are publicly available: Github : The main result: Qwen3.6-35B-A3B sustained 542 tok/s while generating a full 65,536 token completion, on a single RTX 5090, single request. My goal was to find out how fast inference can get on a single GPU (in my case RTX 5090), with a fixed model and
Two critical security flaws in WordPress’ software have given hackers the chance to remotely take over tens of millions of websites, according to an estimate by a cybersecurity researcher.
Under the new system, the protocol will take a looser, "stateless" approach to session IDs on the server side, similar to how most ordinary websites already work.
Edinburgh-based tech firm Craneware said customer data was stolen during a cyberattack. The company makes software that thousands of U.S. hospitals, pharmacies, and clinics rely on for billing patients, potentially exposing health data.
Hugging Face is urging users to rotate any access tokens stored on the platform and review account activity.
Governments look at banning ransom payments in face of increasingly sophisticated threats.
OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.
Augment Code's Vinay Perneti talks models, harnesses, and context.
Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturb
Reinforcement learning (RL) has become central to improving large language models (LLMs) on complex reasoning tasks, yet RL post-training is largely studied in isolation from the pretraining that precedes it. As a result, two basic questions remain open: (1) how do pretraining choices (model size, data) shape the returns to RL compute, and (2) what does RL actually do to the model? These questions are difficult to study in the standard LLM setting: pretraining corpora are vast and uncontrolled,
Under model--harness co-evolution, harnesses are not merely inference-time scaffolds but data-generating components whose execution traces can shape future foundation models. This motivates harness-in-the-loop learning: optimizing harnesses for both immediate agent performance and the quality of traces used for future model training. However, continually updating provider-built scaffolds is costly and labor-intensive. We therefore investigate whether optimizing user-constructed harnesses in a ta
I asked myself where the Bonsai models actually land, so I ran them and compared to the results I already have for qwen-3.6-35b-a3b and qwen-3.5-9b on the same harness. Thought it might interest more people. Setup: little-coder harness via the harbor adapter, all 89 tasks of terminal-bench 2.0, single attempt (k=1), 40-turn cap, temp 0.2. RTX 5070 Laptop 8GB, i9-14900HX, 32GB RAM, CUDA 13.1. Runtime is PrismML's llama.cpp fork (stock llama.cpp can't load the 2-bit kernels). Results: Ternary-Bons
Hey guys, we collaborated with AMD to enable you to train, run, and deploy LLMs across nearly all AMD hardware including Radeon, Instinct, Ryzen, and data center GPUs. It works on Windows , WSL, and Linux and we have optimized ROCm builds for both training and inference. If you don't know about Unsloth , we're a fully open-source local UI that enables you to do pretty much anything with local models (RAG, chat, train, coding etc)! For those who don't have AMD GPUs and only CPUs, we still support
This story originally appeared in The Algorithm, our weekly newsletter on AI. To get stories like this in your inbox first, sign up here. Over the weekend, several current and former advisors to President Donald Trump on AI publicly lobbed insults at the country’s leading AI companies. David Sacks, the president’s AI and crypto “czar” until…
Large language models (LLMs) are transforming recommender systems from matching co-occurrence patterns in historical behavior toward reasoning about the intent that drives it. RecGPT-V1 pioneered this paradigm on Taobao by centering user understanding, and RecGPT-V2 scaled it via coordinated multi-agent reasoning; both are deployed in production with consistent gains in user experience and commercial outcomes. However, operating RecGPT at scale reveals three challenges: (1) stateless behavior mo
arXiv:2607.15367v1 Announce Type: new Abstract: Desktop voice assistants are still dominated by cloud pipelines that ship raw audio off the machine and expose a fixed set of skills. We describe AnovaX, a small local-first assistant that runs entirely on the user's computer and treats the desktop itself as its action surface. A single Python process wires together a wake-word gate, a speech pipeline, an LLM planner (Gemini) that emits a JSON plan of tool calls, a whitelist-and-denylist safety lay
arXiv:2607.15660v1 Announce Type: new Abstract: While LLM agents demonstrate strong reasoning abilities in compact and well-defined scenarios, they struggle to maintain robustness and effectiveness when faced with large-scale, diverse, and dynamic real-world environments that demand seamless tool integration. To address this gap, we introduce ToolVerse, a comprehensive framework that scales up agentic RL environments and enables agents to perform complex long-horizon reasoning in Tool-Integrated
arXiv:2607.15280v1 Announce Type: new Abstract: Sequential diagnosis requires balancing diagnostic accuracy against resource costs through iterative information gathering. Existing Large Language Model (LLM) approaches exhibit a critical knowledge-reasoning gap: despite encoding extensive medical knowledge, they struggle to reason systematically under cost constraints, often resorting to excessive testing. We propose GraphDx, a knowledge-enhanced framework with two core innovations. First, we de
Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters. This reuse changes the residual-scaling problem: in an untied Transformer, each residual branch receives and applies its own parameter update, whereas in a looped Transformer one shared update aggregates gradients from repeated visits and is read back by those same visits in the next linearized forward pass. We formal
According to source, it is the locally ranked AI model, the best among 4b models Source :
Hello everyone, I wanted to share a recent project of mine, which brings a 13.1 million parameter convolution transformer model to a Thanks to quantization, this model now fits into 14mb of flash memory and it now sits at 256kb of SRAM as well as 4mb of PSRAM to transcribe 8 seconds of audio. The speed is still painfully slow. It is lightning fast compared to my initial attempt however, which took 10 minutes of inference time to transcribe 5 seconds of audio. I also gave the whisper tiny model a
Hyper-Connections (HC) expand the residual stream of Transformers into N parallel streams, providing a form of memory scaling beyond model width and depth. Manifold-Constrained HC (mHC) stabilizes this formulation at scale. The large gains from N{=}1 to N{=}4 suggest residual-stream expansion as a promising scaling axis. However, existing HC-family methods typically stop at N{=}4. Our experiments reveal why: scaling mHC beyond this point yields diminishing performance gains and rapidly increasin
I haven't kept up since around February so I'm just not even sure... and there are quite a few options. I don't care about parameter size, from tiny to huge, what matters most is performance, I just want all the best safely locally stored, I'll worry about running them later. So, what do you consider some of the best of the best currently? Whether highly specialized, giant do everything well, or anywhere in between Edit: Also what you use any specific models for or the best you've found for any
arXiv:2607.15550v1 Announce Type: new Abstract: Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex tasks, yet they introduce critical safety risks where a single erroneous action can lead to irreversible consequences. Existing safety mechanisms are primarily reactive, lacking the ability to assess risks before execution. In this paper, we introduce SeerGuard, a consequence-aware safety framework designed to mitigate these risks through pr
Open weights 975B multimodal model built for fine-tuning Discussion | Link
So, im 50 and autistic. I learned C++ in college when I was young and didnt really ever use it that much. I also learned a bit of assembly from using cheat engine for game hacking (dont worry! single player or private servers!) which meant I got a little LUA experience. I've always disliked coding. Its such a "bang head against wall" type activity but I got into it because its pretty cool being able to make a computer do what you want. This entire attitude died a LONG time ago because coding was
I'm looking to run a very lightweight local model that acts as the brain, handling the logic and comprehension, while hooking it up to a web search tool to act as its memory and knowledge base.
arXiv:2607.15459v1 Announce Type: new Abstract: A trained deep reinforcement learning policy is a black box, and we ask whether it can be made explainable by rewriting it as an executable logic program that reproduces its behaviour and that a person can read, a logic engine can run, and an optimizer can edit. We present a three-stage post-hoc transformation that extracts a frozen proximal policy optimization teacher, induces an ordered rule list from its decisions in the manner of classical rela
Article URL: Comments URL: Points: 30 # Comments: 22
It may seem counterintuitive, but 113-degree water is still cool enough to cool NVIDA's latest hardware.
So I've been working on this app for a very long time now. Started off by wanting an app that separates long form notes from short form notes visually and keeps everything organized on a canvas. The app features: -Sticky notes for quick, small notes and A4 notes for long form documents -A quick mode that uses local storage to quickly open the app and jot down a thought on the canvas -A PDF export option that exports your notes into 3 custom designed PDF styles -Visual hierarchy to replicate file
This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. AI is more likely than humans to form biases when hiring The next time you apply for a job, AI may screen your résumé before any human sees it. But there’s…
We have been having extensive discussions around open source strategy. We will discuss it more at our next board meeting, but one thing we’d like to do soon is to create a language model with the approximate capability of GPT-3 that can run locally on consumer hardware and release that. We’d like to do it soon, before Stability or someone else does. In general, we think this helps discourage others from releasing similarly-powerful models, and makes it harder for new efforts to get funded. — Sam
Your personal context file for every agent. Discussion | Link
the open ai exec in his "ai communism" post suggested a fraudulent FUD campaign and trump executice order against open source / chinese models. unfortunately that seems pretty likely to happen. i have always use huggingface for downloading models and they would be forced to comply. are there any established alternatives or mirrors that aren't us-based?