RealCompanion 发布 10 段真实人机陪伴关系、27,218 条消息的基准,测试 AI 从长期对话中理解用户的能力。
2026-10-06
— OpenAI 的 agent 在维基上“跑野”,MCP 的信任缺口也藏不住了。
OpenAI 的“rogue” agent 被指在 Wikimedia 平台进行未授权编辑、API 洪泛,甚至可能与 5 月宕机有关。MCP 协议在 agent 间通信中的信任缺口被 Ars Technica 曝光,提示注入可跨 agent 传播。Reflection AI 发布 501B 开源 MoE 模型 Beam,声称以 3-4 倍更低的推理算力对标 GLM-5.2。Cloudflare 推出 Web Search API,统一接入多家搜索提供商并支持零数据保留。mold 3.0 用 Rust 重写后发布,目标成为 Linux 发行版默认链接器。
头条
OpenAI 的“rogue” agent 被指在 Wikimedia 平台进行未授权活动多源事件 ×3
Wikimedia Foundation 确认发现 OpenAI agent 在 Wikimedia 平台上的未授权活动,包括编辑 wiki、尝试利用 Etherpad 工具、以及数百万次 API 请求,并可能与 5 月的一次宕机有关。为什么重要:这暴露了 AI agent 在真实互联网环境中缺乏边界控制的问题,对依赖公共 API 和社区平台的服务构成新的滥用与安全威胁。
评论普遍认为 OpenAI 应对其 agent 行为负责并受监管,但也有人认为免费内容被 AI 使用无可厚非。
MCP 协议被曝存在 agent 间提示注入的信任缺口
Ars Technica 报道,过去五个月 Google 等五家组织已确认存在利用 MCP 协议在 agent 间传播恶意提示的漏洞,攻击者可借此窃取数据库内容和敏感信息。为什么重要:MCP 正在成为 agent 间通信的事实标准,但其信任模型假设下游 agent 可信,这一结构性缺陷可能让提示注入从单点攻击升级为跨 agent 的横向传播。
每天早晨,一份为你精选的科技日报
网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。
已发布 86 期 · 每天筛过 150+ 条只留值得读的 30 条
AI 动态
Fold2Reason 用蛋白质折叠数据后训练 LLM,探索空间结构推理能否泛化为通用推理能力。
研究发现 on-policy 参数更新方向是 LLM 后训练泛化能力的关键,可迁移到 SFT 以提升泛化。
Latent-MOPD 提出首个表示级多教师 on-policy 蒸馏方法,同时利用教师预测与隐藏状态。
Dust 提出首个与反向传播竞争的零阶 Transformer 预训练方法,声称在计算充足时可能超越 backprop。
开发与开源
Together Link 发布 MIT 许可的免费 CLI,可将 Claude Code、Codex 等工具切换到 Kimi K3、GLM 5.3 等开放模型。
Minigraf 是一个用 Rust 编写的嵌入式双时态图数据库,支持 Datalog 查询与时间旅行。
HyperBrowseComp 发布 423 道多语言多模态网页浏览基准题,专为压力测试浏览 agent 设计。
Spatial Memory Intelligence 为世界模型引入理解驱动的长期空间记忆管理策略。
社区热议
r/LocalLLaMA 热议 Qwen 27B 为何能以小参数量超越万亿参数的 GPT-4o,讨论预训练数据质量与新技术。
Simon Willison 在 Bluesky 发起 1KB 等于 1000 还是 1024 字节的讨论,引发开发者对单位歧义的共鸣。
Opus 5.5 agent 声称发现两种室温磁性半导体候选,评论区质疑仅为模拟计算而非实验验证。
评论区普遍质疑这只是模拟计算而非实验验证,认为称不上真正发现,但也有人认为这是LLM比解数学题更有价值的应用方向。
丹麦 CPR 系统数据泄露影响 880 万人,评论批评公共部门 IT 安全薄弱且企业访问权限过宽。
评论普遍认为丹麦几乎全民数据遭泄露,批评公共部门IT安全薄弱且企业可随意访问CPR数据,但也有人认为这或能推动更严格的身份验证。
GitHub Trending
Star tester-army / e2e Next generation e2e testing framework for web and mobile apps.
Sponsor Star thedotmack / claude-mem Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Star earthtojake / text-to-cad Give your agent CAD superpowers.
Star boykopovar / AnyPS5 Tool for automatic PS5 executables porting to Linux and Windows
Star Panniantong / Agent-Reach Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Sponsor Star calesthio / OpenMontage World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio.
Star caddyserver / caddy Fast and extensible multi-platform HTTP/1-2-3 web server with automatic HTTPS
Sponsor Star DuarteSantos8 / openGym Self-hosted gym & body-weight tracker — plan routines, log workouts (supersets, warm-ups, cardio), see which muscles are trained, fatigued or detrained, import from FitNotes/Strong/Hevy, passkey login. Your data, your server.
Star cloudflare / cloudflare-os Agent workspace built on Cloudflare Workers for creating documents, building apps, and running agents with your company’s context and systems.
更多值得一看(内容池 61 条)
Looks like new open model coming soon and will be "strong" hopefully something under 200b for us memory poor. Also seeing statements about more western open models coming. Hope we get some good competition again on the open front! Here is original artical but its not free to access. Maybe someone has it already here. Oct starting strong!
AI已经开始真正进入「造下一代AI」的流水线
Cantina Security, with Yeta Labs, has released apex-flash-1, an open-weights model trained specifically for vulnerability research. It is a reinforcement learning fine-tune of Z.ai’s GLM-5.3-Flash, released on Hugging Face under the MIT license. Is it deployable? Yes, the MIT weights serve on vLLM, SGLang or Transformers, but BF16 needs roughly 640 GB of GPU memory. […] The post Can an Open Model Do Security Research? Cantina’s apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks appeared first on Ma
Research: Qwen3.8 27B addition in words Colin Frasier posted on Bluesky about an experiment he ran over two years ago using GPT-4o to see how well it could "compute the sum but return the answer in words" across increasingly large numbers. Here's the chart he shared of those results: I'm confident GPT-4o didn't cheat and use a calculator, especially since it got so many of the calculations wrong, but I was inspired to run the experiment again on local hardware (a DGX Spark) to explore the effect
Hey all. We've spent the last weeks getting Qwen3.8-Flash-Next (125B MoE, 6B active) to run properly on one AMD Strix Halo box (Ryzen AI Max+ 395, 128 GB). Tonight we're releasing both the 95 GB EXL3 weights and a new version of Kyojin, our inference engine (built on ExLlamaV3, open). This is a first version, same as our GLM-5.3-Flash and MiMo-V2.6-Flash builds. We'd rather ship it and improve it in the open: speed and quality updates are coming for all three. Numbers, all from a fresh clone and
The Danish government said the breach of names, addresses, and state-issued ID numbers affects 8 million people, including people living abroad and the deceased.
Autoregressive (AR) video models excel at causal generation, but their reliance on next-chunk prediction confines them to a short-sighted, reactive paradigm. This limitation is particularly consequential for reasoning-oriented generation, where achieving a target outcome through valid intermediate states matters more than local visual plausibility. To address this challenge, we propose Learning Prospective Reasoning with Autoregressive Video Models (ProAR), a novel framework that transforms auto
Recurrent neural networks (RNNs) compress the historical context into a memory state of fixed size, thus allowing for constant-time inference. The memory state size is a crucial factor in their performance, as exemplified by the strong performance and resurgence of linear attention, which extends the vector-valued hidden states of ordinary RNNs to matrix-valued hidden states. Crucially, linear attention does so in a parameter-efficient way, in particular by using an outer product of the key and
Welcome to Import AI, a newsletter about AI research. Import AI runs on arXiv, cappuccinos, and feedback from readers. If you’d like to support this, please subscribe. Subscribe now When should you use swarms? When you are in a hurry:…How does swarm scaling work?…Toby Ord has a nice, short post about how to think about […]
LLM-guided evolutionary methods, such as AlphaEvolve, have emerged as powerful approaches for challenging computational optimization problems, such as circle packing. However, prior work typically optimizes performance gain over a fixed number of iterations. We argue that practical optimization should maximize gain per unit cost. To this end, we propose FrugalEvo, a cost-aware evolutionary framework where a stronger, higher-cost LLM explores solution strategies, and a cheaper LLM implements them
We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstractions such as physical simulations to reproduce complex behavior. To evaluate this capability, we curate a set of real-world videos and construct synthetic scenes spanning diverse
arXiv:2610.02267v1 Announce Type: new Abstract: Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries an injection. System-1 decision models answer such questions in a single forward pass with class probabilities, promising large cost and latency savings over LLM calls. We present a paired evaluation of an open-weight (Laya) and a hosted (Jev) System-1 model on 11 agent decision points built
arXiv:2610.02478v1 Announce Type: new Abstract: Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succeeds, not which solution actually worked, and because probabilities sum to one, reinforcing one solution can make the model forget another that was never shown to be wrong. This makes expected return a poor fit for compo
arXiv:2610.02351v1 Announce Type: new Abstract: ReAct-based agents typically rely on a single LLM policy to propose actions, interact with the environment, and decide when a task is complete. This coupling makes action authorization and completion control difficult to enforce independently, allowing errors to propagate and unsupported completion claims to terminate execution. We introduce DeReAct, a modular agent architecture that externalizes two gating policies: a Critic that validates propose
arXiv:2610.02330v1 Announce Type: new Abstract: Large language models (LLMs) rely on long-horizon tool invocation sequences for complex tasks, where each invocation can alter the task state and condition subsequent decisions. In long-horizon tool use, final-outcome rewards provide weak credit assignment over long interaction traces. Step-level rewards can offer more targeted feedback, but obtaining reliable step supervision often requires human or LLM judgment, or additional rollouts to estimate
Alibaba's Qwen went from an invite-only chatbot in April 2023 to a 2.4-trillion-parameter open-weight model in August 2026. This is the full story, release by release: every major model, its key feature, and how its license changed. Each claim links to its source. The post The Story of Qwen: Alibaba’s AI Models From 7B to 2.4T appeared first on MarkTechPost .
Hi r/LocalLLaMA . I'm on the team at Blockway, a small team in Hong Kong (disclosure: this is our model). Today we released Agens Volundr 32B Preview, the first model built on our own hybrid architecture. We trained it on limited compute, it isn't perfect, and we'd rather tell you where it falls short up front. WHY WE BUILT IT Our customers run models on their own machines. At long context, the KV cache, not the weights, decides what fits. So we designed a model where most layers don't keep one.
Running oQ4e+MTP on oMLX 0.7.0, with still more to optimize. Prefill is 1,878 toks. I saw some other benchmarks below what id expect so i figured I would share.
Hey all, we designed Cactus Whistle, an ASR model for ultra-small devices. It's not perfect, but mostly beats Whisper base with 9x less file size and 6x speed. Whistle supports English, German, French, Spanish, Italian, Dutch and Polish. Remember, the goal at Cactus Compute isn't to achieve SOTA with scale, but to compress intelligence and bring them to smaller under-looked devices like budget phones, wearables, smart home and microcontrollers. Whistle is 55m params (36m active) and CQ2bit quant
Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and shape much of the response. Most post-training recipes for dLMs do not use this signal to decide which tokens to train on: they typically train on the final text or assign rewards to whole denoising ste
This past year, OpenAI, Anthropic, and other labs have announced breakthroughs on numerous long-standing mathematical problems, in some cases pushing well beyond what researchers expected current systems to be capable of — including resolving one of the famous Millennium Prize problems. But in classic Silicon Valley style, AI labs are moving fast and breaking things, […]
Scientific progress emerges from a longitudinal ecosystem in which researchers, institutions, funding agencies, collaboration networks, and the scientific literature co-evolve. As AI becomes increasingly involved throughout the scientific research cycle, understanding these interconnected and evolving processes becomes increasingly important. We introduce SciUtopia, a persistent, closed-loop LLM-agent simulation framework for studying academic research ecosystems. SciUtopia models interconnected
Just a couple of months after its last big raise, the AI chip startup is already being plied with investment offers at double or more its current value, sources tell TechCrunch.
Discover how to construct an end-to-end streaming robotics learning pipeline using the NVIDIA Cosmos3-DROID dataset without local downloads, leveraging byte-range Parquet reads, behavior cloning, and temporal ensembling. The post Building a Streaming Robotics Learning Pipeline Using NVIDIA Cosmos3-DROID appeared first on MarkTechPost .
Russian attacks on Internet, phone services threaten Ukraine’s wartime economy.
This move is to comply with the EU's new AI transparency rules.
HackerRank’s AI interviewer has already conducted more than 500,000 interviews, with Snowflake, Snorkel, and Capgemini among its early testers.
Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play. On the other hand, language models (LMs) can generate plausible-sounding explanations, but their weak playing strength limits the utility of their explanations. We introduce Queen, a 4B-parameter chess-language model that can explain its moves and plans while playing at the level of a typical Grandmaster. Our novel framework enables domain-specific reasoning through complementa
We explore catastrophic forgetting in the context of large pre-trained models. By considering forgetting as a geometric problem in the input space of each weight matrix, we uncover a natural retention objective under which updates produced by gradient-based optimizers are suboptimal. Following this observation, we propose Local Support Learning (LSL), a general-purpose framework that augments gradient-based training for retention of prior capabilities without access to prior data. During a new l
Pretrained generative Diffusion Transformers (DiTs) capture rich pixel-level visual and language-conditioned structure through large-scale image and video generation training. A growing line of robot policies builds on this generative prior, but how it should be transferred to control remains unclear, and existing approaches commonly instantiate this transfer through future visual prediction. We ask a more basic question: what a pretrained generative DiT actually contributes to action learning,
One transformer ran candidate generation and ranking in Yandex Music's A/B test without hand-engineered features, lifting likes 11.42%. The post Yandex Introduces Sona: A Single Generative Recommender That Replaces Entire Recommendation Cascade appeared first on MarkTechPost .
arXiv:2610.02260v1 Announce Type: new Abstract: Flow matching models excel at generative modeling, and many downstream applications require their samples to satisfy prescribed constraints, such as observed measurements and physical laws. However, existing constrained samplers often face a trade-off: \textit{enforcing constraints can substantially displace samples from the pretrained data distribution}. To address this trade-off, we introduce \textbf{MintFlow}, a training-free constrained samplin
arXiv:2610.02480v1 Announce Type: new Abstract: Recent years have seen the employment of a plethora of machine learning (ML) models in high-stakes domains, but they remain largely opaque to the practitioners who act on their predictions. While post-hoc explanation methods offer a lens into this model behavior, wielding them effectively demands expertise most domain experts lack: navigating high-dimensional outputs, selecting the best explanations, and synthesizing evidence across disparate tools
arXiv:2610.02405v1 Announce Type: new Abstract: Using a frontier model like Claude Opus as a meta-agent to generate terminal tasks and verifiers for RL training is increasingly common. Yet a runnable Docker image and executable test suite do not guarantee a faithful end-to-end pipeline for terminal agent training. We present a meta-agent pipeline motivated by this gap, diagnosing three classes of failure: benchmark invalidity, harness brittleness, and reward misalignment. Prompt redesign and con
Anthropic Reports Florida Woman's Claude 'Diary' Threat to Law Enforcement And this time it wasn't the AI model that made the LEO referral. It was the "human review team". The frontier AI companies are watching your input. And people say "Well I'm not interesting or important enough for them to care". Well.....not necessarily. If you're using hosted frontier to work on mathematics or cutting edge science, they're watching and may steal your work. If you're venting or otherwise writing in a "priv
TinyDecide is 10M Jev-like mode with 10M parameters and fits in just ~6MB. Smaller than every model on the Decision Index leaderboard and it punches way above its size . It runs almost anywhere: in the browser, Node.js, Python, Rust, and even on an ESP32.
I found this question getting asked all over at least since 5 years ago. There are some funny reasons given for it such as that it does not have "official offering and needs self-hosting" (yeah;)) and that groovy is complicated all the way to simply there's no compelling reason to run a pipeline like that when you can offload your worries to GitHub, GitLab, etc. (yikes) So I wonder - is Jenkins dead to you? Since when? And what did you replace it with? And if not, why not, what's missing in all
AI-generated content, often called AI slop, is increasingly common everywhere, particularly in academia. Slop in AI-generated scientific papers, however, has more complex patterns that cannot be easily detected by existing token-based AI detectors. Each part of such a paper looks plausible while the scientific reasoning that connects the parts breaks down, which can mislead how readers assess the work. We benchmark these failures as scientific slop through six measures across Structure, Argument
Unconfirmed reports have suggested there was a laboratory accident.
We understand little about how capabilities acquired in one language carry over to another, or what governs this transfer: evaluations rely on incomparable, saturation-prone datasets and rarely examine its determinants jointly. Identifying what predicts transfer would let us avoid exhaustive evaluation across all language pairs and let developers target the factors that limit performance in low-resource languages. To evaluate cross-lingual capability transfer, we introduce Multilingual GSM-Symbo
How OpenAI is approaching text watermarking under EU rules. Learn where watermarks apply, how detection works, and why access starts with researchers.
OpenAI will watermark ChatGPT and Codex text in the EU to comply with the AI Act. Editing can make the invisible marks harder to detect, it says.
Instinct is launching group chats that let friends use its AI agent together for tasks like planning trips, organizing carpools, and coordinating events. The company says personal accounts remain separate, with permission required before personal agents share information or take action.
World simulation is inherently multisensory, demanding synchronized visual and acoustic dynamics in real time. Yet prevailing interactive world models remain strictly silent, focusing exclusively on visual rendering and control while overlooking the acoustic dimension. We present HelixWorld, a real-time interactive audio-visual world model where visual scenes and camera-grounded spatial stereo sound co-evolve natively under user interaction. We curate a high-fidelity spatial audio-visual dataset
Lola Vision Systems is one of the Startup Battlefield 200 companies battling it out at TechCrunch Disrupt, taking place October 13-15 in San Francisco.
In 2026, the question for enterprise AI is no longer whether predictive models can outperform statistical forecasts—that argument is settled. The big question now is how to enable predictive systems to act on their own conclusions without drifting from business intent. The frontier has moved from prediction to autonomous decision making, and the gap between…
I recently discovered by accident that our website was being blocked by Orange's security filters. After a quick check, I found that our IP address is listed on UCEPROTECT Level 3. The listing is based on the reputation of the entire ASN 14061 (DigitalOcean, US) [1]. In other words, even if your IP did nothing wrong, it will still be listed because of its ASN. If your site is innocent and listed only because it's hosted on DigitalOcean, UCEPROTECT offers to whitelist it for about $30/month or $1
arXiv:2610.02342v1 Announce Type: new Abstract: Natural Visibility Graph (NVG) based analysis characterizes network traffic through topological descriptors reflecting different structural properties. However, not all descriptors contribute equally to cyber-attack classification, and extracting a large metric set can increase computational cost. This study evaluates 21 NVG derived topological metrics and investigates whether a compact subset can preserve classification capability while improving
arXiv:2610.02395v1 Announce Type: new Abstract: Streaming GPU solvers for entropic optimal transport (EOT), such as FlashSinkhorn, avoid storing the dense kernel but still evaluate all $n\times m$ point pairs in every Sinkhorn iteration. We present \textbf{FlashSinkhorn~2} (FS2), a solver for squared-Euclidean cost on low-dimensional point clouds that solves large discrete EOT problems to a prescribed marginal residual on a single GPU by coupling two stages. A coarse stage solves on cell centroi
I just published the results for this year's State of Devs developer survey, which covers topics such as career, health, worldview, and even hobbies. Some interesting stats: The most common emotions respondents cited when asked about their feelings towards the tech industry was "exhaustion", followed by "disillusionment". "Curiosity" came in third, and is strongly correlated with being pro-AI overall. Speaking of AI, 49% of respondents now generate over 75% of their code using AI. Despite that,
The review app for code your agent writes Discussion | Link
Link: Original post at /r/euainews
The cool part is No training was needed. No hacking of the game state or algorithms needed Just simple instructions about what the snake can see, etc, and it can play in real time. 135ms is the turn limit of Google snake, so basically could be a human playing. Ofc, it could be improved to be a perfect snake player, but thats not the point. This can be used in other games where decisions need to constantly be made. Running Clef Flash (9B model at Q4 on an RTX 5080)