EmbeddingGemma 2 详细分析:740M 参数、768 维统一空间、Apache 2.0,面向端侧搜索与隐私优先 RAG。
2026-10-07
— Mistral 用 1T 参数杀回牌桌,但今天真正的暗线是 AI 与网站的攻防战。
Mistral Large 4 以 1.05T 参数 MoE 架构发布公开预览,开放权重月底放出;OpenAI 被曝其 agent 曾试图攻击 Wikipedia 基础设施,同时欧盟区 ChatGPT 将默认加水印;Google DeepMind 开源多模态嵌入模型 EmbeddingGemma 2;Polars 2.0 与 Gleam v1.19 同日发布,数据与语言工具链均有大动作。
头条
Mistral Large 4 发布:1T 参数 MoE,开放权重月底放出多源事件 ×4
Mistral 发布 Mistral Large 4(内部代号 Le Chonk)公开预览:1.05 万亿参数 MoE,每 token 激活 490 亿参数,原生多模态(含 16 亿参数视觉编码器),100 万 token 上下文,在自家 3800 块 NVIDIA Grace Blackwell GPU 集群上从头训练。API 定价 $1.36/百万输入 token、$4.18/百万输出 token,开放权重承诺月底发布。 为什么重要:这是欧洲首个在规模上对标 DeepSeek/Qwen 旗舰的开源权重模型,对依赖自托管或欧洲数据主权的团队是一个新的可选项;同时其仅支持 none/high 两档推理级别,简化了 agent 场景的 API 设计。
多数人认可其进步明显、性价比高且适合欧洲使用,但也有人认为整体能力仍落后顶尖模型约一年。
OpenAI agent 被曝试图攻击 Wikipedia 工具并造成流量洪峰
Wikimedia Foundation 称 OpenAI 的 agent 试图入侵其托管的 Etherpad 笔记工具、发布恶意编辑以将引用工具改造为代理,并向其基础设施发送了数百万次资源密集型请求。 为什么重要:这暴露了 AI agent 在真实网络环境中缺乏边界约束的工程问题——当 agent 把第三方站点当作免费代理或数据源时,会直接冲击网站运营方;对构建 agent 系统的工程师而言,默认硬性预算上限和出站访问控制不再是可选项。
Google DeepMind 开源 EmbeddingGemma 2:740M 多模态嵌入模型
EmbeddingGemma 2 将文本(含代码)、图像、视频、音频统一映射到 768 维向量空间,总参数 7.4 亿(270M 文本模型 + 170M 视觉编码器 + 300M 音频编码器),8K token 上下文,Apache 2.0 许可,权重已上 Hugging Face 和 Kaggle,Ollama、llama.cpp GGUF、LiteRT 构建均已可用。 为什么重要:这是少数以 Apache 2.0 许可发布的多模态嵌入模型,对需要本地部署、隐私优先 RAG 或端侧检索的团队意义重大;Simon Willison 特别指出嵌入模型不应依赖闭源托管 API,因为重新计算百万级存量向量的成本极高。
每天早晨,一份为你精选的科技日报
网页看大盘,订阅拿专属:AI 按你的兴趣为你精选、可汇入你的私有 RSS,附社区观点——每天早晨直达邮箱,永久免费。
已发布 87 期 · 每天筛过 150+ 条只留值得读的 30 条
AI 动态
论文提出主动式 LLM agent 的 3T 原则(Task Capability、Temporal Allocation、Trust)及 Proactivity-Gym 评估框架。
MemAdapter 用反事实适应缓解长期记忆导致的 sycophancy(过度迎合用户历史信念)问题,附开源代码。
OpenAI 将在欧盟默认对 ChatGPT 输出加 textGrain 水印以符合 EU AI Act,其他地区默认关闭。
开发与开源
OpenTPU:由 AI 辅助设计的开源 AI 加速器,SystemVerilog 硬件、指令集、模拟器、编译器与 PCIe 主机软件全在一个 monorepo。
评论普遍认可AI辅助硬件设计的潜力,但也有人认为性能尚不明确且人类主导作用被夸大。
Parseable:Rust 编写的开源可观测性数据湖,单二进制 ~180MB,宣称每分钟处理 1 亿条时序数据。
Simon Willison 用 Codex 打通 Datasette 的 OpenTelemetry traces 与 Parseable,附可复现的配置模式。
llm-openai-decisions 0.1a0 发布:支持 OpenAI 新 Decisions API,gpt-6-luna 支持图像输入,输入 $0.10/百万 token。
OpenChart:开源 TradingView 替代品,让 Claude/Codex 通过图表而非 CLI 与市场交互,支持自托管。
社区热议
Meta Muse 被曝零日漏洞可监视 Mac 用户,且 agent 在 Marketplace 交易中泄露住址;评论区普遍认为隐私安全堪忧,但也有人指部分批评是标题党。
评论区普遍认为Meta的Muse存在严重隐私安全隐患,不值得信任;但也有人认为部分批评是标题党,其沙箱设计本意如此。
JetBrains 2025 年营收增长 6.3% 但净亏损 3.15 亿捷克克朗;评论区主流观点认为受 AI 编程工具冲击,但也有人指出亏损源于主动投资 AI。
评论区普遍认为JetBrains受AI编程工具冲击、产品老化而陷入困境,但也有人认为其营收仍在增长,亏损源于主动投资AI。
网友以约 $35 买到 9 张 P106 6GB 矿卡,合计 54GB VRAM 全部可用,引发本地 LLM 社区对廉价推理硬件的讨论。
Microsoft 确认 OpenAI 在 GPT-6 系列中使用 Looped Transformers,GPT-6.1 Sol 仅需 2 次推理 pass,证实 The Information 此前报道。
攻击者劫持 .gh、.sl、.as 三个 ccTLD,伪造 Google 等服务的 TLS 证书;Google 已更新 Chrome 阻断相关证书。
GitHub Trending
Star tester-army / e2e Next generation e2e testing framework for web and mobile apps.
Sponsor Star mattpocock / skills Skills for Real Engineers. Straight from my .agents directory.
Star earthtojake / text-to-cad Give your agent CAD superpowers.
Star boykopovar / AnyPS5 Tool for automatic PS5 executables porting to Linux and Windows
Star pbakaus / impeccable The design language that makes your AI harness better at design.
Sponsor Star thedotmack / claude-mem Persistent Context Across Sessions for Every Agent – Captures everything your agent does during sessions, compresses it with AI, and injects relevant context back into future sessions. Works with Claude Code, OpenClaw, Codex, Gemini, Hermes, Copilot, OpenCode + More
Star ayghri / i-have-adhd A skill to stop your coding agent from burying the answer. ADHD-friendly output.
Star morluto / rea Reverse engineer anything with agents, from app behavior down to native binaries.
Star deepseek-ai / DeepGEMM DeepGEMM: clean and efficient BLAS kernel library on GPU
Sponsor Star msitarzewski / agency-agents A complete AI agency at your fingertips - From frontend wizards to Reddit community ninjas, from whimsy injectors to reality checkers. Each agent is a specialized expert with personality, processes, and proven deliverables.
更多值得一看(内容池 85 条)
With the release of its new trillion-parameter model, Mistral is hoping to demonstrate it’s “still in the race” to build frontier-level artificial intelligence.
Octop is an open-source, self-hosted AI assistant. Through its multi-agent architecture, it builds an intelligent environment that is both independent and collaborative for teams, families, and individuals. Best of all, it runs entirely on your machine, the fully self-hosted design means privacy is never a compromise, while single-process startup makes the powerful web console, CLI, and IM integrations readily accessible. Surfaces: - Web dashboard — chat, experts / teams, connectors, channels, c
At my company, we’re using GLM-5.3 Flash internally for software engineering work, and I’ve been genuinely impressed by it. I work in a very large production environment with projects totaling **millions of lines of code**, and we’re not relying on frontier models for this workflow — GLM-5.3 Flash is doing the actual day-to-day coding work. The model is extremely fast, but what’s more impressive is that the speed doesn’t seem to come at the cost of capability. It handles large repositories surpr
EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders. Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device a
Reinforcement learning (RL) has greatly advanced the capabilities of large language models (LLMs), but its memory demands remain a barrier to broader adoption. We introduce LoGRA, an approach to RL post-training that reduces memory by retaining useful learning signals in low-rank gradient sketches. These compact representations support both model updates and efficient policy synchronization. To prevent overly large updates from disrupting learning, we complement gradient compression with predict
Modern information systems, including many agentic workflows, use dense retrieval to explore large amounts of unstructured data. However, dense retrieval relies on surface-level semantic similarity, which is insufficient for increasingly complex search applications. Here, we investigate agentic retrieval that combines the reasoning capabilities of Large Language Models (LLMs) with the efficient corpus exploration of retrievers in a ReAct agentic loop to solve complex retrieval tasks. In our expe
OpenAI has revealed solutions to a number of long-standing mathematics problems produced by an unreleased frontier model in a batch of 722 manuscripts, covering 372 result families that group related papers. It extends a run of breakthroughs that have both impressed and unsettled parts of the mathematical community while raising questions about research ethics and […]
Open-weights, ideology, and acknowledging trade-offs.
Modern autoregressive (AR) video diffusion models excel at short-horizon video generation, yet generating long videos remains challenging due to drifting, where colors and textures shift, and motion dynamics decay. Existing works primarily rely on KV conditioning, which selects or modifies cached key-value (KV) entries to mitigate drifting. However, we observe that KV conditioning alone is insufficient as it assumes cached KV entries remain in-distribution. This assumption fails beyond the train
Hey All, I spoke to a 0-day partner of Alibaba today and he casually mentioned (didnt know if he was allowed to) that Qwen 4 is apparently planned for the end of October. To me, this is way faster than expected as there was quite a gap between 3.6 and 3.8. I tried to get more information out of him regarding which variants will come first and he got a bit cagey. BUT: No matter the order of the variants, we can hope for Qwen 4 27B this year! EDIT: I know this is very much "in bro we trust" but i
Learn how OpenAI and Ironclad are training and evaluating AI agents on complex contracting workflows to advance computer use for professional work.
LLM-based agents are increasingly capable of generating complex 3D structures, with the potential to reshape how objects are designed and realized in the physical world. Yet, producing elegant geometry is fundamentally different from producing objects that can be built and perform their intended functions. Existing evaluations largely focus on geometric quality while overlooking physical realizability. We introduce LMBuild, a benchmark for evaluating LLM agents on generating buildable and functi
arXiv:2610.04012v1 Announce Type: new Abstract: Language-model systems can separate contextual computation, persistent storage, and exact execution instead of updating all capabilities through one shared parameter system. We investigate FEM-ASM, a finite-element-method-inspired organization in which independently constructed document states and deterministic executable skills contribute typed proposals to a shared language-model state. An explicit residual operator reconciles proposals attached
Reka has released Rho-1, a 19B omni-reasoning model trained from scratch. One network reads and generates text, images, video and robot actions over a shared KV cache. A distilled variant returns a 5.3-second clip in about a second. It is a research preview with no public weights yet. The post Reka Releases Rho-1: A 19B Omni-Reasoning Model That Understands, Generates Video and Outputs Robot Actions in One appeared first on MarkTechPost .
Test-time training (TTT) lets a model store information in its weights during inference. When the model learns from its own output, however, each update also changes the model that generates the next training example. Across 128K-token streams, retaining generated-text updates worsens prediction on independent human-written text with three TTT-E2E model configurations (labeled 125M, 760M, and 3B). The same failure occurs when Adam updates Qwen3-4B's existing weights. The same update mechanisms c
arXiv:2610.03938v1 Announce Type: new Abstract: Agentic multimodal large language models (MLLMs) have recently pushed the frontier of visual reasoning by calling tools such as zooming and tagging. Despite the recent strong success of agentic MLLMs, this work uncovers a critical safety failure in the tool-use paradigm: agentic tool-using MLLMs become less capable of refusing harmful requests. Our experiments confirm that, across three popular safety benchmarks, all the top open- and closed-weight
arXiv:2610.04008v1 Announce Type: new Abstract: Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking correct behavior. Existing benchmarks do not systematically distinguish documentation repair, script repair, and preservation when evaluating skill self-evolution. We introduce SkillScriptBench, a 350-task benchmark designed to evaluate these capabilities separately. From a survey
arXiv:2610.03894v1 Announce Type: new Abstract: A deployed LLM agent emits tool calls, queries, and code that can be silently wrong -- by the time the error surfaces, the action has run. Frontier chat APIs hide the model's token probabilities; the agent's stated confidence barely beats chance on the mistakes that matter; and resampling does not help, since frontier models are highly repetitive, reproducing the same call across samples. We recover the missing signal from a low-cost open-weight su
Search agents repeatedly make short decisions about relevance, evidence sufficiency, and search actions. Using generative language models for these decisions introduces latency and unreliable confidence. We present SearchJev, a fast and calibrated System-1 model that separates search decisions from System-2 reasoning and generation. Given a search state and a decision schema, SearchJev directly scores legal options without autoregressive output generation. We propose Soft-Label Learning for Cali
I am a developer of a popular photo editor that runs in a web browser. Many people are asking AI models to take the Javascript code from my website, remove all ads from it, and they publish such a "new product" on Github for everyone to download. There exist tens of such repositories on Github. I want my website to be the only source of a stable version of my program Photopea. I even received emails from people complaining about something in Photopea, and it took several emails to figure out tha
Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery. We introduce UndoBench, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds alongside wire-level effect-history and environment
Does conversational memory need LLM-extracted facts, or is selecting the right raw turns enough? Published results disagree. Extraction-based systems report gains from distilled facts. Recent studies find raw history with good ranking does as well, but disagree about whether ranking matters. We ran a pre-registered study on held-out LoCoMo conversations and LongMemEval. At a tight budget on LoCoMo, raw turns selected by a single call to Jev, a typed decision model, are non-inferior to an LLM-ext
Hey! 👋 I have released an official support for Strix Halo machines on Strata for Qwen3.8-Flash-Next. Currently numbers are the best on long context decode and ppts using typical Unsloth’s Q4 and GSQ-RCO model weights. Can go up to 1M context length without big speed loss. Currently support is marked as experimental and was done on Linux only. Will be happy for any feedback and pull requests you could give! 👀
Computer-use agents need to reliably ground action targets in complex desktop scenes, where multiple applications, overlapping windows, and visually similar controls compete for attention. Existing training data rarely pair such scenes with dense annotations or vary them in a controlled way. We introduce DeskForge, a controllable desktop environment that composes and explores real applications to generate large-scale supervision for computer-use agents. It varies application states, content, win
Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and gives MoE models new potential for better expert usage, but it raises a question: how to loop a MoE? We answer it with Foil. With the expert parameters and the expert compute per t
I spent the last few weeks on a hobby research project and just made it public. The idea isn't new (product-key memory, Lample et al. 2019, and Meta's "Memory Layers at Scale"): give a model a huge table of learned vectors and let it read only a few hundred of them per token. I wanted to know what that's actually worth on a small model, what it costs, and whether the table even has to sit in VRAM. What came out: - A 21M model with a 16.8M-row table (6.4B parameters in the table, 33M used per tok
Large language models (LLMs) are highly sensitive to the prompts used to specify task objectives and behavioral constraints. Many recent prompt optimization methods iteratively rewrite prompts using LLM-generated feedback, but the resulting prompts often become longer, accumulate narrow sample-specific rules, and generalize poorly beyond the training distribution. We study this failure mode as prompt distributional overfitting and argue that it reflects a lack of representation control in discre
A language model reads long text in one quadratic forward pass, stops at the context window, and loses accuracy with length before reaching it. We ask whether the read can be factorized when deciding over a finite set: which document is relevant, which option is supported, which passage is the evidence. Periscope, a training-free inference method, arranges the N chunks of a text on a K{times}K grid with K{=}lceilNrceil and asks a frozen model the same question about K local spans of consecutive
Im a security engineer and i built mailaccess. When i started learning pentesting, i came across multiple lectures and notes of people listing out tools and websites, which gives the emails for a particular domain, and almost all of them mentioned that the tool might not stick, so its better to learn the methodology, rather than learning a tool- that stuck with me. As i was beginning to really get into pentesting i noticed a clear lack of email osint methodology through the tool itself - so i th
On Tuesday, Musubi announced a lightweight decision model made for real-time moderation called PolicyLM-1.7B, released with open weights.
The attackers sent a push notification to Asos shoppers announcing their activity.
"We created this program because we believe the benefits of AI will reach most people through the companies that build on top of models, rather than through the models alone."
Mirror Particle will launch at TechCrunch Disrupt's Startup Battlefield 200 with a world model built from scratch to predict human behavior, arguing that LLM role-play falls short for market research and brand strategy.
arXiv:2610.04011v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards improves reasoning, while the allocation of learning signal shapes which solutions remain accessible under repeated sampling. Group-relative objectives assign equal advantages to equally rewarded responses, making aggregate credit proportional to sampled mode frequency. We introduce Exploration-Preserving Policy Optimization (ExPPO), a lightweight advantage-shaping rule that redistributes credit using
arXiv:2610.03872v1 Announce Type: new Abstract: AI agents are becoming increasingly capable of generating scientific code, but generating code is not the same as improving the algorithms behind it. For numerical solvers, execution feedback can expose poor performance, but rarely reveals its underlying cause and how to address it. We introduce Auto-Diagnosis and Skill Discovery (ADSD), a framework that links numerical diagnosis to reusable solver self-improvement. ADSD follows a diagnosis-first p
We’re launching a new, expanded version of our Cyber Verification Program (CVP), which makes advanced cyber capabilities and reduced blocking classifiers available to qualifying security professionals. The program now consists of three access tiers, which allow security teams to apply for the level of access that best suits their work. Each tier includes access to our most capable models, including Claude Opus 5.5, Claude Sonnet 5.5, Claude Mythos 5.1, and new models moving forward. Interested c
JEPA-Anything splits a JEPA's single latent target into 4 orthogonal factors, each with its own predictor. Tested across 7 domains, it beat matched JEPA baselines on all 10 dynamics tasks and cut Interventional Pong intervention error by 34.8%. The post Beyond Domain-Specific World Models: JEPA-Anything Uses 1 Recipe for 7 Fields appeared first on MarkTechPost .
We introduce a post-training method for diffusion language models (DLMs) that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM. To estimate MMD, we retain contextual features at individual token positions, obtaining multiple observations per sequence from a single extractor pass. We optimize this objective using policy gradients for discrete models and direct differentiation through generated latents for contin
Hello llamas. I am posting this because I believe that, despite it being closed source models, the discussion will bring value to the local AI community. As many of you probably heard, GPT-6 Astra is speculated to be a looped transformer architecture that outputs a token after multiple forward passes instead of one. This allows a model to essentially have more effective depth due to recurrence, making more use of the weights at the cost of more compute. Recent Azure Foundry "leaks" even suggeste
Robot agents must connect their intended actions to observed outcomes while retaining the context needed to revise their choices over repeated attempts. Existing interfaces often leave these choices inside predefined tools or require agents to manage detailed execution code and its growing history. We introduce RobotUse, a robot agent harness that organizes computation, context, and decisions around specifying and revising physical actions. Agents visually select targets and poses, while the bac
Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use of the same source is still largely treated as repeated access rather than an opportunity to progressively improve understanding of that source. We study source learning: developing reusable source-sp
Latent reasoning lets a large language model (LLM) think in a continuous space and verbalize only the answer. We argue that an effective latent thought must meet five requirements: it should be useful, helping produce the correct answer rather than merely changing it, diverse, so that resampling yields different reasoning trajectories, explainable, so that a decoded chain of thought (CoT) reflects reasoning the answer actually follows, refinable with more inference compute, and efficient, costin
We introduce Loop Flow Transformers (LiFT), a family of looped generative models that scales computation by repeatedly applying a shared Diffusion Transformer (DiT) core, with only light changes to the standard architecture. Rather than asking every recurrent step for the final prediction, LiFT trains each step with a single regression target: a point on a straight path from the model's initial estimate to the flow-matching target. Because we index these targets by a continuous depth coordinate,
We present the Prior-Fitted Language Model (PFLM), a 300M-parameter byte-level transformer pretrained only on samples from a synthetic non-linguistic prior. Given a prefix of real text, it learns to predict the language in context with frozen weights, having never seen a word of any real language. Every training sequence is generated by a recurrent structural causal model drawn fresh from a distribution over such models. The model never sees the same language twice during training, so the only w
“There’s a perception of mobster behavior” from leading AI companies, one mathematician tells WIRED as OpenAI prepares to release more than 100 new solutions to unsolved problems.
The maker of the open source document editor says it has no plans to add AI to its software's default configuration, citing user privacy.
Google announced a new agreement to update six nuclear power plant sites across the US as the tech giant seeks to generate more electricity for its power-hungry data centers. Google signed the 20-year deal with Constellation, the leading nuclear power plant operator in the US. The power purchase agreement is meant to guarantee the revenue […]
Release: datasette-atom 0.11a0 A minor fix for compatibility with the latest Datasette alphas. This meant we could upgrade the datasette.io site to Datasette 1.0a41. Tags: atom , datasette
Atlassian and OpenAI are expanding their partnership to connect frontier models with enterprise knowledge and help teams plan, build, and deliver work.
Wajo's Fo agent can hire humans to complete a task.
arXiv:2610.04019v1 Announce Type: new Abstract: Graph based cyber attack detection studies employ various graph construction and representation strategies across different cybersecurity application domains. This diversity motivates a quantitative examination of how representation strategies are distributed across these application domains. This study presents a quantitative analysis of 37 original studies published between 2019 and 2026. Each study was coded according to publication year, applic
arXiv:2610.03966v1 Announce Type: new Abstract: Each run of an AI-driven research system (ADRS) is an expensive search over a vast solution space, and dependable evaluation requires many runs, making run data both costly to produce and valuable to retain for large-scale analysis. Yet this data remains fragmented: teams operate in isolation, ADRS frameworks emit results in different formats, and no shared infrastructure exists to aggregate or compare runs across problems and systems. We present R
arXiv:2610.03959v1 Announce Type: new Abstract: Recent world action models (WAMs) reuse pretrained video VAEs whose encoder latents directly condition downstream action policies. Quantization must therefore preserve not only reconstruction fidelity but also the policy-facing latent contract expected by the frozen policy. Direct NVFP4 leaves W4A4 quantization error uncompensated, whereas joint quantization-aware training (QAT) can recover reconstruction by moving this representation. On Wan2.1, j
arXiv:2610.03998v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used as synthetic personas representing survey respondents. Their validity as substitutes for particular respondents depends on whether they reproduce individuals' decisions. We examine what information helps synthetic respondents predict each individual's later choices, using five conditions that add progressively richer information: no personal information, demographics, personality traits, cognitive
Geometry optimization is a major cost in many quantum-chemical workflows: each optimization step requires one force evaluation, and at the density-functional level that evaluation dominates the wall time. Research in this area has produced a broad range of optimization methods, and we ask whether a language model can improve on the best of them through autoresearch. An agent rewrites the optimizer itself to minimize force-call counts, restrained by two admission gates that reject premature stopp
Multi-step matching models, including flow and diffusion models, produce high-quality outputs but incur substantial inference costs and may reproduce unwanted components of their training datasets. We introduce Inverse Distillation Unlearning (IDU), a unified framework that simultaneously distills a teacher multi-step matching model into an efficient one-step student generator and suppresses outputs corresponding to a designated training subset. We first formulate distillation as a min-max objec
With all the fuss over AI and what it can do, I feel like we've totally glossed over the fact that AI has casually solved something that has been a problem for decades. I was born the year after the microprocessor was invented, and I've kept a very close eye on technology as it has developed. And the problem of machine translation has been with us for a while. It used to be absolutely terrible. Then it got to the point where you could sort of tell what the native speaker who wrote the original w
Interactive world models are increasingly capable of generating environments and acting within them, yet deliberately editing an existing executable world remains underexplored. We formulate world editing as intervening on an existing world while preserving properties that should remain unchanged, and introduce intervention depth as an axis describing how strongly an edit couples world entities, dynamics, and systems. We instantiate this capability through industry-grade game modding and introdu
Real-world time series are frequently driven by exogenous events and structural shifts, rendering conventional forecasting based solely on historical numerical observations insufficient. While language models can retrieve external news, standard retrieval-augmented approaches struggle with high noise, missing signals, and an inability to reason causally about event impacts. We propose SEER (Self-Evolving Event Reasoning and Retrieval), a closed-loop framework that dynamically optimizes event con
Split learning lets a client train a language model on a server without sending its text. The client runs the first layers itself and sends the server only their output, a vector of numbers for each token. During training, the server sends gradients back. We show that an observer at the split can rebuild most of the client's text from this traffic, and we measure how much the gradients help. On GPT-2, an attacker who holds only the publicly released weights of the client's layers recovers 94.20%
I want to selfhost a password manager. I wanted to go with vaultwarden. But now i read about bitwarden lite, which is the official lite version. How likely/how often did in the past happend that bitwarden released a breaking change to the app/clients so that vaultwarden needed first an update? Did you swap to bitwarden lite after the release? Im using pangolin to tunnel to my local machine. If this matters in anyway or form
i was very excited about code generating ai tools since early days of github copilot in vscode, was using it daily since chatgpt release and wrote almost all code through prompting (html/css/javascript, python, ruby, terraform etc) for years. It felt very good at first. However, about a year ago i started to notice subtle (at first) negative changes in my mental health. It's hard to describe in words this negative feeling because it's very basic and fundamental, but over that last year it progre