DigUp 是一款开源 Mac 应用,本地运行 Google DeepMind 的 EmbeddingGemma 2,可跨文本、图像、音频和视频搜索文件内容。
Open Source
Letzte 7 Tage · 100 Einträge
MotherDuck 博客实测 DuckDB 2.0 alpha,重点分析 async I/O 等三项特性带来的速度提升,并给出个人笔记本与 S3 场景下的对比数据。
Nace AI 开源 Drex 1.5,9B 参数决策模型,单次前向传播为每个选项返回概率,Decision Index 0.3.1 得分 58.08,支持最高 128K token 上下文。
社区开源 800M 参数物理基础类型化决策模型 Laya,支持 73K 上下文和图像输入,延续 JEV 架构路线。
评论普遍担忧Bitwarden转向商业许可,认为这是开源项目被风投裹挟的“恶化”信号,但也有人认为只要源码仍开放、自托管可行,影响有限。
Python 3.15.0 已加入 actions/python-versions,现在可在 GitHub Actions 测试矩阵中直接使用 "3.15" 运行测试。
Bitwarden is introducing a dual-licensing model. Starting with the next release, official app store builds will use a commercial license, while GPLv3 versions remain available on GitHub. According to Bitwarden: Existing features remain available in both versions. Self-hosting and forking remain supported. Future features may be exclusive to commercially licensed builds. The free plan remains unchanged. The main concern for me (and probably others in this community) is the potential divergence be
An 11MB local model for typed decisions in one pass Discussion | Link
I remember hearing a while ago that Dwarf Fortress didn't use version control. That fact has been lodged in my head, because given its inherent complexity I have trouble imagining a codebase that would benefit from version control more . That's art. I went looking and I'm sad to report that the days without version control appear to have ended. Quoting Tarn Adams over time: March 2013: "I don't use version control -- I didn't like the feeling of having the code get committed into a black box thi
Anthropic 推出免费 OSS Scanner,用最强模型为开源项目提供周期性安全扫描,但报告无人工审核。
Underdog Saluki 27B is a 7.89 GB, 2-bit GGUF of Qwen3.8-27B under Apache 2.0. It beats the 54 GB original on tool calling but gives up ground on competition math and reasoning. The post Meet the Underdog Saluki 27B: A 2-bit Qwen3.8-27B That Beats the Original at Tool Calling appeared first on MarkTechPost .
Apogee 是 Mozilla Orbit 的本地隐私替代品,基于 WebGPU/WebAssembly 在浏览器内运行 AI 摘要。
It seems Dario's "too powerful for you users" strategy is paying off: we have two open models at the top of the leaderboard, surpassing every single model from Anthropic. Open source prevails. Even Mistral Large 4 is better!
Blog Post : ML Drift: Next-Gen GPU AI/ML Inference at the Edge - Google Developers Blog The Google AI Edge Team is excited to announce the open-source release of ML Drift , our high-performance, cross-platform, on-device GPU compute engine specifically built for on-device AI/ML inference, under the Apache 2.0 license. By abstracting hardware and low-level API complexities of on-device GPUs across OpenGL ES, OpenCL, Metal, and WebGPU, ML Drift empowers developers to build real-time, interactive M
Demo: Link: So after 4 or 5 months of pretty heavy development - I'm barely going back to Gmail, so I think it's time to get some users and with that some feature requests and bugreports 😄 AI First things first: I am Symfony/Webdev for over 10 years. I've started this project with a selfcoded base but fairly quickly leaned more and more into claude code. If you dont feel comfortable with that, I understand, but really cant help you. Motivation Gmail crippled "gmailify" a while ago - external ema
Article URL: Comments URL: Points: 44 # Comments: 8
Open-source desktop AI agent for any model you choose Discussion | Link
Qwen-Image-2.1-Turbo, create and edit images in just 8 denoising steps! Open weights now available! Built on Qwen-Image-2.1, Turbo is an accelerated checkpoint on the same 7B visual generation architecture. Fewer steps does not mean lower quality: it still generates strong 2K images from text, and supports continued creation through natural-language edits, from adding accessories to changing a scene. Start directly with Diffusers: load QwenImage21Pipeline and the checkpoint’s recommended 8-step
Youtu-Parsing-Omni is a compact (5B) omni-modal parsing model. Given a single input — a document page, a natural image, a chart / flowchart, a geometry figure, an audio clip or an audio-visual video — it produces one structured JSON envelope that covers both perception (layout elements, text, tables, formulas, bounding boxes, timestamps, ASR, OCR, acoustic events, camera motion) and cognition (captions, narratives, reports). The output family is selected by the task prompt ( --task in the exampl
Quake 被移植到安全 Rust 并在浏览器中可玩,评论区惊叹流畅度但也有人质疑是 LLM 辅助的 slop 移植。
评论普遍惊叹Quake能在浏览器流畅运行,但也有人认为这是LLM辅助的“slop”移植,可能很快被弃坑。
Debian Linux has put a call out for artist submissions for its next release: Forky.
Alibaba’s Qwen team has released Qwen-Image-2.1-Turbo, an accelerated checkpoint of its open-weight Qwen-Image-2.1 model. It generates and edits images in 8 denoising steps instead of the base model’s 40-step default. For developers, that means 5x fewer denoising steps on the same 7B architecture, plus a hosted API option. TL;DR What is Qwen-Image-2.1-Turbo? Qwen-Image-2.1-Turbo is an […] The post Alibaba Qwen Releases Qwen-Image-2.1-Turbo, an 8-Step 7B Image Model appeared first on MarkTechPost
Article URL: Comments URL: Points: 39 # Comments: 15
Anthropic's offering to help open-source projects track down security vulnerabilities with a new service called OSS Scanner. It says open-source projects that opt-in will get "thorough, periodic security scans by our strongest models at no cost." That could mean open-source projects get alerted about possible security issues sooner, but the trade-off is that OSS Scanner's […]
K10s 用 Go 与 Bubble Tea 打造可点击的 Kubernetes TUI,内置 AI 感知集群上下文。
Article URL: Comments URL: Points: 42 # Comments: 48
Perplexity's pplx-embed-v2-late comes in 2 sizes: a 0.6B model built to run on edge devices, and a 9B model for building high-quality indexes. Its best score is 92.4% on MADQA, and its weakest is 61.2% on ViDoRe v3 Markdown. Both are MIT-licensed and ready to self-host. The post Perplexity AI Releases pplx-embed-v2-late: A 0.6B Edge Model and a 9B Model Scoring 92.4% on MADQA appeared first on MarkTechPost .
anyone has feedback about this one?
Hi all, a bunch of performance improvements have been landed in audio.cpp. The biggest highlight is Higgs Audio TTS, which now runs with around 6 GB VRAM , a 48% reduction in peak memory usage compared to the previous implementation. Thanks to We also made some models significantly faster, especially HTDemucs on GPU and PocketTTS on CPU. No compromises in parity and correctness. Here's a summary of the improvements: Model Peak memory reduction Speedup Higgs Audio TTS 48% VRAM 1.01–1.09× CUDA ACE
Hey everyone, Jovan from UkisAI (Swift Qwen) here! For those who don't know us, UkisAI is a small lab making tiny frontier LLMs, tools and datasets (+doing it open-source!). I'm one of the guys running it aka I train the models and post on Reddit. Our first open-source release is Swift, a series of reasoning-efficient LLMs. It is proof of how penalizing pathological overthinking patterns inside of various LLMs can bring their token usage down -58.3% and speed x1.95 without losing accuracy if RL-
Liquid AI has released Open d1, two open-weight multimodal models in its d1 decision model family. d1-3B reads text and images. d1-omni-600M reads text with an image, or text with audio. Neither model writes text. Each returns calibrated, typed answers in one forward pass with zero output tokens. The target is real-time decisions on the […] The post Liquid AI Releases Open-Weight d1-3B and d1-omni-600M: Multimodal Decision Models With Zero Output Tokens appeared first on MarkTechPost .
Mistral 发布 1 万亿参数开源模型 Large 4(Le Chonk),主打编码与网络防御,预览版已可用。
Potentially big speedup for MoE models that don’t fully fit in VRAM. Are you GPU Poor? Show your speedups ;)
Meta 开源 Rebalancer:C++ 赋值求解器,每天处理约 4000 万次分片/服务器/流量放置问题。
Hi HN, we're Thomas and Olivier from Terse ( ) We've built Durable Actors, an open-source alternative to Cloudflare's Durable Objects. A Durable Object/Actor is a tiny server that handles one request at a time and has its own SQLite database. There's exactly one of each in the world and it is addressed by name. This is the perfect primitive for deploying multiplayer agents. Each agent can have its own Durable Actor, and each user can connect to that Actor via websocket. This is fully horizontall
now you can use GLM 5 Flash MTP locally
What size do you want?
Artcraft 用 AI 逆向工程发布 7 款开源 Adobe 替代应用,复刻 Photoshop、Illustrator 等界面与功能。
Article URL: Comments URL: Points: 54 # Comments: 20
Article URL: Comments URL: Points: 50 # Comments: 19
Explore a comprehensive coding guide to Laya, the open-source zero-shot decision engine. Learn how to implement typed decisions, fit custom temperatures, and build reliable abstention gates using real-world CLINC150 banking data. The post A Developer’s Guide to Laya: Zero-Shot Decisions and Calibration appeared first on MarkTechPost .
Hey Pocket-ID Dev-Team, Hey stonith404 , I just wanted to say thank you for your work. After a couple of months of intensive use, I have to say: Pocket ID is the service that makes using my self-hosted apps so much easier and more comfortable. I came across Pocket ID while searching for an encrypted file transfer service, and I've been following its development ever since. In the beginning, I didn't have much trust that a young developer could build secure and reliable software. Back then, I did
I’ve been pushing TinyStories-style models downward in size, and this is the smallest one so far: MacroStories — 19,969 parameters, 81 KB FP32 For scale: → ~50× smaller than the 1M TinyStories model → ~3,000× smaller than AlexNet → 32-dim hidden state → 378-token vocabulary → one decoder block, recurrently applied 4 times with shared weights It’s obviously not a general-purpose LM, but within its constrained story distribution it can maintain a 100–300 word narrative with a goal, problem, releva
EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders. Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device a
My comment on EmbeddingGemma 2 — Hacker News. I really appreciate that EmbeddingGemma 2 is under the Apache 2.0 license. For embedding models in particular, I don't think it makes sense to use a closed, proprietary, hosted-only model. Most applications of embedding models involve calculating thousands or even millions of embedding vectors and storing them for later comparison. If your model is proprietary, the vendor is likely someday going to decide to stop offering that model. They'll have a b
Octop is an open-source, self-hosted AI assistant. Through its multi-agent architecture, it builds an intelligent environment that is both independent and collaborative for teams, families, and individuals. Best of all, it runs entirely on your machine, the fully self-hosted design means privacy is never a compromise, while single-process startup makes the powerful web console, CLI, and IM integrations readily accessible. Surfaces: - Web dashboard — chat, experts / teams, connectors, channels, c
EmbeddingGemma 2 详细分析:740M 参数、768 维统一空间、Apache 2.0,面向端侧搜索与隐私优先 RAG。
A small win for US open source
EmbeddingGemma 2 is an open multimodal embedding model built by Google DeepMind which maps text (incl. code), images, video, and audio inputs—and combinations thereof—into a single, unified 768-dimensional vector space. The model has 740M total parameters, combining a 270M parameter text model with modular vision (170M) and audio (300M) encoders. Designed to run on consumer hardware such as mobile devices and laptops, EmbeddingGemma 2 delivers low-latency semantic representations for on-device a
OpenTPU:由 AI 辅助设计的开源 AI 加速器,SystemVerilog 硬件、指令集、模拟器、编译器与 PCIe 主机软件全在一个 monorepo。
评论普遍认可AI辅助硬件设计的潜力,但也有人认为性能尚不明确且人类主导作用被夸大。
Parseable:Rust 编写的开源可观测性数据湖,单二进制 ~180MB,宣称每分钟处理 1 亿条时序数据。
OpenChart:开源 TradingView 替代品,让 Claude/Codex 通过图表而非 CLI 与市场交互,支持自托管。
Simon Willison 用 Codex 打通 Datasette 的 OpenTelemetry traces 与 Parseable,附可复现的配置模式。
I spent the last few weeks on a hobby research project and just made it public. The idea isn't new (product-key memory, Lample et al. 2019, and Meta's "Memory Layers at Scale"): give a model a huge table of learned vectors and let it read only a few hundred of them per token. I wanted to know what that's actually worth on a small model, what it costs, and whether the table even has to sit in VRAM. What came out: - A 21M model with a 16.8M-row table (6.4B parameters in the table, 33M used per tok
I am a developer of a popular photo editor that runs in a web browser. Many people are asking AI models to take the Javascript code from my website, remove all ads from it, and they publish such a "new product" on Github for everyone to download. There exist tens of such repositories on Github. I want my website to be the only source of a stable version of my program Photopea. I even received emails from people complaining about something in Photopea, and it took several emails to figure out tha
Im a security engineer and i built mailaccess. When i started learning pentesting, i came across multiple lectures and notes of people listing out tools and websites, which gives the emails for a particular domain, and almost all of them mentioned that the tool might not stick, so its better to learn the methodology, rather than learning a tool- that stuck with me. As i was beginning to really get into pentesting i noticed a clear lack of email osint methodology through the tool itself - so i th
On Tuesday, Musubi announced a lightweight decision model made for real-time moderation called PolicyLM-1.7B, released with open weights.
The maker of the open source document editor says it has no plans to add AI to its software's default configuration, citing user privacy.
Release: datasette-atom 0.11a0 A minor fix for compatibility with the latest Datasette alphas. This meant we could upgrade the datasette.io site to Datasette 1.0a41. Tags: atom , datasette
I want to selfhost a password manager. I wanted to go with vaultwarden. But now i read about bitwarden lite, which is the official lite version. How likely/how often did in the past happend that bitwarden released a breaking change to the app/clients so that vaultwarden needed first an update? Did you swap to bitwarden lite after the release? Im using pangolin to tunnel to my local machine. If this matters in anyway or form
Reflection AI has introduced Beam, its first open-weight model. It is a 501B sparse Mixture-of-Experts model with 23B active parameters, built for coding and agentic work. Reflection says it matches GLM-5.2 on reasoning with 3 to 4x less inference compute. Apache 2.0 weights are due later in October 2026. The post Reflection AI Introduces Beam: A 501B Open-Weight MoE Model With 23B Active Parameters for Coding and Agentic Workloads appeared first on MarkTechPost .
Looks like new open model coming soon and will be "strong" hopefully something under 200b for us memory poor. Also seeing statements about more western open models coming. Hope we get some good competition again on the open front! Here is original artical but its not free to access. Maybe someone has it already here. Oct starting strong!
Reflection is aiming Beam and future models at enterprises and sovereign nations. The pitch is to build “AI factories,” a product that would let institutions build their own customized, local AI system by training Reflection’s AI models on their own proprietary data.
Hey all. We've spent the last weeks getting Qwen3.8-Flash-Next (125B MoE, 6B active) to run properly on one AMD Strix Halo box (Ryzen AI Max+ 395, 128 GB). Tonight we're releasing both the 95 GB EXL3 weights and a new version of Kyojin, our inference engine (built on ExLlamaV3, open). This is a first version, same as our GLM-5.3-Flash and MiMo-V2.6-Flash builds. We'd rather ship it and improve it in the open: speed and quality updates are coming for all three. Numbers, all from a fresh clone and
Minigraf 是一个用 Rust 编写的嵌入式双时态图数据库,支持 Datalog 查询与时间旅行。
Hi r/LocalLLaMA . I'm on the team at Blockway, a small team in Hong Kong (disclosure: this is our model). Today we released Agens Volundr 32B Preview, the first model built on our own hybrid architecture. We trained it on limited compute, it isn't perfect, and we'd rather tell you where it falls short up front. WHY WE BUILT IT Our customers run models on their own machines. At long context, the KV cache, not the weights, decides what fits. So we designed a model where most layers don't keep one.
Hey all, we designed Cactus Whistle, an ASR model for ultra-small devices. It's not perfect, but mostly beats Whisper base with 9x less file size and 6x speed. Whistle supports English, German, French, Spanish, Italian, Dutch and Polish. Remember, the goal at Cactus Compute isn't to achieve SOTA with scale, but to compress intelligence and bring them to smaller under-looked devices like budget phones, wearables, smart home and microcontrollers. Whistle is 55m params (36m active) and CQ2bit quant
Alibaba's Qwen went from an invite-only chatbot in April 2023 to a 2.4-trillion-parameter open-weight model in August 2026. This is the full story, release by release: every major model, its key feature, and how its license changed. Each claim links to its source. The post The Story of Qwen: Alibaba’s AI Models From 7B to 2.4T appeared first on MarkTechPost .
多数人认可 Rust 重写带来性能与依赖精简,但也有人认为 C 版更易早期构建、Rust 并非万能,且质疑是转译而非重写。
TinyDecide is 10M Jev-like mode with 10M parameters and fits in just ~6MB. Smaller than every model on the Decision Index leaderboard and it punches way above its size . It runs almost anywhere: in the browser, Node.js, Python, Rust, and even on an ESP32.
Aleph Alpha 发布 Kolibri:78.1B 参数英德 MoE 模型,仅激活 3.46B,1M token 上下文,Apache 2.0 许可,FP8 权重可在单张 B200/H200 上运行。
评论区普遍认可该优化效果显著,但也有人认为其实现方式或已有讨论值得关注。
Index-Translate is a family of multilingual translation models built on Qwen3.5. The text models cover 150 languages and follow translation instructions such as terminology, formatting, and content-preservation requirements. The family extends this foundation to speech, syllable-controlled translation, and full-document translation. -Index-Translate translates text, structured content, and community expressions. -Index-Echo produces translated subtitles or speech conditioned on the source speake
Been playing around with local TTS with Breeze combined with STT, and the results are amazing. Using Opus 5.5, I can hear the first sound after 500ms if there is no thinking involved, and with thinking on low mode, can be 1-1.5s. I'm using a BLE remote (the kind that are used for taking pics with phones) combined with a wireless microphone. So I can just sit on the couch, and just talk to her. She watches for any claude session that finishes, and sends me the results in a very short, spoken styl
degoog 搜索引擎聚合器发布 v1.0.0 稳定版,面向自托管用户。
浏览器原生的经典 Visual Basic VB6 IDE 项目,在网页中复刻 VB6 开发环境。
Jeden Morgen ein Tech-Digest, für dich kuratiert
Das Web zeigt das große Ganze; Abonnenten bekommen ihr eigenes — nach deinen Interessen kuratiert, dein privates RSS integriert, mit Community-Stimmen, jeden Morgen zugestellt. Dauerhaft kostenlos.
91 Ausgaben erschienen · täglich 150+ Meldungen auf 30 gesiebt