Aleph Alpha releases Kolibri: a 78.1B-parameter English-German MoE model activating only 3.46B, with 1M token context, Apache 2.0 license, and FP8 weights that can run on a single B200/H200.
2026-10-05
— Local large models running on consumer GPUs, open-source agents and KV cache optimization flooding feeds on the same day—today belongs to the engineering crowd.
Another local inference breakthrough: Strata runs a 125B-parameter model at 100T/s on an RTX 4090, while FPGA and small MoE model training also make progress. HeteroFold offers a new scheme for cross-model KV cache migration in multi-agent LLMs, eliminating redundant prefill. Google pauses its open-source vulnerability bounty program due to a surge in AI submissions, as AI slop reports batter the security ecosystem. Rust build optimization Headstart nearly doubles speed, and DeepSeek Harness v0.2 releases an official desktop app.
Schlagzeilen
Strata runs 125B-parameter Qwen3.8 Flash Next at 100T/s on an RTX 4090
The open-source project Strata enables Qwen3.8 Flash Next (125B parameters) to run locally on consumer hardware, achieving a measured generation speed of 100 tokens/s on an RTX 4090, requiring only an NVIDIA or AMD GPU with 12GB+ VRAM, and supporting Windows and Linux. Why it matters: this means large-parameter models no longer depend on servers; developers can run code generation, image understanding, and coding agents on a local PC with data never leaving the machine—a substantive alternative for privacy-sensitive scenarios.
The progress of local large models on consumer hardware is exciting, but some argue that low-bit quantization (IQ3_S) sacrifices output quality, and reliability on long tasks remains questionable.
Google pauses open-source vulnerability bounty program due to surge in AI submissions
Google paused its Open Source Software Vulnerability Reward Program (OSS VRP) on October 1, citing a surge in AI-generated invalid reports and hallucinated content that overwhelmed engineers and maintainers; the company promised an update in Q1 2027. Why it matters: AI slop reports are battering the security research ecosystem, and the pause of a bounty program that serves as a key line of defense for open-source security means real vulnerabilities could be drowned out, requiring security teams to develop new filtering mechanisms.
HeteroFold enables prefill-free KV cache migration across model families
A paper proposes HeteroFold, a prefill-free cross-family KV cache migration method that lets the receiver in a heterogeneous multi-agent LLM system directly reuse the sender's KV cache while both models remain frozen. Why it matters: repeated prefill for text communication in multi-agent systems is a significant overhead; this method eliminates redundancy by aligning model structures and mapping caches, offering direct engineering value for building efficient heterogeneous agent systems.
Headstart nearly doubles Rust build and check speed
The open-source project Headstart makes rustc write out early metadata files immediately after interface checking completes, allowing cargo to start compiling downstream crates earlier; in tests, cargo check and cargo build were up to twice as fast. Why it matters: Rust compilation speed has long been a pain point in developer experience; this optimization works at the dependency graph scheduling level without changing code semantics, significantly improving iteration efficiency for large Rust projects.
The comments generally acknowledge the significant optimization effect, though some note that its implementation approach or similar existing discussions deserve attention.
DeepSeek Harness v0.2 releases official desktop app
DeepSeek released a v0.2 preview of its MIT-licensed open-source agent harness (dsh), adding official desktop apps for macOS (Apple Silicon) and Windows, a built-in plugin manager, file and code change review, scheduled automation tasks, and support for connecting to non-DeepSeek models via OpenAI-compatible endpoints. Why it matters: the agent harness is the key runtime that turns models into executable agents; the official desktop app lowers the barrier to use, while the plugin mechanism and cross-model support let developers build automated workflows more flexibly.
Jeden Morgen ein Tech-Digest, für dich kuratiert
Das Web zeigt das große Ganze; Abonnenten bekommen ihr eigenes — nach deinen Interessen kuratiert, dein privates RSS integriert, mit Community-Stimmen, jeden Morgen zugestellt. Dauerhaft kostenlos.
85 Ausgaben erschienen · täglich 150+ Meldungen auf 30 gesiebt
KI-News
Google Research moves federated learning to TEEs; Gboard's next-word prediction now uses externally verifiable differential privacy.
The top ARC-AGI-3 score on Kaggle jumped from 7% to 56% within 30 days, with small local models in a harness starting to surpass average human performance.
Qwen3.5 architecture 9B/27B INT4 inference implemented on cheap eBay FPGA mining hardware; the 8GB HBM2 SQRL FK33 costs only $280.
Training a 3.87B MoE (1.45B active) model from scratch using only 86.5B tokens, with every layer being MoE, 4096 context, and a Qwen3 tokenizer.
Dev & Open Source
Discussion on why developers don't more actively use native browser platform capabilities; commenters say platform features aren't good enough or flexible, though some attribute it to framework dependency habits.
Commenters generally believe native platform features are not good enough, hard to use, and inflexible, though some think developers are simply used to relying on frameworks and libraries.
A severe IAM vulnerability exposed in Docker Hub could allow impersonating other users via Personal Access Tokens; official remediation advice has been released.
degoog search engine aggregator releases v1.0.0 stable for self-hosted users.
A classic Visual Basic VB6 IDE project native to the browser, replicating the VB6 development environment in a web page.
Community-Themen
A local LLM tinkering history from one 3090 to 20 DGX Sparks, where the home circuit breaker became the first bottleneck; the comments resonate strongly.
Meta's Muse agent system prompt claims users' control over their home unconditionally outranks safety training, sparking heated discussion on r/LocalLLaMA.
The awkwardness of 64GB system memory: running Qwen3.8-27B Q6 plus ComfyUI image inference leaves memory stretched thin.
Micron CEO says memory supply in 2027-2028 will be much tighter than in 2026, with local LLM enthusiasts watching hardware cost trends.
GitHub Trending
Star tester-army / e2e Next generation e2e testing framework for web and mobile apps.
Star pbakaus / impeccable The design language that makes your AI harness better at design.
Sponsor Star coreyhaines31 / marketingskills Marketing skills for Claude Code and AI agents. CRO, copywriting, SEO, analytics, and growth engineering.
Sponsor Star DietrichGebert / ponytail Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.
Star earthtojake / text-to-cad Give your agent CAD superpowers.
Star Panniantong / Agent-Reach Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.
Star getsentry / sentry Developer-first error tracking and performance monitoring
Sponsor Star calesthio / OpenMontage World's first open-source, agentic video production system. 12 production pipelines, 100+ tools, 700+ agent skill and production-knowledge files. Turn your AI coding assistant into a full video production studio.
Star caddyserver / caddy Fast and extensible multi-platform HTTP/1-2-3 web server with automatic HTTPS
Weitere Fundstücke(1 weitere)
How many on the list did you know? Obviously one paper like Attention is All You Need (278k citations) can influence a lot - all the authors are on the list. But still interesting imo. More context: