I've been running Qwen3-0.6B on the M5Stack LLM-8850 card (Axera AX8850 NPU, 24 TOPS, 8GB LPDDR4x) hosted by a Raspberry Pi 5 — as a llama.cpp backend. The problem: the vendor stack requires converting every model through their compiler, and their closed runtime gets 13.5–14.5 t/s. I wanted llama.cpp to just work: GGUF in, tokens out. What I ended up doing: Reverse-engineered the engine format. The vendor's compiled engines (.axmodel) store weights in a blob called npu_params. I decoded it: int8
Proof is in the method honestly. Closed model labs need to constantly reinvent the wheel to keep lead. Open source has a bunch of independent labs practically working somewhat together. Eventually when everyone is just releasing weights and papers on how they did it the closed source secrets just get overrun by having plenty of very good secret sauces to the public. That and NO DOUBT chinese labs are sharing internal secrets amongst each which explains how when any of them makes a big jump the o
There's a lot of capital pouring into the business of giving models away.
Their last version 7.14 was released just a month ago. llama.cpp PR(waiting for approval) for Version 10.0 Hope this version comes with more boost & improvements.
It's allegedly been running dozens of gigantic generators without a permit.
A federal judge ruled the Trump administration illegally labeled Anthropic a supply-chain risk, handing the AI company a victory as its second Pentagon lawsuit continues in Washington.
It’s Time. We’re in the Endgame now.
Built the latest and I'm getting as much as 220 tokens per second and averaging in the 170s, I can't get over it. If anyone on here is on that project, fuckkkin' chapeau man, really incredible job. I can't believe I was able to like double or more my throughput from llama.cpp This is what I set up: command: > ninfer-serve /models/qwen3_8_27b_nvfp4.ninfer --model-id qwen3.8-27b-nvfp4 --host 0.0.0.0 --max-context 240000 --kv-capacity 240000 --max-concurrency 2 --kv-dtype fp8 --host-kv-mib 16384 --
I've got Qwen3.8-Flash-next running on RTX 3090, Ryzen 9 3950X, a PCIe 3.0 motherboard, and 64GB DDR RAM from 2020. IQ4_XS weights, full kvarn5 context, vision on GPU, experts in host RAM, n-grams on disk. MTP works but actually slows decode down even with 80% draft acceptance, as expected since every rejected token eats into the host RAM bandwidth. I get 160 tok/s prefill 16 tok/s decode , which makes it a decent option whenever I know I'll be AFK for at least a couple of hours, but not usable
Just as new data centers face growing backlash from neighboring communities, the US Environmental Protection Agency (EPA) is about to make it harder for people to weigh in on any pollution those centers create. The EPA plans to toss out a federal rule requiring public notice and an opportunity to comment when certain industrial sites […]
Neocloud Lambda has raised $1B in private debt to buy Nvidia AI chips and lease them to Microsoft. It's the latest in a string of loans, underscoring the high cost of the AI boom.
A recent paper argues that AI is often better at doctoring than doctors. Guess who isn't thrilled.
"At Hot Chips 2026, Micron drew a notable comparison: For the same memory capacity, HBM requires approximately three times the wafer area of DDR5." "When asked whether this ratio would improve with newer generations, the Micron Fellow reportedly explained that it definitely would not get better." "According to the data shown at Hot Chips, an HBM4 die, for example, operates with 256 memory banks, while DDR5 is specified with 32. Additional data paths, the power supply, and the Through-Silicon Via
The art portfolio platform Cara, designed for creators who don’t want their work used to train AI, has been under assault by trolls seizing and publishing its data.
The firm, known for its focus on software, is going to start throwing more money at the hardware behind AI.
Hello everyone! Pangolin 1.22 introduces a new resource type: AI Gateway. These resources are identity-aware proxies in front of both cloud model APIs and self-hosted model servers, so coding agents and AI clients call a Pangolin URL. A gateway resource can be keyless by authenticating via the Pangolin client or keyed by minting virtual API keys. We're also moving SSH, RDP, VNC, and private HTTPS resources from Enterprise to Community Edition. Pangolin is an open-source, identity-aware remote ac
Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task can be performed simply by specifying it in the context, without any parameter update. This form of in-context learning (ICL) turns generalization into a problem of task specification. To achieve cross-task generalization, we bring this paradigm to robotic manipulation, and argue that the natural task
arXiv:2608.26109v1 Announce Type: new Abstract: Machine-learning models can predict ICU mortality accurately, but feature-attribution methods alone rarely provide the clinical narrative needed for bedside use. Large language models (LLMs) may bridge this gap, and multi-step agentic pipelines are a plausible extension because they separate data interpretation, guideline checking, and final explanation. This revised feasibility study preserves the original standalone-versus-agentic comparison whil
As a happy user of ds4, I'm very excited about this branch. Ran some prompts and it seems to be working well on my M4 Max 128gb!
Got a desktop and notebook running NixOS. Now I want to get a beefy NAS and run immich, jellyfin, home assistant, some file shares, vaultwarden and papra on it. My plan is to run NixOS on it because i know it and can restore a PC in a few minutes. The named services do exist on NixOS and that’s how I plan to use them. No docker, no Proxmox. Anyone tried it? Anything I need to know? Objections?
Google is now automatically expanding its AI search summaries at the top of the results page for some searches, as reported by Search Engine Roundtable. The change, when it kicks in, pushes the typical list of links from a search much farther down Google's results page; instead of seeing part of an AI Overview with […]
Anthropic refused to support lethal autonomous warfare and mass surveillance.
Modders are trying out an unofficial version of Nvidia's DLSS 5 on Skyrim, Cyberpunk 2077, GTA V, and a bunch of other games after code for the AI upscaling tech appeared in an early-access build of NBA 2K27. Members of the RenoDX modding channel on Discord reportedly found a way to extract the DLSS "Neural […]
The feature, announced this week, allows Brave's users to sign up for websites and other online services without having to share their personal email addresses.
After an affair with a fellow police officer ended, a Georgia cop used Flock to track her movements—and those of a man whose vehicle often showed up near hers, internal investigation records show.
Meta fixes AI glasses to stop recording any time users cover up the safety light.
arXiv:2608.26113v1 Announce Type: new Abstract: We present PICasso, an AI-assisted framework for automated synthesis, verification, and optimization of photonic integrated circuits (PICs) from natural-language specifications. PICasso couples a structured NL -> YAML -> GDS generation pipeline with PDK aware knowledge injection, automated placement and routing, DRC/LVS validation, and SAX-based photonic simulation. To systematically evaluate AI-driven photonic design, we introduce PIC-Set, a bench
Run prompts on every model at once. Score. Version. Ship. Discussion | Link
11 月 20 -21 日,由奇点智能研究院与 CSDN 联合主办的「奇点智能大会北京站」正式举行
arXiv:2608.26107v1 Announce Type: new Abstract: Predicting students' academic risk in online education is crucial for enabling timely interventions that can improve retention and learning outcomes. However, existing models often suffer from limited early detection capability and insufficient interpretability, leading to a "black-box" trust crisis that hinders their adoption in real-world pedagogical settings. To address these challenges, we propose EduRiskX, a neuro-symbolic framework that integ
You can test it out on breezblue's playground or use it locally, its only ~7GB.