DawnSift
Subscribe

Cloud & Infra

Last 7 days · 36 items

2026-10-09 Fri

Architect Financial Technologies has launched Liquid Inference, an LLM router that runs a live auction for every request. Liquid Inference is an LLM inference marketplace from Architect where providers bid to serve each prompt. The buyer pays the lowest offer that meets its rules. For developers, it is quite simple message: swap a base URL, […] The post Architect Launches Liquid Inference, a Real-Time Auction for LLM Inference appeared first on MarkTechPost .

2026-10-08 Thu

Hi HN, we're Thomas and Olivier from Terse ( ) We've built Durable Actors, an open-source alternative to Cloudflare's Durable Objects. A Durable Object/Actor is a tiny server that handles one request at a time and has its own SQLite database. There's exactly one of each in the world and it is addressed by name. This is the perfect primitive for deploying multiplayer agents. Each agent can have its own Durable Actor, and each user can connect to that Actor via websocket. This is fully horizontall

Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughput benefit. We present SlimWise, a serving framework that tailors the expert pool to each inference phase. SlimWise performs prefill with the full model and decode with a pruned mo

I set up borg via borgmatic like a year ago. 3-2-1 strategy. Confirmed it was backing up and did a quick extract test. That was it. That was a year ago. Well I set a vm in Proxmox and because I was still learning I didn’t set up some directories correctly. Later I installed Immich but apparently installed it under a directory owned by Nextcloud. I never updated Nextcloud because it was local and then I decided to make it available remotely via a reverse proxy and all that fun stuff. So I wanted

2026-10-07 Wed

Google announced a new agreement to update six nuclear power plant sites across the US as the tech giant seeks to generate more electricity for its power-hungry data centers. Google signed the 20-year deal with Constellation, the leading nuclear power plant operator in the US. The power purchase agreement is meant to guarantee the revenue […]

2026-10-06 Tue

I found this question getting asked all over at least since 5 years ago. There are some funny reasons given for it such as that it does not have "official offering and needs self-hosting" (yeah;)) and that groovy is complicated all the way to simply there's no compelling reason to run a pipeline like that when you can offload your worries to GitHub, GitLab, etc. (yikes) So I wonder - is Jenkins dead to you? Since when? And what did you replace it with? And if not, why not, what's missing in all

2026-10-05 Mon
2026-10-04 Sun

Here's a product feature which the world is going to need a whole lot more of over the coming months and years: default hard budget caps . I'm talking about the feature of pay-by-usage services and APIs that lets you say "after $X/month, cut this thing off and return errors". These need to be hard limits. Soft caps, "after $X/month, send me a warning email", will not cut it. Coding agents, and personal agents (coding agents wrapped in a less threatening UI), greatly reduce the friction of spinni

FTL 是一个面向云的新型操作系统,用户态 OS 以库形式构建,兼容 Linux 二进制,隔离性优于传统 monolithic kernel。

评论普遍认为FTL思路有趣、作者可信,但也有人认为其定位更像gVisor或微内核,且缺乏与Firecracker等对比。

IQ3_XXS weights are just under 80GB and my slowww DDR4+7900XTX is stabilizing around 45-70/s (sometimes higher while coding depending on mtp). Looking online I'm seeing similar results for users with 12GB and 16GB cards, and significantly faster numbers for owners of DDR5. (In comparison, Llama CPP with tuning was maxing out around 22.5t/s on the same rig. Quality seems reliably superior (I wouldn't recommend the Q2 weights though)) Seriously. Ask to set it up for your specs. If 27B doesnt fit w

2026-10-03 Sat

Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned through local fitting: a router is optimized for a particular query workload and candidate pool, and often requires additional supervision or retraining as the routing environment changes. We ask whether LLM routing can instead be approached from a foundation-model perspect

Every morning, a tech digest curated for you

The web shows the big picture; subscribers get their own — AI curated to your interests, your private RSS folded in, with community takes, delivered each morning. Free forever.

89 issues shipped · 150+ items sifted to 30 worth reading, every day

Every morning, a tech digest curated for you