Architect Financial Technologies has launched Liquid Inference, an LLM router that runs a live auction for every request. Liquid Inference is an LLM inference marketplace from Architect where providers bid to serve each prompt. The buyer pays the lowest offer that meets its rules. For developers, it is quite simple message: swap a base URL, […] The post Architect Launches Liquid Inference, a Real-Time Auction for LLM Inference appeared first on MarkTechPost .
Cloud & Infra
Last 7 days · 36 items
Free SSL/TLS certificate lifetimes reduce to 64 days in February.
Kubernetes co-creators Craig McLuckie and Joe Beda aim to bring agent harnesses fully into the cloud.
Meta 开源 Rebalancer:C++ 赋值求解器,每天处理约 4000 万次分片/服务器/流量放置问题。
Hi HN, we're Thomas and Olivier from Terse ( ) We've built Durable Actors, an open-source alternative to Cloudflare's Durable Objects. A Durable Object/Actor is a tiny server that handles one request at a time and has its own SQLite database. There's exactly one of each in the world and it is addressed by name. This is the perfect primitive for deploying multiplayer agents. Each agent can have its own Durable Actor, and each user can connect to that Actor via websocket. This is fully horizontall
NeMo-DCR 实现万亿参数 agentic RL 的 bit-exact 增量压缩权重同步,将 1T checkpoint 跨区域传输从 87.5 分钟大幅压缩。
Mixture-of-experts (MoE) models activate few experts per token, yet batched decoding can access nearly the entire expert pool, making expert-weight traffic a major bottleneck. Expert pruning reduces this traffic, but conventional approaches also prune compute-bound prefill, sacrificing model quality for little throughput benefit. We present SlimWise, a serving framework that tailors the expert pool to each inference phase. SlimWise performs prefill with the full model and decode with a pruned mo
I set up borg via borgmatic like a year ago. 3-2-1 strategy. Confirmed it was backing up and did a quick extract test. That was it. That was a year ago. Well I set a vm in Proxmox and because I was still learning I didn’t set up some directories correctly. Later I installed Immich but apparently installed it under a directory owned by Nextcloud. I never updated Nextcloud because it was local and then I decided to make it available remotely via a reverse proxy and all that fun stuff. So I wanted
攻击者劫持 .gh、.sl、.as 三个 ccTLD,伪造 Google 等服务的 TLS 证书;Google 已更新 Chrome 阻断相关证书。
Parseable:Rust 编写的开源可观测性数据湖,单二进制 ~180MB,宣称每分钟处理 1 亿条时序数据。
网友以约 $35 买到 9 张 P106 6GB 矿卡,合计 54GB VRAM 全部可用,引发本地 LLM 社区对廉价推理硬件的讨论。
Google announced a new agreement to update six nuclear power plant sites across the US as the tech giant seeks to generate more electricity for its power-hungry data centers. Google signed the 20-year deal with Constellation, the leading nuclear power plant operator in the US. The power purchase agreement is meant to guarantee the revenue […]
Discover how to construct an end-to-end streaming robotics learning pipeline using the NVIDIA Cosmos3-DROID dataset without local downloads, leveraging byte-range Parquet reads, behavior cloning, and temporal ensembling. The post Building a Streaming Robotics Learning Pipeline Using NVIDIA Cosmos3-DROID appeared first on MarkTechPost .
Russian attacks on Internet, phone services threaten Ukraine’s wartime economy.
I found this question getting asked all over at least since 5 years ago. There are some funny reasons given for it such as that it does not have "official offering and needs self-hosting" (yeah;)) and that groovy is complicated all the way to simply there's no compelling reason to run a pipeline like that when you can offload your worries to GitHub, GitLab, etc. (yikes) So I wonder - is Jenkins dead to you? Since when? And what did you replace it with? And if not, why not, what's missing in all
从 1 张 3090 到 20 台 DGX Spark 的本地 LLM 折腾史,家庭电闸先成了瓶颈,评论区共鸣强烈。
Docker Hub 曝出严重 IAM 漏洞,可能通过 Personal Access Token 冒充其他用户,官方已发布修复建议。
多数人认为数据中心用水相对农业等并不夸张,但也有人认为其耗电耗水仍值得警惕。
Under the One Big Beautiful Bill Act, data center projects in rural areas could be eligible for major tax benefits starting next year. Some hyperscalers do not seem eager to take the free cash.
Here's a product feature which the world is going to need a whole lot more of over the coming months and years: default hard budget caps . I'm talking about the feature of pay-by-usage services and APIs that lets you say "after $X/month, cut this thing off and return errors". These need to be hard limits. Soft caps, "after $X/month, send me a warning email", will not cut it. Coding agents, and personal agents (coding agents wrapped in a less threatening UI), greatly reduce the friction of spinni
FTL 是一个面向云的新型操作系统,用户态 OS 以库形式构建,兼容 Linux 二进制,隔离性优于传统 monolithic kernel。
评论普遍认为FTL思路有趣、作者可信,但也有人认为其定位更像gVisor或微内核,且缺乏与Firecracker等对比。
IQ3_XXS weights are just under 80GB and my slowww DDR4+7900XTX is stabilizing around 45-70/s (sometimes higher while coding depending on mtp). Looking online I'm seeing similar results for users with 12GB and 16GB cards, and significantly faster numbers for owners of DDR5. (In comparison, Llama CPP with tuning was maxing out around 22.5t/s on the same rig. Quality seems reliably superior (I wouldn't recommend the Q2 weights though)) Seriously. Ask to set it up for your specs. If 27B doesnt fit w
评论普遍质疑Cloudflare借隐私之名集中流量,但也有人认为OHTTP分离身份与请求的设计确有价值。
The CEO of Amazon Web Services tried to push back against widespread suspicion of data centers.
Amazon praised for ending NDAs but slammed for downplaying data center pollution.
让每一次请求选对模型,让每一次反馈都成为下一次更优、更省的选择
Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned through local fitting: a router is optimized for a particular query workload and candidate pool, and often requires additional supervision or retraining as the routing environment changes. We ask whether LLM routing can instead be approached from a foundation-model perspect
Every morning, a tech digest curated for you
The web shows the big picture; subscribers get their own — AI curated to your interests, your private RSS folded in, with community takes, delivered each morning. Free forever.
89 issues shipped · 150+ items sifted to 30 worth reading, every day