If you scroll down from the countdown at , you see a big model card with a bunch of sections: Highlights, Model Overview, Quickstart, Best Practices, Citation, etc! No benchmarks on this yet as far as I can tell. We'll still need to wait another 5.5 hours for those I reckon. Edit: Ladies and gentlemen, the model is live. Let the testing begin!
Interestingly, the 3.8 version has exactly the same architecture - meaning all the capability gains come from training improvements! See the diff (0 changes) here!
With your experiments, Qwen 3.8 27B most close which frontier model? And please specify which quantization you run. I will post to comments my tests and experience too.
I didn't think we'd get here so quickly. I can run this shit on a computer I spent less than $2k for (back before prices exploded). Crazy world Source: Artificial Analysis Intelligence Index v4.1.1
Release: llm-gemini 0.33 It's been a while since the last llm-gemini release. This version of the plugin adds support for today's Gemini 3.7 Flash release, plus gemini-3.6-flash , gemini-3.5-flash-lite and two embedding models gemini-embedding-2 and gemini-embedding-001 . The plugin is also upgraded for compatibility with LLM 0.32, which means you can now see reasoning traces and you can also enable server-side tools using this pattern: llm -m gemini-3.7-flash -T CodeExecution \ 'use python to c
Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefined injection locations, limiting scalable security research across broader domains. To bridge this gap, we propose **ToolHazard**, a scalable adversarial environment synthesis framework that reduces human engineering an
The Administrative Office of the U.S. Courts told TechCrunch that it will start disclosing how many times judges authorized the use of spyware to wiretap suspected criminals.
Over the last couple of weeks, hackers have targeted and broken into the systems of several water plants in the United States. Here’s what we know and don’t know about this wave of attacks allegedly carried out by the Iranian government.
US groups release cheaper models after new challenges to their trillion-dollar ambitions.
We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately before full attention layers, forming pre-attention spikes (PAS), and can persist through intervening linear attention layers, giving rise to inter-spike plateaus (ISP). As full attention becomes denser, successive PAS become increasingly connected through ISP, ultimately recovering the stable MA morphology of full
Recent optimization studies on Retrieval-Augmented Generation (RAG) have exploited chunk-level KV cache reuse to avoid processing long retrieved contexts for higher efficiency, while significant information redundancy and noise still remain in the coarse-grained chunks. This paper optimizes the Pareto frontier under low prefill latency constraints while maximizing accuracy by proposing CoinRAG (Contextualized Information Nugget KV Cache Reuse for Long-Context RAG). The name metaphorically reflec
The "thinking-with-images" paradigm equips multimodal LLMs with active visual operations such as crop-and-zoom. However, models using these operations often achieve only marginal or negative gains over direct inference at substantially higher token cost. They may also repeatedly crop irrelevant regions and fail on questions that direct inference answers correctly. We ask whether the returned visual evidence causally affects the answer. To answer this question, we formulate visual tool-use as a c
Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet existing methods largely rely on fixed, handcrafted heuristics and costly expert trial and error. As pruning objectives, budgets, and model architectures diversify, manually navigating the expanding design space becomes increasingly difficult. This paper aims to build an AI4AI framework for visual-token pruning by addressing a natural question: Can large language models automatically
Here is my self-hosted setup; below is a list of the hardware and apps I use daily. Overall, my setup has been rock solid and runs like a well-oiled machine; it's a very light touch, and I try to keep things simple. The hardware Server, Proxmox VE: Supermicro X14SBM-TP4F, Xeon 6-6521P, 256GB DDR5 Boot: 2x 1TB NVMe mirror VM storage: 4x 8TB WD_BLACK SN850X, RAIDZ1 Media: 10x 18TB WD Red Pro, RAIDZ2 plus a hot spare LSI 9305-16i in IT mode, ICY DOCK 4-bay M.2 cage over MCIO GPUs: Tesla T4 16GB and
My homelab has two disks. One with the OS and files, and the other as a backup. Every night, I run dd to copy everything to the second disk. The idea was simple: if one disk fails, I can just switch to the other one. I've managed to rescue my data several times this way. But not today. If one disk gets corrupted and then dd copies that corrupted data to the other disk, well... now both disks are corrupted. What I lost: All my recipes in Mealie My unchecked-out code in Gitea, including my diary t
As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers eval
Tim O’Reilly built a publishing empire that AI is helping to destroy. Yet he loves AI—as long as it’s open source.
Article URL: Comments URL: Points: 44 # Comments: 18
What is a personal agent?
What's happening in China? Kimi K3-2.8T Qwen3.8-2.4T DeepSeek-V4-Pro-0813-1.6T GLM-5.3-743B They’re all less than a month old!
少花一半成本,部署效率提升80%!机器人研发可以不用重复造轮子了。
Apple has recently sent out notifications, warning users that they'd been targeted by mercenary spyware attack.
Don't get too hyped + take with a grain of salt as there have been an endless amount of quantization schemes with big promises that never really became a thing. Tim Dettmers is a pretty well known researcher though, so maybe something will come of this. Time shall tell. Another tweet about the method, DS4 Pro on a single B300 (288 GB VRAM):
Release: sqlite-utils 4.2.1 Fixes a crashing bug in sqlite-utils 4.2 . I'd introduced code that looks like this: from typing_extensions import Self It turned out the typing-extensions package was not listed as a dependency for sqlite-utils - it was installed by one of the other dependencies in the dev dependency group , but when you uvx sqlite-utils directly you don't get those dependencies. As part of fixing this I figured out how to run a smoke test to ensure the CLI tool still works even with
Composable agent harness where everything is a plugin Discussion | Link
Judge warns pro se litigants are using chatbots wrong and getting desperate.
Aurora Innovation and Kodiak AI, two companies developing self-driving trucks, have received permits from the California Department of Motor Vehicles.
The invisible SynthID will remain.
Natural gas prices could triple in some parts of the U.S., which could saddle hyperscalers with massive bills to power their AI data centers.
Meta released Glimmer this week, an open-weight AI model anyone can download and run on their own hardware — a contrast to Muse Spark, the company’s more powerful model that stays locked behind its own APIs. The release landed alongside a letter from Mark Zuckerberg arguing AI should be “for everyone” rather than controlled by a handful of labs, but as Equity’s […]
Facing any issues? Chat Template is fine? Looping issue? Too much reasoning thing? How's MTP with this one? Any other issues faced by Qwen3.6-27B & Qwen3.5-27B during release time? If I missed any other items, please mention in your comments. AND Share comparison with Qwen3.6-27B. On Memory & t/s stats How much memory takes for this model if you use full 256K context + unquantized KVCache + MTP? For Q4 & above quants. Particularly Q8 please, want to know it's possible to hold this in 32GB VRAM.
Not directly LOCALLlama related but I thought it was interesting since Mistral and Z.ai are competitors, and more surprisingly they are pricing it (GLM-5.2) even cheaper than their current flagship model Mistral Medium 3.5. Does this suggest a pivot in Mistral's strategy? Are they going to abandon frontier model development and instead focus on selling compute while developing smaller specialized models like Shieldstral?
Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth. Dense attention gives each token broad horizontal access to the past, but the vertical feedback channel between decoding steps remains narrow: only the sampled token returns to the bottom of the stack, while the top-layer hidden state is discarded. We introduce the full-bandwidth transformer, which widens this channel with latent feedback: at each decoding step, the previ