Drama swirls around OpenAI’s legendary mathematical milestone
OpenAI says it found a solution to a major math problem that has remained unsolved for around 90 years, as reported earlier by The New York Times and Wired. In a blog post on Tuesday, OpenAI announced that it discovered a solution to the Navier-Stokes problem - which relates to the flow of liquid and […]
Kimi K3 (2.8T) at 1 token/s on a MacBook Pro, streamed from four SSDs
Latest open artifacts (#24): Motif-3, GLM-5.3, Hy4-preview and open model licenses
The open model ecosystem continues to expand in its breadth
OpenAl Says It Has Cracked One of Math's “Millennium Problems” (Navier-Stokes) [N]
As reported by the New York Times: OpenAI’s announcement:
Beneath the Surface of Chains-of-Thought: A Mechanistic Interpretation of Reasoning Operations in LLMs
Reasoning in large language models unfolds through diverse functional operations, such as problem formulation, goal decomposition, and deduction. Although these operations are explicitly distinguished in text, little is known about how they are geometrically organized in representation spaces. To this end, we investigate whether distinct reasoning operations exhibit corresponding geometric structure in hidden representations. We find that operations are separable in held-out representations, wit
llm 0.35
Release: llm 0.35 New OpenAI model: gpt-6-astra for GPT-6 Astra . Tags: openai , llm , gpt-6-astra
OpenBMB Releases MiniCPM5-2B: A 2.52B Dense Model Averaging 53.9 Across 34 Benchmarks and Built to Run On Device
OpenBMB has released MiniCPM5-2B, a dense causal language model with 2,516,756,480 parameters and a native 131,072 token context. It averages 53.9 across the 34 benchmarks in its model card, ahead of Qwen3.5-4B at 51.1, with its clearest leads in tool use, coding agents and long-context retrieval. Post-training pairs 400B tokens of deep-thinking SFT with RL teachers and on-policy distillation that merges 16 expert models into one checkpoint. The weights ship under Apache 2.0 alongside the pre-tr
Ling-3.0-flash-VL inherits the language, reasoning, and long-context capabilities of Ling-3.0-flash, while extending them with native image and video understanding. The model has 124B total parameters, with only 5.5B parameters activated per token, and supports a context window of up to 1M tokens. The architecture of Ling-3.0-flash-VL is designed to integrate visual information into real-world reasoning and agentic workflows. A ViT visual encoder extracts features from images and videos, while a
Google DeepMind Releases AlphaGenome Atlas With Precomputed Molecular Effect Predictions and AVI Scores for 9 Billion Human DNA Variants
DeepMind's AlphaGenome Atlas maps every single-letter change in the human genome with 1 impact score per variant. The post Google DeepMind Releases AlphaGenome Atlas With Precomputed Molecular Effect Predictions and AVI Scores for 9 Billion Human DNA Variants appeared first on MarkTechPost .
Chrome is now shipping updates every 2 weeks as AI changes the security landscape
Google is speeding up Chrome’s release schedule to ship security patches and new features faster.
Google Cloud races to catch up in the AI deployment wars with Accenture deal
Google Cloud expands its enterprise AI push with Accenture, betting on forward-deployed engineers to drive adoption and overcome deployment bottlenecks.
The Work Now Within Reach
Explore how more capable, affordable AI can expand the work people and businesses can accomplish—and make growth more economical.
I tested 10 model/harness combinations on the same Three.js task
Show HN: Wg-admin – web UI for an existing WireGuard host
Reads /etc/wireguard, lets you add/edit/remove peers, applies with wg syncconf (no interface bounce). Does not install WireGuard or rewrite your PostUp/NAT. That's all Comments URL: Points: 31 # Comments: 11
Tao: Open math problems being non-renewably mined by AI
Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference
Layer dropout (a.k.a. stochastic depth) has been shown to enable faster training, higher accuracy, and robustness to zero-shot layer pruning in both language and vision transformers. However, as models and datasets have scaled, dropout - particularly layer dropout - has largely disappeared from large language models (LLMs) pre-training recipes. While some prior work has reported that dropout can degrade accuracy, no comprehensive study has quantified, let alone mitigated, this effect. In this st
Spatial canvas IDE for AI coding agents Discussion | Link
Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system keeps only a small fixed slice of that pool. Which frames survive that slice is usually treated as a preprocessing detail; we test whether it should be. Published selectors make the comparison hard because they change the frame scorer, the prompt boundary, the resolution policy, and the answering model all at once. We hold each fixed and vary one decision at a time: selection, spa
The VMs Powering Mobile Agents (Instinct, Claude Code)
1Password increases engineering productivity 21% with Codex
Engineers at 1Password use Codex to rapidly build new features and internal tools, reaching production-readiness while maintaining rigorous security policies.
DeepSeek Flash 4.1 is already being tested via API and rolling out.
Translation: "Internal beta testing for an intermediate version of DeepSeek V4.1 Flash is now open; you are welcome to try it out. It adopts a new model architecture featuring native multimodal support, stronger capabilities, faster speeds, and lower costs. Keep your base_url unchanged and set the model name to deepseek-v4.1-flash-expires-on-0910 to call the API. Current pricing is identical to deepseek-v4-flash, with a rate limit of 20 concurrent requests per account." From Chubby on 𝕏:
My business partner sent a 5K vibe-coded PR that he didn't even test
Qwen3-0.6B (400 MB) on a Samsung Note 8 (2017) phone drives a real desktop Chrome
Up front: I'm one of the people building the page-perception layer used here. We started by testing small local models. The result turned out to be more interesting than the original test. 12 small models, 3 verifiable tasks, logs, and offline replay. Setup: Galaxy Note 8 (2017, Android 9, 6 GB), llama.cpp in Termux, Qwen3-0.6B Q4_K_M. A laptop with Chrome open, not headless. The phone drives the browser through our relay. What the model does: it gets a structured representation of the page (her
For Strix Halo - Official llama.cpp isn't ideal and how to highest possible throughput
I've been making a lot of comments about optimal setup for Strix Halo (gfx1151) and from my observation, 90% of our community is using offcial llama.cpp for it, which is NOT optimized for Strix Halo at all, official llama.cpp is having extremely hard time to reach 50% hardware theory, wasting the silicon of this device. Here's alternatives that can bring the speed of Strix Halo to a totally different world, I will link to user's sastifaction comment to prove that the result is real: - ~50t/s dec
Show HN: Copperhead – Cursor for circuit boards
AI power users claim Anthropic duped them with subscriptions, and they’re taking it to court
Anthropic says power users are key to its business - it's prioritized them even when it means cutting off other popular applications, like OpenClaw. But some of these same customers say Anthropic misled them into believing they'd get more out of a top-tier pricing subscription than they did. In an expanded class action lawsuit filed […]
Cognition hits $48B valuation, signaling investors believe AI coding is far from a winner-take-all market
Cognition's valuation multiple is higher than Cursor's was before selling to SpaceX.
What's going on with OpenAI and the Navier-Stokes controversy?
Artificial intelligence has solved a major mathematics problem, but credit for the accomplishment is murky.
Meta debuts its Muse AI agent. Will consumers trust it?
Meta's new personal AI agent Muse wants access to users' email, calendars, payments, health services, and more — making the company's biggest consumer AI bet yet a major test of whether people still trust Meta with their data.
Meta reveals its AI agent that can shop, send emails and plan trips on your behalf
Muse, Meta's newly revealed AI agent, can perform tasks autonomously, if you trust it to do that.
Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic
Wrote up my thoughts on the whole OpenAI Navier–Stokes Millennium Prize Problem story, and how it highlights the still c…
Wrote up my thoughts on the whole OpenAI Navier–Stokes Millennium Prize Problem story, and how it highlights the still confusing question of what using my data "to improve model performance" actually means simonwillison.net/2026/Sep/8/o...
FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow
We present FlashRender, a few-step generative rendering framework that retakes a source video along a target camera trajectory in seconds. We identify sampling-step-dependent camera control as a prominent manifestation of discretization error in existing multi-step generative rendering models and show that resolving this inconsistency substantially lowers denoising trajectory curvature, facilitating subsequent step distillation. To this end, we introduce Representation Transformation and Alignme
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization
Large language models (LLMs) are increasingly used to formulate optimization models from natural-language problem descriptions, yet realistic operations research (OR) requests are often incomplete: missing objectives, constraints, or business rules can change the resulting mathematical program. Existing evaluations largely assume a complete specification and therefore overlook whether an agent knows when clarification is needed before modeling. We introduce OR-Clarify, a benchmark for pre-formul
Reducto Releases r-1: A Single Pass Document Parsing Model That Cuts Errors 20% at 1 Cent Per Page
We look at r-1, the document parsing model Reducto released on September 1, 2026. We walk through how it folds OCR, layout detection, tables, formatting and grounding into one full page pass, replacing the multi stage agentic pipeline it ships alongside. We break down the two numbers that matter for a migration decision: a reported 20% error reduction and a flat 1 cent per page rate against the legacy 3 to 6 cents. The post Reducto Releases r-1: A Single Pass Document Parsing Model That Cuts Err
Large language models develop novel social biases through adaptive exploration
C*: Unifying Programming and Verification in C (2025)
OUI-1: world's first model for Generative UI
I don't think anyone posted about this here, but Qwen released a finetuned version of 3.5 4 for driving. The full Bf16 checkpoint is 9B. This is a very interesting development of Chinese AI labs tackle self driving next with open weight models. Edit: the HF repo links to the github repo, which in the citation links to a 40 page technical report . Here's the abstract: We present Qwen-Drive-1.0, an initial step towards a vision-language foundation model for autonomous driving. Qwen-Drive-1.0 retai
NeurIPS desk-rejected 178 papers for being "AI-generated". The detector flagged the track chairs' own papers at 24-69% [N]
hey all. the NeurIPS Position Paper Track just used a proprietary AI detector (Pangram) to desk-reject 18.4% of all submissions. no human review, no appeal process, just out. there's been a lot of noise about this, so i went through the actual conference statements and Pangram's technical docs to see how this actually went down. the reality is wildly worse than just "the AI detector made a mistake." here are the receipts: The track chairs would have failed their own test. Independent researchers
GPU guide (GB per dollar, bandwidth)
First plot: GB / $ Second plot: bandwidth (spec on paper, not t/s) Third plot (bandwidth / price) in the comment. Hope that helps, my script uses the GPUs most discussed on the LocalLLaMA, LowEndLocalAI, and LocalLLM subs. At first, I tried to include more, but it became unreadable. Prices were collected by ChatGPT (so may contain inaccuracies). New prices were used where available, second hand otherwise. And I understand this is a basic comparison, but it's better than nothing. For example, you
Qwen 3.8 27b with PI agent - pushed to its 3D graphic game limits
I was inspired by Bijan Bowen video - Subway FPS Wondered how far I can push Qwen 3.8 27b so I used a plan made by Fable 5.1 DESIGN.md which has 267 KB! ( 26K of design line for a game ... LOL ) So I gave that design.md to my qwen 3.8 27b q4xl (llama-server) working on PI agent with 120k context + vision on CPU ( offroad ) + MTP ( for speed ) .... read 11M tokens and write 3.2 M tokens ( worked 12 hours ) .... than that is result. That is insane what we can do locally on own computer !
Which local model is actually good at knowing when to stop and ask you a question?
I’ve been thinking about this after using more agentic/local coding models. A lot of the newer models are surprisingly good at continuing on their own. But sometimes that seems like the problem. If a requirement is ambiguous, I’d rather the model stop and ask: “Do you mean A or B?” instead of spending 10 minutes reasoning, making an assumption, calling tools and then confidently building the wrong thing. I don’t see this behavior discussed much in benchmarks either. We measure coding, reasoning,
OpenAI fought dirty on career-making math problem, says NYU mathematician
There is a $1 million bounty for the first person providing a solution to the Navier-Stokes existence and smoothness problem.
Muse – Meta’s personal AI agent
Muse, Meta’s New Personal AI Agent, Needs You to Trust It
Designed to compete with OpenClaw and Instinct, the company says Muse can do everything from sell your car to book you a plane ticket.
The rise in drive prices is being driven by AI and even HDDs can't hide
SSDs are preferable for many reasons, but with storage prices skyrocketing, there are some scenarios where opting for an HDD makes more sense.
Top chipmakers embrace ASML’s $400M machines, agree to crucial chipmaking change
Chipmaking changes could boost productivity of new ASML machines by 40 percent.
ENEAS: Embedding-guided Neural Ensemble for Adaptive Segmentation
We present ENEAS, a unified, text-promptable method for instance tracking and semantic discovery. Text-promptable segmentation models, including the latest foundation models such as SAM 3, still suffer from temporal hallucinations, spatial fragmentation, and semantic misclassification: they fail to report target absence when an object leaves the field of view, segment local textures instead of the complete object during extreme close-ups, and prioritize visual features over ontological reality,
Mistral raises €3B as sovereign AI becomes big business
The French AI lab has raised €3 billion at a €21 billion valuation in a Series D round led by Samsung, Scaleup Europe, and PSG Equity.
Sunshine, the open source alternative to discontinued Nvidia GameStream, has started to follow Plex steps.
Installed it on a new PC and was met with this screen. Upon some investigation figured out that they now basically push users into purchasing a paid driver to use gamepads (do you really believe they will continue to support free fallback forever?) and suggest purchasing it for keyboard and mouse too, which suggests that eventually free mouse/keyboard support will also be dropped/degraded to less useable state. Issues are already being marked as resolved by redirecting "fixes" into paid version
Connecting the machines
Percolation Dynamics in Optimization : Variance Cascades and Discrete Scale Invariance
We study the dynamics of Stochastic Gradient Descent (SGD), which is known to steer deep neural networks toward invariant sets that correspond to simpler subnetworks. How this steering unfolds over time remains poorly understood. We answer this by modeling the stochastic gradient flow (SGF) as a percolation process, in which architectural symmetries force subnetworks to merge in discrete simultaneous blocks rather than one at a time. These structural transitions register as variance spikes in a
Do you think it happened? Research stolen from their Codex private chats
The Cloud for AI Agents Discussion | Link
Voice conversations between Gemma4 12B and E2B on GPU and Jetson Orin
Gemma 4 12B runs on an RTX PRO 4500 Blackwell. Gemma 4 E2B run on a Jetson Orin NX 16GB; similar performance is expected on a Jetson Orin Nano Super 8GB. Both systems use a reSpeaker Flex 4-mic array and a 3W speaker. Inference is handled by Cortexist Little Gemma, a small LLM engine written in C for CUDA devices. On Jetson Orin it is faster than llama.cpp, and no degradation after long voice prompt. The pipeline supports lip sync, expressions, and gestures. Everything is open source. They talk
Docker vs Podman
Hi, is there anyone here who's running Podman instead of Docker? I tried Podman for the first time (on Ubuntu, with quadlets), and it's a nightmare. Just for LibreChat with code interpreter, I have like 10 quadlets instead of one or two Docker Compose stacks, with no reasonable GUI, just programmatically with questionable Cockpit Podman. Am I just a masochist, or does anyone here really deploy Podman quadlets instead of Docker, and why? Thanks.