DawnSift
Subscribe

Dev Tools

Last 7 days · 127 items

2026-10-09 Fri

JetBrains released Mellum2.1, an Apache 2.0, 12B mixture-of-experts thinking model with 2.5B active parameters. RL in real repositories lifted its SWE-bench Verified score from 2.0 to 47.0. The post JetBrains Releases Mellum2.1: A 12B MoE Open Model for Coding Agents appeared first on MarkTechPost .

We introduce SWE-Game, a benchmark of 247 tasks grounded in 41 executable reference Godot games spanning 13 gameplay categories in 2D and 3D. Five task types cover development from a brief, implementation from a game design document, skeleton completion, repair of 83 injected-fault cases, and Godot-to-Unity porting. Reference materials specify the intended gameplay, while a shared instrumentation interface lets evaluator-owned drivers and probes execute actions and observe independently implemen

arXiv:2610.08900v1 Announce Type: new Abstract: Agentic coding makes code generation cheap, but reliable completion remains difficult: the agent that writes the code is a weak judge of whether it is done. We present Humanize, a multi-agent orchestration workflow for agentic coding built around judgement engineering: explicit, mechanically enforced decisions at the boundaries between planning, implementation, review, and learning. A human approves a plan contract, a builder agent implements it in

Just noticed this today when I went to run the built-in "UPDATE" script and git failed because there was no common ancestor. Looked into why, and apparently every historical commit has been re-written to strip the "Co-Authored by Claude" text from the descriptions. Personally I think that's pretty gross. I'm struggling to think of any reason to do this other than an intention to be dishonest about the origins of the project.

Whistle 发布 16.9 MB 语音识别模型,CPU 运行、无依赖,支持 7 种语言转录与词级时间戳。

普遍认可英语识别准确且体积小、CPU可跑,但也有人认为非英语及嘈杂场景效果差、不如Parakeet。

ttok 0.4Simon Willison2 minDev ToolsAI

Simon Willison 更新 ttok 0.4,修复 Click 警告并新增 --list-models 命令。

Opus5.5, Sol6.1, Fable, and Astra have all proven they can and the scene has exploded this past week. Part of that is from the tools and feedback loops maturing though. Are any open weight models (at all, so including K3, GLM5.3, Qwen3.8-Max, and Mimo-2.6) able to do this? Can the larger models this sub regularly runs (GLM 5.3-Flash, Qwen 3.8-Next-Flash, V4.1-deepseek Flash..) handle a simpler one (GBA and PSP having smaller roms and mature pipelines)?

Recent game design agents have made substantial progress in generating playable games. However, program correctness does not ensure an enjoyable experience for players. We present Recursive Game Creator, an experience-oriented harness to advance agentic game development from rough game prototypes into entertaining games. Recursive Game Creator organizes recursive development around four components: Designer, Builder, Player, and Reviewer. The Designer translates user instructions and Reviewer's

Computer programming is, fundamentally, about two things: Problem-solving using computers Learning to control complexity while solving these problems I have a hard time imagining a future where knowing how to solve problems with computers and how to control the complexity of those solutions is less valuable than it is today, so I think it will continue to be a viable career even with the advent of AI tools. — Carson Gross Tags: computer-science , carson-gross , careers , ai

arXiv:2610.08808v1 Announce Type: new Abstract: Satisfiability Modulo Theories (SMT) solvers are foundational to software verification, program analysis, and compiler testing, particularly over the theory of Quantifier-Free Floating-Point (QF_FP). While recent optimization-based SMT solvers have successfully applied gradient descent to continuous relaxations of logical formulas, they are fundamentally bottlenecked by gradient domination, a phenomenon where a small subset of difficult clauses hij

The other week I posted about Jev vs. Kev compared and since then, OpenAI released the decisions endpoint, Cloudflare released Clef and many here asked about Laya as well. This time we compared six popular decision models by making them play Pac-Man: kev 1.13, Kev 4B, Clef, Clef Flash, GPT-6 Luna and Laya. Since they respond within ms it works for them to play the game in real time. We published a leaderboard and the repo is open-source, so anyone can run their own decision model, like your own

2026-10-08 Thu

EmbeddingGemma 2 came out this week. It maps images and text into one 768-dim space, so you can search photos by describing them. I ported its text and vision towers to ruNNtime, a WebGPU inference library in TypeScript, and made a small photo gallery where search runs entirely on your GPU in the browser. ruNNtime also supports plenty of other vision-like models, and you can play with them in the interactive docs source:

Hi HN! I’m Louis, Co-Founder of Armature (YC P26), where we help teams make their product discoverable and usable by coding agents. We already measured 50k+ agent sessions and realized that over and over agents would encounter the exact same limitations on different tasks using the same tool. So we wondered why these weren’t fixed. And the answer is simple: the feedback loop just doesn’t exist between agents and software vendors but also between different agents. Humans can share their experienc

Docker 发布 docker-agent CLI 插件,用 YAML 声明式配置构建多 agent 协作,无需写代码。

评论区普遍质疑Docker Agent定位模糊、与Docker关联不明,认为其像跟风的agent框架,但也有人认为它对安全可复现的容器化开发流程有潜在价值。

ReikaProduct Hunt1 minDev ToolsAI

A coding agent CLI designed around small local models first Discussion | Link

Language model agents are deployed with a harness, the software around the model that manages its context, tools, and feedback. When such an agent is distilled into a smaller one, the harness stays in place, so the student mainly needs the teacher-specific abilities that the harness cannot provide, such as acting correctly on harness information. Standard distillation, however, imitates the teacher's full outputs and treats the harness as part of the input. We propose Harness-Aware Distillation

I set up borg via borgmatic like a year ago. 3-2-1 strategy. Confirmed it was backing up and did a quick extract test. That was it. That was a year ago. Well I set a vm in Proxmox and because I was still learning I didn’t set up some directories correctly. Later I installed Immich but apparently installed it under a directory owned by Nextcloud. I never updated Nextcloud because it was local and then I decided to make it available remotely via a reverse proxy and all that fun stuff. So I wanted

2026-10-07 Wed

arXiv:2610.03984v1 Announce Type: new Abstract: Autonomous coding agents solve repository issues by reading code, running commands, editing files, and submitting patches. Extra inference-time compute yields gains only when it produces a useful repair and supplies reliable evidence for choosing one. Three behaviors decide both, and we argue they are teachable rather than byproducts of scale, so a policy can carry them instead of a scaffold. Location diversity remains narrow, since attempts return

At my company, we’re using GLM-5.3 Flash internally for software engineering work, and I’ve been genuinely impressed by it. I work in a very large production environment with projects totaling **millions of lines of code**, and we’re not relying on frontier models for this workflow — GLM-5.3 Flash is doing the actual day-to-day coding work. The model is extremely fast, but what’s more impressive is that the speed doesn’t seem to come at the cost of capability. It handles large repositories surpr

llm-openai-decisions 0.1a0 发布:支持 OpenAI 新 Decisions API,gpt-6-luna 支持图像输入,输入 $0.10/百万 token。

llm-mistral 0.16Simon Willison2 minAIDev Tools

Release: llm-mistral 0.16 Adds support for reasoning models, such as the newly released Mistral Large 4 . Tags: llm , mistral , llm-reasoning

Scrimshaw JukeboxSimon Willison1 minAIDev Tools

Tool: Scrimshaw Jukebox I wanted to see if Claude Opus 5.5 could compose music, so I tried this : I want you to write some computer game music for me. First design simple text based format for the music and build an artifact that can play it out loud - include some example tracks in that artifact I am looking for music of the quality of the original secret of Monkey Island It leaned a lot harder into the Monkey Island theme than I had intended, but the results are surprisingly good. I wonder if

arXiv:2610.04008v1 Announce Type: new Abstract: Executable Agent Skills combine natural-language instructions and scripts into reusable packages for LLM agents, and revising them requires fixing errors without breaking correct behavior. Existing benchmarks do not systematically distinguish documentation repair, script repair, and preservation when evaluating skill self-evolution. We introduce SkillScriptBench, a 350-task benchmark designed to evaluate these capabilities separately. From a survey

Hey! 👋 I have released an official support for Strix Halo machines on Strata for Qwen3.8-Flash-Next. Currently numbers are the best on long context decode and ppts using typical Unsloth’s Q4 and GSQ-RCO model weights. Can go up to 1M context length without big speed loss. Currently support is marked as experimental and was done on Linux only. Will be happy for any feedback and pull requests you could give! 👀

Tool-using AI agents are increasingly deployed across enterprise software systems, yet widely used benchmarks primarily evaluate nominal task completion, conflating baseline planning competence with operational fault recovery. We introduce UndoBench, a benchmark spanning 36 base workflows and 36 fault scenarios across 8 enterprise domains, decoupling task competence from recovery capability via counterfactual paired trials under identical seeds alongside wire-level effect-history and environment

arXiv:2610.03872v1 Announce Type: new Abstract: AI agents are becoming increasingly capable of generating scientific code, but generating code is not the same as improving the algorithms behind it. For numerical solvers, execution feedback can expose poor performance, but rarely reveals its underlying cause and how to address it. We introduce Auto-Diagnosis and Skill Discovery (ADSD), a framework that links numerical diagnosis to reusable solver self-improvement. ADSD follows a diagnosis-first p

i was very excited about code generating ai tools since early days of github copilot in vscode, was using it daily since chatgpt release and wrote almost all code through prompting (html/css/javascript, python, ruby, terraform etc) for years. It felt very good at first. However, about a year ago i started to notice subtle (at first) negative changes in my mental health. It's hard to describe in words this negative feeling because it's very basic and fundamental, but over that last year it progre

Release: datasette-atom 0.11a0 A minor fix for compatibility with the latest Datasette alphas. This meant we could upgrade the datasette.io site to Datasette 1.0a41. Tags: atom , datasette

arXiv:2610.03966v1 Announce Type: new Abstract: Each run of an AI-driven research system (ADRS) is an expensive search over a vast solution space, and dependable evaluation requires many runs, making run data both costly to produce and valuable to retain for large-scale analysis. Yet this data remains fragmented: teams operate in isolation, ADRS frameworks emit results in different formats, and no shared infrastructure exists to aggregate or compare runs across problems and systems. We present R

JetBrains 2025 年营收增长 6.3% 但净亏损 3.15 亿捷克克朗;评论区主流观点认为受 AI 编程工具冲击,但也有人指出亏损源于主动投资 AI。

评论区普遍认为JetBrains受AI编程工具冲击、产品老化而陷入困境,但也有人认为其营收仍在增长,亏损源于主动投资AI。

2026-10-06 Tue

The "old" version of Cowork runs model inference in the cloud, executing tool calls in an Anthropic-provided VM we shipped to your computer. We added the VM for capability, safety, and security reasons - mapping in just the data you explicitly added to your session. People loved what they were able to do with Claude but didn't love the disk, battery, and performance cost of running the VM locally. Also, people didn't love that closing your laptop means the work stops. The "new" version of Cowork

LLM-guided evolutionary methods, such as AlphaEvolve, have emerged as powerful approaches for challenging computational optimization problems, such as circle packing. However, prior work typically optimizes performance gain over a fixed number of iterations. We argue that practical optimization should maximize gain per unit cost. To this end, we propose FrugalEvo, a cost-aware evolutionary framework where a stronger, higher-cost LLM explores solution strategies, and a cheaper LLM implements them

arXiv:2610.02405v1 Announce Type: new Abstract: Using a frontier model like Claude Opus as a meta-agent to generate terminal tasks and verifiers for RL training is increasingly common. Yet a runnable Docker image and executable test suite do not guarantee a faithful end-to-end pipeline for terminal agent training. We present a meta-agent pipeline motivated by this gap, diagnosing three classes of failure: benchmark invalidity, harness brittleness, and reward misalignment. Prompt redesign and con

I found this question getting asked all over at least since 5 years ago. There are some funny reasons given for it such as that it does not have "official offering and needs self-hosting" (yeah;)) and that groovy is complicated all the way to simply there's no compelling reason to run a pipeline like that when you can offload your worries to GitHub, GitLab, etc. (yikes) So I wonder - is Jenkins dead to you? Since when? And what did you replace it with? And if not, why not, what's missing in all

评论普遍质疑Cloudflare做搜索中间层的价值,认为直接调用或自建更划算,但也有人认为其统一接口和ZDR承诺对代理场景有用。

I just published the results for this year's State of Devs developer survey, which covers topics such as career, health, worldview, and even hobbies. Some interesting stats: The most common emotions respondents cited when asked about their feelings towards the tech industry was "exhaustion", followed by "disillusionment". "Curiosity" came in third, and is strongly correlated with being pro-AI overall. Speaking of AI, 49% of respondents now generate over 75% of their code using AI. Despite that,

ReviuProduct Hunt1 minDev ToolsAI

The review app for code your agent writes Discussion | Link

2026-10-05 Mon

DeepSeek has released official macOS and Windows desktop apps for DeepSeek Harness v0.2, its MIT-licensed agent harness. The preview adds a plugin manager, file and code-change review, and scheduled Automation Tasks. It also supports non-DeepSeek models through OpenAI-compatible endpoints. The post DeepSeek Harness v0.2 Brings Official Desktop Apps to Its Open-Source Agent Harness appeared first on MarkTechPost .

So I tried that miracle engine everyone is talking about. Asked the IQ3_S model to express its opinion on a post from this sub to measure the tps on a long-ish generation: Can you help with the following problem? So Kimi K2 is outdated, and so is GPT OSS 120b. Which of the modern open weights models can boast the least sycophancy? I need this both for creative/research assistant usage (sycophancy led me down blind alleys of my own bad ideas many times) and agentic coding (more sycophancy less bu

ApertureProduct Hunt1 minAIDev Tools

An AI code editor that checks its own work Discussion | Link

讨论为何开发者不更积极地使用浏览器原生平台能力,评论区认为平台功能不够好、不灵活,但也有人归因于框架依赖习惯。

评论区普遍认为平台原生功能不够好、难用且不灵活,但也有人认为开发者只是习惯依赖框架和库。

2026-10-04 Sun

pi pod runs sessions of the pi coding agent in isolated sandboxes ("pods") on a server you run, in composable environments. ---- Since moving my company towards AI-native work, I have been really frustrated by the state of "agentic engineering" environments. Products by the labs (claude code, codex) lock you into a single provider for your tokens. Agnostic solutions (factory, devin, arguably cursor) make you pay per-token costs. None of these products allow you to fully customize the harness, an

EDIT: Same with the website. EDIT 2: New commits added on clarifying what you can and can't do. "Apache 2.0 in 2029" -> "Apache 2.0" now changed on the website. Also, u/jotkaPL (creator of Dockhand) wrote something along the lines of: "No worries, it will convert to Apache 2.0 in 2029" but later deleted the comment here . Original post: The most important lines have been changed: Previously New Change Date: January 1, 2029. Change License: Apache License, Version 2.0 Change Date: Four years from

I love Qwen3.8 27B so much that I decided to show my gratitude to the Alibaba ecosystem by building a dedicated RAG/search node using a cheap Mini PC from AliExpress. Turns out, my ecosystem loyalty got rewarded with an absolute masterpiece of fraud: Promised: Intel N150 + DDR4/DDR5 Delivered: Core i3-7020U (2018 Kaby Lake, 2C/4T) + DDR3 1600MHz The Scam: The seller literally hardcoded New_N150 into the BIOS release string ( HSHW_M6_DDR3_EC_Intel_Com_New_N150_K001 ). So now my Qwen3.8 RAG stack

CookTrace is a self-hosted recipe manager, pantry and shopping list, an alternative to Mealie, Tandoor and Paprika. AGPL-3.0, single Docker container, native Android app and a Wear OS app, no telemetry and no cloud sync. Your recipes are a SQLite file on your own machine. Part of the TraceApps family: NutriTrace (nutrition), CookTrace (recipes / pantry / shopping), LiftTrace (strength / lifting), NoteTrace (notes / tasks / reminders). New here? It keeps your recipes and cooks them with you. Impo

It feels like we are coming full circle to the times of the Silo Engineer, but on steroids. Back in the day when social media feeds were chronological (if they even existed), the Silo Engineer would retreat into his cubicle to Create (it was usually a he). He was Creating and Creating, and then a few weeks later the results came back. If you were lucky, you got what you asked for or something close to it. No one knew what the Silo Engineer was making, whether it was a sound solution from the arc

2026-10-03 Sat

Zig v0.17.0 发布,重写构建系统并引入 Build Server Protocol,206 位贡献者参与。

评论普遍认可 Zig 设计与目标支持,但也有人认为其核心成员态度强硬、AI 政策转向引发争议。

《Agentic Coding 的四骑士》引发热议,评论普遍担忧智能体编程侵蚀代码质量与团队协作,但也有人认为问题源于使用方式。

评论普遍担忧智能体编程侵蚀团队协作与代码质量,但也有人认为问题源于使用方式而非工具本身。

Hi HN, I'm Justin. Breadcrumb records everything you do on your Mac (screen + meetings + AI transcripts + what you and your AI decided) and turns it into memory your AI can search. It's local and encrypted. You can also teach it rules by talking to it and it makes sure the right rules turn up in the right context. Works with Claude Code / Codex / Cursor / opencode. All of this is exposed to your AI as 30+ MCP tools (here's the definitions): I started it in June because I wanted to understand wha

Retrieval-based speculative decoding (SD) drafts tokens by copying continuations from existing text, which suits coding agents that repeatedly reproduce code, logs, and earlier attempts. Yet existing methods fall short in agent pipelines: much of the reusable text is missing from their corpora or stored in a form that differs from what the agent emits, and their draft lengths ignore that accept length varies across agents and drifts over turns. We present AgSpec, a framework that supplies the co

arXiv:2610.00025v1 Announce Type: new Abstract: Agent harnesses increasingly want to run small language models (SLMs) on the microtasks around a frontier large language model (LLM) planner: auto-approving shell commands, writing memory, selecting tools, ranking past turns. We ask whether off-the-shelf SLMs meet practitioner-defined thresholds and, when they fail, why, and whether quantization changes the answer. We build a benchmark of 4 such microtasks with fixed prompts and automatic metrics,

Every morning, a tech digest curated for you

The web shows the big picture; subscribers get their own — AI curated to your interests, your private RSS folded in, with community takes, delivered each morning. Free forever.

89 issues shipped · 150+ items sifted to 30 worth reading, every day

Every morning, a tech digest curated for you