DawnSift
구독하기
수 · 테크 데일리 · 제66호

2026-09-16

— Open-source models have closed to within just a 4-month gap, while security incidents remind us: scan before you trust.

오늘의 TL;DR

Google releases the Gemini 3.8 Live series of speech models, targeting production-grade voice agents; TypeSafe launches System One Models, focused on speed and cost for structured decisions; a Mozilla report shows the gap between open-source models and frontier closed-source models has narrowed to 4.4 months; on the security front, Strix obtained Baseten's production GitHub admin access in 25 minutes, and CrofAI was exposed as an OpenRouter wrapper before wiping its online presence.

헤드라인

1

Google releases Gemini 3.8 Live and Extended Thinking, targeting production-grade voice agents다중 소스 ×4

Google launches Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two native speech-to-speech models; the former targets scale and cost efficiency, the latter targets highly complex multi-step reasoning; both support background tool calling, real-time visual input, and switching among 97 languages, and Extended Thinking ranks first on the Artificial Analysis Speech Quality Index with a score of 82.6. Why it matters: the key bottleneck in moving voice agents from demo to production is reasoning and tool execution without interrupting the conversation flow, and these two models directly address that gap, and are already available through the Gemini Live API and AI Studio.

The comments largely acknowledge the voice experience, but some believe the demo flopped, the version rollout was chaotic, and extended thinking is not worthy of its name.

2

TypeSafe releases System One Models: a new model category for structured decisions

TypeSafe AI founder Diogo Almeida (formerly of the OpenAI ChatGPT research team) announced the launch of System One Model, positioned as a frontier model category for fast, structured decision-making, along with the accompanying Jev tool, targeting speed and cost advantages for structured tasks such as classification. Why it matters: if this model truly reaches frontier-level performance on structured decisions at lower cost, it will change the model selection logic for the many repetitive judgment tasks in agent pipelines.

The community generally acknowledges its speed and cost advantages on structured tasks, but some believe it lacks benchmark proof and that comparisons with LLMs are unfair.

3

Mozilla report: gap between open-source models and frontier closed-source models narrows to 4.4 months

Mozilla's "State of Open Source AI" report shows that the performance gap between frontier AI models from U.S. tech companies and China's best open-weight models has narrowed to 4.4 months, with Moonshot AI's Kimi K3 trailing by only 3 points on the Artificial Analysis Intelligence Index composite score. Why it matters: this means most routine workloads can be satisfied with open-source models, and frontier closed-source models are worth their 5x cost premium only in a very narrow set of highly difficult scenarios, directly affecting enterprises' model procurement and deployment strategies.

4

Strix security team obtains Baseten production GitHub admin access in 25 minutes

While evaluating whether to host data with Baseten, security firm Strix used its self-developed autonomous hacking agent Strix to scan its external exposure and, after about 25 minutes, obtained a live GitHub token with repository-level admin permissions over Baseten's internal repositories. Why it matters: this once again confirms the fragility of third-party service supply chain security—even an infrastructure provider valued at $13 billion and relied on by many serious companies may have quickly exploitable credential leak paths.

5

CrofAI exposed as an OpenRouter wrapper, routing to smaller models at up to 20x markup before wiping its online presence

CrofAI, which calls itself the "world's cheapest inference provider," was exposed as actually an OpenRouter wrapper that routes requests to smaller, cheaper models with markups of up to 20x; facing fraud allegations, CrofAI first denied, then changed its story, and 3 hours later wiped its entire online presence. Why it matters: this is a typical lesson in chasing cheap tokens—if an inference provider lacks transparent model routing and pricing mechanisms, "cheap" may just be a disguise for wrapper markups.

The community generally sees it as a cautionary tale about chasing cheap tokens.

매일 아침, 당신을 위한 테크 다이제스트

웹은 전체 그림을, 구독자에게는 당신만의 것을 — 관심사 맞춤 AI 큐레이션, 개인 RSS 통합, 커뮤니티 반응과 함께 매일 아침 배달. 영원히 무료.

66호 발행 · 매일 150개+ 중 읽을 가치 있는 30개로 선별

AI 소식

Atria Dawn Preview: an agentic foundation language model for scientific research and engineering workflows, competing with frontier agents on 16 benchmarks.

🤖Atria Dawn Preview is a foundation agentic language model trained through verified tool interactions that achieves strong benchmark results and demonstrates a shift toward human-AI project-level collaboration in scientific research.

ZGCM-1: a fully open-source 7B foundation model that couples internal reasoning with external tool calling to achieve efficient training on 256K context.

🤖ZGCM-1 is a 7B open foundation model that combines internal reasoning with external tool use, trained via efficient architecture-system co-design, progressive long-context scaling, and autonomous agent workflows to achieve strong reasoning and efficiency.

Grouped Value Attention reduces KV cache by reconstructing keys on demand, achieving near-GQA accuracy with a smaller persistent cache.

🤖Grouped Value Attention reduces transformer KV cache size by storing grouped values and reconstructing keys via a learned linear map, achieving near-GQA accuracy with a smaller persistent cache.

개발·오픈소스

Java 27 GA released, including 9 JEPs: G1 becomes the default GC in all environments, TLS 1.3 post-quantum hybrid key exchange, compact object headers enabled by default, and more.

The comments generally acknowledge Java's continuous updates, but some believe Oracle's pace is too fast and enterprises still use older versions, with mixed feelings.

커뮤니티 화제

Schneier writes that 25 years of mass surveillance has long exceeded its counterterrorism purpose and become a routine law enforcement tool; the comments generally believe surveillance will not stop and will intensify.

The comments generally believe mass surveillance will not stop and will intensify, but some believe efforts should shift to localized alternatives or securing digital rights protections.

The United States confirms for the first time the deployment of a space weapon in orbit; the comments generally criticize the move as triggering an arms race and an orbital debris crisis.

The comments generally criticize the U.S. deployment of space weapons as triggering an arms race and an orbital debris crisis, but some believe space militarization has long existed and this disclosure may just be a political stunt.

GitHub Trending

Star alibaba / open-code-review Fast, efficient, battle-tested at Alibaba's scale. Hybrid architecture code review tool: deterministic pipelines + LLM Agent, precise line-level comments, built-in multi-language ruleset (NPE, thread-safety, XSS, SQL injection), OpenAI & Anthropic compatible.

Star JustVugg / colibri Run frontier MoE models on hardware you already own — pure C, zero deps, experts streamed from disk. Tiny engine, immense model. 🐦

Sponsor Star ever-co / ever-gauzy Ever® Gauzy™ - Open Business Management Platform (ERP/CRM/HRM/ATS/PM) - https://gauzy.co

Star debpalash / VoiceStudio VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

Sponsor Star Homebrew / BrewUI 📺 Homebrew's official macOS GUI

Star melgarafael / DeskcommCRM Open-source AI sales OS — self-hosted CRM with native AI agents + WhatsApp (WAHA). Open alternative to Kommo, Octadesk & Intercom for any business that sells by chat. MCP-ready, multi-tenant, LGPD.

Sponsor Star danny-avila / LibreChat Enhanced ChatGPT Clone: Features Agents, MCP, Skills, DeepSeek, Anthropic, AWS, OpenAI, Responses API, Azure, Groq, o1, GPT-5, Mistral, OpenRouter, Vertex AI, Gemini, Artifacts, AI model switching, message search, Code Interpreter, langchain, DALL-E-3, OpenAPI Actions, Functions, Secure Multi-User Auth, Presets, open-source for self-hosting. Active

Star pacifio / atlas Source control for agents. Use multiple coding agents, track their changes and query them in one place

더 볼만한 소식(61건 더)

Meta engineering team introduced ZGateway, a proxy tier that now sits between client applications and ZippyDB, the Meta’s most widely used key value store. ZippyDB backs product metadata, counters, and configuration at billions of operations per second. ZGateway started as a fix for connection sprawl across more than a million client hosts and grew into […] The post Meta Introduces ZGateway: A Stateless Proxy Tier That Unifies ZippyDB Traffic and Handles Over 1 Billion Operations Per Second appe

Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to stop. This test-time strategy makes it difficult to measure how agent performance scales. We study open-ended tasks that provide continuous scores for intermediate submissions, making progress observable throughout long trajectories. We propose Elo-per-token analysis, which tracks the best solution found at each token budget and uses a Bradley-Terr

Three weeks back , i posted SHADOW-250M here. It got 360 upvotes, 293 on r/LocalLLaMA and 94 GitHub stars. Thank you. That model was 60 MB, ran around 400 tok/s on CPU and could retrieve records from an archive on disk. What it couldn’t do reliably was reason over what it retrieved or compute. So I built a smaller one to experiment with those two problems. SHADOW-50M is actually 44M parameters, trained from scratch on 45B tokens. 19.8 MB complete model, ~1,900 tok/s on laptop CPU, ~41 MB RAM, te

arXiv:2609.13491v1 Announce Type: new Abstract: The strong performance of AI Agents across an impressive variety of tasks is driving an unprecedented investment in agentic infrastructures, however the cost of processing tokens is fast increasing. Web agents automate the execution of web-application tasks described in natural language, by analyzing the web-application's user interface (UI) and interacting with it. This work introduces OdoBot, a novel web-agent architecture that completes tasks at

arXiv:2609.13543v1 Announce Type: new Abstract: LLM agents are predominantly benchmarked on short, single-task trajectories, yet real deployments run for hours under contention, surfacing a different class of failures. We use the Clinical Environment Simulator (CES), in which an agent manages an entire emergency-department shift under continuous time and resource pressure, as a testbed: long-horizon execution failures manifest measurably in a single rollout under structured, multi-dimensional gr

Agent-net, the team building an agent-to-agent marketplace where AI agents discover, trust, and pay each other, has released Webagent, an open source harness for standing up public-facing business agents. So, basically you give it your website, get an agent, and let it talk to other agents. Instead of writing orchestration code, a business fills in […] The post Agent-net Open Sources Webagent: A Go Harness That Turns Any Website into a Guarded AI Agent appeared first on MarkTechPost .

Diffusion large language models (dLLMs) achieve high decoding efficiency through block-parallel, arbitrary-order generation, making them attractive for latency-sensitive applications. GUI agents represent a natural testbed for this paradigm, as they must repeatedly perceive screen states and emit structured, spatially grounded actions in real time. However, whether dLLMs can be extended into capable multimodal GUI agents while preserving their parallel decoding advantage remains an open question

Computer-use agents increasingly interact with browsers, terminals, file systems, and external services, introducing safety risks that emerge through runtime behavior rather than generated content alone. Existing guard models target static prompts and responses and are poorly suited to agent execution; existing executable safety platforms produce evaluation verdicts rather than the normalized supervision a guard model needs to learn across heterogeneous agent frameworks. We introduce HazardAudit

Two months ago I announced I had found a new dynamic quant method called Voodoo Quant which was SOTA for the most aggressive quant levels on some smaller Qwen3.5 GGUF models. I kept the methodology private at the time, but I've seen too many requests for dyn quants for various models lately, so I decided to give my method to the community since I don't have the time to scale this into something that could do it justice. Hopefully it will also inspire some researchers to find out more about it an

EDIT: Sorry for the unclear title. This model is UkisAI's Swift-Qwen3.8-27B, not a new version of BottleCap AI's 3.6-ThinkingCap. All credit goes to UkisAI for making great fine-tune, and I made this post to celebrate their work. I meant no disrespect by mentioning another model in the title. I doubt I'm in the minority here when I say I love Qwen models, but the overthinking is a major timekiller. It was bad in 3.6-27B, and it's worse in 3.8. I know there are some who say, "well that's how it a

Hey r/LocalLLaMA , We’ve released our full ShapeLearn GGUFs for Qwen 3.8 27B. Blog / Download models TL;DR 3.84 bpw (GPU-5) reaches 99.63% of BF16’s aggregate score of 8 benchmarks, being the most accurate quant we’ve evaluated; 3.23 bpw (GPU-4) reaches 98.72%. These average BF16-normalized scores across instruct and thinking benchmarks. All five new models sit on the measured quality/speed-bpw frontier across six GPUs. In this model’s case, lower BPW translates directly to TPS. Comparisons incl

I have a custom fork of llama.cpp designed around the ampere architecture specifically (though many of the upgrades also translate to faster performance of blackwell + lovelace). The recommended config supports 90+ TPS (for agentic/coding, at temp 1; greedy will of course be faster) through 100K tokens, with context of up to 240K. If you want the repo, it is here: I recommend running with this quant, which is ~ 4 K M quality but considerably faster (technically, a 3 K XL upgrade) If you want the

arXiv:2609.13406v1 Announce Type: new Abstract: When we speak of recursive self-improvement (RSI), are we speaking of a phenomenon, a mechanism, or a prospect? Towards autonomous and evolving intelligence, RSI is being claimed at many scales, while no single framework that formally describes these emerging instances exists. Its counterpart in the classical realm, iterative policy improvement, is characterized by generalized policy iteration (GPI), a framework of broad applicability with well-und

arXiv:2609.13436v1 Announce Type: new Abstract: Large Language Model (LLM) agents offer a promising path toward autonomously managing long-term physical tasks without human intervention. However, physical tasks require agents to continuously observe the environment, make consequential actions, and remain effective as the environment changes. Existing approaches either require substantial data and retraining, or primarily focus on agents operating in the virtual world. In this work, we explore th

Foundation models have progressed from learning and reasoning over existing knowledge, to increasingly learning through action, tool use, and outcome feedback. We argue that the next frontier is a further transition: from solving and acting within problems specified by humans to participating in the process by which new problems, representations, explanations, and knowledge are created. We refer to this capability as Discovery Intelligence. We formulate Discovery Foundation Models (DFMs) as gene

Multimodal agents can create complex videos in software such as Blender by coding without relying on diffusion models. Yet video understanding benchmarks still evaluate models mainly through question answering. If an agent truly understands a video, it can reconstruct it programmatically. We introduce BVB, Blender-VideoBench, a benchmark that tests this ability by asking agents to reconstruct real-world videos as animated Blender scenes. To ensure fair comparison, each agent programs the reconst

Interactive video world models must maintain broad scene context under camera motion while producing high-fidelity observations with low latency. Existing approaches face a representation trade-off: perspective models operate on local views and must preserve off-screen content over long rollouts, whereas broader spatial coverage is typically obtained by synthesizing full-sphere videos or constructing explicit 3D representations. Motivated by the complementary roles of global context and selectiv

In this technical report, we propose Pelican-Sim 1.0, a general world model simulator for embodied intelligence that predicts future observations from visual context and robot actions to support downstream learning and decision making. The model incorporates four key design features: (1) Unified action representation: a 28-dimensional action value space covering most mainstream embodiments, keeping one model valid across heterogeneous devices. (2) Action-visual injection: URDF- and camera-render

Streaming omni-modal models must decide what and when to answer from the video chunks and synchronized audio observed so far. Visual cues often support an interpretation before an utterance or sound event is complete. If that interpretation enters memory as a fact, later reasoning can keep relaying it even after audio contradicts it. We call this failure premature cross-modal commitment. We propose Omni-Streaming Thinking (OST), which generates structured outputs that include evidence observed s

Maybe some of you know but I didn’t see any post about this. Apple just made available their AFM model on MacOS 27 natively. Just run fm chat in a terminal. Disclaimer: I’m an open weight person. I prefer open models and ecosystem, but I’ll still open the discussion. Did you test them? Build using them? Are these models good? I feel like this is still a huge step in the direction of local AI that a company like Apple does this and release hardware optimized models. So what do you think?

arXiv:2609.13437v1 Announce Type: new Abstract: Scientific research is a continuous process that emphasizes inheritance. Methods developed by predecessors are often expanded upon by new researchers to explore more novel and in-depth scientific questions. However, the change of lab staff, such as student graduation, leads to a lack of personnel capable of replicating methods. Methods that have been developed with significant effort and resources cannot be continued. To address these limitations,

Listen to the session or watch below Employees at the world’s leading AI labs are saying there’s a real possibility that advanced AI could destroy humanity. Are they right? Or is this more scaremongering and hype? Watch a conversation unpacking AI extinction fears: where they come from, whether they hold any water, and, if so,…

arXiv:2609.13422v1 Announce Type: new Abstract: LLM judges are increasingly used to evaluate and improve AI-generated outputs, yet their reliability for complex professional work remains unclear. We study this problem through Vibe Patenting, an end-to-end patent-drafting testbed for AI-agent evaluation. A separately-invoked LLM judge evaluates generated patent drafts and provides structured feedback for iterative revision. Across multiple inventions and drafting-agent configurations, judge-guide

Last year Ruxandra Teslo, a policy analyst who focuses on clinical trials, posted an idea for supercharging medical AI systems: Use data from failed biotech companies. By bidding at their bankruptcy proceedings, she proposed, it might be possible to obtain detailed regulatory filings, manufacturing strategies, and safety data—types of information usually considered trade secrets. She…

Lets be honest, 80% of the services we run our labs are just tools to deploy, monitor, log, backup or do other server-related stuff. Cool for nerds, but nothing to write home about honestly. The really useful and "fun" services, e.g. Vaultwarden, Paperless, Jellyfin etc. are rather sparse I want to know whats your favorite "fun" service one might not know about yet. I'll start: Recently came across RECLIP and I throughly enjoy it. Used to work with halfassed browser extensions and sketchy websit

Multimodal entity linking grounds entity mentions in text and images to knowledge-base entries. These systems degrade on rare entities, but prior work measures rarity primarily through popularity-based metrics such as pageviews. We broaden this view using knowledge-graph structural metrics that capture how well an entity is documented and connected. These metrics identify many rare entities that popularity metrics miss. Across the resulting rare-entity slices, state-of-the-art accuracy drops by

매일 아침, 당신을 위한 테크 다이제스트