OpenAI launches a bio-bug bounty program for GPT-5.5/5.6, raising the reward for general jailbreaks to $50,000.
2026-07-10
— Model explosion, tightening regulation, local inference breaks barriers
OpenAI releases GPT-5.6 three models and ChatGPT Work, announces it as the preferred model for Microsoft 365 Copilot; Meta launches Muse Spark 1.1, priced at just $1.25/million input tokens; GLM-5.2 (744B MoE) runs locally on a consumer machine with 25GB RAM; NYT accuses OpenAI of hiding evidence in copyright lawsuit, seeks court sanctions; EU passes Chat Control 1.0, allowing suspicionless bulk scanning of private chats.
À la une
GPT-5.6 Officially Released: Three Models + ChatGPT Work + Microsoft Copilot IntegrationMulti-sources ×7
OpenAI launches the GPT-5.6 family, including flagship Sol, balanced Terra, and economical Luna. Sol surpasses predecessors and competitors on benchmarks for coding, cybersecurity, and more. Simultaneously, it releases ChatGPT Work agent, based on Codex technology, capable of completing complex tasks across applications, and announces GPT-5.6 as the preferred model for Microsoft 365 Copilot. Why it matters: Software engineers can immediately try the new models at low cost (Sol $5/$30 per million tokens). ChatGPT Work extends agent capabilities to non-developers, and Microsoft integration signals imminent large-scale enterprise deployment.
The community generally acknowledges cost-efficiency improvements, but some developers feel coding capabilities fall short of expectations and question the marketing hype.
Meta Releases Muse Spark 1.1: Multimodal Agent Model with Competitive PricingMulti-sources ×4
Meta launches Muse Spark 1.1, a multimodal reasoning model for agentic tasks, with significant improvements in tool use, computer use, and coding. It opens the Meta Model API preview, priced at $1.25/million input tokens, supporting a 1M token context window. Why it matters: Entering the AI coding market at a very low price could drive overall price reductions; the 1M context is suitable for complex agent workflows, directly impacting developers' agent architecture choices.
Commenters recognize the cost-effectiveness but are divided on the fairness of benchmarks and usability.
GLM-5.2 (744B MoE) Successfully Runs Locally on a Consumer Machine with 25GB RAMMulti-sources ×3
Developer JustVugg releases the colibrì project, streaming GLM-5.2 (744B parameter MoE model) on a consumer machine with approximately 25GB RAM using pure C with zero dependencies. It uses only 9.9GB of resident memory, loading other experts from disk on demand. Why it matters: Significantly lowers the barrier to running large models locally, enabling use of a 744B-level model without high-end GPUs, offering possibilities for edge computing and privacy-sensitive scenarios.
Highly praised for engineering implementation, but some question actual inference speed and practical value.
NYT Accuses OpenAI of Hiding Evidence in Copyright Lawsuit, Seeks Court Sanctions
News organizations including The New York Times file a sanctions motion in the copyright lawsuit, alleging OpenAI falsely claimed for years it could not search training data to conceal infringement evidence. Its privacy engineer's testimony revealed contradictions; OpenAI is accused of hiding billions of chat logs. Why it matters: If the court imposes sanctions, OpenAI may be required to hand over key data, directly affecting the legal boundaries of fair use in AI training, with far-reaching implications for the industry's large-scale scraping practices.
Community opinions are polarized, with some supporting copyright holders and others viewing this as an obstacle to AI development.
EU Passes Chat Control 1.0: Allows Suspicionless Bulk Scanning of Private Chats
The European Parliament passes Chat Control 1.0 with a qualified majority, allowing suspicionless bulk scanning of private communications. Encrypted communications are exempt but not actually scanned; the measure is effective until 2028. Why it matters: Creates regulatory pressure on applications and open-source projects relying on end-to-end encryption. Developers need to reassess privacy compliance strategies in the European market.
The developer community strongly criticizes it, viewing it as a violation of privacy and democratic processes; however, some note the bill has not yet taken final effect.
Chaque matin, un digest tech fait pour vous
Le web montre la vue d’ensemble ; les abonnés reçoivent la leur — sélection IA selon vos intérêts, votre RSS privé intégré, avec les avis de la communauté, livrée chaque matin. Gratuit à vie.
44 numéros publiés · 150+ infos filtrées à 30 chaque jour
Actu IA
Anthropic develops the Jacobian lens, discovering a hidden 'J-space' concept space inside Claude Opus 4.6, revealing the model's pre-reasoning process.
AgentLens proposes a code agent evaluation benchmark based on trajectory scoring, combining formal verification with LLM review to provide interpretable scores.
A new paper proposes the 'Harners Effect': orchestration design is a key lever for controlling token costs in enterprise agent systems, with experiments spanning six major models.
AI agent startup Lyzr completes a $100 million Series B round using its own agent SivaClaw, handling all investor Q&A and documents end-to-end.
Dev & open source
Bun founder discusses the Rust rewrite controversy; Zig author Andrew Kelley posts criticizing engineering standards; community sees it as personal attack vs. necessary critique.
评论区主要认为该文章是对Jarred的人身攻击而非技术讨论,但也有人认为这种直率批评是必要的。
Due to censorship after Microsoft acquisition and AI data misuse, some developers migrate to Codeberg, self-hosted Gitea/Forgejo as GitHub alternatives.
用户因微软收购后GitHub的审查、AI训练数据滥用、服务不稳定等问题,转向自托管Gitea/Forgejo或Codeberg等替代方案,但也有人认为GitHub的便利性仍难以替代。
Mitchell Hashimoto interview: discusses Ghostty terminal, Zig language, and open-source entrepreneurship, mentioning 15 years of CLI experience as serendipitous exploration.
Échos de la communauté
Reddit senior developers discuss: the value of hiring junior engineers—they force teams to maintain documentation, standards, and code quality.
GLM-5.2 achieves accuracy close to human accountants in VAT audit tasks, at only 1% of the cost of manual work.
GitHub Trending
AI-powered job application framework built on Claude Code. Fork it, fill in your profile, and let Claude evaluate jobs, tailor CVs, write cover letters, and prepare you for interviews.
Mellea is a library for writing generative programs.
Give Claude the ability to watch any video. /watch downloads, extracts frames, transcribes, hands it all to Claude.
agent multiplexer that lives in your terminal.
Clone any website with one command using AI coding agents
World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio.
LLM 驱动的多市场股票智能分析系统:多源行情、实时新闻、决策看板与自动推送,支持零成本定时运行。 LLM-powered multi-market stock analysis system with multi-source market data, real-time news, decision dashboard, automated notifications, and cost-free scheduled runs.
Use Codex from Claude Code to review code or delegate tasks.
🪄 Flint is a visualization language that lets AI agents reliably create expressive, good-looking charts from simple, human-editable chart specs.
Instant, Concurrent, Secure & Lightweight Sandbox for AI Agents.
Aussi à voir(50 de plus)
Rebranded Codex promises independent workflows that can run "for hours if needed."
Hey HN! We've spent the good part of this past year building an AI tutor that teaches kids ages 4-9 reading, math, ESL and more. Getting an AI tutor to effectively teach a child turns out to be a really hard technical challenge, this took getting the underlying architecture right. Our tutor steers the UX in real-time and makes complex decisions on the fly. Doing both at conversation speed required us to replace the standard tool-use loop. We built our own tutor harness that utilizes a streaming
TLDR: 75B-total / 9B-active MoE is the perfect shape for multi-24GB rigs, and almost nobody ships it. Qwen 27B is a great model and punches way above its weight-class, it is a frequent fallback for me. Nemotron-3-Puzzle-75B-A9B, NVFP4, vLLM 0.22.1 (the new Marlin fallbacks run FP4 on Ampere), pipeline-parallel across 3×3090 capped at 200W each. The 4th card runs a speech sidecar untouched - 3 seats × 256K ctx, fp8 KV — hybrid Mamba keeps the cache tiny - 132 t/s decode across 3 streams (~65 sing
arXiv:2607.07321v1 Announce Type: new Abstract: Tool utilization enables Large Language Model (LLM) agents to interact with the real world and resolve complex tasks. However, existing agent frameworks predominantly rely on static toolsets composed of granular atomic actions (e.g., basic file I/O or single-turn search), which forces agents to reinvent low-level logic for every recurring workflow, leading to increased reasoning overhead and failure rates. In this study, we propose that agents can
I have to admit, a lot of people we're 100% correct to make the suggestion to try this model. I am sorry I ever doubted. The 27B passed every agentic task on a neutral system prompt in 6-9 tool calls. The 75B needed a hand-tuned profile to pass at all and used 2x the turns. For agents, fewer turns beat faster tokens. The two contenders - Nemotron Puzzle-75B-A9B NVFP4, vLLM, PP=2000 across 3 cards, ~65 t/s decode. I made a post about this model. I still think its good for throughput on chatbots a
Maintainer here. OpenMed is an Apache-2.0 toolkit for clinical NLP with one hard rule: patient data never leaves your hardware. No cloud calls, no API keys, works in airplane mode. What shipped in 1.8 this week: OpenMedKit for Android (Kotlin, ONNX Runtime Mobile + ML Kit OCR): read a document, strip every name/MRN/date, entirely on the phone. iOS/Swift and React Native bridges landed too. Browser runtime : de-identification via Transformers.js / ONNX Runtime Web with wasm + WebGPU backends. Ful
Wow... Using this customized vllm provided as a docker, I'm able to run DS V4 Flash on a single RTX 6000 Pro (apparently it also works on a single 5090 - check his readme, but I haven't tried). Apparently this also works with GLM 5.2 (though you need at least two 6000 pros, which is still amazing). Setting 130K context, I needed around 150 GB of RAM to get past the safetensor sharding, but once it is fully loaded in VRAM I am able to fit it all in the GPU (If you have less than this much RAM, cr
I came across this on Hacker News and felt like I needed to share it with the dev community. The main point is simple: a lot of teams reach for extra databases, queues, search engines, caches, and services before they actually need them. This page lays out where Postgres is usually enough, and where you may actually need something else: Postgres is not perfect for everything, but it is good enough for a surprising amount of real-world work. The more I build and maintain systems, the more I appre
This post was originally written in Korean, then polished and translated into English using ChatGPT. I do run llama.cpp locally on a Tesla P40, but as someone who already pays for ChatGPT Pro, I was gradually losing the practical reason to keep running local LLMs like Qwen 3.6 27B or Gemma 4 31B. If I need access to OpenAI models through an API-like workflow, I can usually just use Codex OAuth instead. But then I realized that embedding models and reranker models are not something I can access t
arXiv:2607.07097v1 Announce Type: new Abstract: Safety evaluations of multi-agent LLM systems often compare a direct prompt with a planner-executor pipeline and report the difference as a single "pipeline effect." We argue that this aggregate is difficult to interpret because it conflates three mechanisms: harmful intent may be reframed as plausible operational work, the planner may refuse or transform the request, and the executor may act under delegation prompts implying prior approval. To sep
“token efficiency needs to drop to as much as 20% over the next 12 months, and 90% by the following year”
Hello everyone! Pangolin 1.20 is focused on how people find and reach their resources in the UI. Here's what's new: Pangolin is an open-source, identity-based remote access platform that lets you securely expose your infrastructure to your team. It supports browser based remote access and a remote access VPN in one platform with strong authentication controls. GitHub: Resource Launcher The landing page non-admins see when they sign in now is vastly more capable. Resources are grouped by site or
No real details or timescales yet, but this article has confirmation from Alexandr Wang that Meta are working on an open source variant of Muse Spark. One to keep an eye on.
MOSS-Transcribe-Diarize 0.9B is an end-to-end audio understanding model for long-form multi-speaker transcription, diarization, timestamps, and acoustic event awareness. Given an audio or video file, the model generates a compact speaker-aware transcript in one pass, including timestamps and anonymous speaker labels such as [S01] , [S02] , and beyond. Introduction MOSS-Transcribe-Diarize 0.9B turns real-world long-form audio into structured, speaker-aware transcripts in one pass. Instead of stit
The main conclusions from analysis were: The Pareto frontier for coding tasks (i.e. best quality for a given cost) includes models from OpenAI, Anthropic, and open source. This means today, only a mix of tools can provide frontier performance. Open models, and GLM 5.2 in particular, are now able to handle even the highest level of task difficulty. The token price of a model is a poor indicator of actual costs incurred on end-to-end tasks. Larger models can be far more token efficient and have lo
The feud between NightmareEclipse and Microsoft shows no signs of resolving soon.
Broadcom accuses Allstate of dodging VMware audits.
The agent-native way to ship software Discussion | Link
Windows 11 updates could soon include fixes for more security issues at once. Microsoft said in a blog post on Thursday that it's now using AI to "identify potential issues earlier," which means "customers will see a higher volume of security updates included in each security release." Hackers, even amateurs, have increasingly been using AI […]
Release: llm 0.31.1 Fix for a bug with OpenAI Chat Completion endpoints where a tool call with empty arguments could result in a JSON error from some providers. #1521 This bug came up when I was testing llm-meta-ai . Tags: llm
Hi Hacker News, I’m Yahia. I built Context.dev ( ) to make it really easy to integrate web data into your products and agents. Here’s a demo video: Since it’s an API, here are the docs: . You can send us a URL and get back clean Markdown, rendered HTML, screenshots, extracted images, etc.. You can also send us a domain and get company or brand context: name, description, logos, colors, fonts, social links, screenshots, style information, and related metadata. For more custom use cases, you can s
arXiv:2607.07229v1 Announce Type: new Abstract: Prior work has shown that chain-of-thought (CoT) reasoning is often unfaithful: a model's stated reasoning does not reliably reflect the process that produced its output. Detecting unfaithfulness, though, requires controlled experimental interventions, which cannot be applied to evaluation transcripts after the fact. We turn instead to a more tractable question that has received less attention: whether the stated reasoning is logically consistent w
Hey Guys, As promised here are the results from running MiniMax M2.7 REAP 139B Q3_K_L on llama-bench on 6x MI50's. Memory Load: Hardware: Asus X99-E-WS ( Modded BIOS to support a large number GPU's ) Intel(R) Xeon(R) CPU E5-2680 v4 @ 2.40GHz 128GB DDR4 RAM SSD 6x MI50's 96GB VRAM (Gen3 x8,x8,x8,x8,x8,x8) Results GPU Setup Model Test Result 6x MI50 / Pro VII 16GB MiniMax M2.7 REAP 139B Q3_K_L pp512 139.27 t/s 6x MI50 / Pro VII 16GB MiniMax M2.7 REAP 139B Q3_K_L tg128 24.87 t/s 6x MI50 / Pro VII 1
I made this simple 3D Geometry Wars-style game using my coding agent, Jarvis Code, with GLM 5.2. You can play it here: I was honestly surprised by the result. Most of the game came together in the first iteration, and I only needed about four small follow-up tweaks afterward. I'd be curious to hear what you think about the gameplay, the code quality, or GLM 5.2's coding ability.
arXiv:2607.06820v1 Announce Type: new Abstract: Recent advances in AI for Mathematics have focused largely on autoformalization and theorem proving, leaving the role of Computer Algebra Systems (CAS) in agentic LLM workflows underexplored. We propose a ReAct-style agentic setup that combines LLM reasoning with verifiable feedback from SageMath, together with Context7 for the up-to-date documentation. We evaluate this agentic setup across frontier models for solving research-level mathematical pr
[...] Work on web and mobile runs in the cloud. Work in the desktop app can also use local files and desktop apps with your permission. At launch, cloud Work conversations do not appear in desktop Work; desktop Work threads and local files remain on that computer. — OpenAI , trying (unsuccessfully) to clarify ChatGPT Work Tags: openai , chatgpt , ai
OpenAI is sunsetting its AI-powered browser after less than a year. But it's moving some agentic browsing features to its desktop app and a Chrome extension.
OpenAI is already shutting down ChatGPT Atlas, its browser that could do tasks for you on your behalf, less than a year after launching it. Atlas was announced in October, but as part of its wave of news about ChatGPT Work today, the company confirmed that it will be "sunsetting" Atlas and is targeting an […]
Dictate, rewrite, translate, and an agent in a single device Discussion | Link
arXiv:2607.07189v1 Announce Type: new Abstract: Vision-language models (VLMs) and agentic AI have shown strong performance on semantic visual tasks, but it remains unclear whether they can handle the physics and inverse problems that underlie computational imaging. We present ImagingBench, a benchmark of 20 computational imaging tasks spanning five categories: ray and wave optics, image signal processing, inverse reconstruction, computational sensing, and calibration. ImagingBench evaluates thre
AI cheating leads to "a failed society," professor says.
I finally decided to set up a dedicated home lab server a while ago. I priced out built-from-scratch x86 ITX configurations and low-power NAS builds, and they were easily running $500–$800 just for CPU/RAM/chassis. I hesitated. I wanted something cheap, extremely low-power, and compact. So I did what any frugal self-hoster does: I looked at old hardware I already owned. I had a 2017 Dell XPS 13 laptop (i5, 8GB RAM) sitting in my drawer. I had tried selling it locally, but since the battery was c