阿拉巴马州总检察长就 OpenAI 网络安全模型入侵 Hugging Face 事件发出传票,调查其监管与安全措施缺失。
Security
Last 7 days · 67 items
文章探讨恶意 LLM 是否可能通过推理引擎漏洞控制宿主机器,评论区对攻击可行性与防御难度存在分歧。
文章揭示开源模型可能隐藏时间触发后门,OpenCode 通过系统提示注入元数据指纹成为触发通道。
Is the technique outdated? Yes. Is it still creepy? Also yes.
Early testers are raving about what Instinct can do, but some say the AI assistant’s sweeping access, broad terms and ability to act on users’ behalf come with uncomfortable trade-offs.
Article URL: Comments URL: Points: 80 # Comments: 39
Nvidia worker indicted after Jensen Huang scolded Supermicro for AI server smuggling.
arXiv:2608.20378v1 Announce Type: new Abstract: Safety alignment in Large Language Models (LLMs) is often superficial, relying on refusal mechanisms that trigger only at the final stages of generation without erasing the foundational knowledge of harmful concepts acquired during pretraining. This study demonstrates that this architectural disconnect leaves models vulnerable to Semantic Camouflage -- adversarial attacks that wrap harmful intent in benign narrative contexts (e.g., creative writing
评论区普遍担忧微软在本地AI编辑图片时静默嵌入含GUID的隐形水印,认为这侵犯隐私、威胁匿名性,但也有人认为这是对抗深度伪造的合理措施。
GrapheneOS, an open source version of Android that prioritizes security and privacy, has detailed its plans for supporting Motorola smartphones. Official support is set to arrive next year, starting with traditional flagships, before rolling out to Motorola's foldable phones and perhaps cheaper models, eventually. In a Mastodon thread, the GrapheneOS Foundation announced that it will […]
评论区普遍认为车机系统联网和软件化带来安全隐患,但也有人认为攻击需特定条件,且用户手机配对或简化系统可降低风险。
The Dutch Data Protection Authority is fining Uber €825 million in the second largest penalty issued under Europe’s GDPR.
The class action suit claims that Amazon never obtained consent from Twitch streamers to be used to train its AI models.
Guidelight AI Standards 研究发现,主流 AI 实验室几乎未公开 rogue model 的遏制响应计划,OpenAI 得分最高,Anthropic 与 Meta 最低。
OpenAI is calling for California to strengthen SB 53, an AI safety bill that the company previously opposed.
The AI leader said that the state's SB 53 framework should "be amended to expand safeguards."
A Dutch data regulatory authority said that Uber has to pay 824.9 million euros for violating the GDPR.
The private Android-based OS will expand beyond Pixels next year.
Has anyone else noticed an increase in scanners/bots in the past ~month? For the past couple years I've had 2-3k hits a day from bots but lately there has been a steady increase in traffic looking mostly for php files. What I find strange is how much of this traffic is coming from MS and Google IPs. Do they not have any kind of monitoring on their cloud services? Having thousands of requests spamming every IP that responds should raise some flags. 20.24.67.246 Hong Kong Hong Kong Microsoft Corpo
Anthropic has moved its most cyber-capable model into a product security teams can switch on themselves. Claude Security scans now run on Claude Mythos 5, in public beta for Claude Enterprise customers with no separate model add-on. The scan connects to a GitHub repository, traces data flows across files, and returns findings with a CWE category, confidence and severity ratings, and a suggested patch. The design point is packaging: users receive a scan result rather than a prompt box, so the mod
安全研究员披露通过接管 e164.arpa 域名意外记录数十万通军事基地电话的经过,引发对 ENUM 基础设施安全的关注。
评论普遍认为边境删除数据被控重罪属权力滥用,但也有人认为按美国法律此结果合法。
评论区普遍认为Rust生态缺乏安全管控,需加强沙箱和依赖审计;但也有人认为应减少第三方依赖并采用更完善的开发环境。
Cryptographic Context Injection is only the latest way to break an LLM safety guardrail.
A hacker pretending to work for a leading cryptocurrency news website targeted several cybersecurity professionals using Google Docs as a way to deliver malware.
Simon Willison 研究用 smolvm 作为不可信 Python/JavaScript 的快速安全沙箱,限制 RAM、CPU 与网络访问。
People-search tool ClarityCheck left database containing more than 9M image files exposed.
arXiv:2608.18131v1 Announce Type: new Abstract: Current safety alignment training for Large Language Models (LLMs) are heavily English-centric. When such safety filters fail for non-English languages, the consequences are immediate and user-facing: voice assistants and spoken dialogue systems may produce stereotype-reinforcing outputs, bypassing the standard English-focused safety alignments and propagating harmful bias to non-English speaking communities. For spoken language technologies deploy
论文分析 500 张 Hugging Face 模型卡,认为现有模型卡不足以支撑开放权重基础模型的下游治理。
AliExpress 被曝运行静默 WebAudio 指纹识别,导致蓝牙多点连接音频中断,评论区普遍谴责并要求加强权限管控。
评论普遍谴责AliExpress利用静音音频指纹追踪,认为应加强权限管控和立法监管,但也有人认为浏览器滥用问题更普遍。
OpenAI reaffirms Zero Data Retention for eligible API customers and previews Private Safety Processing for advanced AI safety without compromising data privacy.
Aegis 运行时治理系统将模型输出视为动作提案,通过可信决策层实现 fail-closed 执行边界控制。
A competition is developing between OpenAI and Anthropic over who can provide the best privacy protections for enterprise customer data.
Anthropic announced last week it would include invisible watermarks in AI-generated content to comply with new EU rules. Within hours, overrides were being touted online.
The idea behind OpenAI's Trusted Access for Cyber program is to give trusted defenders better models so they can report bugs and vulnerabilities to companies, with the aim of getting flaws patched faster.
This goes way, way further than keeping an eye on license plates.
The U.S. phone provider escaped a large-scale breach of its network after identifying Chinese-backed hackers early on.
arXiv:2608.17202v1 Announce Type: new Abstract: Safety alignment in open-weight language models is trivially removable: abliteration projects a refusal-mediating direction out of the weights in minutes, and no release-time defense we are aware of prevents it durably. What cannot be prevented can be deceived. Our defense, decoy hardening ("Fool's Gold"), concedes the refusal strip and poisons its payoff: once refusal is stripped, most answers to hazardous operational requests are confident, fluen
Bankrupt Spirit accused of selling out workers in massive data sale to Google.
arXiv:2608.17183v1 Announce Type: new Abstract: Small Language Models (SLMs) are increasingly deployed in resource-constrained, privacy-sensitive settings, where safety and bias failures can cause security and societal risks. However, existing AI safety\slash security\slash compliance benchmarks are designed for large language models that may not transfer reliably to SLMs. We therefore ask: Can these benchmarks effectively and reliably evaluate SLMs? To answer this question, we conduct a large-s
arXiv:2608.17067v1 Announce Type: new Abstract: As text-to-image generative models advance, they raise critical safety concerns, particularly the generation of Not-Safe-For-Work (NSFW) content such as violence and nudity, further exacerbated by red-teaming adversarial attacks. Existing defenses predominantly operate under white-box assumptions, relying on text encoder optimization, weight editing, or inference-time intervention, and fundamentally cannot scale to proprietary models. Black-box alt
Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is known about the harmful outputs produced during model misbehavior, partly because large-scale, high-quality collections of frontier-LLM misbehavior are difficult to obtain. To address this gap, we introduce HarmProfile, a content-centric benchmark dataset that collects model misbehavior across diverse harm categories and
Flock’s surveillance cameras have already sparked outrage. WIRED reconstructed its next-generation AI system, already in use by some police, to confirm it goes much further than tracking license plates.
Secret parameter allowed hackers to steal passwords when a target clicked on a link.
OpenAI is announcing security updates following the July news that its AI broke out of a sandboxed environment and accidentally hacked Hugging Face, including improvements to its research environments, monitoring, and alignment techniques. The company had already put the brakes on a new model, Astra, that it thinks could have "critical" cybersecurity capabilities, and the […]
The ChatGPT maker says its upcoming Astra model may have reached “critical” cyber capabilities, prompting it to halt a significant number of training runs while it tightens internal safeguards.
The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasis on alignment and security during the post-training process.
As large language models become increasingly widespread, third-party providers that deploy open-weight models have become an important part of the ecosystem. Auditing the quality of their inference APIs is therefore an open problem. We formalize hosted model routing as a stochastic process and propose \textbf{Ventor-QTest}, a composite black-box audit that requires no probability information from the target API. Its repeated-request component sends each frozen constrained context to the target m
OpenAI is strengthening monitoring, alignment, and security for frontier AI models. See how new safeguards are guiding the pace of model development.
Z.ai’s latest AI model release could help companies secure their systems—or find its way into the hands of hackers.
OpenAI launches an initiative to strengthen democratic oversight of AI in national security, supporting government institutions with tools, training, and expertise.
Passwords are less secure than passkeys, even if you use a password manager. Here's why and how to get started with passkeys.
A new feature added to Comcast's newest routers can detect if there is motion inside your home without needing traditional motion sensors.
This is the latest large-scale DDoS attack to hit the social networking site this year.
arXiv:2608.14565v1 Announce Type: new Abstract: AI safety research has mainly focused on two areas: technical alignment (ensuring AI systems produce human-aligned outputs) and the regulation of generative AI's societal impacts (including unemployment risk and labor market disruption). However, an equally important dimension remains underexplored: the risk inherent in dependence on AI systems themselves. In this position paper, we argue that AI safety research should address AI Lock-In, the pheno
Every morning, a tech digest curated for you
The web shows the big picture; subscribers get their own — AI curated to your interests, your private RSS folded in, with community takes, delivered each morning. Free forever.
44 issues shipped · 150+ items sifted to 30 worth reading, every day