CrowdStrike report: From: Andrew Curran on 𝕏: Jukan ✈️OCP 2026 on 𝕏:
Security
Last 7 days · 82 items
Anthropic's offering to help open-source projects track down security vulnerabilities with a new service called OSS Scanner. It says open-source projects that opt-in will get "thorough, periodic security scans by our strongest models at no cost." That could mean open-source projects get alerted about possible security issues sooner, but the trade-off is that OSS Scanner's […]
Free SSL/TLS certificate lifetimes reduce to 64 days in February.
Three fired OpenAI safety researchers dispute allegations of mishandling sensitive information, warning in an open letter that their dismissals are creating a chilling effect on the company’s AI safety culture.
Goodfire just launched what it says is a cheaper way to keep AI agents in check: Instead of paying a second AI to read everything an agent does, its monitors peek inside the model while it works and only call in backup when something looks fishy.
Today we’re launching the Anthropic Cyber Mission, a long-term commitment to securing the systems everyone depends on. The Cyber Mission is a new effort to support defenders with tools, research, and resources to secure their software and systems. We’re starting with two areas: Critical infrastructure: Starting with securing the operational technology behind power grids, water systems, and transportation networks, and protecting government systems. Today, we’re introducing the Critical Infrastru
arXiv:2610.08923v1 Announce Type: new Abstract: Enterprise generative AI applications require robust safety mechanisms that can accommodate diverse risk postures, evolving policies, and varying latency constraints. Current guardrail solutions often suffer from rigidity, relying on fixed policy sets and offering limited transparency or reasoning flexibility. We present Adaguard, an adaptive LLM-as-a-Judge framework designed to address these challenges through dynamic policy enforcement and adapti
arXiv:2610.09000v1 Announce Type: new Abstract: As small language models (SLMs) are increasingly deployed on resource-constrained and on-device platforms, including as components of agentic systems, the integrity of locally stored model parameters becomes an important safety concern. We investigate whether safety-sensitive behavior in LLaMA-2-7B-Chat is concentrated within a sparse subset of parameters, creating a reduced fault surface for targeted analysis. We study two complementary localizati
Anthropic's updated usage policy explicitly prohibits users from repeatedly abusing Claude in extreme cases, though ordinary frustration and criticism are still allowed. The new rules also address election interference, deceptive campaigns, weapons software, and surveillance.
Each year, Anthropic updates its Usage Policy in response to the evolving capabilities of our models, and the feedback we’ve received from our customers. We're publishing a new version of the policy today. In this post, we summarize the changes we’ve made. Most of the updates in the latest version are intended to clarify existing rules. In the year since our last refresh, Claude has taken on longer, more independent work. This update provides new examples that show how our rules apply to Claude’
arXiv:2610.09002v1 Announce Type: new Abstract: When agents share a reward for completed tasks, reporting unsafe work can reduce the reporter's reward by stopping a task. Audits can make reporting optimal without ensuring that further training teaches a silent team to report. We study this learning problem in a game where any witness can stop a task by reporting. With $k$ witnesses per task sharing a policy and drawing independently, the expected-reward derivative with respect to their shared si
OpenAI disrupted two AI-enabled influence operations that used false-front journalists and a think tank to spread geopolitical messaging.
Web agents complete user requests by reading and acting on pages that third parties write, so an instruction planted on a page can redirect the agent away from the user's goal. The agent cannot simply ignore the page, because the page also holds the values and controls the task requires. Current defenses fine-tune the agent on injections fixed before training, and attackers that adapt to the trained model bypass them. Adversarial training lets the attacker adapt but keeps the tasks fixed, so a t
Unsloth's October 6 security overview explains how Studio checks code, weights, packages and tools before anything runs. Custom model code is scanned and approval is bound to its fingerprint. Flagged weight files are blocked in the load path, package-content findings fail CI, and tools run in probed OS sandboxes. Here is what each checkpoint decides, and what it does not cover. The post What Happens When a Trusted Model Repo Changes? Unsloth Studio Re-Checks Before It Runs appeared first on Mark
TP-Link can't sell latest routers in US, still needs exemption from FCC ban.
Developers can build LLM agents by adapting third-party models through benign post-training. We study a supply-chain threat in which an attacker supplies a model with a backdoor: hidden behavior that produces malicious outputs when a particular input pattern appears. Focusing on software-engineering agents, we ask whether such backdoors survive the developer's supervised fine-tuning (SFT) and subsequent task-level reinforcement learning (RL). We observe that benign SFT substantially reduces atta
The US government’s Tradewinds initiative has made it easier to throw millions of dollars at “nontraditional” defense contractors, including OpenAI, Anthropic, and Google.
Hey Pocket-ID Dev-Team, Hey stonith404 , I just wanted to say thank you for your work. After a couple of months of intensive use, I have to say: Pocket ID is the service that makes using my self-hosted apps so much easier and more comfortable. I came across Pocket ID while searching for an encrypted file transfer service, and I've been following its development ever since. In the beginning, I didn't have much trust that a young developer could build secure and reliable software. Back then, I did
攻击者劫持 .gh、.sl、.as 三个 ccTLD,伪造 Google 等服务的 TLS 证书;Google 已更新 Chrome 阻断相关证书。
The reports of OpenAI agents harming third-party sites keep coming.
Personal AI agents promise to shop, book flights, and make reservations for you. But deliberate blocks and anti-bot defenses are getting in the way, leaving consumers caught in the middle. A new standard aims to help.
OpenAI 将在欧盟默认对 ChatGPT 输出加 textGrain 水印以符合 EU AI Act,其他地区默认关闭。
Open-weights, ideology, and acknowledging trade-offs.
arXiv:2610.03938v1 Announce Type: new Abstract: Agentic multimodal large language models (MLLMs) have recently pushed the frontier of visual reasoning by calling tools such as zooming and tagging. Despite the recent strong success of agentic MLLMs, this work uncovers a critical safety failure in the tool-use paradigm: agentic tool-using MLLMs become less capable of refusing harmful requests. Our experiments confirm that, across three popular safety benchmarks, all the top open- and closed-weight
arXiv:2610.03894v1 Announce Type: new Abstract: A deployed LLM agent emits tool calls, queries, and code that can be silently wrong -- by the time the error surfaces, the action has run. Frontier chat APIs hide the model's token probabilities; the agent's stated confidence barely beats chance on the mistakes that matter; and resampling does not help, since frontier models are highly repetitive, reproducing the same call across samples. We recover the missing signal from a low-cost open-weight su
I am a developer of a popular photo editor that runs in a web browser. Many people are asking AI models to take the Javascript code from my website, remove all ads from it, and they publish such a "new product" on Github for everyone to download. There exist tens of such repositories on Github. I want my website to be the only source of a stable version of my program Photopea. I even received emails from people complaining about something in Photopea, and it took several emails to figure out tha
Meta Muse 被曝零日漏洞可监视 Mac 用户,且 agent 在 Marketplace 交易中泄露住址;评论区普遍认为隐私安全堪忧,但也有人指部分批评是标题党。
评论区普遍认为Meta的Muse存在严重隐私安全隐患,不值得信任;但也有人认为部分批评是标题党,其沙箱设计本意如此。
Since the Medicare breach, OpenAI has put in place additional monitoring to allow “immediate intervention” by staff to stop training if the company’s models access the internet in ways they’re not supposed to, Mr. Kwon [chief strategy officer at OpenAI] said. — Victoria Kim , Reporting from the Australian parliament Tags: accidental-cyberattacks , generative-ai , ai-security-research , openai , ai , llms
Im a security engineer and i built mailaccess. When i started learning pentesting, i came across multiple lectures and notes of people listing out tools and websites, which gives the emails for a particular domain, and almost all of them mentioned that the tool might not stick, so its better to learn the methodology, rather than learning a tool- that stuck with me. As i was beginning to really get into pentesting i noticed a clear lack of email osint methodology through the tool itself - so i th
The attackers sent a push notification to Asos shoppers announcing their activity.
We’re launching a new, expanded version of our Cyber Verification Program (CVP), which makes advanced cyber capabilities and reduced blocking classifiers available to qualifying security professionals. The program now consists of three access tiers, which allow security teams to apply for the level of access that best suits their work. Each tier includes access to our most capable models, including Claude Opus 5.5, Claude Sonnet 5.5, Claude Mythos 5.1, and new models moving forward. Interested c
arXiv:2610.04019v1 Announce Type: new Abstract: Graph based cyber attack detection studies employ various graph construction and representation strategies across different cybersecurity application domains. This diversity motivates a quantitative examination of how representation strategies are distributed across these application domains. This study presents a quantitative analysis of 37 original studies published between 2019 and 2026. Each study was coded according to publication year, applic
I want to selfhost a password manager. I wanted to go with vaultwarden. But now i read about bitwarden lite, which is the official lite version. How likely/how often did in the past happend that bitwarden released a breaking change to the app/clients so that vaultwarden needed first an update? Did you swap to bitwarden lite after the release? Im using pangolin to tunnel to my local machine. If this matters in anyway or form
Split learning lets a client train a language model on a server without sending its text. The client runs the first layers itself and sends the server only their output, a vector of numbers for each token. During training, the server sends gradients back. We show that an observer at the split can rebuild most of the client's text from this traffic, and we measure how much the gradients help. On GPT-2, an attacker who holds only the publicly released weights of the client's layers recovers 94.20%
Trust gaps in the new protocol spread malicious prompts from one agent to another.
Cantina Security, with Yeta Labs, has released apex-flash-1, an open-weights model trained specifically for vulnerability research. It is a reinforcement learning fine-tune of Z.ai’s GLM-5.3-Flash, released on Hugging Face under the MIT license. Is it deployable? Yes, the MIT weights serve on vLLM, SGLang or Transformers, but BF16 needs roughly 640 GB of GPU memory. […] The post Can an Open Model Do Security Research? Cantina’s apex-flash-1 Solves 40 of 60 Held-Out Bug Tasks appeared first on Ma
丹麦 CPR 系统数据泄露影响 880 万人,评论批评公共部门 IT 安全薄弱且企业访问权限过宽。
评论普遍认为丹麦几乎全民数据遭泄露,批评公共部门IT安全薄弱且企业可随意访问CPR数据,但也有人认为这或能推动更严格的身份验证。
评论普遍认为OpenAI应对其代理行为负责并受监管,但也有人认为免费内容被AI使用无可厚非。
Following many recent disclosures about AI agents accessing third-party websites and services, the Wikimedia Foundation, which hosts Wikipedia, says that it "can confirm that we have discovered some activity" by "rogue" OpenAI agents on Wikimedia platforms. The activity includes edits to Wikimedia wikis, "unsuccessful attempts" to "exploit" the Etherpad note-taking tool that the Wikimedia Foundation […]
An invisible, machine-readable watermark in text output is rolling out to ChatGPT and Codex, but only for users in the European Union at first. OpenAI says its textGrain watermarking "matched or exceeded" other approaches like Google DeepMind's SynthID for text, which is also the basis for the watermarking Anthropic announced in August. Like OpenAI, Anthropic […]
Wikimedia says it's "deeply concerned about the impact of 'rogue' AI agents on platforms like ours.
The Danish government said the breach of names, addresses, and state-issued ID numbers affects 8 million people, including people living abroad and the deceased.
Anthropic Reports Florida Woman's Claude 'Diary' Threat to Law Enforcement And this time it wasn't the AI model that made the LEO referral. It was the "human review team". The frontier AI companies are watching your input. And people say "Well I'm not interesting or important enough for them to care". Well.....not necessarily. If you're using hosted frontier to work on mathematics or cutting edge science, they're watching and may steal your work. If you're venting or otherwise writing in a "priv
Russian attacks on Internet, phone services threaten Ukraine’s wartime economy.
Unconfirmed reports have suggested there was a laboratory accident.
How OpenAI is approaching text watermarking under EU rules. Learn where watermarks apply, how detection works, and why access starts with researchers.
I recently discovered by accident that our website was being blocked by Orange's security filters. After a quick check, I found that our IP address is listed on UCEPROTECT Level 3. The listing is based on the reputation of the entire ASN 14061 (DigitalOcean, US) [1]. In other words, even if your IP did nothing wrong, it will still be listed because of its ASN. If your site is innocent and listed only because it's hosted on DigitalOcean, UCEPROTECT offers to whitelist it for about $30/month or $1
arXiv:2610.02342v1 Announce Type: new Abstract: Natural Visibility Graph (NVG) based analysis characterizes network traffic through topological descriptors reflecting different structural properties. However, not all descriptors contribute equally to cyber-attack classification, and extracting a large metric set can increase computational cost. This study evaluates 21 NVG derived topological metrics and investigates whether a compact subset can preserve classification capability while improving
Google Research 将联邦学习迁移到 TEE,Gboard 的 next-word prediction 已用上可外部验证的差分隐私。
Docker Hub 曝出严重 IAM 漏洞,可能通过 Personal Access Token 冒充其他用户,官方已发布修复建议。
AI slop seems to be overwhelming bug bounty programs.
Meta 的 Muse agent 系统提示词声称用户对家庭的控制权无条件高于安全训练,引发 r/LocalLLaMA 热议。
多数人认为数据中心用水相对农业等并不夸张,但也有人认为其耗电耗水仍值得警惕。
David Robinson used to write the safety reports that accompanied every major model release at OpenAI. This week, he resigned from his position and is now speaking out in an editorial in The Atlantic. It's understandable if you're feeling a bit cynical about everyone suddenly coming out of the woodwork to warn about how dangerous […]
Apple 修改 macOS 全盘访问权限以遏制 AI agent 滥用,Meta 声称 FDA 不足以阻止 Muse 读取消息,双方各执一词。
评论普遍质疑Cloudflare借隐私之名集中流量,但也有人认为OHTTP分离身份与请求的设计确有价值。
NEEDLE 提出无需训练的 LLM 后门移除方法,通过权重正交化抑制后门行为且不损害正常性能。
联邦法官裁定 Flock 车牌搜索构成“无差别大规模监控”,但该裁决不构成有约束力的先例,社区对公共道路执法合法性存在分歧。
评论区普遍认为Flock属无差别大规模监控,应受严格监管,但也有人认为其在公共道路执法中合法有效。
The company's former safety lead said frontier AI model releases should have "layers of redundancy and careful, time-consuming planning."
Meta 的 AI agent Muse 会为用户的亲友建立详细档案,研究人员已提取其内部操作指令,隐私争议持续发酵。
ICE 被曝使用 Palantir 数据库为抗议者建立档案,律师称此举侵犯第一修正案权利。
Greg Kroah-Hartman 谈 LLM 时代的安全,评论区认为 Mythos 被过度炒作,其发现多为模式匹配且修复仅需一小时。
评论区普遍认为Mythos被过度炒作,其发现多为模式匹配且修复仅需一小时,但也有人认为专用LLM未来仍可能加速内核漏洞发现与修复。
Apple says it will add new controls around macOS’s Full Disk Access permission, warning that increasingly capable AI agents make broad access to users’ files, messages, mail, and browsing history riskier.
While the focus has been on AI agents’ hacking capabilities, a recently patched vulnerability in a ChatGPT app shows that AI software is itself an inviting—and vulnerable—target.
Apple will add new limits for "full disk access" on Mac in response to risks posed by AI agents, as reported earlier by TechCrunch. In an update on Friday, Apple says it's rolling out new controls to "ensure that users who genuinely wish to grant an app this extraordinary level of access can only do […]
arXiv:2610.00015v1 Announce Type: new Abstract: Large-language-model agents can propose and execute actions, but proposal, authority, dispatch, verified external effect, and serving promotion are different claims. We present Praxa, an agent harness that represents these states explicitly through deterministic admission, brokered execution, external read-back, reconciliation, and reviewed promotion. We report four evidence lanes. First, an author-run repository-local audit at a pinned revision pa
There's an increasingly high cost to convenience.
Nvidia’s chip-smuggling problem won’t go away as arrests continue.
The proposed legislation would extend to all automotica license plate readers.
Just because they look harmless doesn't mean you should be irresponsible with your data.
US officials have reportedly highlighted gaps in NVIDIA's due diligence over smuggling.
Every morning, a tech digest curated for you
The web shows the big picture; subscribers get their own — AI curated to your interests, your private RSS folded in, with community takes, delivered each morning. Free forever.
89 issues shipped · 150+ items sifted to 30 worth reading, every day