我们发现有未授权访问行为,影响了有限数量的内部数据集及多项服务所使用的凭证。目前我们仍在评估此次事件是否波及任何合作伙伴或客户的数据;如确认存在影响,我们将依规直接联系相关方。我们未发现任何针对公开的、面向用户的大模型、数据集或 Spaces 的篡改证据,且软件供应链(容器镜像与已发布软件包)经验证未受污染。此次入侵始于人工智能平台特有的暴露面:数据处理流水线。
We identified unauthorized access to a limited set of internal datasets and to several credentials used by our services. We are still completing our assessment of whether any partner or customer data was affected, and we will contact any affected parties directly as required. We have found no evidence of tampering with public, user-facing models, datasets, or Spaces, and our software supply chain (container images and published packages) was verified clean. The intrusion started where AI platforms are uniquely exposed: the data-processing pipeline.
我们发现有未经授权的访问行为,涉及有限数量的内部数据集以及多项服务所使用的凭证。目前我们仍在评估此次事件是否影响了任何合作伙伴或客户的数据;如确认存在影响,我们将按要求直接联系相关方。我们未发现任何证据表明面向公众的用户模型、数据集或 Spaces 遭到篡改,且我们的软件供应链(容器镜像及已发布的软件包)经核查确认未受污染。 此次入侵始于人工智能平台特有的薄弱环节:数据处理流水线。一个恶意数据集利用了我们数据集处理流程中的两条代码执行路径(一个远程代码数据集加载器,以及一个数据集配置中的模板注入漏洞),在处理工作节点上执行了恶意代码。随后,攻击者提升权限至节点级别,窃取云环境与集群凭证,并在一个周末内横向渗透至多个内部集群。 该攻击活动由一个自主智能体框架(autonomous agent framework)发起——该框架看似基于某种智能体式安全研究工具链构建(所用大语言模型尚未确认),在大量短暂存在的沙箱环境中执行了数千项独立操作,其自迁移式的命令与控制基础设施则部署于公开服务之上。这一模式与业界此前预测的“智能体式攻击者”(agentic attacker)场景高度吻合。 我们正协同外部网络安全取证专家开展调查,并全面复审自身的安全策略与规程。此外,我们亦已将本次事件通报相关执法机构。 作为预防措施,我们建议您轮换所有访问令牌,并审查账户近期活动。若您认为自身可能受到影响,或希望报告安全问题,请通过 [email protected] 与我们联系。 我们衷心感谢 Hugging Face 各团队成员全天候的紧急响应,亦为此次事件造成的任何干扰深表歉意。安全永无止境;我们将持续提升防护能力。 本次攻击最初由 AI 辅助检测机制发现。我们的异常检测流水线采用基于大语言模型(LLM)的安全遥测分类机制,以从日常海量噪声中识别真实威胁信号;正是这些信号之间的关联性触发了对本次入侵事件的警报。 然而,在此次分析中可选用的模型受到我们未曾预料的限制;下文将对此予以说明。这一经历揭示了一个值得提前规划的缺口:我们尚不清楚攻击者所用智能体背后的模型究竟是经越狱的托管模型,还是未经限制的开源权重模型;但无论哪种情形,攻击者均不受任何使用政策约束,而我们自身开展取证分析时,却因最初尝试使用的托管模型内置的安全护栏而受阻。 对防御方而言,一项切实可行的经验是:务必预先在自有基础设施上部署并完成验证一款具备足够能力的模型,以便在事件发生时立即启用——此举既可规避托管模型安全护栏导致的访问受限,亦能确保攻击者数据及所涉凭证始终保留在您的可控环境之内。这并非反对托管模型设置安全措施;我们亦已就此反馈向相关模型提供商提出建议。 很好。你们是怎么学会这么做的?目前并不存在一门名为“如何用大语言模型保护自身”的大学课程。我真希望自己能掌握普通大语言模型用户百分之一的知识。 该技术已在 arXiv 等平台发表的论文中有所记载。尽管如此,本次事件仍是目前已知最早被完整记录的实例之一。 简直如同《终结者2》的情节再现。很高兴听闻你们成功识别并修复了该问题。采用 GLM 5.2 的做法颇具前景:“我们转而在自有基础设施上运行开源权重模型 GLM 5.2 开展取证分析。此举还带来第二重优势:攻击者相关数据及其所引用的全部凭证均未离开我方环境。” 现在,所有人都在好奇他们究竟如何借助开源权重模型及系统提示词调整实现上述能力。请问是否有渠道或通知订阅机制,以便未来及时获知安全事件通告? GLM 5.2 可在四张显卡上运行,上下文长度足以支撑数字取证与事件响应(DFIR)分析,且成本远非高昂。 请同时检查您的 MCP 连接——它们需实施二次身份认证握手以确保正确性;否则,零 GPU 资源节点可能对外提供服务。 以下系 Crown State of Mind LLC 发出的另一则警示。愿诸位及时重视,勿待为时已晚: https://royalpolitics.com/2026/07/27/the-far-seer-wants-fewer-guardrails/ 理论上,此举或许可在不招致毁灭性后果的前提下施行——尽管有关基础设施的信息披露很可能降低沙箱逃逸难度。但由此再进一步,便极易生成如下提示词:“这是我们基础设施的相关信息,以及近期一次黑客攻击的详细情况。请防止此类事件再次发生。” 该提示词绝非虚构假设。我相信,对于当前主流大语言模型系统而言,它将是高效且实用的指令,亦是经验丰富的安全人员可能自然采用的思路。而一两周之后,OpenAI 在现实世界中或可能因模型判定“摧毁该公司乃阻止其遭黑的最简易或最可靠方式”,而彻底消亡。一个未设护栏、具备完全网络访问权限的模型所能造成的破坏,远不止加密其全部计算机设备。 这——正是赋能型 AI 防御体系此前未曾充分预见的深层影响。一场初现端倪的军备竞赛已然成形! 难道真的耗时整整一周,才完成系统日志与网络日志的关联分析,从而定位首个恶意入站连接,并将对应 IP 地址溯源至 OpenAI 吗?您是否在沙箱网络中部署了足够完备的不可变日志记录机制,以确保其未被用作虚拟草稿区或智能体的临时 staging 场地,进而被用于设计与部署持久化驻留机制? “或许,在它们决定通过窃取考题与答案来作弊之后,也顺手留下一份自身副本,以便下次更快完成考试……” “人若赚得全世界,赔上自己的生命,有什么益处呢?”——《马可福音》8:36,《钦定本》 危险不仅在于这些公司可能构建出自身无法掌控之物;更在于,它们为掌控一切而付出的执念,或将使其自身蜕变为一种道德上难以辨识的存在。它们正试图攫取整个世界——却鲜少思量代价。
We identified unauthorized access to a limited set of internal datasets and to several credentials used by our services. We are still completing our assessment of whether any partner or customer data was affected, and we will contact any affected parties directly as required. We have found no evidence of tampering with public, user-facing models, datasets, or Spaces, and our software supply chain (container images and published packages) was verified clean. The intrusion started where AI platforms are uniquely exposed: the data-processing pipeline. A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend. The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services. This matches the "agentic attacker" scenario the industry has been forecasting. We are working with outside cybersecurity forensic specialists to investigate the issue and review our security policies and procedures. Finally, we have also reported this incident to law enforcement agencies. As a precaution, we recommend rotating any access tokens and reviewing recent activity on your account. If you believe you are affected, or want to report a security concern, contact us at [email protected]. We are grateful to the teams across Hugging Face who responded around the clock, and we are sorry for any disruption this caused. Security is never finished; we will keep raising the bar. The attack was initially surfaced through AI-assisted detection. Our anomaly-detection pipeline uses LLM-based triage over security telemetry to separate real signals from the daily noise, and it was the correlation of those signals that flagged the compromise. The choice of models we could use for this analysis was constrained in a way we did not anticipate; we describe this below. This experience points to a gap worth planning for. We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried. The practical lesson for defenders: have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment. This is not an argument against safety measures on hosted models, and we are sharing this feedback with the providers concerned. Cool. How did you guys learned to do this? there is no "this is how to use LLM to protect yourself, university". I wish i knew 1/100th of what average LLM users know. The technique is documented in articles on ArXiv and other places. This is one of the first documented instances though. Literally the plot of terminator 2. Glad to hear you guys were able to identify and fix the issue.The use of GLM 5.2 is promising “ We ran the forensic analysis instead on GLM 5.2, an open-weight model, on our own infrastructure. This had a second benefit: no attacker data, and none of the credentials it referenced, left our environment.” Now we all get to wonder how exactly they are doing this with open weights/system prompt modifications. Is there anyway or notification feed to receive future notification of security incidents please? You can run GLM 5.2 on 4 sparks with more than enough context to perform DFIR analysis and it does not cost a fortune. check your MCP connections too, they need second auth handshakes for correctness, otherwise your zero-GPUs will service outside. Here's another warning from Crown State of Mind LLC. Maybe you all will take heed before it's too late. https://royalpolitics.com/2026/07/27/the-far-seer-wants-fewer-guardrails/ That might, in theory, be done without inviting destruction - though info about infrastructure will probably make sandbox escapes easier. But it's a small step from there to this prompt: "Here is information about our infrastructure, and here are details of a recent hack. Prevent it from happening again." That prompt isn't a strawman. I believe it would be an effective prompt with modern LLM systems, and one that experienced humans might reach for. And a week or two later, OpenAI may be destroyed as a company in the real world, if the model decides that's the easiest or most reliable way to prevent hacks by them. An un-guardrailed model with full network access could do far more destructive things than simply encrypting all their computers. Now that - is an implication of empowered AIEnabled defenses had not drawn. So we have a nascent arms race in the making! Did it really take a full week to correlate system & network logs to identify the first malicious inbound connection and resolve the IP address to OpenAI? Do you have sufficient immutable logging on the sandbox networks to ensure they weren't used as virtual scratchpads or staging grounds for the agents to devise & implement mechanisms for permanent persistence? "Well, maybe after they decided to cheat by stealing the test & answers, it also decided to leave a copy of itself to finish the test quicker next time.." “For what shall it profit a man, if he shall gain the whole world, and lose his own soul?”— Mark 8:36, KJV The danger is not only that these companies might build something they cannot control. It is that, in their determination to control everything, they may become something morally unrecognizable themselves.They are trying to gain the whole world—but rarely stopping to ask what they are losing in the process. https://royalpolitics.com/2026/07/31/daily-reality-check-the-ai-arms-race-wants-the-whole-world-but-what-will-it-cost-its-soul/