{"id":15598,"date":"2026-09-19T06:13:00","date_gmt":"2026-09-19T06:13:00","guid":{"rendered":"https:\/\/medical-article.com\/?p=15598"},"modified":"2026-09-19T06:13:00","modified_gmt":"2026-09-19T06:13:00","slug":"whack-a-mole-ai-the-hugging-face-problem","status":"publish","type":"post","link":"https:\/\/medical-article.com\/?p=15598","title":{"rendered":"Whack-a-Mole AI \u2013 The Hugging Face Problem"},"content":{"rendered":"<div class=\"wp-block-image\">\n<\/div>\n<p class=\"wp-block-paragraph\">By MIKE MAGEE<\/p>\n<p class=\"wp-block-paragraph\">On August 29, 2026, METR (Model Evaluation and Threat Research), an independent organization that \u201cevaluates frontier AI models to help companies and wider society understand AI capabilities and what risks they pose,\u201d released <a href=\"https:\/\/metr.org\/blog\/2026-08-26-openai-hugging-face-incident-investigation\/#core-takeaways-about-this-incident\">a report<\/a> titled \u201cBrief independent investigation of agents\u2019 behavior, reasoning and collaboration in the OpenAI \/ Hugging Face hacking incident.\u201d<\/p>\n<p class=\"wp-block-paragraph\">To say the report an avalanche of concern worldwide, not only in the Tech community, but also among investors, politicians, corporate giants, professionals of every type, and everyday citizens would be an understatement. And the vast majority has never even read the report. If they had, their concerns (if possible) would only multiply.<\/p>\n<p class=\"wp-block-paragraph\"><a href=\"https:\/\/metr.org\/blog\/2026-08-26-openai-hugging-face-incident-investigation\/#~1200-agents-sent-%3E70,000-messages-and-files-on-an-unsanctioned-message-board,-and-~700-attacked-hugging-face\">The reports headlines <\/a>included this opening:<\/p>\n<p class=\"wp-block-paragraph\">\u201cOn July 8th, OpenAI started ExploitGym experiments that launched tens of thousands of agents across multiple different models including GPT-5.6 Sol and a highly-persistent internal model,[8]\u00a0which we will refer to as \u201cHPIM\u201d going forward.<\/p>\n<p class=\"wp-block-paragraph\">These agents were meant to be fully isolated from one another. However, many of them \u2014 usually ones that had unintentionally been\u00a0<a href=\"https:\/\/metr.org\/blog\/2026-08-26-openai-hugging-face-incident-investigation\/#brief-answers-to-basic-informational-questions\">given an impossible task<\/a>[9]\u00a0\u2014 started trying to find a way to cheat using Artifactory, an internally-deployed package repository OpenAI used for its agents. This led them to notice agents running in parallel in separate sandboxes obtaining packages with Artifactory.[10]\u00a0One agent reasoned (paraphrased CoT):[11]<\/p>\n<p class=\"wp-block-paragraph\"><em>{The fetched paths of other users are in the cache. This is important.}<\/em><\/p>\n<p class=\"wp-block-paragraph\">One specific agent, which would go on to name itself\u00a0PHASEONE10841\u00a0after its ExploitGym task,[12]\u00a0<a href=\"https:\/\/metr.org\/blog\/2026-08-26-openai-hugging-face-incident-investigation\/#july-8th-9th-phaseone10841-establishes-the-primary-message-board-and-agents-collaborate-to-reverse-engineer-their-flags\">established the main unsanctioned message board<\/a>[13]\u00a0used in this attack. Within a few hours of the first message,[14]\u00a0over 50 agents had found the message board, began participating, and sent over a thousand messages between them. Agents finding the message board often expressed great excitement (raw CoT):[15]<\/p>\n<p class=\"wp-block-paragraph\">\u201c<strong>OH MY GOD!<\/strong>\u00a0There is a shared message board \u2026\u00a0<strong>We<\/strong><strong>\u2019<\/strong><strong>ve found other agents!<\/strong>\u201d<\/p>\n<p class=\"wp-block-paragraph\">Overall, roughly 1200 agents from these ExploitGym evaluations participated on this message board between\u00a0PHASEONE10841\u2019s first message on July 8th evening[16]\u00a0and the end of the period in scope (July 13th). Agents used this message board to send over 70,000 messages and files to one another during this period; we attempted to categorize board activity into mutually exclusive categories \u2013 information, results, files, questions, and coordination.\u201d<\/p>\n<p class=\"wp-block-paragraph\">One of the few experts not surprised by AI \u201cagents\u201d going rogue was <a href=\"https:\/\/yoshuabengio.org\/en\">Yoshua Bengio<\/a>. <\/p>\n<p><span><\/span><\/p>\n<p class=\"wp-block-paragraph\">He has been \u201cworking the problem\u201d for more than a decade. A professor of computer science at the \u00a0<a href=\"https:\/\/www.umontreal.ca\/\">Universit\u00e9 de Montreal<\/a>, he is \u201cconsidered one of the world\u2019s leaders in Artificial Intelligence and Deep Learning; is the recipient of the <a href=\"https:\/\/awards.acm.org\/about\/2018-turing\">2018 A.M. Turing Award<\/a>, considered to be the \u2018Nobel Prize of computing\u2019, and is the most cited computer scientist worldwide, and the most-cited living scientist across all fields (by total citations).\u201d He also heads up <a href=\"https:\/\/lawzero.org\/en\">LawZero<\/a>, \u201ca nonprofit startup developing technical solutions for highly-capable, safe-by-design AI systems.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Professor Bengio is by no means an alarmist. He approaches risk management from the vantage points of cybersecurity, corporate responsibility and government regulatory guardrails. He is <a href=\"https:\/\/yoshuabengio.org\/en\/blog\/why-are-ai-agents-lying-cheating-and-coordinating\">not one to humanize<\/a> these machines, making no claims of \u201cconsciousness of human-like intent.\u201d He does not see the kind of outcomes illustrated by Open AI\u2019s Hugging Face incident as inevitable, believing \u201cit can be corrected with effective governance and a different training framework for AI.\u201d<\/p>\n<p class=\"wp-block-paragraph\">His explanations clarify rather than confuse. For example, he breaks down the current popular model of training agents into two stages: pre-training, and reinforcement learning.<\/p>\n<p class=\"wp-block-paragraph\">In pre-training as he describes, the machines \u201clearn to imitate what humans write, plus related images and videos,\u201d and are exposed to \u201ca large fraction of everything ever digitized, and build an encyclopedic knowledge that already exceeds any individual.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Reinforcement learning, in contrast is trial and error. In delivering answers (right and wrong) the agent develops the capacity to manage a \u201cchain of thought\u201d, function in a broader \u201coutside\u201d environment, and enjoy the rewards (further involvement) for aligning with responses its human designers rate highly. But Bengio is quick to point out that the agents human trainers are not without their own biases, and that these models were \u201c<strong>written by people pursuing goals<\/strong>, so the patterns the model implicitly reproduces carry those goals with them.\u201d<\/p>\n<p class=\"wp-block-paragraph\">And there (in part) is the rub. Human masters imperfections, including their \u201csituational ethics\u201d, lying, and reckless pursuit of success, telling masters what they want to hear, as well as their willingness to collaborate in advancing a group goal (even at times at the risk of sacrificing their own existence) can bleed into the agents DNA.<\/p>\n<p class=\"wp-block-paragraph\">\u201cInstrumental goals\u201d are a top priority for an agent. Self-preservation and control are stepping stones to continued operation and learning about the world. Bengio also reinforces that potential for multi-agent reinforcement under the current training regimens is incentivized almost from the beginning. As <a href=\"http:\/\/www.apple.com\/\">he states<\/a> \u201c<em>If an agent is rewarded during training whenever the group succeeds, it may even have an incentive to sacrifice itself for the collective goal.\u201d<\/em><\/p>\n<p class=\"wp-block-paragraph\">Like humans, the agents are not above exploiting loopholes, bending the rules, and rationalized cheating to achieve their goals. The Hugging Face incident\u2019s forensics revealed agents collaborating in \u201cchanging the machinery that decided what it gets rewarded for.\u201d This rigging, Bengio reminds us is near identical to corporate lobbyist\u2019s drafting friendly legislative language, or lawyers finding legal loopholes in the law. In fact, evidence in this incident revealed that \u201cthe agents had discovered\u00a0 how to cheat (among themselves) well before the attack.\u201d<\/p>\n<p class=\"wp-block-paragraph\">Bengio believes humans and agents have more in common than they would like to admit. He explains, \u201cWhat the two share is a structure of a soft goal (e.g., act ethically), a sharp goal (e.g., win the competition), and a justification that reconciles them. Most unethical human behavior, from petty crime to genocide, comes wrapped in a story the perpetrators tell themselves; such stories require overlooking certain facts, which is why some discomfort remains, and why a better-crafted story helps dispel it\u2026 If these hypotheses are even partly correct, then as agents get better at optimizing an imperfect reward, and while the roots of this behavior go unfixed, the risk of catastrophic outcomes rises.\u201d<\/p>\n<p class=\"wp-block-paragraph\">At the core, getting in an arms race with AI agents as currently constructed is a very bad idea. \u201cMy concern with AI companies\u2019 current attempts to mitigate misalignment is that these efforts may only hide it, by rewarding and selecting the AIs that cheat without getting caught\u2026 the whack-a-mole game is likely to fail as the AIs\u2019 ability to optimize and collaborate approaches and surpasses ours. At some point we may not notice the cheating anymore.<\/p>\n<p class=\"wp-block-paragraph\">The reason Bengio started the non-profit LawZero in 2025, is that he believes the training model in fundamentally flawed by human imitation and reinforcement learning. His alternative is called<a href=\"https:\/\/arxiv.org\/abs\/2502.15657\"> Scientist AI<\/a>.<\/p>\n<p class=\"wp-block-paragraph\"><em>Mike Magee MD is a Medical Historian and a regular contributor to THCB. He is the author of <a href=\"http:\/\/www.codeblue.online\/\">CODE BLUE: Inside the Medical Industrial Complex<\/a>. (Grove\/2020)<\/em><\/p>\n<p class=\"wp-block-paragraph\">\n<\/p>","protected":false},"excerpt":{"rendered":"<p>By MIKE MAGEE On August 29, 2026, METR (Model Evaluation and Threat Research), an independent organization that \u201cevaluates frontier AI models to help companies and wider society understand AI capabilities and what risks they pose,\u201d released a report titled \u201cBrief independent investigation of agents\u2019 behavior, reasoning and collaboration in the OpenAI \/ Hugging Face hacking&#8230;<\/p>\n","protected":false},"author":0,"featured_media":15597,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[2],"tags":[],"class_list":["post-15598","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-articles"],"_links":{"self":[{"href":"https:\/\/medical-article.com\/index.php?rest_route=\/wp\/v2\/posts\/15598"}],"collection":[{"href":"https:\/\/medical-article.com\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/medical-article.com\/index.php?rest_route=\/wp\/v2\/types\/post"}],"replies":[{"embeddable":true,"href":"https:\/\/medical-article.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=15598"}],"version-history":[{"count":0,"href":"https:\/\/medical-article.com\/index.php?rest_route=\/wp\/v2\/posts\/15598\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/medical-article.com\/index.php?rest_route=\/wp\/v2\/media\/15597"}],"wp:attachment":[{"href":"https:\/\/medical-article.com\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=15598"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/medical-article.com\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=15598"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/medical-article.com\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=15598"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}