{"id":35033,"date":"2026-07-13T14:11:30","date_gmt":"2026-07-13T18:11:30","guid":{"rendered":"https:\/\/kapdec.com\/blog\/?p=35033"},"modified":"2026-08-09T12:50:53","modified_gmt":"2026-08-09T16:50:53","slug":"new-ai-auditing-technique-accurately-detects-models-adapted-for-harmful-content","status":"publish","type":"post","link":"https:\/\/kapdec.com\/blog\/new-ai-auditing-technique-accurately-detects-models-adapted-for-harmful-content\/","title":{"rendered":"New AI Auditing Technique Accurately Detects Models Adapted for Harmful Content"},"content":{"rendered":"<span class=\"span-reading-time rt-reading-time\" style=\"display: block;\"><span class=\"rt-label rt-prefix\">Reading Time: <\/span> <span class=\"rt-time\"> 2<\/span> <span class=\"rt-label rt-postfix\">minutes<\/span><\/span><p>The rapid rise in generative artificial intelligence (AI) technology has brought numerous open-source models online, allowing users to adapt them for tasks like creating product renderings in specific artistic styles. However, this accessibility also poses risks, as some individuals exploit these models to produce illegal content, such as hate speech or child sexual abuse material (CSAM). This issue is becoming increasingly severe; the National Center for Missing and Exploited Children recorded over 1.5 million reports of AI-generated CSAM in 2025, a significant rise from 67,000 cases in 2024.<\/p>\n<p>To combat this growing problem, a team of researchers from MIT and the nonprofit organization Thorn has developed a new auditing method to determine whether AI models have been adapted to produce CSAM. This technique focuses on analyzing the model&#8217;s internal mechanisms rather than generating outputs, thus avoiding legal and ethical complications associated with producing illegal content.<\/p>\n<p>Led by MIT graduate student Vinith Suriyakumar and associate professors Ashia Wilson and Marzyeh Ghassemi, in collaboration with Thorn, the researchers created a method that inspects the modifications made by an algorithm called low-rank adaptation (LoRA) during the model&#8217;s fine-tuning process. This technique, known as Gaussian probing, uses random data inputs to study the model&#8217;s internal structure changes, identifying potential harmful specializations without generating any outputs. Testing has shown the method to be 100% accurate in detecting models adapted to create CSAM.<\/p>\n<p>The innovative approach offers a scalable and cost-effective solution for platforms hosting open-source models, enabling them to identify and remove harmful adaptations swiftly. It also provides a new tool for law enforcement to address AI safety issues that have broad negative impacts.<\/p>\n<p>The study encompassed variations of three model types and compared results to known data from LoRA adaptors capable of generating harmful and safe content. The high accuracy of this technique suggests it could become a vital tool for preventing the dissemination of harmful AI-generated content.<\/p>\n<p>Future plans include expanding the evaluation of this method on a broader variety of model variations and investigating its ability to detect harmful capabilities in base models before they undergo adaptation. The researchers believe their work could have a significant impact on child safety and AI governance, addressing a critical issue affecting children globally.<\/p>\n<p>This research was partially funded by the Bridgewater AIA Labs Research Fellowship and was presented at the \u201cTrustworthy AI for Good\u201d workshop at the International Conference on Machine Learning.<\/p>\n<hr>\n<p>\n<strong>Source:<\/strong> MIT News<br \/>\n<strong>Read Original:<\/strong><br \/>\n<a href=\"https:\/\/news.mit.edu\/2026\/new-method-keeps-kids-safe-from-illegal-ai-generated-content-0713\" target=\"_blank\" rel=\"noopener\">https:\/\/news.mit.edu\/2026\/new-method-keeps-kids-safe-from-illegal-ai-generated-content-0713 <\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p><span class=\"span-reading-time rt-reading-time\" style=\"display: block;\"><span class=\"rt-label rt-prefix\">Reading Time: <\/span> <span class=\"rt-time\"> 2<\/span> <span class=\"rt-label rt-postfix\">minutes<\/span><\/span>Researchers developed an auditing technique to test generative AI models for malicious capabilities, without prompting them for illegal outputs.<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"om_disable_all_campaigns":false,"_monsterinsights_skip_tracking":false,"footnotes":""},"categories":[226],"tags":[484,826,572],"class_list":["post-35033","post","type-post","status-publish","format-standard","hentry","category-ai-digest","tag-ai-in-education","tag-smart-and-modern-learning","tag-stem-education"],"acf":[],"_links":{"self":[{"href":"https:\/\/kapdec.com\/blog\/wp-json\/wp\/v2\/posts\/35033","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/kapdec.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/kapdec.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/kapdec.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/kapdec.com\/blog\/wp-json\/wp\/v2\/comments?post=35033"}],"version-history":[{"count":1,"href":"https:\/\/kapdec.com\/blog\/wp-json\/wp\/v2\/posts\/35033\/revisions"}],"predecessor-version":[{"id":35483,"href":"https:\/\/kapdec.com\/blog\/wp-json\/wp\/v2\/posts\/35033\/revisions\/35483"}],"wp:attachment":[{"href":"https:\/\/kapdec.com\/blog\/wp-json\/wp\/v2\/media?parent=35033"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/kapdec.com\/blog\/wp-json\/wp\/v2\/categories?post=35033"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/kapdec.com\/blog\/wp-json\/wp\/v2\/tags?post=35033"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}