You Wouldn’t Scrape the Internet to Make an LLM: Law and Policy of Scraping the Ago of AI

No ratings

Presented at Shmoocon 2024 by

AI projects are built on large language models, developed by ingesting massive amounts of data used to train on billions of parameters, often scraped from the internet. The growth of AI in the public consciousness has led to spirited debates about the legality and policy issues with gathering and processing this data–the application of privacy, free expression, copyright, and anti-hacking law. This intense focus on LLMs has led to a multitude of cases challenging the AI models that will shape the legal landscape, including online scans, searches and scrapes used for cybersecurity and threat intelligence, whether enhanced with AI or not. At the same time, security professionals may be charged with technically protecting privacy or personal data against being ingested, at least without permission. We will discuss the state of the law on online scraping, examining key legal cases that paved the way for search engines built on scraping–web, images, and books–and the cases that limited or allowed access to some private sites, as well as the open tech policy issues as the courts and Congress attempt to balance innovation, safety, and privacy. Finally, we will tie this together with the potential impact on cybersecurity.