But robots.txt doesn't effectively address content shared across platforms, nor does it give individual creators a way to easily communicate consent when publishing on third-party sites or when others reuse their work. Today’s AI systems scrape vast amounts of content from the open web, including websites like Wikipedia, news outlets such as The Guardian and The New York Times (which is now suing OpenAI), public domain and pirated books, code from platforms like GitHub, and public forums like Reddit.