Protecting your website content from AI
A client in the charity sector recently raised a query regarding how they could protect their website content and in particular photographs of young people being scraped by AI. They were wondering if there was a way to stop AI from taking screenshots or downloading the photos. One suggestion was to distort the photos if there was an AI crawler visiting the site.
My response
You cannot.
AI is incredibly efficient at scraping websites. Any sort of solution such as disabling downloads or distorting the photos would not only be technically complex but would also be quickly circumvented by any determined crawler.
Whatever’s on your website, assume it’s already been stolen.
Policy matters
The only real solution or possible navigation of this issue is to have policies in place with a basic workflow before the images are uploaded on your website. Some of the steps you can take include:
- Remove image metadata, such as location and date, before uploading images onto your website
- Have an internal policy that sets out guidelines on what should or shouldn’t be included in images
- Anyone who submits images (parents, museums, schools, etc) should have signed off on the privacy policy and terms and conditions
The very knowledgeable Mike Ellis directed me to a great section on the Kids in Museums privacy policy regarding AI:
Once photos are published online, we cannot guarantee that they will not be downloaded and/or misused by others. It is very easy for AI bots to scrape material from our website and social media channels. While Kids in Museums technically remains the data controller, there may be situations where it is close to impossible for the organisation to fulfil that role. This is true for any organisation that shares images publicly.
We will take steps to protect images of children and vulnerable adults, such as ensuring that no identifying information appears in the photo or is used in captions. This is a rapidly evolving area, and we will continue to review practical solutions to better protect images on our website.
The standout point is that this is a rapidly evolving area but efforts will be made to review practical solutions.
But what about robots.txt?
Theoretically you can use your websites robots.txt to deny access to AI crawlers which are scraping your site and gathering content.
This comes with its own caveats:
- You are presuming that AI has not already crawled and stolen all your content or images
- You are presuming the crawlers actually respect your robots.txt
- You are potentially blocking good traffic
Furthermore the number of AI crawlers increases daily so you are fighting an uphill battle trying to stop your data being visible.
Conclusion
Always take appropriate steps to protect people who may appear in images on your website as a pre-measure. Presume everything you publicly upload to your website can and will be stolen.
I’m a Leeds based Web Developer and Designer with 10+ years experience in building all manner of websites, from large-scale e-commerce platforms to sites for charities and non-profit organisations.
