AI, Web Scraping and Personal Data: New European Guidance Raises Questions for Danish Companies 

Written by

in

European data protection regulators are sharpening their focus on one of the most important questions surrounding artificial intelligence: what happens when companies collect huge amounts of information from the internet to develop or train AI systems? 

On 8 July 2026, the European Data Protection Board, known as the EDPB, adopted new guidelines on web scraping in the context of generative AI. At the same time, it adopted separate guidance intended to clarify when information can genuinely be considered anonymous. 

The developments matter for Danish companies developing AI products, buying AI services or using large datasets. Information being publicly accessible online does not automatically mean that businesses are free to collect and reuse it without considering the General Data Protection Regulation, or GDPR. 

For companies following technology and data protection developments through Lead Roedl, the new guidance highlights an increasingly important intersection between AI development, privacy law and ordinary business compliance. 

What Is Web Scraping? 

Web scraping is the automated extraction of information from websites and other online sources. 

Businesses have used scraping technologies for years. They can collect product prices, market information, public records, reviews and other online material. 

Generative AI has dramatically increased the importance of the practice. 

AI models can require enormous amounts of training information. Developers may therefore use automated tools to collect text, images and other material available across the internet. 

The difficulty arises when that material contains information about identifiable people. 

A website may publicly display someone’s name, photograph, professional history or comments. Social networks and forums can contain even more personal information. 

The fact that someone can view that information online does not automatically remove it from GDPR protection. 

GDPR Can Apply to Web Scraping 

The EDPB makes an important point in its new guidance: GDPR applies to web scraping when the activity involves processing personal data. 

Processing can include activities such as: 

  • Collecting personal information 
  • Storing scraped information 
  • Organising datasets 
  • Retrieving information 
  • Using personal data for AI training 
  • Combining information from multiple sources 

This means companies cannot assume that publicly available information is automatically available for unrestricted AI use. 

Instead, the organisation needs to identify whether personal data is involved and, if so, determine how GDPR requirements apply. 

Companies Need a Lawful Basis 

One of the fundamental GDPR requirements is that processing personal data needs a lawful basis. 

For AI developers, legitimate interests may sometimes be considered as a possible legal basis for web scraping. 

But legitimate interest is not an automatic permission. 

An organisation relying on it generally needs to identify a legitimate interest, show that processing is necessary for that interest and balance its interests against the rights and freedoms of the people whose information is being processed. 

The context can make a major difference. 

Scraping professional information from a corporate website may create a different privacy impact from collecting highly personal posts from a health-related forum. 

Companies should therefore examine what information they collect, where it comes from and what individuals could reasonably expect to happen to it. 

Publicly Available Does Not Mean Risk-Free 

This distinction may be one of the most important lessons for businesses. 

Internet users often publish information for a particular purpose. 

A person might place their professional details online so potential employers can find them. Someone might participate in a public discussion because they want to communicate with other members of a community. 

Neither situation necessarily means that the individual expects the information to become part of a large AI training dataset. 

The EDPB’s approach encourages companies to consider the original context and purpose of the information rather than treating the public internet as an unrestricted data source. 

For Danish businesses developing AI systems, the source of training information therefore deserves careful attention. 

Special Categories of Data Create Greater Risk 

Web scraping becomes even more complicated when sensitive information is involved. 

Under GDPR, certain information receives additional protection. These special categories can include data revealing: 

  • Racial or ethnic origin 
  • Political opinions 
  • Religious or philosophical beliefs 
  • Trade union membership 
  • Genetic information 
  • Biometric information used for identification 
  • Health information 
  • Information concerning a person’s sex life or sexual orientation 

Processing these categories is generally prohibited unless a specific exception under Article 9 of GDPR applies. 

The EDPB emphasises that there is no general AI or web scraping exemption from these requirements. 

If special-category information is scraped, a business needs both an appropriate lawful basis under Article 6 and a valid Article 9 condition. 

That can make indiscriminate collection particularly risky. 

Data Minimisation Matters for AI Training 

Another important GDPR principle is data minimisation. 

Companies should collect personal data that is adequate, relevant and limited to what is necessary for the intended purpose. 

This principle can appear difficult to reconcile with the way some AI systems are developed. 

A common assumption in machine learning has been that more training data is better. GDPR pushes organisations to ask a different question: 

Do we actually need all of this personal information? 

The EDPB recommends measures designed to limit unnecessary collection. 

Depending on the project, companies could consider excluding particular websites, categories of information or data sources that create excessive privacy risks. 

The goal is not simply to collect as much information as technology makes possible. 

Accuracy Also Matters When Scraping the Web 

Online information is not always accurate. 

Profiles become outdated. Websites repeat incorrect information. Posts can be misleading. Information can be copied from one website to another without verification. 

If inaccurate information enters an AI training dataset, the problem can spread further. 

The EDPB therefore recommends that organisations scraping information use reliable sources, record when information was collected and validate data before using it for AI training. 

This connects web scraping to another fundamental GDPR principle: personal data should be accurate and, where necessary, kept up to date. 

For companies, documenting where training information originated can therefore be valuable from both a technical and compliance perspective. 

Transparency Can Be Difficult, But It Still Matters 

Imagine an AI developer scraping information concerning millions of people from thousands of websites. 

Personally contacting every individual could be extremely difficult. 

GDPR recognises that there can be circumstances where providing information individually is impossible or would involve disproportionate effort. 

However, that does not mean transparency can simply be ignored. 

Businesses still need to consider appropriate ways of explaining their processing. 

This could include clear public privacy information describing: 

  • What types of information are collected 
  • Where information comes from 
  • Why it is collected 
  • How it is used for AI 
  • How long it is retained 
  • What rights individuals have 
  • How individuals can object or request action 

The precise requirements depend on the circumstances, but transparency should be built into the project rather than considered only after scraping has taken place. 

Anonymisation Is Not Simply Removing a Name 

The EDPB’s separate 2026 anonymisation guidance is also important for companies working with AI datasets. 

Businesses sometimes assume that deleting names, email addresses or identification numbers automatically makes information anonymous. 

The legal position is more demanding. 

According to the EDPB, information is anonymous when it does not relate to an identified or identifiable natural person. 

Whether someone remains identifiable depends on the context and the means reasonably likely to be used to distinguish that individual. 

The EDPB proposes three practical criteria for assessing anonymisation: 

  • No record isolation: an individual should not be singled out within the dataset 
  • No linkage: records should not be capable of being linked to information about the same person 
  • No inference: information about an individual should not be capable of being inferred 

If all three criteria are satisfied, the EDPB says the information can safely be considered anonymous. 

If one or more are not satisfied, further analysis is required. 

Pseudonymous and Anonymous Data Are Different 

This distinction is particularly important for businesses. 

Replacing a person’s name with a code or identifier may reduce privacy risks, but the resulting information can still be personal data. 

If somebody can reconnect the code with the person using additional information, the dataset may be pseudonymous rather than anonymous. 

GDPR generally continues to apply to pseudonymous personal data. 

True anonymisation can place information outside GDPR because it no longer relates to an identifiable individual. 

Companies should therefore avoid describing datasets as anonymous without examining whether individuals could realistically be identified through linkage, inference or other available information. 

What Should Danish Companies Review? 

The new guidance provides a useful reason for businesses to examine how they obtain and use information for AI. 

Companies involved in AI development could consider: 

  • Identifying which datasets contain personal information 
  • Recording where scraped information comes from 
  • Establishing a lawful basis for processing 
  • Reviewing whether sensitive information could be collected 
  • Limiting collection to information genuinely needed 
  • Assessing whether particular websites should be excluded 
  • Reviewing transparency notices 
  • Providing appropriate ways for individuals to exercise their rights 
  • Testing whether supposedly anonymous information can be reidentified 
  • Documenting decisions concerning legitimate interests and anonymisation 

Businesses buying AI technology from external suppliers also have questions to ask. 

They may want to understand how providers obtained training data, how personal information is handled and what contractual protections apply. 

AI Compliance Is Becoming a Wider Business Issue 

The EDPB’s 2026 guidance demonstrates that AI regulation does not exist in isolation. 

A company might comply with requirements under the EU AI Act while still facing separate obligations under GDPR. 

Copyright, employment law, consumer protection, contractual obligations and cybersecurity requirements may also be relevant depending on the technology and how it is used. 

For businesses following European regulatory developments through Lead Roedl, this overlap is becoming increasingly important. 

Companies developing or deploying AI need to consider not only whether the technology works, but also where its information comes from and whether that information has been processed lawfully. 

The new EDPB guidance does not mean that web scraping for AI is automatically prohibited. 

It does mean that large-scale collection should not be treated as legally invisible simply because the information was available on the internet. 

For Danish companies, one of the most useful questions to ask before building or buying an AI system may therefore be a very simple one: 

Where did the data come from? 

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *