AI glossary

General Data Protection Regulation (GDPR)

The General Data Protection Regulation (GDPR) is a European Union law that sets rules for how organizations collect, store, and process the personal data of people in the EU. It is formally Regulation (EU) 2016/679.

What GDPR is and when it applies

The European Parliament adopted GDPR on April 14, 2016, and it became enforceable on May 25, 2018, after a two-year transition period. Unlike many privacy laws that apply only to entities physically located within their jurisdiction, GDPR applies to any organization that processes the personal data of people in the EU, regardless of where the organization itself is based. This extraterritorial scope is why it affects AI products and companies well outside Europe.

For software developers and data scientists, this means that if your model processes user data to generate insights, recommendations, or decisions for EU residents, you are subject to these rules. The regulation covers both personal data and anonymized data that can be re-identified. It establishes strict guidelines on data minimization, purpose limitation, and user consent, forming the legal backbone for how modern AI systems must handle the individuals whose data trains them.

Article 22 and automated decisions

One of the most critical provisions for AI teams is GDPR’s Article 22. It gives individuals the right not to be subject to a decision based solely on automated processing, including profiling, when that decision produces legal effects concerning them or similarly significantly affects them.

This provision is directly relevant to AI systems used for things like automated loan approval, hiring screening, or insurance pricing, since those are exactly the kinds of significant, automated decisions Article 22 is aimed at. If a machine learning model decides to deny a credit application or flag a candidate for rejection without human oversight, it triggers this right.

There are exceptions, including when the automated decision is necessary for a contract with the person, is authorized by law, or is based on the person’s explicit consent; in those exception cases, the person still generally has the right to obtain human intervention, to express their point of view, and to contest the decision. This means your system cannot be entirely “black box” if it impacts livelihoods. You must have a mechanism to escalate the decision to a human reviewer who can override the algorithm.

The “right to explanation” debate

GDPR’s recitals and related articles (13, 14 and 15) require that people be given “meaningful information about the logic involved” in automated decisions covered by Article 22, which many commentators refer to informally as a “right to explanation,” although GDPR itself never uses that exact phrase.

Exactly how much detail “meaningful information about the logic” requires, especially for complex models like large neural networks, has been an ongoing subject of legal and academic debate rather than a single settled standard. For a simple linear regression model, explaining the logic is straightforward. For a deep learning model with millions of parameters, it is significantly harder.

Practically, this often leads teams to incorporate techniques from Explainable AI to generate feature importance scores or counterfactual explanations. The goal is not necessarily to explain the entire mathematical internal state of the model, but to provide enough context for the user to understand why the decision was made and whether it was reasonable.

Penalties

Non-compliance is not a minor administrative issue. GDPR allows fines of up to €20 million or 4% of a company’s total worldwide annual turnover from the preceding financial year, whichever amount is higher, for the most serious violations.

For large technology companies, 4% of global turnover can amount to billions of dollars. This financial risk is one reason AI ethics and governance are treated as engineering requirements, not just legal formalities. It is not just about avoiding a fine; it is about maintaining trust with users who are increasingly aware of how their data is being used by automated systems.

What it means for AI teams in practice

Because GDPR restricts fully automated significant decisions and requires some ability to explain automated outcomes, it is one of the reasons interpretability and human-review steps are commonly built into AI systems that make consequential decisions about individuals in the EU.

GDPR’s data minimization and purpose-limitation principles are also part of why some teams turn to techniques like synthetic data or federated learning, which can reduce how much real personal data needs to be centrally collected or stored for training. Synthetic data allows teams to train Machine learning models without exposing raw user records, while federated learning enables model training across decentralized devices holding local sample data, reducing the need to transfer personal data to a central server.

For technical product managers and engineers, this means designing systems with privacy by default. You need to decide where the human-in-the-loop sits, how you will document the logic for regulators, and how you will handle data deletion requests. Compliance is not a one-time audit; it is an ongoing engineering challenge that affects model architecture, data pipelines, and user interface design.

FAQ

Does GDPR apply to non-EU companies?

Yes. If you process the personal data of people in the EU, regardless of where your company is based, GDPR applies to you.

What is the “right to explanation” in GDPR?

It is an informal term for the requirement that individuals receive meaningful information about the logic involved in automated decisions, though the exact text of GDPR does not use this specific phrase.

What are the penalties for GDPR violations?

Fines can reach up to €20 million or 4% of a company’s total worldwide annual turnover, whichever is higher.