Large language models (LLMs) demonstrate remarkable prowess in areas such as code-writing, summarizing complex concepts, and even tackling mathematics problems. One question, however, arises amidst all these capabilities: have these models been programmed to ‘forget’? This question could hold substantial implications for issues relating to the General Data Privacy Regulation (GDPR).
The GDPR propagates the “right to be forgotten,” or the right of erasure, which seems difficult for developers of LLMs to comply with, except through the deletion of their entire models. Essentially, language models trained on publicly available personal data currently lack an mechanism for “forgetting” or erasing specific parts of their training set. This erasure might otherwise require the model to be completely retrained from scratch, a process which is practically and computationally extensive.
This quandary poses serious questions as to whether these AI models can indeed comply with the GDPR’s right to be forgotten. As we further advance in the field of AI and machine learning, we will have to contemplate whether there exists a machine ‘unlearning’ alternative, a path to make these language models GDPR-compliant.
For deeper insight into this pressing issue, read more in the detailed LegalTech News article.