25 August 2026

On 20 July 2026, the Personal Data Protection Commission (“PDPC”) published the Advisory Guidelines on Use of Personal Data in Generative AI (“Guidelines”). The Guidelines clarify how the Personal Data Protection Act 2012 (“PDPA”) applies in the context of organisations incorporating generative artificial intelligence (“GenAI”) into their work processes and provide guidance on the use of personal data in GenAI models and systems.

Organisations can refer to these Guidelines to better understand how to responsibly collect and use personal data for GenAI development, particularly in scenarios involving web-scraping and the re-use of data originally provided for purposes other than GenAI. The Guidelines also set out the data responsibilities of key GenAI stakeholders and provide guidance on how organisations should respond to individuals’ requests regarding the processing of their personal data for GenAI.

The Guidelines build on and should be read together with the Advisory Guidelines on Use of Personal Data in AI Recommendation and Decision Systems and Advisory Guidelines on Key Concepts in the PDPA.

The Guidelines are organised according to the typical stages of the GenAI lifecycle. Key recommendations are set out below.

Development: Collecting and using personal data to develop GenAI models

Publicly Available Exception

Where personal data forms part of online data that is publicly accessible without any restrictions, organisations can rely on the Publicly Available Exception under the PDPA, in lieu of seeking consent, to web-scrape and collect the data to develop a GenAI model.

However, where personal data forms part of online data that is placed behind a digital barrier, such as paywalls and authentication mechanisms, organisations must assess that the data is publicly available before they can rely on the Publicly Available Exception. Relevant factors to be considered include:

  • the purpose of the digital barrier (e.g. to enable data monetisation);
  • the effect of the digital barrier (e.g. whether the online data remains accessible to the public at large or only a specific group of persons);
  • the steps needed (e.g. number and complexity) to access the personal data; and
  • whether the personal data can be accessed without any restrictions from other online sources.

Examples of what PDPC considers to be publicly available data even though placed behind digital barriers can be found in the Guidelines.

An organisation that relies on the Publicly Available Exception even though it is reasonably arguable that personal data behind a digital barrier is not publicly available must explain its assessment and reasoning by way of a Data Protection Impact Assessment or other written record(s). This documentation must be produced by the organisation if it is required by PDPC to justify its reliance on the Publicly Available Exception. This ensures that organisations properly assess whether data is publicly available and keep proper records.

Organisations and/or individuals who make personal data available online but do not intend, or have not obtained consent, for their personal data to be web-scraped should implement appropriate digital barriers.

Consent and notification obligations

A key source of data used to develop GenAI models is personal data provided by an individual to an organisation, or personal data about an individual created in the course of or as a result of the individual’s use of the organisation’s products or services (“User Data”).

Where User Data is used for GenAI model development, and exceptions to consent (such as the business improvement and research exceptions) do not apply, organisations must obtain consent by providing “AI-Specific Notifications” that articulate the purpose of use, what data will be used and how the data will be used, and how individuals can decline or withdraw consent.

Anonymisation of data

To minimise unnecessary risks, PDPC encourages organisations to anonymise their datasets as much as possible and practise data minimisation when developing GenAI models.

Deployment: Processing personal data in deployed GenAI models and/or systems

The Guidelines clarify the key roles and responsibilities of the following GenAI stakeholders in safeguarding personal data:

  • Model Providers who develop and make available GenAI models for distribution and use. Model Providers must comply with all PDPA obligations when they process data to develop and deploy GenAI models. This extends to situations where Model Providers collect and use personal data from downstream systems for model development. Where Model Providers preserve training data which include identifiable personal data to develop future models, they are encouraged to develop and make available a data retention policy that explains the rationale for retaining data for a longer period of time. The personal data should be reviewed regularly to determine if it is still needed. When processing data on behalf of downstream users, Model Providers are encouraged to document and share information on model-level safeguards.
  • System Providers who are engaged to develop bespoke, customisable GenAI systems or who develop and retail commercial off-the-shelf systems. Where such System Providers process personal data as part of their own datasets to develop systems, they are organisations. Where they process data on behalf of downstream deployers (e.g. to customise systems for specific use cases, in delivering their Software as a Service, they are data intermediaries. In their role as data intermediaries, System Providers are expected to periodically review the need for additional security arrangements as they develop and make available new types of systems, and they are encouraged to share information on the system-level safeguards they have implemented with downstream deployers.
  • System Deployers who use and/or enable the use of GenAI systems under their authority. System Deployers bear primary responsibility for ensuring that the GenAI systems they have chosen to use meet their obligations under the PDPA and that the personal data processed by their chosen system is for relevant purposes. System Deployers must also safeguard personal data in their possession or under their control, including new categories of data sources that are collected through their systems, and they are encouraged to develop clear written policies and document processes in relation to the safeguards adopted. Further, System Deployers should regularly review the sufficiency of the safeguards, especially for agentic AI systems.

Post-deployment: Addressing individuals’ requests about personal data

Individuals can request to access and correct their personal data even after it has been used in GenAI development (“Access and Correction Obligations”).

Personal data under an organisation’s control includes data transferred to a data intermediary and organisations must take such data into account when responding to access and correction requests. Data intermediaries may facilitate requests or forward them to controlling organisations, but they are not obliged to do so under the PDPA.

In the GenAI context, the Access and Correction Obligations apply to personal data which has been collected, used, or disclosed for model or system development or deployment. Organisations must accede to individual requests unless an exception under the PDPA applies.

Despite the challenges to facilitating GenAI-related access and correction requests (e.g. due to massive amounts of data used to develop GenAI models and/or systems, and training data not being stored in a traditional repository), organisations are expected to adopt the following best practices to support compliance with their Access and Correction Obligations as reasonable and appropriate in the circumstances:

  • Adopt upstream data handling measures such as (i) verifying data accuracy at the point of collection; (ii) implementing data cleaning techniques such as de-duplication and outlier detection; and (iii) maintaining data provenance records to document the lineage of training data;
  • Review access and correction requests on a case-by-case basis and accede where reasonable (e.g. where the request refers to personal data stored in a Retrieval-Augmented Generation database). Organisations are encouraged to ensure that personal data, including inaccurate data, is removed from training datasets before they undertake future AI training runs; and
  • Track the maturity of and progressively adopt appropriate technical measures to remove inaccurate personal data from models and systems. In the interim, organisations can consider output filters and other safeguards to minimise the likelihood of models or systems producing inaccurate data as outputs.

Background

From 2 June 2026 to 1 July 2026, PDPC conducted a public consultation on a draft version of the Guidelines. PDPC published its response to feedback received from the public consultation on 20 July 2026.

Reference materials

The following materials are available on the PDPC website www.pdpc.gov.sg: