Can I Use Large Language Models and Other AI (such as ChatGPT, Google Gemini, etc.) with SOEP Data?
General Statement: Currently, large language models (LLMs) and other AI tools may not be used to manage, process, or analyze SOEP data.
Under SOEP Data Use Agreements, researchers are forbidden to distribute data or other materials we supply (apart from codebooks and metadata, described below) to other members, organizations, or individuals. This means, that use of LLMs is a violation under our present SOEP Data Use Agreements.
For purposes of this policy, LLMs are classified into three categories:
Type 1: LLMs that retain user-provided data for any purpose, including training the LLM (e.g., GPT, Llama)
Type 2: LLMs that are licensed by an institution and have conditions of use that do not permit the retention of user-provided data
Type 3: Type 2 LLMs that are isolated within a secure network with no access to the Internet
|
LLM Type |
SOEP Data Use |
Reason |
|
Type 1 |
None |
Type 1 LLMs ingest and make use of the data. This counts as redistributing the data to the company operating the LLM, so it is not permitted. |
|
Type 2 |
None |
Type 2 LLMs do not retain or make use of the data, so this does not count as redistribution. However, they are not isolated from broader networks or the Internet, so they would not comply with data security plans for Restricted-Use data. At present, we are also not accepting requests for this kind of data use for Public-Use data. |
|
Type 3 |
None |
Type 3 LLMs do not retain or make use of the data and are operated on individual machines or within secure networks. However, current evidence shows that complete isolation cannot yet be reliably guaranteed, as unintended access beyond such controlled environments (“sandbox escapes”) remains possible. Given these unresolved security risks, we currently do not accept requests to use Type 3 LLMs with SOEP microdata. |
Thanks to ICPSR and Sebastian Karcher at Syracuse University for originating this taxonomy of LLMs.
It is permissible to use LLMs with our public-facing documentation, codebooks, study-level metadata, and aggregated group or population estimates. However, person-level or household-level microdata must not be provided to or processed by LLMs. This restriction does not apply to other machine-learning methods, including generative machine-learning methods, provided that their use complies with applicable data protection and data-use requirements.
Q: Is it acceptable to use an LLM to write analysis code (e.g. Stata, Python, or R scripts) based data?
A: Using an LLM to write analysis code is acceptable as long as no SOEP microdata are upload to the LLM.
Q: Is it acceptable for derived analytic output (e.g., model coefficients or summary statistics) shared with an LLM?
A: Uploading analytic output to an LLM is acceptable as long as no SOEP microdata are upload the LLM.
Q: Does the policy apply to all editions of SOEP data (e.g., teaching or international edition)?
A: Yes. The policy applies to SOEP editions.
Q: Does the policy also apply to SOEP data that are not distributed through the SOEP Research Data Center (SOEP-FDZ), but can be accessed in a data-protection-compliant manner through other channels, such as international databases (e.g., LIS, LWS) or harmonized data resources (e.g., Cross National Equivalent File (CNEF))?
A: Yes. The policy applies to SOEP files across all distribution platforms.
Q: Is it acceptable to use AI-assisted development tools that automatically index files in the working directory (e.g., AI features built into code editors) with SOEP data files?
A: No. Users must disable this feature or store SOEP microdata outside any directory to which such a tool has access.
We gratefully acknowledge the Health and Retirement Study (HRS) at the University of Michigan, whose AI/LLM Use Policy served as the basis for this document.