Journal of Intellectual Property (J Intellect Property; JIP)

KCI Indexed
OPEN ACCESS, PEER REVIEWED

pISSN 1975-5945
eISSN 2733-8487
Research Article

Transparency Without Exposure: A Selective Verifiability Framework for Copyright and Trade Secrets in Generative AI Training Data

Professor, Department of Global Business and Trade, Joongbu University, Republic of Korea

Correspondence to Heungsok Cha (heungsok@joongbu.ac.kr)

Volume 21, Number 3, Pages 277-297, September 2026.
Journal of Intellectual Property 2026;21(3):277-297. https://doi.org/10.34122/jip.2026.21.3.277
Received on July 11, 2026, Revised on August 06, 2026, Accepted on September 04, 2026, Published on September 30, 2026.
Copyright © 2026 Korea Institute of Intellectual Property.
This is an Open Access article distributed under the terms of the Creative Commons Attribution-NonCommercial-NoDerivatives (https://creativecommons.org/licenses/by-nc-nd/4.0/) which permits use, distribution and reproduction in any medium, provided that the article is properly cited, the use is non-commercial and no modifications or adaptations are made.

Abstract

Governance of generative artificial intelligence (AI) faces a structural dilemma. Copyright owners and regulators need evidence of training-data acquisition, licensing, rights reservations, and processing, while unrestricted disclosure may expose valuable dataset compilations and methods. This article systematically compares AI transparency duties in Korea, the European Union, the United States/California, and Japan by examining disclosure and information-provision requirements, labeling, obligated actors and triggers, exceptions and safeguards, and enforcement and sanctions. Copyright fair use and text-and-data-mining limitations are treated separately: they govern specified uses of protected subject matter and do not generally displace independent transparency duties covering copyrighted and non-copyright information. The comparison finds that public summaries improve orientation but generally cannot establish whether a particular work was included, on what legal basis it was processed, or in which model version it was used. Therefore, the article proposes “selective verifiability”: the capacity for an authorized actor to test a bounded compliance claim against lifecycle-linked records without receiving unrestricted corpus access. The framework has three tiers: a standardized public training-content summary; a confidential, source-level provenance manifest available to regulators or accredited auditors; and a claim-triggered verification channel for rights holders. W3C PROV, SPDX, machine-learning bills of materials, digital signatures, and append-only commitments can support the architecture but cannot independently prove legal validity or record completeness. For Korea, the article recommends a phased pilot, explicit institutional allocation, purpose-limited data processing, proportionate recordkeeping duties, reviewable trade-secret claims, and evidentiary incentives that do not immunize infringement.
Keywords

generative AI, training data, copyright, trade secrets, transparency, data provenance, selective verifiability

Notes

Conflicts of Interest

No potential conflict of interest relevant to this article was reported.

Funding

The author received manuscript fees for this article from Korea Institute of Intellectual Property.

Section