Japan panel backs ‘Principle Code’ to boost AI training data transparency
Japan panel approves voluntary ‘Principle Code’ to promote AI training data transparency; firms asked to disclose model design, data types, training methods and IP safeguards.
The Cabinet Office expert panel on artificial intelligence on August 18, 2026, gave broad approval to a draft "Principle Code" urging AI training data transparency for developers of generative AI. The code would ask domestic and foreign companies operating in Japan to disclose key information about the datasets and methods used to train their systems. The move is voluntary and carries no legal penalties, reflecting the government’s choice of a soft-law approach to manage AI risks while supporting technological development.
Panel decision and scope of the draft
The expert panel agreed in principle to the voluntary framework at its meeting on August 18, 2026, after months of discussion that began in late 2025. The draft frames transparency as a set of public-facing disclosures that companies would publish on their websites. It makes clear that compliance is left to each provider’s discretion, and that the code itself would not create binding legal obligations.
The panel described the measure as a basic set of principles to guide information disclosure by developers and providers of broadly capable generative AI models. Where companies cannot publish particular items, the draft requires them to explain the reasons for non-disclosure. The approach is intended to balance creators’ rights and public interest with incentives for continued innovation.
Required disclosures under the draft code
Under the proposed principles, participating firms would be expected to reveal the design specifications of the AI models they offer and the general categories of data used in training. The draft lists model architecture descriptions, the types of content used (such as text, images, audio, or code), and the training methods as core items for public disclosure. It also calls for explanation of steps taken to avoid scraping from pirate or infringing websites and for details on technical measures aimed at preventing intellectual property violations.
The emphasis in the draft is on providing sufficient information to creators, rights holders and the public so they can assess how copyrighted material and other sensitive content may have been used. The disclosures are framed as summaries or overviews rather than itemized lists of every dataset, reflecting concerns about trade secrets and operational burdens.
Why Japan chose a voluntary "soft-law" route
Japanese officials and panel members said they opted for a soft-law mechanism to avoid stifling ongoing research and commercial development by imposing strict legal requirements. The Cabinet Office has been working on transparency rules since the previous autumn, and the draft reflects consultations with industry, creators and rights management bodies. Proponents argue voluntary disclosure can raise minimum transparency standards while leaving room for competitive considerations and legitimate confidentiality.
Critics warn that without enforceable sanctions the code may have limited force, and that voluntary adoption could be uneven across domestic and international providers. The draft attempts to mitigate this by recommending public posting of disclosures and requiring rationales when items cannot be made public, but it stops short of mandatory reporting or penalties.
International context and regulatory contrasts
The Japanese proposal arrives amid a growing global push for disclosure rules for large AI models. The European Union’s AI Act requires providers of general-purpose models to publish summaries of the content used in training, with noncompliance subject to significant fines — up to €15 million or 3% of global annual turnover, whichever is higher. In the United States, California enacted a law in 2026 requiring generative AI developers to disclose sources of training data and related information for systems offered in the state.
Panel members cited these international moves when designing the draft, aiming to align Japanese expectations with global norms without imposing identical legal sanctions. The government has signaled a preference for coordination with international standards while retaining flexibility for Japanese industry practices.
Reactions from creators, rights holders and industry
Creative professionals and rights organizations have long voiced concern that their works can be incorporated into AI training datasets without clear notice or recourse. The panel’s draft was presented against that backdrop, and spokespeople for creators welcomed steps toward clearer disclosure while indicating the need for stronger protections. Industry representatives expressed cautious support, saying transparency could build public trust if implemented in a manner that protects legitimate business secrets.
The draft leaves open questions about enforcement, the granularity of required disclosures and how to reconcile transparency with proprietary model development. Rights holders and some lawmakers may press for tougher measures in future legislation if voluntary practices fail to satisfy creators’ demands.
Next steps for the Cabinet Office include public consultation and further refinement of the text before the panel issues final recommendations. The government will then consider whether to promote the code as a national standard for providers operating in Japan.
The panel’s decision on August 18, 2026, marks a first regulatory step toward clarifying how generative AI systems are trained, but substantial policy debate remains over the balance between transparency, intellectual property protection and innovation incentives.