3 Health data governance and standards

Key takeaways

  • Health data receive special legal protection as long as they are not anonymized, requiring distinct strategies beyond general AI approaches.

  • Techniques to preserve privacy, such as federated learning, can become competitive advantages, while enabling collaboration.

  • Data governance strategies must consider both regulatory compliance and health system interoperability requirements. The latter may require consideration of coordinated IP, technical and policy approaches, as appropriate.

  • Sharing cross-border health data calls for coordinated legal, technical and policy approaches.

  • The adoption and use of technical standards enables ecosystem growth while preserving competitive differentiation.

  • Public health partnerships can accelerate validation and deployment while creating sustainable impact.

  • Health data governance frameworks vary significantly. Innovators targeting deployment in low- and middle-income countries may benefit from engaging early with local data protection authorities and community stakeholders.

3.1 Data rights and privacy frameworks

Data rights and privacy regulation are important considerations to support AI-enabled health innovation, shaping not only legal compliance but also IP strategy, the freedom to operate and commercial scalability. Unlike patents, data sets are generally not protectable as such, although some jurisdictions have specific protections that may apply. Value is preserved through a combination of regulatory compliance, contracts, trade secret protection and database rights, where available. The interaction between data protection and IP strategies is relevant in considering AI-enabled health innovation, particularly where data governance may affect the long-term exploitation of AI models trained on health data. This chapter outlines some selected data governance frameworks that may be relevant for innovation and IP strategies. Emerging tools, including the European Health Data Space, may have a bearing on this discussion in the future.

3.1.1 European Union: General Data Protection Regulation

The 2016 General Data Protection Regulation (GDPR), effective as of May 2018 establishes one of the world's most stringent regimes for processing personal health data. The material scope of the regulation is set out under Article 2(1). Further definitions of personal data are in Article 4(1) and Recital 26. Provisions apply to personal data, meaning information relating to an identified or identifiable natural person and will only apply if data are not anonymized. In the case of anonymization, health data do not fall under the GDPR. If the GDPR does apply, developers must identify and document a lawful basis for processing,  under Article 6, as well as, where health or other sensitive data are involved, a condition permitting the processing of special categories of personal data under Article 9. Depending on the context, these may include Article 6(1)(a) and explicit consent under Article 9(2)(a); public interest in public health under Article 9(2)(i); or scientific research exemptions under Article 9(2)(j). In addition, controllers must implement data protection by design and by default (Article 25) and, where processing is likely to result in high risk to individuals' rights and freedoms,conduct data protection impact assessments (Article 35).

Other GDPR provisions that may be relevant to AI-enabled health innovation include:

  • Special category data (Article 9): Health data receive heightened protection requiring explicit consent or specific exemptions. AI training on health data must satisfy both general lawfulness (Article 6) and special category conditions (Article 9).

  • Genetic data (Article 9(1)): These data are explicitly listed among the special categories of personal data, alongside health data and other sensitive information, as categories considered specifically sensitive and subject to stricter data processing requirements.

  • Automated decision-making (Article 22): Individuals may not be subject to decisions made exclusively by automated processing where such decisions have legal consequences or similarly significant effects. In healthcare systems, this may apply in particular to determinations related to medical diagnoses or treatments. Health systems using AI must incorporate human oversight or obtain explicit consent.

  • Transparency requirements (derived from Articles 13, 14 and 16 as well as Recital 71): Data subjects have the right to obtain meaningful information about the logic involved in automated processing. This may impact AI model opacity and confidential business information or trade secret protection.

  • Data portability (Article 20): In certain circumstances individuals can request their data in machine-readable format, potentially requiring AI developers to provide patient-level data even when aggregated for model training.

  • Right to erasure (Article 17): Data subjects have the right to have their personal data deleted where certain conditions are met. This can create tensions, especially where data sets are valuable assets. While IP rights may protect the structure or compilation of a data set (e.g., database rights), they do not override data protection rights. In practice, this means that AI developers and data holders must design systems that allow for compliance with erasure requests, even if this is technically complex or economically inconvenient. This right is not absolute, however, and may be limited by competing interests, such as legal obligations or public interest (Article 17(3)).

3.1.2 European Union: Data Governance Act

The 2022 European Union Data Governance Act (DGA), effective as of September 2023, complements the GDPR by creating structured mechanisms for lawful data sharing, including:

  • Data intermediaries (Articles 10–12): Trusted entities that facilitate data-sharing between data holders and data users without accessing the data themselves. Data intermediaries must maintain neutrality and cannot use shared data for their own purposes.

  • Data altruism organizations (Articles 16–24): Not-for-profit entities that collect data based on consent or permission to make it available for objectives of general interest (e.g., health research, public health or official statistics). Registered data altruism organizations receive a European Union-recognized label.

  • Public sector data reuse frameworks (Articles 3–9): Conditions for the reuse of certain categories of public sector health data that are protected for reasons such as commercial confidentiality, statistical confidentiality, intellectual property rights, or personal data protection. This may be relevant for health-related public-sector datasets (e.g., hospital records, disease registries or health surveys). Access remains subject to applicable EU and national rules, as well as technical, organizational, confidentiality, and privacy-preserving safeguards.

Importantly, the DGA does not create new IP rights in data; instead, it clarifies how existing rights (database rights, trade secrets, confidential information) may be exercised in data-sharing environments. Article 5(10) explicitly states that reuse of protected data shall be without prejudice to IP rights, and that data intermediaries must respect trade secret protection for commercially sensitive information shared through their platforms.

3.1.4 United States: Health Insurance Portability and Accountability Act

In the United States, the 1996 Health Insurance Portability and Accountability Act (HIPAA) governs the use of protected health information by health care entities and their business associates, including AI developers. They are required to enter into business associate agreements; implement administrative, technical and physical safeguards; and follow prescribed de-identification protocols via either the safe harbor or expert determination methods (45 CFR § 164.514(a)–(b)). HIPAA compliance has direct implications for AI development, since data sets that are inadequately de-identified may retroactively become protected health information subject to breach notification requirements. This may undermine trade secret protections and trigger regulatory penalties. Several states, including California, Washington and Nevada, have enacted health privacy laws exceeding HIPAA requirements. AI-enabled health companies operating in the United States must comply with the most stringent applicable standard at both the federal and state levels.

3.1.5. Canada: Personal Information Protection and Electronic Documents Act

Canada's 2000 Personal Information Protection and Electronic Documents Act, amended in 2015, establishes a consent-based framework for private-sector data processing. Bill C-27 (the Digital Charter Implementation Act, introduced in 2022 but no longer under consideration) would have launched the Consumer Privacy Protection Act and the Artificial Intelligence and Data Act, with enhanced requirements for AI transparency, algorithmic impact assessments and automated decision-making disclosures. The federal Government is considering new privacy reforms and AI regulation. The Office of the Privacy Commissioner has issued guidance on AI and data analytics (2021), emphasizing transparency obligations when using health data for machine learning.

3.1.6 Japan: Act on the Protection of Personal Information

Japan's 2003 Act on the Protection of Personal Information, with a major revision in 2017 and a further amendment in 2020, effective as of 2022, permits the processing of anonymized information and pseudonymized information with specific safeguards. The 2020 amendments introduced explicit provisions for cross-border data transfers and created a new category of “pseudonymous information” that enables broader research use while maintaining privacy protections. That said, the Act requires consent for processing, cross-geographic transfer and third-party receipt of personal information. Many forms of consent may be required. AI-enabled health innovators will have to determine how to manage these requirements. Japan's Personal Information Protection Commission has issued guidance on AI and automated profiling (2021).

3.1.7 India: Digital Personal Data Protection Act

India's Digital Personal Data Protection Act, enacted in August 2023, introduces consent-centric data governance modeled partially on GDPR principles but with differences (e.g., no data localization requirements and streamlined compliance for start-ups). The Act requires “verifiable consent” for data processing. It grants access, correction and erasure rights to individuals. Sectoral rules for health data processing were under consultation as of early 2026, with implementation rules expected to specify additional safeguards for sensitive personal data, including health information and genetic data. The Data Protection Board of India is being established to enforce the Act.

3.1.8 Brazil: Lei Geral de Proteção de Dados (General Data Protection Law)

The 2018 Lei Geral de Proteção de Dados (General Data Protection Law), effective as of 2020, establishes comprehensive data protection requirements closely aligned with GDPR principles, including lawful bases for processing, data subject rights (access, correction, deletion and portability), data protection impact assessments and accountability requirements. Brazil's National Data Protection Authority has issued guidance on health data processing (2021) and AI systems (2023). The law requires explicit consent for processing sensitive personal data, including health information, with provisions for research exceptions under ethical oversight.

3.1.9 China: Personal Information Protection Law and Data Security Law

China's data governance framework for AI-health innovation is governed by two laws enacted in 2021.

The Personal Information Protection Law establishes comprehensive requirements for the processing of personal information, including health data, with obligations closely paralleling those of the GDPR in structure but with distinct Chinese characteristics. Health information is classified as sensitive personal data under the law, requiring separate and explicit consent for processing. The cross-border transfer of personal data requires either a security assessment administered by the Cyberspace Administration of China, certification by an accredited institution or execution of standard contractual clauses issued by the Cyberspace Administration. These provisions create significant compliance infrastructure requirements for international AI-health collaborations involving Chinese patient data.

The Data Security Law introduces a national data classification system based on the importance of data to national security and public interest, with health data potentially qualifying as “important data” subject to heightened protection and mandatory security assessments before a cross-border transfer. For AI health innovators, these frameworks create meaningful data localization pressures. Training data sets derived from Chinese health institutions may be subject to restrictions that preclude transfers to overseas model development environments, making federated learning architectures particularly relevant in the Chinese market.

3.1.10 Australia: Privacy Act and My Health Records Act

Australia's privacy framework for health data operates across two instruments. The Privacy Act 1988, incorporating the Australian Privacy Principles (APPs), establishes baseline requirements for handling personal information, including health records. APP 11 imposes specific security obligations and APP 8 governs cross-border disclosure. Health information receives heightened protection as a sensitive information category. The My Health Records Act 2012 establishes a distinct national framework for the digital health record system. It imposes strict limitations on the secondary use of records for purposes including AI training, requiring specific legislative authorization for research use. Australia's Privacy Act is currently undergoing significant reform. A 2023 review proposed strengthening individual rights, introducing a direct right of action and expanding the definition of personal information in ways that would affect AI training data practices. AI-enabled health innovators targeting the Australian market should monitor reform developments closely.

3.1.11 Singapore: Personal Data Protection Act

Singapore's 2012 Personal Data Protection Act provides a mature and frequently updated framework for personal data governance. In 2020, amendments introduced mandatory data breach notifications, enhanced accountability obligations, and, importantly for AI and health contexts, new provisions on deemed consent through legitimate interests. The last element provides a lawful basis for data processing, where the individual would reasonably expect use and it does not adversely affect them. The act's research exception permits the processing of personal data without consent for research purposes, where results are not published in individually identifiable form. This is relevant to AI model training on health data sets. Singapore's Personal Data Advisory Committee has issued specific guidance on AI governance through the Model AI Governance Framework (2020, updated). It addresses data quality, algorithmic transparency and human oversight in AI systems, making it one of the more developed non-EU frameworks for AI-specific data governance in the health context.

3.1.12 African regional frameworks

National data protection laws across Africa reflect principles established by the GDPR and included in the 1995 EU Data Protection Directive. The latter, which precedes the GDPR, includes data subject rights and the recognition of special categories of data and adequacy determination. The African Union adopted the Convention on Cyber Security and Personal Data Protection (Malabo Convention, 2014), which establishes GDPR-aligned principles, including lawfulness, purpose limitation, data minimization and accountability. As of 2026, 14 African Union member states had ratified the Convention.

National implementation shows strong convergence with GDPR standards. South Africa's 2013 Protection of Personal Information Act became fully effective in 2021. It closely mirrors the GDPR structure with special protections for health data and the creation of an information regulator. Kenya's 2019 Data Protection Act establishes data subject rights, requires impact assessments for high-risk processing and restricts international data transfers to countries with inadequate protection, mirroring GDPR's adequacy requirements. Nigeria's 2019 Data Protection Regulation introduced GDPR-aligned principles, later strengthened through the 2023 Nigeria Data Protection Act, which created the Nigeria Data Protection Commission with powers similar to EU supervisory authorities. Rwanda's Law No. 058/2021 established heightened protections for health data, mandatory breach notifications and cross-border transfer restrictions.

3.2 Standards and interoperability

Standardization inside Standard Development Organizations (SDOs) and Standard Essential Patents (SEPs) have played central roles in the information and communication technology sector, where interoperability is critical for widespread adoption (see: Box 2 Experiences with SEPs).

Box 2 Experiences with SEPs

SEPs protect inventions that are used when implementing technical standards. The IPR policies of SDOs usually set forth requirements that such patents are declared and their holders undertake to license them either royalty free or on FRAND terms. In the health technologies sector, FRAND licensing typically applies to interoperability standards such as DICOM for medical imaging, HL7/FHIR for health data exchange, and communication protocols for connected medical devices. FRAND undertakings are not characteristic to pharmaceutical products, as drug patents are not subject to technical standardization requirements. The WIPO report on Valuation Methods in Licensing Standard Essential Patents(1)https://www.wipo.int/web-publications/frand-economics-valuation-methods-in-licensing-standard-essential-patents/en/index.html offers a practical reference on the economic methods used to assess FRAND licensing terms in the field of SEPs.

To balance incentivizing participation in developing standards and broad implementation, and to mitigate possible competition law concerns, SDO IPR policies generally include requirements to disclose potential SEPs and offer licenses on fair, reasonable, and non-discriminatory (FRAND) terms. ITU's Common Patent Policy for ITU-T/ITU-R/ISO/IEC and the European Telecommunications Standards Institute (ETSI) IP policy illustrate this approach. (2)Examples of different IP policies are available at https://www.wipo.int/en/web/patents/topics/sep.

In the health sector, however, exclusivity strategies may reflect a range of specific considerations, including market positioning and regulatory requirements . Accordingly, FRAND licensing will not be equally relevant across all AI-enabled health technologies. Its role is most likely to arise in relation to standardized technical specifications and interoperability-enabling interfaces – such as data formats and communication protocols – rather than safety critical clinical decision logic or proprietary algorithmic architectures that form the competitive core of medical AI products. The latter are not necessarily the kinds of interfaces or specifications that a standard requires all implementers to use.

Case study: Owkin – federated learning for drug discovery

Owkin, a start-up founded in 2016 in France and the United States, specializes in applying AI and federated learning to medical research, enabling multiple institutions to collaborate on AI model training without sharing raw patient data. Its federated learning platform distributes model training across participating hospitals, research centers and pharmaceutical companies. Each site trains a local model on its own data and shares only model updates with a central aggregator, which combines them into a global model for further refinement. This approach to preserving privacy addresses one of the fundamental challenges in health care AI: training robust models on large, diverse data sets while respecting patient privacy and institutional data governance requirements.

Owkin's IP strategy combines patents, trade secrets and selective open-source contributions. Its core European patent application (EP 4256758A1) covers the administration and orchestration of federated learning networks. It protects key architectural elements, including distributed training protocols, secure aggregation mechanisms, privacy-preserving safeguards and techniques for managing data heterogeneity across participating institutions. By grounding claims in system-level architecture and concrete technical solutions, rather than abstract algorithmic concepts, the patent establishes a defensible position targeting the practical implementation challenges of federated learning in health care environments.

In tandem, the company maintains trade secrets covering proprietary data sets, model architectures, training procedures optimized for federated settings and operational deployment configurations. It also contributes selectively to open-source federated learning frameworks to build research community credibility without exposing production implementation or clinical validation methods.

Owkin has established significant pharmaceutical partnerships that demonstrate the commercial viability of its platform. In 2021, Sanofi made a USD 180 million equity investment to collaborate on oncology drug discovery and patient stratification, representing one of the largest AI–pharma collaborations in Europe. Additional partnerships with Bristol Myers Squibb and Johnson & Johnson cover immunology, oncology and surgical outcomes prediction. As of March 2025, Owkin has raised over USD 304 million in total funding and achieved unicorn status with a valuation exceeding USD 1 billion.

Key lessons for innovators

  • Federated learning addresses a critical market need. It enables AI training on distributed health care data without centralization, thereby respecting privacy regulations and institutional autonomy.

  • A mixed IP strategy is essential. Patents protect the core technical architecture, trade secrets safeguard operational know-how and partnerships, and selective open-source contributions build ecosystem trust.

  • Strategic partnerships with and investment by pharmaceutical companies, provide both funding and access to real-world clinical data, regulatory expertise and commercialization pathways.

  • European data governance frameworks (i.e., GDPR and DGA) can be a competitive advantage when properly navigated. They enable compliant multi-institutional collaborations that would be difficult to replicate in less regulated environments.

Sources: Owkin company website (owkin.com); Sanofi press releases (2021, 2024); Reuters reporting on Sanofi investment; Forbes interviews with Thomas Clozel; EPO (EP 4256758A1).