From March 2029, per the European Health Data Space (EHDS) regulation, hospitals across the EU will progressively face new obligations to make defined categories of electronic health data available for cross-border primary use and respond to research requests coordinated through national Health Data Access Bodies. Hospitals already perform multiple mandatory reporting and data-submission processes today. These processes are often fragmented, request-driven, largely manual/semi-automated, with data repeatedly re-curated for each reporting destination, effectively “curate many times, use once". EHDS adds regulatory pressure to make data available, faster and more extensively across borders, without specifying how to improve existing suboptimal processes.
AIDAVA tested the opposite approach: curate once, use many times. Clinical data are cleaned, structured and validated upstream into a reusable digital representation of the patient record. From this single curated resource, different outputs can then be generated automatically without repeating the underlying data preparation.
In four hospitals in Estonia, Austria and the Netherlands, working in three languages with patients and local data stewards, AIDAVA demonstrated that a single curated record could generate three unrelated outputs without further manual preparation: a European patient summary of the kind required under the EHDS from 2029, a dataset for an interoperable, federated breast cancer registry, and a cardiovascular risk score for the treating doctor. Generating the breast cancer registry dataset took under one minute compared with approximately 20 minutes using the conventional approach. The cardiovascular risk score was generated in under two seconds, compared with approximately five seconds previously. The consortium also modelled the potential economic impact across the EU-27. The conventional, request-driven approach was estimated to cost approximately €6.6 billion over five years. An upstream curation approach was estimated at approximately €1.04 billion, representing a potential 84% reduction. These figures are scenario-based estimates, not observed EU-wide savings, and the consortium stresses that they require validation against real-world national data.
AIDAVA is a research prototype tested at four sites in two disease areas and the results also highlight important remaining challenges. Automated extraction of medical concepts from free text achieved 72% precision and 73% recall across the three languages in scope. Natural Language Processing remains the weakest link, human intervention is still required for difficult or ambiguous cases. The project therefore does not present AI as a replacement for data curation, but as a way of moving much of the repetitive work upstream and making data reuse more cost-effective and AI-ready.
The project's closing message to policymakers is clear: meeting the EHDS implementation deadlines and extracting lasting value from the European Health Data Space are two different objectives. Compliance alone will not deliver the full potential of Europe's health data infrastructure. The way data are prepared and curated at hospital level will determine whether the EHDS becomes primarily a compliance burden or a reusable European digital asset. AIDAVA therefore calls for five actions:
- Launch implementation pilots in three to five Member States to test scalable approaches under real-world conditions.
- Establish a shared European semantic foundation, governed as a public asset, to enable EU-wide interoperability and scalability of AIDAVA-like solutions.
- Invest in skills and open-source tools to strengthen hospitals’ capacity for AI-assisted data curation and semantic interoperability.
- Create sustainable funding mechanisms so that the costs of preparing data for European reuse are shared across the actors that benefit from it.
- Coordinate the AI Act, Medical Device Regulation and emerging legislation such as the Biotech Act to provide clear and proportionate pathways for certifying AI-enabled health-data infrastructure