Data distribution impacts the performance and generalisability of contrastive learning-based foundation models of electrocardiograms

Gul Rukh Khattak , Konstantinos Patlatzoglou, Joseph Barker, Libor Pastika, Boroumand Zeidaabadi, Aidan R. Birdi, Jiayu Huo, Ahmed El-Medany, Hesham Aggour, Yixiu Liang, Antonio H. Ribeiro, Jeffrey Annis, Antonio Luiz Pinho Ribeiro, Junbo Ge, Daniel B. Kramer, Jonathan W. Waks, Evan Brittain, Nicholas Peters, Fu Siong Ng  , Arunashis Sau 

Abstract

Contrastive learning is a widely adopted self-supervised pretraining strategy, yet its dependence on cohort composition remains underexplored. We present Contrasting by Augmented Patient Electrocardiograms (CAPE) foundation model and pretrain on four cohorts (n = 5,203,269), from diverse populations across three continents (North America, South America, Asia).

Introduction

Electrocardiograms (ECGs) provide a non-invasive and widely accessible means of recording the heart’s electrical activity, capturing myocardial depolarisation and repolarisation across multiple leads positioned on the skin. Despite the potential to reveal important physiological insights, the clinical utility of ECGs remains in part constrained by the need for expert interpretation. Additionally, recent studies have demonstrated the potential for artificial intelligence-enhanced ECG (AI-ECG) analysis to identify undiagnosed disease and predict risk of adverse events better than human experts 

Results

This study utilises a diverse set of large-scale electrocardiogram (ECG) cohorts from five countries across four continents: the Beth Israel Deaconess Medical Centre (BIDMC) [3] and Vanderbilt University Medical Centre (VUMC) [30] from the United States, the Clinical Outcomes in Digital Electrocardiography (CODE) cohort [31] from Brazil, the Shanghai Zhongshan Hospital (SHZS) cohort from China, the UK Biobank (UKB) [32] from the United Kingdom, and the Physikalisch-Technische Bundesanstalt (PTB-XL) dataset [22] from Germany.

Discussion

To our knowledge, this is the first systematic exploration of how different pretraining cohorts influence learned feature representations, and how varying downstream cohorts affect predictive performance on target labels. 

Conclusions

This study emphasises the central importance of generalisation in contrastive learning for clinical applications. We find that the ability of a model to perform well across diverse patient cohorts depends critically on the composition of the pretraining data.

Materials and methods

All relevant ethical permissions have been obtained for all cohorts explored in the current study. The BIDMC cohort approval is provided by the Beth Israel Deaconess Medical Centre Committee on Clinical Investigations, IRB protocol # 2023P000042. The CODE study is approved by the Research Ethics Committee of the Universidade Federal de Minas Gerais, protocol 49368496317.7.0000.5149. The Institutional Research Board of Zhongshan Hospital (No. 2023-253R) approved the use of SHZS data with a waiver of patient consent.

Acknowledgments

The authors would also like to thank the InSIGHT Core in the Center for Healthcare Delivery Science at Beth Israel Deaconess Medical Center for assistance in obtaining primary data.

Citation: Khattak GR, Patlatzoglou K, Barker J, Pastika L, Zeidaabadi B, Birdi AR, et al. (2026) Data distribution impacts the performance and generalisability of contrastive learning-based foundation models of electrocardiograms. PLOS Digit Health 5(9): e0001623. https://doi.org/10.1371/journal.pdig.0001623

Editor: Jie-Zhi Cheng, Shanghai United Imaging Intelligence Co Ltd, CHINA

Received: January 31, 2026; Accepted: July 14, 2026; Published: September 2, 2026

Copyright: © 2026 Khattak et al. This is an open access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited.

Data Availability: A subset of the CODE dataset, comprising 15% of patients, is publicly available at https://zenodo.org/records/4916206. The PTB-XL dataset is open access and can be obtained from https://physionet.org/content/ptb-xl/1.0.1/. Access to the UK Biobank data is available upon approved application through https://www.ukbiobank.ac.uk. Due to ethical and legal restrictions, the remaining datasets used in this study are not publicly available. The results underlying the figures and tables in this manuscript will be made available by the authors upon request, where possible.

Funding: British Heart Foundation (BHF), UK: programme grant funding to FSN and NSP (RG/F/22/110078), clinical research training fellowship to AS, JB and AEM (FS/CRTF/21/24183 and FS/CRTF/24/24624), BHF Centre of Research Excellence funding to FSN, AS and NSP (RE/18/4/34215 and RE/24/130023). Medical Research Council (MRC), UK funding to LP (MR/Y000803/1). AS is also funded by the Academy of Medical Sciences (SGCL033\1014), Cardiomyopathy UK (C2502) and an NIHR Academic Clinical Lectureship. YL is funded by Shanghai Municipal Key Clinical Specialty (shslczdzk01701) and Fudan University AI4S Project (FudanX24AI056). The authors are supported by the National Institute for Health Research Imperial Biomedical Research Centre. FSN and ALPR received funding from the British Council through the Brazil–UK Research in DiGital HEalth and AI-ECG (BRIDGE-ECG) project (Grant No. RE2025-4590).

Competing interests: I have read the journal’s policy and the authors of this manuscript have the following competing interests: J.W.W. and D.B.K. were previously on the advisory board for Heartcor Solutions LLC, for whom they remain independent consultants. J.W.W. also reports research funding from Anumana and is a consultant for HeartBeam SInc. F.S.N. reports speaker fees from GE Healthcare and is on the advisory board for J&J Medtech. A.S., L.P., B.Z., and F.S.N. declare inventorship on a patent application relating to AI-ECG methods. A.S., L.P., B.Z., F.S.N., D.B.K., J.W.W., and N.S.P.S hold equity shares in Cardiovolt.ai Limited. The remaining authors declare no conflicts of interest.

https://journals.plos.org/digitalhealth/article?id=10.1371/journal.pdig.0001623