Abstract
Objectives
This study aims to investigate the potential relationship between radiomic features extracted from 18F-fluorodeoxyglucose (18F-FDG) positron emission tomography (PET)/computed tomography (CT) images of patients with breast cancer and molecular markers including estrogen receptor (ER), progesterone receptor (PR), human epidermal growth factor receptor 2 (HER-2) and Ki-67.
Methods
Radiomic features were obtained via an open-source software (LIFEx) from 18F-FDG PET/CT images of 162 patients with histopathologically confirmed breast cancer. In order to examine the association between each molecular marker and radiomic features, machine learning models were developed seperately using the Python software. After the data were divided into test and training subsets, features were selected and scaled. Using the selected features, five different machine learning models (Random Forest, XGBoost, Support Vector Machine, Logistic Regression and Naive Bayes) were established and their performance was evaluated based on accuracy, sensitivity, specificity, F1 score, balanced accuracy, MCC and area under the curve (AUC) the receiver operating characteristic curve.
Results
For HER-2 status, the model yielded an AUC of 0.76, a balanced accuracy of 0.76, a sensitivity of 62.5% and a specificity of 90.0%. For ER, PR and Ki-67 status, AUC values were 0.59, 0.55, 0.60, balanced accuracies were 0.65, 0.61, 0.62, sensitivities were 96.6%, 81.8%, 68.4% and specificities were 33.3%, 40.0%, 55.2% respectively.
Conclusion
In this study, radiomic features derived from 18F-FDG PET/CT images demonstrated poor or non-discriminatory performance in the assessment of ER, PR and Ki-67 status. Although a preliminary signal of an association was observed for HER-2 status, this finding should be considered hypothesis-generating only due to the low number of HER-2-positive cases in the test subset and the lack of independent external validation. In larger cohorts prospective, multicenter studies are needed to validate this potential association.
Introduction
Breast cancer is the most prevalent malignancy diagnosed among women across the globe, as reported by the World Health Organization (1). Owing to its high prevalence and heterogeneous biology, accurate molecular characterization and appropriate therapeutic stratification remain crucial.
Conventionally, breast cancer is classified into several histological including invasive ductal carcinoma, invasive lobular carcinoma, tubular carcinoma, and mucinous carcinoma (2). Over time, it has become evident that tumors sharing the same histological characteristics can display distinct biological behaviors (3). Breast cancer constitutes a heterogeneous malignancy characterized by diverse biological and histological properties, clinical presentations, behaviors and therapeutic responses (4, 5). Such heterogeneity emphasizes the need to individualize therapeutic approaches according to the histological, molecular, and genetic features of the tumor, thereby highlighting the increasing significance of personalized therapy (6, 7).
In the current approach, the European Society for Medical Oncology guidelines for early breast cancer classify the disease into four molecular subtypes: “Luminal A,” “Luminal B,” “human epidermal growth factor receptor 2 (HER-2)-positive,” and “basal-like.” This classification is mainly determined by the evaluation of estrogen receptor (ER), progesterone receptor (PR), HER-2, and Ki-67 levels (8).
18F-fluorodeoxyglucose positron emission tomography/ computed tomography (18F-FDG PET/CT) is a functional imaging modality that reflects tumor glucose metabolism and may reveal biological heterogeneity within tumors beyond conventional visual analysis. Radiomic analysis converts imaging features into quantitative parameters, thereby enabling the investigation of associations between imaging phenotypes and molecular characteristics.
As in breast cancer, most imaging techniques used in medicine primarily rely on shades of gray, which represent transitions between black and white. Although imaging methods are capable of generating a wide spectrum of gray tones, the human eye can only distinguish a limited portion of them. Consequently, when clinicians evaluate medical images, the results often remain constrained by the perceptual limits of human vision. The necessity of computational support in analyzing the vast amount of information embedded within these gray tones is therefore evident. Through radiomic data analysis, the relationships between the smallest image pixels are converted into quantitative data, enabling information beyond the visual perception of the human eye to be processed in greater detail with the aid of machine learning algorithms (9, 10, 11, 12).
In our study, we aim to investigate whether radiomic features extracted from 18F-FDG PET/CT images of patients with breast cancer are associated with molecular markers (ER, PR, Ki-67, and HER-2) commonly used to determine molecular subtypes. Although tissue sampling is the gold standard for determining these markers, in certain situations, such as when intratumoral heterogeneity limits representation of the whole lesion or in advanced disease when sampling every lesion is not feasible and targeted tracers are unavailable, imaging-based approaches may provide exploratory information about tumor heterogeneity and potential imaging-molecular associations. This study was designed as an exploratory, hypothesis-generating analysis, and the findings should not be interpreted as supporting clinical implementation without further multicenter validation.
Materials and Methods
Study Population
For this retrospective study, we analyzed data from female patients over the age of 18 with histopathologically confirmed breast cancer based on surgical excision specimens and from available 18F-FDG PET/CT images acquired between January 2015 and January 2022. Molecular marker profiles, including ER, PR, HER-2, and Ki-67, were determined exclusively from surgical pathology specimens, which were considered the histopathological reference standard. The maximum time interval between PET/CT imaging and histopathological evaluation was one month.
The study initially included 477 patients. Patients were excluded if histopathology reports were unavailable. Analyses were performed separately for each molecular marker, and patients without an available result for a specific marker were excluded only from the corresponding endpoint analysis. For cases with HER-2 scores of +2, FISH correlation is recommended; patients without corresponding FISH results were excluded.
Furthermore, patients who had received any local or systemic treatment for breast cancer prior to 18F-FDG PET/CT imaging or who had undergone excisional biopsy before imaging were excluded from the study.
As a result, 162 patients meeting all inclusion criteria were enrolled. Among these, ER and PR data were available for 160 patients each, HER-2 data for 139 patients, and Ki-67 data for 158 patients.
Classification thresholds for ER, PR, HER-2, and Ki-67 were selected according to clinical criteria and standards reported in the literature. ER and PR expression was considered negative when the result was 0% and positive when ≥1% (13). HER-2 status was determined by immunohistochemistry, with scores of 0 and +1 classified as negative and +3 classified as positive. Cases with a +2 score were further evaluated using FISH; those with a negative FISH result were categorized as negative, whereas those with a positive FISH result were categorized as positive (14). Ki-67 expression was considered negative for values <20% and positive for values ≥20% (15).
This study received ethical approval from Çanakkale Onsekiz Mart University Clinical Research Ethics Committee (decision number: 2023/01-17, date: 04.01.2023) and informed consent was not required due to the study’s retrospective nature.
PET/CT Acquisition Procedure
18F-FDG PET/CT images were acquired using a hybrid PET/CT imaging system (Gemini TF16 PET/CT; Philips Medical Systems). Imaging was performed 60±5 minutes after intravenous administration of 350-550 MBq of 18F-FDG labeled 18F-FDG in patients who had fasted for 4-6 hours and had a blood glucose level <150 mg/dL. Low-dose, non-contrast CT images (120 kVp, 60-150 mA, 5 mm slice thickness) covering the vertex to the proximal thigh were obtained first; PET acquisition in 3D mode at 3 minutes per bed position was performed immediately afterward. PET images were reconstructed using the line-of-response row-action maximum-likelihood algorithm (LOR-RAMLA; Philips Astonish TF), producing transverse, coronal, and sagittal image planes.
Feature Extraction
For the purpose of segmenting breast cancer lesions and extracting radiomic features from PET/CT images, the Java-based Local Image Feature Extraction (LIFEx, version 7.1.9) software (www.lifexsoft.org) was employed (16). The software allows for standardized texture analysis, and all texture parameters used in this study conform to the guidelines established by the Image Biomarker Standardization Initiative, ensuring reproducibility and consistency in radiomic feature quantification (17).
PET/CT images in DICOM format were exported from the workstation. For texture analysis, a three-dimensional (3D) spherical volume of interest (VOI) encompassing the entire target lesion was delineated on the PET images. Using a semi-automated approach, a fixed threshold corresponding to 40% of the maximum standardized uptake value within the VOI was applied to segment the lesion. The fixed 40% SUVmax threshold is a commonly applied method in PET-derived radiomics studies due to its reproducibility (18).
The image-processing stage following segmentation is critically important. To ensure standardization of tissue analysis and facilitate comparison with other studies, it is essential to define appropriate image processing parameters before calculating radiomic features. Experience in this area has increased as the number of studies has grown (19). In the present study, all segmented lesions were spatially resampled to a voxel size of 4 mm along the x, y, and z axes. Absolute resampling, as recommended for intensity rescaling in PET imaging, was applied (20). Since SUV values in 18F-FDG-based oncological imaging typically range between 0 and 20, these values were set as the minimum and maximum for intensity rescaling. A fixed bin width of 0.3125 was used, resulting in 64 gray levels within the mentioned SUV range. The workflow of the radiomic analysis is illustrated in Figure 1.
In the subsequent step, 45 features — conventional, histogram-based, and texture — were extracted from the segmented lesion. Histogram-based features represent first-order statistics, while texture features correspond to second-order statistics (21). The texture features were primarily derived from the gray-level co-occurrence matrix (GLCM), gray-level run length matrix (GLRLM), neighborhood gray-level difference matrix (NGLDM), gray-level zone length matrix (GLZLM), and their respective subgroups (22). The full list of features used to develop machine learning models is presented in Table 1.
Machine Learning Model Establishment
Machine learning models were developed separately for the exploratory assessment of associations between radiomic features and ER, PR, HER-2, and Ki-67 molecular profiles.
Due to differences in sample size and class distribution among molecular markers, the train-test split was determined separately for each endpoint. While an 80/20 split was used for ER, PR, and HER-2, a 70/30 split was applied to Ki-67 to ensure adequate class representation in the test set (random seed =42). Class distributions between training and test subsets were comparable for all endpoints. To prevent information leakage, all preprocessing steps are applied exclusively to the training subset.
For each molecular marker, 45 radiomic features were initially considered. The features were scaled using a normalization method to map their minimum and maximum values to a range between 0 and 1. This step was performed only on the training subset, and once the parameters were learned, they were applied to the test subset.
To determine the most relevant features for use in machine learning models, the sequential feature selection (SFS) method was applied only to the training subset. Feature selection was guided by k-scores, which represent the mean cross-validation performance of each candidate feature subset. Through this approach, the test subset was not accessed during feature scaling or feature selection.
Subsequently, the selected features were used to develop five machine learning models. The models implemented in this study included Random Forest, XGBoost, Support Vector Machine, Logistic Regression, and Naive Bayes. These models were trained on the training subset and evaluated on the held-out test subset. As model comparison and selection of the reported model were performed on the same test subset, the overall workflow does not represent a fully nested validation strategy. Accordingly, the reported performance estimates should be interpreted within an exploratory, hypothesis-generating framework and may be optimistically biased.
Statistical Analysis
Data analysis was performed using Python software (version 3.9.13). The performances of the five machine learning models were evaluated and compared using performance metrics including accuracy, precision, recall, F1 score, and support. Additionally, balanced accuracy and the matthews correlation coefficient (MCC) were calculated to account for class imbalance. Confidence intervals for sensitivity and specificity were calculated using binomial proportion methods. Based on the exploratory comparisons on the held-out test subset, one model for each molecular marker was selected. The confusion matrix of the selected model was further examined and its sensitivity and specificity were calculated. Additionally, ROC curves were plotted for all models, and the corresponding area under the curve (AUC) values were determined.
Results
The distribution of patients, the machine learning model selected for detailed reporting for each molecular marker, and the corresponding model performance metrics are presented in Tables 2 and 3.
ER Prediction
For the assessment of ER status, the most relevant 7 features out of 45 radiomic features were SUVmin, SUVq1, SUVq2, skewness, GLCM Contrast, GLRLM_LRE and GLRLM_SRLGE. Although the accuracy of Random Forest, XGBoost, Support Vector Machine, and Logistic Regression models was 0.91, this performance metric is misleading because of a marked class imbalance, with ER-positive cases comprising 89% of the cohort. When the Random Forest model was used to predict ER expression from radiomic data, the sensitivity, specificity, and AUC were 96.6%, 33.3%, and 0.59, respectively. The combination of high sensitivity and low specificity indicates inadequate discrimination of ER-negative tumors. The balanced accuracy was 0.65, and the MCC was 0.37, further confirming the weak discriminatory performance despite high sensitivity.
PR Prediction
For the assessment of PR status, the five most relevant radiomic features were identified as SUVstd, SUVq2, SUVpeakSphere 0.5ml, GLZLM_SZE and GLZLM_ZP. The Random Forest model achieved an accuracy of 0.69, with sensitivity, specificity, and AUC of 81.8%, 40.0%, and 0.55, respectively. The balanced accuracy was 0.61 and the MCC was 0.23. These findings suggest that the model performs similarly to a majority-class classifier. No discriminatory performance was observed for PR status.
HER-2 Prediction
The ten most relevant radiomic features for HER-2 status assessment were SUVmean, SUVmax, SUVpeak-sphere-1 mL, sphericity, compacity, GLCM dissimilarity, GLRLM_LRLGE, GLRLM_RLNU, NGLDM Contrast and GLZLM_GLNU. Among the models evaluated, Random Forest achieved the highest balanced accuracy of 0.76. When this model was used to predict HER-2 expression from radiomic data, sensitivity and specificity were 62.5% and 90.0%, respectively, and the highest AUC was 0.76 (Figure 2). These results indicate a preliminary discriminatory signal for HER-2 status. However, given the limited number of HER-2 positive cases in the test subset and the exploratory model selection framework, these findings should be considered hypothesis-generating rather than evidence of robust predictive capability.
Ki-67 Prediction
For the assessment of Ki-67 status, the six most relevant radiomic features were identified as SUVmax, SUVq3, compacity, GLRLM_LRE, GLZLM_LGZE and GLZLM_GLNU. The highest accuracy for Ki-67 prediction was 0.60, achieved using both XGBoost and Logistic Regression models. The F1 scores were 0.58 for XGBoost and 0.42 for Logistic Regression. Using the XGBoost model, Ki-67 prediction yielded sensitivity, specificity, and AUC of 68.4%, 55.2%, and 0.60, respectively. No discriminative performance was observed for Ki-67 expression relative to the radiomic features in this study.
Discussion
The present study aimed to explore whether radiomic features extracted from 18F-FDG PET/CT images were associated with ER, PR, HER-2, and Ki-67 status in breast cancer. Radiomic analysis may provide information regarding tumor tissue heterogeneity which is presumed to be reflected in but remains imperceptible to the human eye (10).
In this context, radiomic analysis may be capturing tumor metabolic heterogeneity rather than the molecular marker directly.
In the present exploratory analysis, model performance for ER, PR, and Ki-67 status was weak or nondiscriminatory. For HER-2 status, a preliminary signal of discrimination was observed, with the Random Forest model achieving a balanced accuracy of 0.76 and an AUC of 0.76. Nevertheless, this finding should be interpreted cautiously because of the limited number of HER-2-positive cases in the test subset and the exploratory nature of the model selection framework.
Because model comparison was performed on the held-out test set, the reported performance estimates may be optimistically biased. Therefore, the results should be interpreted as exploratory and not be considered sufficient to support robust prediction or clinical decision-making. Also, the results require external validation.
For studies in the literature to be comparable and reproducible, it is essential to examine not only the imaging protocols but also the parameters used for texture feature extraction and the statistical methods applied. Various software tools are available for extracting texture features. However, the nomenclature of radiomic features often differs across platforms. Moreover, the choice of parameters during the feature extraction process significantly influences the results. Therefore, studies should provide detailed descriptions of how these parameters are defined to enable meaningful comparisons across studies. A review of the literature reveals that many studies fail to report the texture analysis parameters in sufficient detail, which hampers the reliability of cross-study comparisons. In the present study, the parameters used for feature extraction are described comprehensively. Establishing standardized parameters for texture analysis is crucial to ensure consistency and reproducibility in future research. Recent years have witnessed a substantial increase in texture analysis studies.
In the literature, studies investigating the relationship between ER status and radiomic features, such as those by Ha et al. (23), Antunovic et al. (24), Moscoso et al. (25) and Araz et al. (26) have reported no discriminatory performance. However, Acar et al. (27) reported discriminatory performance for ER using GLRLM_GLNU, GLRLM_RLNU, GLZLM_GLNU and GLZLM_ZLNU features while Soussan et al. (28) reported an association using the HGRE feature. In the present study, despite the high observed sensitivity for ER status, the low specificity, the balanced accuracy of 0.65, and the AUC of 0.59 indicate weak discriminatory performance, particularly for identifying ER-negative tumors.
Likewise, the studies in the literature investigating PR status, such as those by Ha et al. (23), Antunovic et al. (24), Acar et al. (27) and Lemarignier et al. (29) have not demonstrated discriminatory performance. However, studies by Soussan et al. (28) for HGRE feature and Moscoso et al. (25) reported potentially relevant findings for PR status. The present study yielded an AUC of 0.55, a balanced accuracy of 0.61, and an MCC of 0.23. These findings indicate no discriminatory performance with respect to PR status.
For HER-2 status, studies by Ha et al. (23), Soussan et al. (28) and Lemarignie et al. (29) with limited sample sizes did not demonstrate discrimination between HER-2 status and radiomic features, whereas other studies including those by Moscoso et al. (25) and Chen et al. (30), reported varying degrees of discriminatory performance. In the present study, the Random Forest model achieved an AUC of 0.76, a balanced accuracy of 0.76, a sensitivity of 62.5%, and a specificity of 90.0%. These findings suggest a preliminary imaging-molecular association with HER-2 status. Due to the limited number of HER-2-positive cases in the test subset and the absence of independent external validation, these results should be interpreted as hypothesis-generating rather than evidence of clinically reliable predictive utility. Therefore, these findings require confirmation in larger, prospective, multicenter cohorts and independent external validation.
From a biological perspective, while 18F-FDG uptake reflects higher glycolytic activity, which is influenced by cellular proliferation, hypoxia, and oncogenic signaling pathways, HER-2 overexpression is associated with increased tumor aggressiveness and activation of downstream pathways such as PI3K/Akt/mTOR. This pathway is known to regulate cell growth, proliferation and metabolism and may promote glycolysis (31, 32, 33). Consequently, HER-2 positive tumors may demonstrate increased metabolic activity and heterogeneous spatial metabolic patterns on 18F-FDG PET imaging (25, 34).
Notably, some conventional features, such as SUVmean, SUVmax, and SUVpeak, which are among the most relevant, are also selected for HER-2 status. This may suggest that both global metabolic activity and intratumoral heterogeneity contribute to HER-2 imaging patterns. As conventional PET features reflect overall tumor metabolism, radiomic features may contribute by quantifying intratumoral spatial heterogeneity. However, this interpretation remains hypothesis-generating and should be specifically investigated in future studies.
Regarding Ki-67 status, previous studies by Antunovic et al. (24) and Acar et al. (27) have not demonstrated that radiomic features meaningfully discriminate Ki-67 expression. However, Ha et al. (23) reported a significant association (p=0.006). In the present study, the AUC of 0.60, sensitivity of 68.4%, and specificity of 55.2% indicate a lack of discriminatory performance for Ki-67 status.
In this study, several factors, including tumor grade, size, distribution of intrinsic molecular subtypes, and uptake time, may confound 18F-FDG uptake and radiomic features. Although all efforts were made to standardize imaging protocols, the possibility of some residual variability cannot be entirely excluded.
Study Limitations
This study has several limitations. The retrospective, single-center design and the relatively limited sample size restrict the generalizability of these findings. In particular, the small number of HER-2 positive cases in the test subset may have limited the reliability of performance estimates for predicting HER-2 status.
Although feature scaling and SFS were performed exclusively within the training subset, the same held-out test subset was used both to compare candidate models and to select the final model. Therefore, the present workflow does not constitute independent validation; consequently, the performance estimates may be optimistically biased.
Furthermore, lesion segmentation was performed by a single observer, and interobserver variability was not assessed. This may limit the reproducibility of the extracted radiomic features. In addition, a fixed 40% SUVmax threshold was applied to extract radiomic features. Although this is a commonly used method in PET-derived radiomic studies, segmentation based on a fixed threshold may affect the robustness of radiomic features, especially in small or heterogeneous tumors.
This was a single-center study, and no external validation was performed. Therefore, the present findings should be interpreted strictly within an exploratory, hypothesis-generating framework. Prospective, multicenter studies with independent external validation are needed to further confirm and generalize these findings.
Conclusion
In this exploratory study, PET-derived radiomic features demonstrated weak or no discriminatory performance in assessing ER, PR, and Ki-67 status. A preliminary imaging-molecular association for HER-2 status was observed. However, these findings should be interpreted as hypothesis-generating because of the limited number of HER-2-positive cases in the test subset, the model selection framework, and the absence of independent external validation. Therefore, these results do not support robust predictive capability or clinical implementation of 18F-FDG PET radiomics for molecular marker assessment. Larger, prospective, multicenter studies with independent external validation are required to confirm the potential association between 18F-FDG PET radiomic features and HER-2 status.


