Bin-Wise Quantized CatBoost for Energy-Efficient Drug-Review Rating Prediction in Biomedical E-Commerce
DOI:
https://doi.org/10.56979/1102/2026/1477Keywords:
Artificial Intelligence, Green Machine Learning, Gradient Boosting, CatBoost, Feature Bin Reduction, Carbon Footprint, Drug Review Prediction, Sentence Embeddings, Equivalence Testing, Biomedical E-CommerceAbstract
Artificial intelligence is increasingly applied on biomedical e-commerce platforms such as online drug-review aggregators, and machine learning provides most of its practical capabilities. Ma-chine learning models for predicting patient-generated ratings on such platforms are typically optimized for accuracy alone, while the energy and carbon cost of training is rarely reported. This study examines whether bin-wise quantization in CatBoost. reducing the feature bin count. can lower the computational and carbon cost of training such models without degrading predictive accuracy. Two public drug-review datasets are used: the UCI Drugs.com corpus (50,000 reviews, rating 1–10) and the WebMD corpus (50,000 reviews, satisfaction 1–5). Each review is represented by a 384-dimensional sentence embedding of its text, eight text-length and style statistics, a sentiment score, and two tabular features. CatBoost is used as the primary model because it produced markedly more stable generalization than the other libraries, and its bin count is directly exposed to the practitioner. Five CatBoost-based models are first compared per dataset under five-fold cross-validation: Ridge regression, standard CatBoost, and CatBoost with the feature bin count fixed at 64, 128, and 255. Hyperparameters are held constant across CatBoost variants to isolate the effect of the bin count. Accuracy equivalence is assessed with two one-sided tests (TOST) at a predefined margin of ±0.01 R², and training time and CO₂ emissions are tracked with CodeCarbon. The test R² values of the bin-reduced and standard CatBoost are statistically equivalent across both datasets and all three bin counts (TOST, all p < 0.001), while 64-bin reduction lowers training time and CO₂ emissions by around 48%. To test whether the effect is specific to CatBoost, LightGBM and XGBoost are evaluated under the same protocol; bin reduction lowers their training cost by roughly 25–60% while remaining statistically equivalent in accuracy (TOST, all p < 0.002), con-firming that the effect is general to histogram-based gradient boosting. SHAP analysis shows the review-text embedding is the dominant predictor, with sentiment and engagement as secondary signals. The work is presented as a sustainability-oriented empirical study: explicit control of the bin-count parameter exposes a favorable accuracy–efficiency trade-off applicable to biomedical review marketplaces.
Downloads
Published
How to Cite
Issue
Section
License
This is an open Access Article published by Research Center of Computing & Biomedical Informatics (RCBI), Lahore, Pakistan under CCBY 4.0 International License




