Ding T., Larrea-Gallegos G., Busio F., Marvuglia A., Schaubroeck T.
Rsc Sustainability, 2026
The rapid expansion of registered chemicals, coupled with persistent data gaps, poses a major challenge for toxicity assessment in life cycle assessment (LCA) and Safe and Sustainable by Design (SSbD). This study proposes a data-driven framework to directly predict toxicity characterization factors (CFs) from the molecular simplified molecular input line entry system (SMILES), using the environmental footprint (EF) v3.1 database as the training benchmark. We evaluate five machine-learning and deep-learning approaches—random forest, XGBoost, Gaussian process, deep neural networks, and graph neural networks via message-passing neural networks (MPNNs)—across three molecular representations: Mordred descriptors, molecular graphs, and large-scale pretrained molecular embeddings (GROVER). Predictive performance is strongly target-dependent, with ecotoxicity CFs showing consistently higher predictability (R<sup>2</sup> = 0.47–0.67) than human toxicity CFs (R<sup>2</sup> = 0.44–0.56). Mordred-based models, particularly XGBoost, shows robust and better performance across multiple targets. Graph-based MPNNs achieved competitive performance, with graph-only multi-target MPNNs showing the clearest benefit over single-target training, especially for human toxicity targets. Adding Mordred descriptors to graph-based models generally improved human toxicity prediction, but can sometimes reduce performance for ecotoxicity prediction. GROVER embeddings provided advantages in specific clusters (e.g., the highest mean R<sup>2</sup> over three runs for one cluster is 0.70) and offer a promising alternative to handcrafted descriptors. The framework further integrates applicability domain analysis and chemical clustering to enable domain-consistent prediction. A textile-sector case study shows that incorporating predicted CFs for previously uncovered chemicals can lead to substantial underestimation of toxicity impacts, by up to about 60%.
