Data-Driven Framework for Rational Design of Gold Nanoclusters (AuNC)

Negar Molavi and Farnaz Heidar-Zadeh

Department of Chemistry, Queen's University, Kingston, Ontario, Canada

Gold nanoclusters (AuNCs), especially those stabilized by N-heterocyclic carbenes (NHC), are stable metal-based compounds that have attracted significant attention due to their tunable photophysical properties and potential applications in targeted cancer therapeutics. One of the main challenges in using NHC-stabilized AuNCs for light-mediated cancer therapy is the limited penetration of light through biological tissues due to scattering and absorption. As a result, identifying compounds with favorable photophysical properties in the near-infrared (NIR) region is a key challenge, as NIR light exhibits deeper tissue penetration and reduced scattering. Hence, researchers generally seek to develop compounds with maximized emission wavelengths. However, the exploration and optimization of NHC ligands through traditional experimental and computational methods can be both time-consuming and resource-intensive. To overcome these challenges, this project aims to develop a data-driven framework for chemical property prediction and accelerated discovery of NHC-AuNCs. We have created a dataset of more than 2500 NHC ligands and fragments derived from literature and synthesized compounds. These structures are then converted into Self- Referencing Embedded Strings (SELFIES), a robust molecular representation that guarantees the generation of chemically valid molecules by enforcing valence constraints. SELFIES are superior to the typically used Simplified Molecular Input Line Entry System (SMILES) representations, because not all of the SMILES strings correspond to a chemically valid structure. We use SELFIES as input to train a variational autoencoder (VAE), which learns a continuous latent representation of the ligand structures. During training, the molecular rep- resentations are compressed into low-dimensional latent vectors that capture the underlying structural features of the ligands. We then combine the learned latent space representations with the Bayesian optimization to identify ligands predicted to exhibit the highest emission wavelengths. This strategy allows us to leverage existing experimental and computational data to efficiently navigate the chemical space and identify promising candidate ligands for further experimental investigations.

Back to List of Abstracts