Sgueglia, Gianmattia (2024) Novel machine learning-based approaches for metalloprotein design and classification. [Tesi di dottorato]
|
Documento PDF
thesis_modified_hd_gianmattia_sgueglia.pdf Visibile a [TBR] Amministratori dell'archivio Download (17MB) | Richiedi una copia |
| Tipologia del documento: | Tesi di dottorato |
|---|---|
| Lingua: | English |
| Titolo: | Novel machine learning-based approaches for metalloprotein design and classification |
| Autori: | Autore Email Sgueglia, Gianmattia gianmattia.sgueglia@unina.it |
| Data: | 12 Dicembre 2024 |
| Numero di pagine: | 193 |
| Istituzione: | Università degli Studi di Napoli Federico II |
| Dipartimento: | Scienze Chimiche |
| Dottorato: | Scienze chimiche |
| Ciclo di dottorato: | 37 |
| Coordinatore del Corso di dottorato: | nome email Lombardi, Angelina alombard@unina.it |
| Tutor: | nome email Lombardi, Angelina [non definito] Chino, Marco [non definito] |
| Data: | 12 Dicembre 2024 |
| Numero di pagine: | 193 |
| Parole chiave: | Protein design, deep learning, artificial intelligence, language models |
| Settori scientifico-disciplinari del MIUR: | Area 03 - Scienze chimiche > CHIM/03 - Chimica generale e inorganica |
| Informazioni aggiuntive: | 37° Ciclo di Dottorato in Scienze Chimiche |
| Depositato il: | 20 Gen 2026 19:20 |
| Ultima modifica: | 12 Ago 2026 05:38 |
| URI: | https://www.fedoa.unina.it/id/eprint/16528 |
Abstract
Metalloproteins are among nature's most versatile and powerful tools, routinely carrying out exceptionally challenging reactions that chemists can only approximate in controlled lab settings. These proteins operate with remarkable specificity and efficiency, catalysing complex processes such as nitrogen fixation, oxygen transport, and electron transfer—reactions essential to life and foundational to fields like energy production and environmental sustainability. Their ability to bind metal ions and harness their reactivity makes metalloproteins uniquely capable of performing chemical transformations that are otherwise completely unfeasible in normal conditions. Protein design seeks to capture this power, providing a theoretical framework and advanced computational tools to re-create or even improve upon these natural capabilities. By understanding and manipulating the principles governing metalloprotein structure and function, custom proteins can be designed to fold into precise structures, bind specifically to metal ions, and interact predictably with substrates or other chemical entities. Such tailor-made proteins hold immense potential for diverse applications, from sustainable catalysis and drug design to biosensing and synthetic biology. Ultimately, protein design promises to expand the possibilities of metalloproteins beyond natural evolution, opening doors to ad hoc proteins engineered to meet the specific demands of modern science and industry. In this context, the aim of this thesis is to expand the repertoire of tools available to design metalloproteins using state-of-the-art methodologies based on machine learning, which have already been applied to the field of protein design, in addition to many others, with great success. This work recapitulates the efforts to develop and study different models for specific applications related to metalloproteins, in addition to also including some work based on classical methodologies in an attempt to overcome some of the limitations of machine learning models. Chapter 2 dives into the creation of models to perform geometry and coordination number classification of metal sites both in proteins and small molecule complexes. To this end, large structural databases like the Cambridge Structural Database (CSD) and the Protein Data Bank (PDB), are leveraged to accumulate the data required to train an artificial neural network to recognize and classify the geometry of a metal site. This chapter describes the rigorous process of data curation and the preparation of input features to capture spatial patterns relevant to the task of metal site classification, showing that this approach can be successfully translated from abiogenic sites to bioinorganic ones. The identification and role of metal-site distortions is also further investigated and related to the observed performance of the model. Chapter 3 shifts to the development of fine-tuned protein language models that generate metalloprotein sequences. The potential advantages and disadvantages of using synthetic data in the context of protein language models are discussed and the effect of different strategies for data generation is investigated in relation to the type and quality of sequences obtained by the models. This chapter reveals how capable such models are at introducing meaningful variability, enabling the generation of new protein sequences resembling natural ones, but with modifications which may have relevant functional implications. Chapter 4 presents a case study applying the current state-of-the-art computational methods to build a pipeline for protein design. Here, an alpha-helical peptide is designed to bind a cobalt-oxo cubane cluster, a complex with potential applications in catalysis. This chapter covers the preliminary experimental validation of the designed peptide, which was synthesized, purified, and characterized spectroscopically. The results demonstrate the peptide’s improved physico-chemical properties compared to its predecessors and its capacity to bind the cluster, although the desired cluster-peptide assembly is not the only species formed when the two react together.
Downloads
Downloads per month over past year
Actions (login required)
![]() |
Modifica documento |


