Sgueglia, Gianmattia (2024) Novel machine learning-based approaches for metalloprotein design and classification. [Tesi di dottorato]

[thumbnail of thesis_modified_hd_gianmattia_sgueglia.pdf] Documento PDF
thesis_modified_hd_gianmattia_sgueglia.pdf
Visibile a [TBR] Amministratori dell'archivio

Download (17MB) | Richiedi una copia
Tipologia del documento: Tesi di dottorato
Lingua: English
Titolo: Novel machine learning-based approaches for metalloprotein design and classification
Autori:
Autore
Email
Sgueglia, Gianmattia
gianmattia.sgueglia@unina.it
Data: 12 Dicembre 2024
Numero di pagine: 193
Istituzione: Università degli Studi di Napoli Federico II
Dipartimento: Scienze Chimiche
Dottorato: Scienze chimiche
Ciclo di dottorato: 37
Coordinatore del Corso di dottorato:
nome
email
Lombardi, Angelina
alombard@unina.it
Tutor:
nome
email
Lombardi, Angelina
[non definito]
Chino, Marco
[non definito]
Data: 12 Dicembre 2024
Numero di pagine: 193
Parole chiave: Protein design, deep learning, artificial intelligence, language models
Settori scientifico-disciplinari del MIUR: Area 03 - Scienze chimiche > CHIM/03 - Chimica generale e inorganica
Informazioni aggiuntive: 37° Ciclo di Dottorato in Scienze Chimiche
Depositato il: 20 Gen 2026 19:20
Ultima modifica: 12 Ago 2026 05:38
URI: https://www.fedoa.unina.it/id/eprint/16528

Abstract

Metalloproteins are among nature's most versatile and powerful tools, routinely carrying out exceptionally challenging reactions that chemists can only approximate in controlled lab settings. These proteins operate with remarkable specificity and efficiency, catalysing complex processes such as nitrogen fixation, oxygen transport, and electron transfer—reactions essential to life and foundational to fields like energy production and environmental sustainability. Their ability to bind metal ions and harness their reactivity makes metalloproteins uniquely capable of performing chemical transformations that are otherwise completely unfeasible in normal conditions. Protein design seeks to capture this power, providing a theoretical framework and advanced computational tools to re-create or even improve upon these natural capabilities. By understanding and manipulating the principles governing metalloprotein structure and function, custom proteins can be designed to fold into precise structures, bind specifically to metal ions, and interact predictably with substrates or other chemical entities. Such tailor-made proteins hold immense potential for diverse applications, from sustainable catalysis and drug design to biosensing and synthetic biology. Ultimately, protein design promises to expand the possibilities of metalloproteins beyond natural evolution, opening doors to ad hoc proteins engineered to meet the specific demands of modern science and industry. In this context, the aim of this thesis is to expand the repertoire of tools available to design metalloproteins using state-of-the-art methodologies based on machine learning, which have already been applied to the field of protein design, in addition to many others, with great success. This work recapitulates the efforts to develop and study different models for specific applications related to metalloproteins, in addition to also including some work based on classical methodologies in an attempt to overcome some of the limitations of machine learning models. Chapter 2 dives into the creation of models to perform geometry and coordination number classification of metal sites both in proteins and small molecule complexes. To this end, large structural databases like the Cambridge Structural Database (CSD) and the Protein Data Bank (PDB), are leveraged to accumulate the data required to train an artificial neural network to recognize and classify the geometry of a metal site. This chapter describes the rigorous process of data curation and the preparation of input features to capture spatial patterns relevant to the task of metal site classification, showing that this approach can be successfully translated from abiogenic sites to bioinorganic ones. The identification and role of metal-site distortions is also further investigated and related to the observed performance of the model. Chapter 3 shifts to the development of fine-tuned protein language models that generate metalloprotein sequences. The potential advantages and disadvantages of using synthetic data in the context of protein language models are discussed and the effect of different strategies for data generation is investigated in relation to the type and quality of sequences obtained by the models. This chapter reveals how capable such models are at introducing meaningful variability, enabling the generation of new protein sequences resembling natural ones, but with modifications which may have relevant functional implications. Chapter 4 presents a case study applying the current state-of-the-art computational methods to build a pipeline for protein design. Here, an alpha-helical peptide is designed to bind a cobalt-oxo cubane cluster, a complex with potential applications in catalysis. This chapter covers the preliminary experimental validation of the designed peptide, which was synthesized, purified, and characterized spectroscopically. The results demonstrate the peptide’s improved physico-chemical properties compared to its predecessors and its capacity to bind the cluster, although the desired cluster-peptide assembly is not the only species formed when the two react together.

Downloads

Downloads per month over past year

Actions (login required)

Modifica documento Modifica documento