"Detection, evolutionary history and structure-function relationships of orphan genes..."

"Detection, evolutionary history and structure-function relationships of orphan genes in plant parasitic nematodes"

25 September 2026

2:00 p.m Sophia Antipolis - INRAE PACA - A010

Ercan Seçkin, doctoral student in the GAME team. will defend her thesis on Friday 25 September 2026 at 2.00 pm, room « Euler Violet » au Centre Inria d'Université Côte d'Azur à Sophia Antipolis

President of the jury :                       

  • To be decided

Rapporteurs :                     

  • Prof. Erich Bornberg-Bauer, PR, Universität Münster | Dr. Hamed Khakzad, CPJ, Inria          

Examiners :

  • Dr. Jean-Luc Pellequer, DR, IBS | Dr. Richard Copley, DR, IMEV
  • Dr. Mathilde Carpentier, MdC, ISYEB | Dr. Anna Grandchamp, MdC, TAGC  

Thesis Directors :

  • Dr. Etienne GJ Danchin, DR, ISA | Dr. Edoardo Sarti, CR, Inria 
  • Dr. Dominique Colinet, MdC, ISA

Abstract :

Orphan genes are genes that lack detectable homologs outside a restricted taxonomic group. They are widespread in eukaryotic genomes and may contribute to lineage-specific innovation, but their origins, evolution and functions remain difficult to study because classical homology-based approaches cannot be used. This thesis investigates orphan genes in root-knot nematodes of the genus Meloidogyne, which are major plant parasites with unusual genome biology and significant economic impact. The aims were to characterize their prevalence, evolutionary origin, molecular properties, structural features and possible functions.

First, a comparative genomics framework was developed to identify transcriptionally supported orphan genes across the Meloidogyne genus. Orphan genes accounted for at least 16% of coded proteins and their emergence showed several phases of acceleration along the phylogeny. By combining homology searches, ancestral sequence reconstruction and synteny analysis, nearly 20% of orphan genes were inferred to have emerged de novo from previously non-genic regions, whereas another 20% likely diverged from preexisting genes beyond detectable sequence homology. Orphan genes tended to encode short proteins, in which signal peptides for secretion were frequent. Furthermore, these genes were preferentially expressed at the pre-parasitic juvenile stage, suggesting that some contribute to the secreted parasitic arsenal of root-knot nematodes. Machine learning classifiers further showed that orphan versus non-orphan proteins, and de novo versus highly diverged orphan proteins, can be reliably distinguished using nucleotide and protein-level sequence features.

Second, the thesis evaluated whether transformer-based models can provide reliable structural information for orphan proteins. Using the curated Meloidogyne orphan dataset, predictions from state-of-the-art 3D structure predictors were compared between orphan and non-orphan proteins. 3D structure predictions for orphan proteins had lower confidence and poor agreement between methods, regardless of multiple sequence alignment usage. This limitation could not be explained solely by intrinsic disorder, as orphan proteins were not generally more disordered than the other proteins. Predicted 3D structures for orphan proteins also rarely matched known structures, further limiting their interpretation. However, secondary structure content was more consistently recovered across methods, indicating that orphan proteins may retain usable local structural signals even when full three-dimensional folds cannot be reliably inferred.

Finally, this observation motivated an alternative strategy based on secondary structure representations and profile hidden Markov models (pHMMs). Instead of comparing amino acid sequences or uncertain tertiary structures, proteins were represented as ordered sequences of secondary structure elements (SSEs), encoded as vectors containing information about each element. pHMMs were adapted to work with these SSE vectors, trained from secondary structure alignments of known structural domain families, and calibrated with a dedicated null model and e-value framework. This approach aimed to test whether fold-level signals can be detected in orphan proteins without relying on sequence homology or fully reliable 3D models. Preliminary results showed that the model classifies known domain folds with high performance. When applied to orphan proteins, the results are promising but still require further exploration.

Overall, this thesis shows that orphan genes are abundant, evolutionarily heterogeneous and biologically important components of root-knot nematode genomes. It also demonstrates that they represent a stringent test case for current homology detection and structure prediction methods. By combining comparative genomics, synteny, transcriptomic and proteomic evidence, machine learning, structure prediction and HMM-based modelling, this work provides a framework for studying evolutionary novelty beyond the limits of sequence similarity.

Keywords :

Orphan genes, de novo gene birth, evolution, phylogeny, multi-omics, protein structure, machine learning, transformers

Contact: animisa@inrae.fr