Liam M. Longo

Specially Appointed Associate Professor

Earth-Life Science Institute

Institute of Science Tokyo(formerly Tokyo Institute of Technology)

Associate Research Scientist

Blue Marble Space Institute of Science

Recent News

Contrastive learning unites sequence and structure in a global representation of protein space [ Link ]
Guy Yanai, Gabriel Axel, Liam M. Longo*, Nir Ben-Tal*, and Rachel Kolodny*

Establishing a coherent mapping of the relationships among all known proteins is crucial for elucidating processes of protein emergence and evolution. Yet the capacity to fully capture relationships of protein similarity is complicated by the nonstraightforward interplay between sequence and structure; indeed, proteins with unrelated sequences can adopt similar structures, and, conversely, proteins with similar or identical sequences can manifest radically different structures. Here, we introduce Contrastive Learning Sequence–Structure (CLSS), a contrastive protein language model (PLM) trained to coembed sequence and structure information in a self-supervised manner, facilitating a holistic representation of protein relatedness. CLSS represents the structures and sequences of full domains and domain subsequences as vectors in the same high-dimensional latent space. We show that this approach yields meaningful shared representations, which recapitulate the extensive structure- and sequence-based knowledge encoded in human-curated hierarchical protein classification systems (ECOD and CATH). Moreover, the representations generated by CLSS outperform those generated by alternative state-of-the-art PLMs in downstream classification tasks. Notably, we show that even the far larger space of domain subsequences is successfully coembedded, establishing a PLM tailored to these evolutionarily meaningful objects. CLSS embeddings produce informative representations of the protein universe without further downstream processing, as we demonstrate by analyzing preferential associations between protein architectures and ligand types across protein space.

Primitive Basic Amino Acids Promote Mineral-Catalyzed Electrochemical Reduction of H+ and CO2
Siang Chen, Tatsuya Corlett, Norio Kitadai, Ryuhei Nakamura, Masahiro Miyauchi, Liam M. Longo*, and Akira Yamaguchi*

Early in the history of life, mineral surfaces, such as the NiFeS mineral violarite (FeNi2S4), may have been the primary catalysts. How these primitive chemical systems could have undergone robust chemical evolution remains unknown. Here, we show that simple amino acids, such as diaminobutyric acid (DAB), can accelerate H+ and/or CO2 electrochemical reduction on violarite. Reactions performed with structurally similar compounds demonstrate that both amino groups of DAB participate in HER catalytic enhancement. Spectroscopic analyses suggest that the role of DAB is to deliver protons to the mineral surface. The observation that simple basic amino acids can enhance surface-supported electrochemistry reveals that amino acids could have produced a positive feedback that promoted metabolic complexification even before their eventual incorporation into complex peptide or protein catalysts.

The Borderlands of Foldability: Lessons from Simplified Proteins
Koh Seya, Alfie-Louise R. Brownless, Shina Caroline Lynn Kamerlin*, and Liam M. Longo*

Proteins make complex life possible, yet our understanding of their emergence remains limited. What are the informational limits of protein folding, and how did the first proteins emerge? Protein simplification studies—in which contemporary folds are built from limited alphabets, symmetrized, fragmented, or shortened—have provided key insights into these questions. These studies use design constraints to address the discoverability of, and connectedness between, protein folds. By considering various environments, such as high salt concentrations or peptide–nucleic acid coacervates, the role of context in the emergence of folded domains is explored. Taken together, these studies support the early emergence of protein folds and reveal the existence of highly connected and readily traversable regions of sequence–structure space.

A universal polyphosphate kinase powers in vitro transcription [ Link ]
Ryusei Matsumoto^, Takayoshi Watanabe^, Eishin Yamazaki, Ako Kagawa, Liam M. Longo*, and Tomoaki Matsuura*

Polyphosphate kinases (PPKs) catalyze phosphoryl transfer between polyphosphates and nucleotides. Polyphosphates are a cost-effective source of phosphorylating power, making PPKs attractive enzymes for nucleotide production. However, at present, applications that require the simultaneous utilization of diverse nucleotides are not possible due to the restricted substrate profiles of PPKs. Here, we present a universal PPK capable of efficiently phosphorylating all eight common ribonucleotides (purines and pyrimidines, monophosphates and diphosphates) to triphosphates. Under optimal conditions, ~70% triphosphate conversion was observed for each substrate. To demonstrate the biotechnological potential of a universal PPK, we developed a one-pot assay for PPK-powered in vitro transcription. Primitive biology likely relied on enzyme promiscuity to support nascent metabolism with a compact proteome. This work highlights how applying the same principle to synthetic biology can facilitate the construction of complex in vitro reaction systems.

RNA binding and coacervation promote preservation of peptide form and function across the heterochiral-homochiral divide [ Link ]
Manas Seal, Ilan Edelstein, Yosef Scolnik, Orit Weil-Ktorza, Norman Metanis, Yaakov Levy*, Liam M. Longo*, and Daniella Goldfarb*

Recent evidence suggests that peptide-RNA coacervates may have buffered the emergence of folded domains from flexible peptides. As primitive peptides were likely composed of both L- and D-amino acids, we hypothesized that coacervates may have also supported the emergence of chiral control. To test this hypothesis, we compared the coacervation propensities of an isotactic (homochiral) peptide and a syndiotactic (alternating chirality) peptide, both with an identical sequence derived from the ancient helix-hairpin-helix (HhH) motif. Using electron paramagnetic resonance (EPR) spectroscopy and atomistic molecular dynamics (MD) simulations, we found that the syndiotactic peptide does not form stable dimers with high α-helicity in solution, unlike the isotactic peptide. However, both peptides do coacervate with RNA, albeit with distinct reentrant phase behaviors. Coacervation in each case is facilitated by oligomer formation, likely dimerization, upon RNA binding that promotes RNA cross-linking. Additionally, RNA cross-linking and coacervation of the syndiotactic peptide may involve α-helical conformations, according to atomistic MD simulations. Coarse-grained MD simulations indicate that the differences in reentrant phase behavior of isotactic and syndiotactic peptides are associated with differences in dimer flexibility and stability, which modulate the strength of peptide-peptide and peptide-RNA interactions and, consequently, the effectiveness of RNA cross-linking. These results illustrate how RNA binding and/or coacervation by early proteins could have promoted the transition of flexible, heterochiral peptides into folded, homochiral domains.

Exonize: a tool for finding and classifying exon duplications in annotated genome [ Link ]
Marina Herrera Sarrias, Christopher W. Wheat, Liam M. Longo*, and Lars Arvestad*

The protein-coding regions of eukaryotic genes are fragmented into exons that, like the genes within which they are situated, can be duplicated, deleted, or reorganized. Cataloging and organizing within-gene exon similarities is necessary for a systematic study of exon evolution and its consequences. To facilitate the study of exon duplications, we present Exonize, a computational tool that identifies and classifies coding exon duplications in annotated genomes. Exonize implements a graph-based framework to handle clusters of related exons resulting from repeated rounds of exon duplication. The interdependence between duplicated exons or groups of exons across transcripts is classified. By identifying duplication events between exonic and intronic regions, Exonize can detect unannotated or degenerate exons. To aid in data parsing and downstream analysis, the Python module exonize_analysis is provided. The application of Exonize to 20 eukaryote genomes identifies full-exon duplications in at least 4% of vertebrate genes, with more than 900 human genes having a full-exon duplication event.

How Nature Chooses Phosphate to Make Life Possible
Shina Caroline Lynn Kamerlin (PI), Liam M. Longo (Co-I), Marie Skepö (collaborator)

If you are an experimentalist interested in studying protein histories as a graduate student in the Longo Lab at ELSI, please get in touch!