z-logo
open-access-imgOpen Access
PhaMMseqs: a new pipeline for constructing phage gene phamilies using MMseqs2
Author(s) -
Christian H. Gauthier,
Steven G. Cresawn,
Graham F. Hatfull
Publication year - 2022
Publication title -
g3 genes genomes genetics
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 1.468
H-Index - 66
ISSN - 2160-1836
DOI - 10.1093/g3journal/jkac233
Subject(s) - biology , pipeline (software) , computational biology , genetics , evolutionary biology , programming language , computer science
The diversity and mosaic architecture of phage genomes present challenges for whole-genome phylogenies and comparative genomics. There are no universally conserved core genes, ∼70% of phage genes are of unknown function, and phage genomes are replete with small (<500 bp) open reading frames. Assembling sequence-related genes into "phamilies" ("phams") based on amino acid sequence similarity simplifies comparative phage genomics and facilitates representations of phage genome mosaicism. With the rapid and substantial increase in the numbers of sequenced phage genomes, computationally efficient pham assembly is needed, together with strategies for including newly sequenced phage genomes. Here, we describe the Python package PhaMMseqs, which uses MMseqs2 for pham assembly, and we evaluate the key parameters for optimal pham assembly of sequence- and functionally related proteins. PhaMMseqs runs efficiently with only modest hardware requirements and integrates with the pdm_utils package for simple genome entry and export of datasets for evolutionary analyses and phage genome map construction.

The content you want is available to Zendy users.

Already have an account? Click here to sign in.
Having issues? You can contact us here
Accelerating Research

Address

John Eccles House
Robert Robinson Avenue,
Oxford Science Park, Oxford
OX4 4GP, United Kingdom