Introduction
nf-core/proteinfamilies is a bioinformatics pipeline that generates protein families from amino acid sequences and/or updates existing families with new sequences. It takes a protein fasta file as input, clusters the sequences and then generates protein family Hidden Markov Models (HMMs) along with their multiple sequence alignments (MSAs). Optionally, existing family HMMs (and their seed and/or full MSAs) can be given in order to update those families with new sequences in case of matching hits.
Check quality and pre-process
Generate input amino acid sequence statistics with (SeqFu) and pre-process them (i.e., gap removal, convert to upper case, validate, filter by length, replace special characters such as /, and remove duplicate sequences) with (SeqKit)
Create families
- Cluster sequences (
MMseqs2) - Perform multiple sequence alignment (MSA) (
FAMSAormafft) - Optionally, clip gap parts of the seed MSA (
ClipKIT), recalculating the sequence coordinates - Generate family HMMs and fish additional sequences into the family (
hmmer) - Optionally, remove redundant and/or merge similar families by comparing family representative sequences against family models with (
hmmer) - Optionally, from the remaining families, remove in-family redundant sequences by strictly clustering with (
MMseqs2) and discarding non-cluster representatives - If in-family redundancy was not removed, reformat the
.stofull MSAs to.fas, for downstream analyses compatibility, with (HH-suite3) - Optionally, infer sequence phylogeny, by calculating the maximum parsimonious likelihood estimation trees for the final full MSAs with (
CMAPLE) - Present statistics for remaining/updated family size distributions and representative sequence lengths (
MultiQC)
The new iterative family generation algorithm (i.e., mgnifam), can now be used in place of steps 2-4.
In a nutshell, mgnifam iterates through these steps to refine the protein family alignments and models,
for up to three rounds, or until the family model has converged. For more information, see usage.md
Update families
- Pool the input sequences with the members of existing full MSAs, if given, and find which families to update by comparing them against existing family models with (
hmmer) - Optionally, remove in-family redundant hits by strictly clustering with (
MMseqs2) and keeping cluster representatives - Rebuild each hit family like a new one, following steps 2-4 of
Create familieswith thestandardalgorithm (always used for updates): new seed MSA and HMM, and new full MSA recruited from the same pool, reformatted to.fas(HH-suite3) - Input fasta sequences that no updated family holds (including those recruited by a rebuilt HMM) continue in the
Create familiesparagraph above; members of existing full MSAs never do
Prepare downstream samplesheets
Optionally, prepare the downstream samplesheets for the nf-core/proteinfold and nf-core/proteinannotator pipelines.
Usage
If you are new to Nextflow and nf-core, please refer to this page on how to set-up Nextflow. Make sure to test your setup with -profile test before running the workflow on actual data.
First, prepare a samplesheet with your input data that looks as follows:
samplesheet.csv:
id,fasta,existing_hmms,existing_seed_msas,existing_full_msasCONTROL_REP1,input/mgnifams_input_small.faa,,,Each row contains a fasta file with amino acid sequences (gzipped or uncompressed).
Optionally, a row may contain existing families’ HMMs (a tar.gz archive, or one HMM library such as a previous run’s <id>.lib.gz), and optionally tar.gz archives of their seed and/or full MSAs, in order to be updated.
Each HMM NAME is a family (unique, case-sensitive), and each seed or full MSA file must be named after one (same base filename, not the extension).
Hit families will be updated, while input fasta sequences in no updated family will create new families (members of existing full MSAs never do).
Every run also archives each sample’s final families under archives/<id>/, in the same tar.gz shape, so they can be updated again later.
Now, you can run the pipeline using:
nextflow run nf-core/proteinfamilies \ -profile <docker/singularity/.../institute> \ --input samplesheet.csv \ --outdir <OUTDIR>Please provide pipeline parameters via the CLI or Nextflow -params-file option. Custom config files including those provided by the -c Nextflow option can be used to provide any configuration except for parameters; see docs.
For more details and further functionality, please refer to the usage documentation and the parameter documentation.
Pipeline output
To see the results of an example test run with a full size dataset refer to the results tab on the nf-core website pipeline page. For more details about the output files and reports, please refer to the output documentation.
Credits
nf-core/proteinfamilies was originally written by Evangelos Karatzas.
We thank the following people for their extensive assistance in the development of this pipeline:
Contributions and Support
If you would like to contribute to this pipeline, please see the contributing guidelines.
For further information or help, don’t hesitate to get in touch on the Slack #proteinfamilies channel (you can join with this invite).
Citations
If you use nf-core/proteinfamilies for your analysis, please cite the article as follows:
nf-core/proteinfamilies: A scalable pipeline for the generation of protein families.
Evangelos Karatzas, Martin Beracochea, Fotis A. Baltoumas, Eleni Aplakidou, Lorna Richardson, James A. Fellows Yates, Daniel Lundin, nf-core community, Aydin Buluç, Nikos C. Kyrpides, Ilias Georgakopoulos-Soares, Georgios A. Pavlopoulos & Robert D. Finn
GigaScience. 2026 Jan. doi: 10.1093/gigascience/giag009.
You can cite the nf-core/proteinfamilies zenodo record for a specific version using the following doi: 10.5281/zenodo.14881993.
An extensive list of references for the tools used by the pipeline can be found in the CITATIONS.md file.
You can cite the nf-core publication as follows:
The nf-core framework for community-curated bioinformatics pipelines.
Philip Ewels, Alexander Peltzer, Sven Fillinger, Harshil Patel, Johannes Alneberg, Andreas Wilm, Maxime Ulysse Garcia, Paolo Di Tommaso & Sven Nahnsen.
Nat Biotechnol. 2020 Feb 13. doi: 10.1038/s41587-020-0439-x.
