Introduction

nf-core/proteinfamilies is a bioinformatics pipeline that generates protein families from amino acid sequences and/or updates existing families with new sequences. It takes a protein fasta file as input, clusters the sequences and then generates protein family Hidden Markov Models (HMMs) along with their multiple sequence alignments (MSAs). Optionally, existing family HMMs (and their seed and/or full MSAs) can be given in order to update those families with new sequences in case of matching hits.

nf-core/proteinfamilies workflow overviewnf-core/proteinfamilies workflow overview

Check quality and pre-process

Generate input amino acid sequence statistics with (SeqFu) and pre-process them (i.e., gap removal, convert to upper case, validate, filter by length, replace special characters such as /, and remove duplicate sequences) with (SeqKit)

Create families

  1. Cluster sequences (MMseqs2)
  2. Perform multiple sequence alignment (MSA) (FAMSA or mafft)
  3. Optionally, clip gap parts of the seed MSA (ClipKIT), recalculating the sequence coordinates
  4. Generate family HMMs and fish additional sequences into the family (hmmer)
  5. Optionally, remove redundant and/or merge similar families by comparing family representative sequences against family models with (hmmer)
  6. Optionally, from the remaining families, remove in-family redundant sequences by strictly clustering with (MMseqs2) and discarding non-cluster representatives
  7. If in-family redundancy was not removed, reformat the .sto full MSAs to .fas, for downstream analyses compatibility, with (HH-suite3)
  8. Optionally, infer sequence phylogeny, by calculating the maximum parsimonious likelihood estimation trees for the final full MSAs with (CMAPLE)
  9. Present statistics for remaining/updated family size distributions and representative sequence lengths (MultiQC)

The new iterative family generation algorithm (i.e., mgnifam), can now be used in place of steps 2-4. In a nutshell, mgnifam iterates through these steps to refine the protein family alignments and models, for up to three rounds, or until the family model has converged. For more information, see usage.md

Update families

  1. Pool the input sequences with the members of existing full MSAs, if given, and find which families to update by comparing them against existing family models with (hmmer)
  2. Optionally, remove in-family redundant hits by strictly clustering with (MMseqs2) and keeping cluster representatives
  3. Rebuild each hit family like a new one, following steps 2-4 of Create families with the standard algorithm (always used for updates): new seed MSA and HMM, and new full MSA recruited from the same pool, reformatted to .fas (HH-suite3)
  4. Input fasta sequences that no updated family holds (including those recruited by a rebuilt HMM) continue in the Create families paragraph above; members of existing full MSAs never do

Prepare downstream samplesheets

Optionally, prepare the downstream samplesheets for the nf-core/proteinfold and nf-core/proteinannotator pipelines.

Usage

Note

If you are new to Nextflow and nf-core, please refer to this page on how to set-up Nextflow. Make sure to test your setup with -profile test before running the workflow on actual data.

First, prepare a samplesheet with your input data that looks as follows:

samplesheet.csv:

id,fasta,existing_hmms,existing_seed_msas,existing_full_msas
CONTROL_REP1,input/mgnifams_input_small.faa,,,

Each row contains a fasta file with amino acid sequences (gzipped or uncompressed). Optionally, a row may contain existing families’ HMMs (a tar.gz archive, or one HMM library such as a previous run’s <id>.lib.gz), and optionally tar.gz archives of their seed and/or full MSAs, in order to be updated. Each HMM NAME is a family (unique, case-sensitive), and each seed or full MSA file must be named after one (same base filename, not the extension). Hit families will be updated, while input fasta sequences in no updated family will create new families (members of existing full MSAs never do). Every run also archives each sample’s final families under archives/<id>/, in the same tar.gz shape, so they can be updated again later.

Now, you can run the pipeline using:

nextflow run nf-core/proteinfamilies \
-profile <docker/singularity/.../institute> \
--input samplesheet.csv \
--outdir <OUTDIR>
Warning

Please provide pipeline parameters via the CLI or Nextflow -params-file option. Custom config files including those provided by the -c Nextflow option can be used to provide any configuration except for parameters; see docs.

For more details and further functionality, please refer to the usage documentation and the parameter documentation.

Pipeline output

To see the results of an example test run with a full size dataset refer to the results tab on the nf-core website pipeline page. For more details about the output files and reports, please refer to the output documentation.

Credits

nf-core/proteinfamilies was originally written by Evangelos Karatzas.

We thank the following people for their extensive assistance in the development of this pipeline:

Contributions and Support

If you would like to contribute to this pipeline, please see the contributing guidelines.

For further information or help, don’t hesitate to get in touch on the Slack #proteinfamilies channel (you can join with this invite).

Citations

If you use nf-core/proteinfamilies for your analysis, please cite the article as follows:

nf-core/proteinfamilies: A scalable pipeline for the generation of protein families.

Evangelos Karatzas, Martin Beracochea, Fotis A. Baltoumas, Eleni Aplakidou, Lorna Richardson, James A. Fellows Yates, Daniel Lundin, nf-core community, Aydin Buluç, Nikos C. Kyrpides, Ilias Georgakopoulos-Soares, Georgios A. Pavlopoulos & Robert D. Finn

GigaScience. 2026 Jan. doi: 10.1093/gigascience/giag009.

You can cite the nf-core/proteinfamilies zenodo record for a specific version using the following doi: 10.5281/zenodo.14881993.

An extensive list of references for the tools used by the pipeline can be found in the CITATIONS.md file.

You can cite the nf-core publication as follows:

The nf-core framework for community-curated bioinformatics pipelines.

Philip Ewels, Alexander Peltzer, Sven Fillinger, Harshil Patel, Johannes Alneberg, Andreas Wilm, Maxime Ulysse Garcia, Paolo Di Tommaso & Sven Nahnsen.

Nat Biotechnol. 2020 Feb 13. doi: 10.1038/s41587-020-0439-x.