Introduction
nf-core/proteinfamilies is a bioinformatics pipeline that generates protein families from amino acid sequences and/or updates existing families with new sequences. It takes a protein fasta file as input, clusters the sequences and then generates protein family Hidden Markov Models (HMMs) along with their multiple sequence alignments (MSAs). Optionally, existing family HMMs (and their seed and/or full MSAs) can be given in order to update those families with new sequences in case of matching hits.
Samplesheet input
You will need to create a samplesheet with information about the samples you would like to analyse before running the pipeline. Use this parameter to specify its location. It has to be a comma-separated file with 2 mandatory and 3 optional columns, and a header row as shown in the examples below.
--input '[path to samplesheet file]'id,fasta,existing_hmms,existing_seed_msas,existing_full_msasCONTROL_REP1,amino_acid_sequences_input.faa,,,CONTROL_REP2,amino_acid_sequences_extra.faa.gz,existing_hmms.tar.gz,,CONTROL_REP3,amino_acid_sequences_extra.faa.gz,existing_hmms.tar.gz,existing_seed_msas.tar.gz,existing_full_msas.tar.gz| Column | Description |
|---|---|
id |
Custom sample name. Only letters, digits, dots (.), underscores (_) and dashes (-) are allowed, since the sample name is used to build output paths and to parse family names back out of them. |
fasta |
Full path to amino acid fasta file. Allowed extensions are “.faa”, “.fasta” and “.fa”, with or without a following “.gz” for gzipped files. |
existing_hmms |
(Optional) Full path to a “.tar.gz” archive of HMM files, or to one HMM library (“.hmm”, “.hmm.gz”, “.lib” or “.lib.gz”, e.g. this pipeline’s <id>.lib.gz). A sample with existing HMMs is updated: its sequences are searched against these families, and only the sequences without hits go on to create new families. |
existing_seed_msas |
(Optional, needs existing_hmms) Full path to a “.tar.gz” archive with seed MSAs (aligned FASTA, as this pipeline writes them; optionally gzipped), each named after its family’s HMM NAME (as file name stem). |
existing_full_msas |
(Optional, needs existing_hmms) Full path to a “.tar.gz” archive with full MSAs (aligned FASTA, as this pipeline writes them, or Stockholm; optionally gzipped), each named after its family’s HMM NAME. Their members are pooled with the input sequences and searched again, so families keep the old members that still hit. |
Input sequences named <sequence>/<start>-<end> (Pfam convention) are treated as slices of <sequence>: family members cut from them are named in the parent sequence’s coordinates (a hit on residues 3-180 of seqA/10-200 becomes seqA/12-189). Any other name is taken as a full protein.
Updating existing families
Each model in existing_hmms is an existing family, identified by its NAME (case-sensitive); HMM files may hold one or more models, and their file names do not matter. Seed and full MSAs must be named after that NAME, without extensions (e.g. fam_1.aln for NAME fam_1). Not every family needs an MSA, but every MSA file needs a family.
hmmsearch reports hits by HMM NAME, while MSAs are matched to their family by file name. The pipeline stops if two models share a NAME, if two MSA files give the same family, or if an MSA file is not named after an existing HMM NAME. It also stops if an existing family is named like the families this run creates for the sample (<id>_<number>..., e.g. after a previous run with the same id): use a new id for the update run, such as <id>_r2.
HMM and MSA archives written by nf-core/proteinfamilies already follow these rules and can be used as they are.
The input sequences, together with the members of any existing_full_msas (gaps removed; a member seq/<start>-<end> is skipped if its region lies inside an input sequence of the same protein seq, where a name without a range is the whole protein, or inside another member; partial overlaps are kept), are searched against the existing HMMs.
Each family’s hits are then rebuilt like a newly created family: optionally made non-redundant, aligned and trimmed into a new seed MSA, built into a new HMM, and used to recruit the new full MSA from the same pool (the new seed MSA serves as the full MSA with --skip_additional_sequence_recruiting).
With --skip_update_refinement, the existing HMMs are kept instead: each one aligns its hits into the new full MSA (hmmalign), and its existing_seed_msas file, if given, passes through unchanged. Seed MSAs are never searched, so sequences only found in a seed MSA must also be in the fasta or in a full MSA to stay in their family.
Families without any hit, or whose rebuilt HMM recruits nothing, are left unchanged (their existing HMM, seed and full MSA pass through) and listed with the reason in update_families/unchanged_families/<id>_unchanged_existing_families.tsv.
Input sequences that end up in no updated family go to family creation (a sample without any hit sends all of them); members of existing full MSAs that no family holds anymore are dropped and never create new families.
Every run writes each sample’s final families to archives/<id>/<id>_{hmms,seed_msas,full_msas}.tar.gz (see output). To update them later, give the three archives in the existing columns, with the new sequences in fasta and a new id (the families created for <id> are named <id>_<number>, so reusing it would clash):
id,fasta,existing_hmms,existing_seed_msas,existing_full_msass1_r2,new_sequences.faa.gz,results/archives/s1/s1_hmms.tar.gz,results/archives/s1/s1_seed_msas.tar.gz,results/archives/s1/s1_full_msas.tar.gzMigrating from v2 to v3
- Samplesheet: rename the columns
sample→id,existing_hmms_to_update→existing_hmmsandexisting_msas_to_update→existing_full_msas, and add anexisting_seed_msascolumn (may be left empty). MSA archives are now optional; a row with MSAs but no HMMs is rejected. - Existing HMMs may be a
.tar.gzarchive or one HMM library; families are identified by HMMNAME, and MSA files must be named after it (see Updating existing families). - Parameters:
--skip_msa_trimmingis now--skip_seed_msa_trimming;--clipkit_out_format,--save_update_families_pre_clipped_fastaand--save_update_families_clipped_fastaare removed (updated families use the samesave_*parameters as created ones);--skip_update_refinementis new. - Outputs:
clipkit/folders are nowtrimmed/(FASTA.aln). Updated families are published like created ones underupdate_families/{seed_msa,hmm,full_msa}/raw/<tool>/<id>/, with full MSAs reformatted to aligned FASTA (full_msa/raw/hhsuite_reformat/, next to the hmmalign Stockholm ones) instead ofupdate_families/full_msa/<tool>/andupdate_families/fasta/. Existing families left unchanged are listed inupdate_families/unchanged_families/, and every sample’s final families are archived underarchives/<id>/for later updates.
Parameter specifications
Here we provide guidance regarding some parameter choices.
clustering_tool[“cluster”, “linclust”]: The mmseqs algorithm used for clustering. Theclusteroption is slower but more sensitive, and is recommended where there are sufficient compute resources available and a more sensitive search is called for. It tends to produce fewer and larger clusters thanlinclust. Thelinclustoption is less sensitive, but extremely fast for clustering larger datasets.cluster_cov_mode[0, 1, 2]: The default bidirectional value for coverage mode (cluster_cov_mode= 0) automatically sets the MMseqs2 clustering mode to greedy cluster set. However, users can opt to override this parameter either indirectly, by changing the coverage mode, or directly, by setting the--cluster-modeargument in the modules configuration file.alignment_tool[“famsa”, “mafft”]: Multiple Sequence Alignment (MSA) options. Thefamsaoption is generally recommended as the best time-memory-accuracy combination. Themafftoption offers various alignment strategies, but in general is slower and less sensitive thanfamsa.trim_ends_only: Flag to either clip seed MSA gaps throughout the alignment, or only at the ends. Only used ifskip_seed_msa_trimmingis off. Full MSAs are never trimmed. The pipeline authors strongly recommend keepingtrim_ends_onlyon (default): gaps inside the sequences may still carry evolutionary significance, and only end trimming keeps row coordinates correct.
Trimmed MSA rows that lost residues are renamed <sequence>/<start>-<end> to the residues they still hold, recalculated from the residues removed at the alignment ends; rows that lost none keep their name.
With --trim_ends_only false, residues removed from interior columns are not reflected, so a row’s range spans more residues than the row contains and no longer maps back to its exact source residues.
Only turn it off if you need interior trimming and do not rely on row coordinates downstream.
Family generation algorithms
family_generation_algorithm [“standard”, “iterative”] selects how clusters become family models. Both produce the same kinds of outputs and feed the same redundancy removal, merging and reporting steps, so the choice is invisible downstream.
standard (default)
Each cluster is chunked into its own FASTA file and processed by a chain of tools orchestrated by Nextflow: FAMSA or mafft aligns the cluster into a seed MSA, ClipKIT optionally trims it, hmmbuild builds the family HMM, and hmmsearch optionally recruits further members from the sample’s sequence pool in a single pass, which hmmalign then aligns into the full MSA.
iterative
Whole chunks of clusters, clusters_per_chunk at a time, are handed to mgnifam, which builds a family from each cluster by looping: build an HMM, recruit members from the sequence pool, realign the expanded membership, and repeat up to three times or, until the family converges or the cluster is discarded. A single task therefore emits many families.
mgnifam performs each step in-process with its own libraries rather than by calling the pipeline’s tools:
Because of that, the parameters below are honoured only by the standard algorithm. They are ignored when the iterative path creates or merges families, which always behaves as stated:
| Parameter | Behaviour of the iterative algorithm |
|---|---|
alignment_tool |
Always FAMSA, through pyfamsa |
skip_seed_msa_trimming |
Trimming is always applied, through pytrimal |
trim_ends_only |
ClipKIT is not used; pytrimal trims by column gap occupancy (gap_threshold) |
skip_additional_sequence_recruiting |
Recruitment is always performed, and repeated until convergence |
hmmsearch_write_target, hmmsearch_write_domain, save_hmmsearch_results |
Searching is in-process, so no hmmsearch report files exist |
Updating existing families (samplesheet entries with existing HMMs) always runs the standard update path, whichever algorithm is selected, so the standard parameters above (e.g. alignment_tool, skip_seed_msa_trimming, trim_ends_only, gap_threshold, skip_additional_sequence_recruiting) apply to updated families.
The parameters both algorithms share are mapped onto their mgnifam equivalents:
| Parameter | mgnifam option |
|---|---|
cluster_size_threshold |
applied while chunking clusters |
min_seq_length |
--discard_min_rep_length |
max_seq_length |
--discard_max_rep_length |
cluster_seq_identity_for_redundancy |
--max_seq_identity |
gap_threshold |
--max_gap_occupancy |
hmmsearch_evalue_cutoff |
--recruit_evalue_cutoff |
hmmsearch_query_length_threshold |
--recruit_hit_length_percentage |
mgnifam’s remaining options (--discard_min_starting_membership, --max_seed_seqs, --batch_size, --prefetch_targets) keep their tool defaults and can be set through ext.args, as described in Custom Tool Arguments.
clusters_per_chunk (default 1000) trades parallelism against scheduling overhead: smaller chunks give more tasks, better load balancing and finer-grained -resume, at the cost of more per-task startup. Set save_iterative_family_metadata to publish mgnifam’s family roster, metadata, converged, successful and discarded records, its per-family HMM consensus match states (rf), representative sequence fasta, and its log.
Running the pipeline
The typical command for running the pipeline is as follows:
nextflow run nf-core/proteinfamilies --input ./samplesheet.csv --outdir ./results -profile dockerThis will launch the pipeline with the docker configuration profile. See below for more information about profiles.
Note that the pipeline will create the following files in your working directory:
work # Directory containing the nextflow working files<OUTDIR> # Finished results in specified location (defined with --outdir).nextflow_log # Log file from Nextflow# Other nextflow hidden files, eg. history of pipeline runs and old logs.If you wish to repeatedly use the same parameters for multiple runs, rather than specifying each flag in the command, you can specify these in a params file.
Pipeline settings can be provided in a yaml or json file via -params-file <file>.
Do not use -c <file> to specify parameters as this will result in errors. Custom config files specified with -c must only be used for tuning process resource specifications, other infrastructural tweaks (such as output directories), or module arguments (args).
The above pipeline run specified with a params file in yaml format:
nextflow run nf-core/proteinfamilies -profile docker -params-file params.yamlwith:
input: './samplesheet.csv'outdir: './results/'<...>You can also generate such YAML/JSON files via nf-core/launch.
Updating the pipeline
When you run the above command, Nextflow automatically pulls the pipeline code from GitHub and stores it as a cached version. When running the pipeline after this, it will always use the cached version if available - even if the pipeline has been updated since. To make sure that you’re running the latest version of the pipeline, make sure that you regularly update the cached version of the pipeline:
nextflow pull nf-core/proteinfamiliesReproducibility
It is a good idea to specify the pipeline version when running the pipeline on your data. This ensures that a specific version of the pipeline code and software are used when you run your pipeline. If you keep using the same tag, you’ll be running the same version of the pipeline, even if there have been changes to the code since.
First, go to the nf-core/proteinfamilies releases page and find the latest pipeline version - numeric only (eg. 1.3.1). Then specify this when running the pipeline with -r (one hyphen) - eg. -r 1.3.1. Of course, you can switch to another version by changing the number after the -r flag.
This version number will be logged in reports when you run the pipeline, so that you’ll know what you used when you look back in the future. For example, at the bottom of the MultiQC reports.
To further assist in reproducibility, you can use share and reuse parameter files to repeat pipeline runs with the same settings without having to write out a command with every single parameter.
If you wish to share such profile (such as upload as supplementary material for academic publications), make sure to NOT include cluster specific paths to files, nor institutional specific profiles.
Core Nextflow arguments
These options are part of Nextflow and use a single hyphen (pipeline parameters use a double-hyphen)
-profile
Use this parameter to choose a configuration profile. Profiles can give configuration presets for different compute environments.
Several generic profiles are bundled with the pipeline which instruct the pipeline to use software packaged using different methods (Docker, Singularity, Podman, Shifter, Charliecloud, Apptainer, Conda) - see below.
We highly recommend the use of Docker or Singularity containers for full pipeline reproducibility, however when this is not possible, Conda is also supported.
The pipeline also dynamically loads configurations from https://github.com/nf-core/configs when it runs, making multiple config profiles for various institutional clusters available at run time. For more information and to check if your system is supported, please see the nf-core/configs documentation.
Note that multiple profiles can be loaded, for example: -profile test,docker - the order of arguments is important!
They are loaded in sequence, so later profiles can overwrite earlier profiles.
If -profile is not specified, the pipeline will run locally and expect all software to be installed and available on the PATH. This is not recommended, since it can lead to different results on different machines dependent on the computer environment.
test- A profile with a complete configuration for automated testing
- Includes links to test data so needs no other parameters
docker- A generic configuration profile to be used with Docker
singularity- A generic configuration profile to be used with Singularity
podman- A generic configuration profile to be used with Podman
shifter- A generic configuration profile to be used with Shifter
charliecloud- A generic configuration profile to be used with Charliecloud
apptainer- A generic configuration profile to be used with Apptainer
wave- A generic configuration profile to enable Wave containers. Use together with one of the above (requires Nextflow
24.03.0-edgeor later).
- A generic configuration profile to enable Wave containers. Use together with one of the above (requires Nextflow
conda- A generic configuration profile to be used with Conda. Please only use Conda as a last resort i.e. when it’s not possible to run the pipeline with Docker, Singularity, Podman, Shifter, Charliecloud, or Apptainer.
-resume
Specify this when restarting a pipeline. Nextflow will use cached results from any pipeline steps where the inputs are the same, continuing from where it got to previously. For input to be considered the same, not only the names must be identical but the files’ contents as well. For more info about this parameter, see this blog post.
You can also supply a run name to resume a specific run: -resume [run-name]. Use the nextflow log command to show previous run names.
-c
Specify the path to a specific config file (this is a core Nextflow command). See the nf-core website documentation for more information.
Custom configuration
Resource requests
Whilst the default requirements set within the pipeline will hopefully work for most people and with most input data, you may find that you want to customise the compute resources that the pipeline requests. Each step in the pipeline has a default set of requirements for number of CPUs, memory and time. For most of the pipeline steps, if the job exits with any of the error codes specified here it will automatically be resubmitted with higher resources request (2 x original, then 3 x original). If it still fails after the third attempt then the pipeline execution is stopped.
To change the resource requests, please see the max resources and customise process resources section of the nf-core website.
Custom Containers
In some cases, you may wish to change the container or conda environment used by a pipeline steps for a particular tool. By default, nf-core pipelines use containers and software from the biocontainers or bioconda projects. However, in some cases the pipeline specified version maybe out of date.
To use a different container from the default container or conda environment specified in a pipeline, please see the updating tool versions section of the nf-core website.
Custom Tool Arguments
A pipeline might not always support every possible argument or option of a particular tool used in pipeline. Fortunately, nf-core pipelines provide some freedom to users to insert additional parameters that the pipeline does not include by default.
To learn how to provide additional arguments to a particular tool of the pipeline, please see the customising tool arguments section of the nf-core website.
nf-core/configs
In most cases, you will only need to create a custom config as a one-off but if you and others within your organisation are likely to be running nf-core pipelines regularly and need to use the same settings regularly it may be a good idea to request that your custom config file is uploaded to the nf-core/configs git repository. Before you do this please can you test that the config file works with your pipeline of choice using the -c parameter. You can then create a pull request to the nf-core/configs repository with the addition of your config file, associated documentation file (see examples in nf-core/configs/docs), and amending nfcore_custom.config to include your custom profile.
See the main Nextflow documentation for more information about creating your own configuration files.
If you have any questions or issues please send us a message on Slack on the #configs channel.
Running in the background
Nextflow handles job submissions and supervises the running jobs. The Nextflow process must run until the pipeline is finished.
The Nextflow -bg flag launches Nextflow in the background, detached from your terminal so that the workflow does not stop if you log out of your session. The logs are saved to a file.
Alternatively, you can use screen / tmux or similar tool to create a detached session which you can log back into at a later time.
Some HPC setups also allow you to run nextflow within a cluster job submitted your job scheduler (from where it submits more jobs).
Nextflow memory requirements
In some cases, the Nextflow Java virtual machines can start to request a large amount of memory.
We recommend adding the following line to your environment to limit this (typically in ~/.bashrc or ~./bash_profile):
NXF_OPTS='-Xms1g -Xmx4g'