nf-core/proteinfamilies
Generation and updating of protein families
metagenomicsprotein-families
Version history
Changed
- #192
- Updated nf-core modules and subworkflows to latest.
mgnifam4.0.0 now also writes a<chunk>_mgnifam_stats.jsonrun summary, saved with the existing--save_iterative_family_metadata, and renames and reorders the columns of<chunk>_metadata.csv. Since mgnifam 3.0.0 a chunk in which a family crashed internally exits3, which stops the pipeline. (by @vagkaratzas) - Added missing Font Awesome icons to the parameters and parameter groups of
nextflow_schema.json(#185). (by @vagkaratzas)
- Updated nf-core modules and subworkflows to latest.
- #189 - nf-core tools template update to 4.1.0. (by @vagkaratzas)
Fixed
- #193
- Family numbering and member order no longer follow the unstable row order of the MMseqs2 clustering TSV: clusters are sorted by representative ID and members by ID, so repeated runs give identically named families and representatives. (by @vagkaratzas)
- Fixed the Nextflow head job running out of heap (
java.lang.OutOfMemoryError: Java heap space) on large samples: eachMERGE_FAMILIES:MERGE_SEEDStask now stages only the seed MSAs of its own pool instead of every seed MSA of the sample (#191). (by @vagkaratzas) - Fixed sample and family names containing dots (e.g. sample
run.v2, or Pfam-stylePF00069.29families in the update tarballs) being cut at their first dot when family IDs are read from file names. This crashed family redundancy removal and the iterative algorithm, kept redundant families, paired each updated family’s new hits with other families’ MSAs, and wrote wrong family IDs to the members TSVs and the MultiQC family table. Family IDs now drop only.gzand the last extension. (by @vagkaratzas)
Dependencies
| Tool | Previous version | New version |
|---|---|---|
| mgnifam | 2.0.0 | 4.0.0 |
Added
- #182 - Added an alternative, iterative family generation algorithm, selectable with
--family_generation_algorithm iterative. It hands whole chunks of clusters (--clusters_per_chunk) tomgnifam, which repeats HMM building, sequence recruitment and realignment per cluster until each family converges or is discarded. Applies to both newly created and merged families. (by @vagkaratzas)
Fixed
- #182 - Fixed
MERGE_FAMILIES:MERGE_SEEDSusing the first sample’s seed MSAs for every sample: pooled families are now paired with their own sample’s seeds. (by @vagkaratzas) - #177
- Fixed
OSError: [Errno 36] File name too longinMERGE_FAMILIES:MERGE_SEEDSwhen a pool combines many families:merged_idnow collapses to a stable hash once the readable form would exceed the filesystem name-length limit. (by @vagkaratzas, bug reported by @Yixuan39) - Fixed
FAMSA_ALIGNaborting with exit status 134 on merged families:MERGE_SEEDSnow drops all-gap records, whichCLIPKITcan leave behind when a member’s residues all fall in trimmed columns, and which FAMSA cannot realign. (by @vagkaratzas, bug reported by @Yixuan39)
- Fixed
Changed
- #186 - Raw family HMMs are now published under a subfolder named after the tool that built them,
hmm/raw/hmmer_hmmbuild/orhmm/raw/mgnifam/, matching how the seed and full MSA outputs are already organised. (by @vagkaratzas, suggested by @nickp60) - #186 - Restricted samplesheet sample names to letters, digits, dots, underscores and dashes. Family names are parsed back out of file names with
split("<sample>_"), which takes a regex, so a sample name carrying a regex metacharacter could mis-split; the sample name also builds output paths. (by @vagkaratzas, raised by @erikrikarddaniel) - #182 - Updated
CHUNK_CLUSTERSandMERGE_SEEDSto optionally outputTSV, required by the iterative family generation subworkflow. (by @vagkaratzas) - #179 - nf-core tools template update to 4.0.3. (by @vagkaratzas)
Dependencies
| Tool | Previous version | New version |
|---|---|---|
| seqkit | 2.9.0 | 2.13.0 |
| mgnifam | - | 2.0.0 |
| gunzip | - | 1.13 |
Changed
- #170 - Updated local modules to versions topic output. (by @vagkaratzas)
- #171
- Updated nf-core modules and subworkflows to latest, removing all remaining ch_versions. (by @vagkaratzas)
- Updated pipeline-level nf-schema to 2.7.2. (by @vagkaratzas)
Dependencies
| Tool | Previous version | New version |
|---|---|---|
| multiqc | 1.34 | 1.35 |
Added
- #159 - Added functionality to generate a HMM library file (compressed) with its respective final protein families HMMs, per each input sample row. Family library files can be found at
hmm/library(by @juanfmx2) (Hackathon 2026) - #154 - Added optional save parameters for
update_familiesmode:--save_update_families_pre_clipped_fasta, and--save_update_families_clipped_fasta(with gaps removed) to save FASTA files from updated family MSAs at various stages of the subworkflow (update_families/fasta/pre_clipped/,update_families/fasta/pre_clipped_non_redundant_sequences/, andupdate_families/fasta/post_clipped/). (by @eparisis) (Hackathon 2026)
Changed
- #162 - nf-core tools template update to 4.0.2. (by @vagkaratzas)
- #154 - Moved
update_familiesnon-redundant sequence FASTA output frommmseqs/update_families/non_redundant_sequences/toupdate_families/fasta/pre_clipped_non_redundant_sequences/${meta.id}/(now controlled by--save_update_families_pre_clipped_fasta). (by @eparisis) (Hackathon 2026) - #153 - Typo fixes and unused boilerplate removal. (by @vagkaratzas)
- #152
- Removed an always-truthy if condition. (by @vagkaratzas)
- Moved a wrongly placed
.first()to now properly report all sample representative sequences on the MultiQC report. (by @vagkaratzas)
- #150 - Pipeline logos and workflow maps updated. (by @vagkaratzas)
- #149 - Updated the nf-core/proteinfamilies article citation to the GigaScience publication. (by @vagkaratzas)
- #147 - Modules, subworkflows and pipeline code updates to fix linting warnings and errors. (by @vagkaratzas)
Dependencies
| Tool | Previous version | New version |
|---|---|---|
| clipkit | 2.4.1 | 2.11.4 |
| multiqc | 1.33 | 1.34 |
Added
- #143
- Added the
cmaplemodule for optional phylogenetic tree inference for final family full MSAs. (by @vagkaratzas) - Added extra nf-tests for the
REMOVE_REDUNDANCYsubworkflow. (by @vagkaratzas)
- Added the
- #140 - Using the new workflow output syntax to publish the downstream
nf-core/proteinannotatorsamplesheet. (by @vagkaratzas)
Changed
- #143 - Updated metro-maps, citations, README.md and output.md to include
cmaplephylogenetic trees. (by @vagkaratzas) - #141 - Aligning the default output directory of the new output syntax with the publishing directory of the pipeline. (by @vagkaratzas)
- #140
- Updated metro-map to also depict the optional creation of downstream samplesheets for
nf-core/proteinfoldandnf-core/proteinannotator. Also swappedseqkit/rmdupandseqkit/replacemodules to their proper execution sequence. (by @vagkaratzas) - Updated the
test_fullprofile time and memory requirements to avoid AWS failures on release. (by @vagkaratzas)
- Updated metro-map to also depict the optional creation of downstream samplesheets for
Fixed
- #143
- Fixed a bug in
REMOVE_REDUNDANCYsubworkflow, where the combination of these skip flags--skip_sequence_redundancy_removal true,--skip_additional_sequence_recruiting trueand--skip_additional_sequence_recruiting false, would executeHHSUITE_REFORMAT_FILTEREDto reformat Stockholm alignments while they were already in fasta (or clipkit) format. (by @vagkaratzas) - Fixed a bug in
REMOVE_REDUNDANCYsubworkflow, where unmerged similar families would be removed along with redundant ones, when--skip_family_mergingwastrueand--skip_family_redundancy_removalwasfalse. (by @vagkaratzas)
- Fixed a bug in
Dependencies
| Tool | Previous version | New version |
|---|---|---|
| multiqc | 1.32 | 1.33 |
| cmaple | - | 1.1.0 |
Added
- #133 - Using the new workflow output syntax to publish the downstream
nf-core/proteinfoldsamplesheet. (by @vagkaratzas) - #132 - Added optimized memory and time resources for
testandtest_fullprofiles. (by @vagkaratzas)
Changed
- #136 - Based on protein family reproducibility benchmarks,
clusteris now the default MMseqs2 mode, due to its increased sensitivity compared tolinclust. (by @vagkaratzas) - #135 - nf-core tools template update to 3.5.1. (by @vagkaratzas)
Dependencies
| Tool | Previous version | New version |
|---|---|---|
| mmseqs | 17.b804f | 18.8cc5c |
Special Thanks
To @jfy133 and @JoseEspinosa for their assistance with all queries regarding chaining nf-core/proteinfamilies to nf-core/proteinfold.
Added
- #124
- Added new subworkflow
MERGE_FAMILIESthat can optionally merge similar (but not redundant) generated protein families. (by @vagkaratzas) - Added new functionality to the local module
IDENTIFY_REDUNDANT_FAMSwhich now also detects and outputs the identifiers of similar families that can optionally be merged downstream. These identifiers are written to “/remove_redundancy/<samplename>/similar_fam_ids.txt”, and the corresponding family pairwise similarity scores to “/remove_redundancy/<samplename>/similarities.csv”. (by @vagkaratzas) - Added new local module
POOL_SIMILAR_COMPONENTSthat generates family clusters, from a family-similarity edgelist. (by @vagkaratzas) - Added new local module
MERGE_SEEDSthat merges seed alignments of similar families, before restarting the family generation subworkflow. (by @vagkaratzas)
- Added new subworkflow
- #118
- Added preprint citation to the repo. (by @vagkaratzas)
- Added separate metro map files for dark and light browser modes. (by @vagkaratzas)
- Added new local module
EXTRACT_FAMILY_MEMBERSwhich outputs a two-column TSV file containing the final family identifiers and their corresponding member sequence identifiers. The file is saved at “/family_reps/<samplename>/<samplename>.tsv”. (by @vagkaratzas)
- #117
- Added
SEQKIT_SEQfor optional sequence preprocessing in the quality check subworkflow. (by @vagkaratzas) - Added
SEQKIT_REPLACEfor optional sequence name parsing in the quality check subworkflow. (by @vagkaratzas) - Added
SEQKIT_RMDUPfor optional removal of duplicate names and sequences in the quality check subworkflow. (by @vagkaratzas)
- Added
Changed
- #128 - nf-core tools template update to 3.4.1.
- #124
- Conditional workflow flags switched to their
skipopposites;--trim_msato--skip_msa_trimming,--recruit_sequences_with_modelsto--skip_additional_sequence_recruiting,--remove_family_redundancyto--skip_family_redundancy_removal,--remove_sequence_redundancyto--skip_sequence_redundancy_removal. (by @vagkaratzas)
- Conditional workflow flags switched to their
- #118
- Swapped the local
CHECK_QUALITYsubworkflow with the new nf-core oneFAA_SEQFU_SEQKIT. (by @vagkaratzas) - Based on protein family reproducibility benchmarks (i.e., computationally reproducing manually curated protein family resources), the
cluster_seq_identityandcluster_coverageparameter default values have been updated to0.3and0.5(down from0.5and0.9) respectively. (by @vagkaratzas)
- Swapped the local
- #117 - Swapped the local
SEQKIT_STATSand the localSEQKIT_STATS_TO_MQCmodules with theSEQFU_STATSone, which runs a bit faster and produces a MultiQC-ready output without the need for manual parsing. (by @vagkaratzas)
Dependencies
| Tool | Previous version | New version |
|---|---|---|
| seqfu | - | 1.20.3 |
| multiqc | 1.30 | 1.31 |
Deprecated
- #124 - Deprecated
--trim_msa,--recruit_sequences_with_models,--remove_family_redundancyand--remove_sequence_redundancy. (by @vagkaratzas)
Special Thanks
To @jfy133, @erikrikarddaniel and @chrisAta for this version’s PR code reviews.
Fixed
- #112 - Fixed a bug in
EXTRACT_FAMILY_REPS, where all sequences were pasted into the family representative one, and updated the relevant local nf-test. (by @vagkaratzas)
Changed
- #106 - Swapped the local
EXECUTE_CLUSTERINGsubworkflow with the new nf-coreMMSEQS_FASTA_CLUSTERone. (by @vagkaratzas)
Dependencies
| Tool | Previous version | New version |
|---|---|---|
| multiqc | 1.29 | 1.30 |
Changed
- #104 - Pulling
paramsfrom local subworkflows into main workflow. - #103 - Parallelized execution for the
EXTRACT_FAMILY_REPSlocal module and changed its input fromfull_msatofasta. - #100 -
CAT_CATmodule replaced withFIND_CONCATENATEto avoid large scaleArgument list too longerrors. - #98 - nf-core tools template update to 3.3.2.
Added
- #105 -
CHECK_QUALITYsubworkflow added at the start of the pipeline. It utilizes theseqkit/statsnf-core module to generate aMultiQC-ready report with statistics for the input amino acid sequences. The metro-map has been updated to reflect this change.
Added
- #93
- Added nf-test and
meta.ymlfile for local subworkflowGENERATE_FAMILIES. - Added nf-test and
meta.ymlfile for local subworkflowREMOVE_REDUNDANCY. - Added nf-test and
meta.ymlfile for local subworkflowUPDATE_FAMILIES.
- Added nf-test and
- #88
- Added nf-test and
meta.ymlfile for local moduleBRANCH_HITS_FASTA. - Added nf-test and
meta.ymlfile for local moduleFILTER_NON_REDUNDANT_FAMS. - Added nf-test and
meta.ymlfile for local moduleIDENTIFY_REDUNDANT_FAMS. - Added nf-test and
meta.ymlfile for local moduleEXTRACT_FAMILY_REPS. - Added the default pipeline end-to-end nf-test.
- Added nf-test and
Changed
- #81 - nf-core tools template update to 3.3.1.
Fixed
- #80 - Fixed a bug where, due to a missing check for equal family sizes, non-redundant families were erroneously marked as redundant through transitive relationships and were removed
Changed
- #77 - Default branch changed from
mastertomain. - #73 - Changed the fasta parsing library of the
CHUNK_CLUSTERSlocal module, frompyfastxback to the latest version ofbiopython, and parallelized its writing mechanism, achieving decreased execution time.
Dependencies
| Tool | Previous version | New version |
|---|---|---|
| biopython | 1.84 | 1.85 |
| pyfastx | 2.2.0 |
Removed
- #73 - Deprecated
pyfastxmodule version ofCHUNK_CLUSTERS, since it was struggling performance-wise with larger datasets.
Added
- #69 - Added the
hhsuite/reformatnf-core module to reformat.stoalignments to.faswhen in-family sequence redundancy is not removed. Also added the option to save intermediate and final family fasta files throughout the workflow with varioussaveparameters. - #58 - Added nf-test and
meta.ymlfile for local moduleREMOVE_REDUNDANCY_SEQS(Hackathon 2025) - #56 - Added nf-test and
meta.ymlfile for local moduleFILTER_RECRUITED(Hackathon 2025) - #55 - Added nf-test and
meta.ymlfile for local moduleCHUNK_CLUSTERS(Hackathon 2025) - #54 - Added nf-test for local subworkflow
ALIGN_SEQUENCES(Hackathon 2025) - #53 - Added nf-test for local subworkflow
EXECUTE_CLUSTERING(Hackathon 2025) - #51 - Added nf-test and
meta.ymlfile for local moduleCALCULATE_CLUSTER_DISTRIBUTION(Hackathon 2025) - #34 - Added the
EXTRACT_UNIQUE_CLUSTER_REPSmodule, that calculates initialMMseqsclustering metadata, for each sample, to print withMultiQC(Id,Cluster Size,Number of Clusters)
Fixed
- #69 - Fixed a bug where redundant family alignments were not published properly, if intra-family redundancy removal mechanism was switched off #68
- #65 - Fixed a bug in
CHUNK_CLUSTERS, where pipeline would crash if the module filtered out all clusters, due to a high membership threshold #64 - #35 - Fixed a bug in
remove_redundant_fams.py, where comparison was between strings instead of integers to keep larger family - #33 - Fixed an always-true condition at the
filter_non_redundant_hmms.pyscript, by adding missing parentheses - #29 - Fixed
hmmalignempty input crash error, by preventing theFILTER_RECRUITEDmodule from creating an empty output .fasta.gz file, when there are no remaining sequences after filtering thehmmsearchresults #28
Changed
- #69 - Changed the publish directory architecture for HMMs, seed MSAs, full MSAs and family FASTA files, to make it more intuitive.
REMOVE_REDUNDANT_FAMSlocal module converted toIDENTIFY_REDUNDANT_FAMSto extract redundant family ids which will then be used downstream.FILTER_NON_REDUNDANT_HMMSlocal module converted toFILTER_NON_REDUNDANT_FAMSand reused four times (HMM, seed MSA, full MSA, FASTA). Changed the output format of theEXTRACT_FAMILY_REPSandREMOVE_REDUNDANT_SEQSlocal modules from.fato.faa. Metro map updated with newhhsuite/reformatmodule. - #57 - slight improvements of
nextflow_schema.json(Hackathon 2025) - #57 - slight improtmenets of
assets/schema_input.json(Hackathon 2025) - #34 - Swapped the
SeqIOpython library withpyfastxfor theCHUNK_CLUSTERSmodule, quartering its duration - #32 - Updated
ClipKIT2.4.0 -> 2.4.1, that now also allows ends-only trimming, to completely replace the customCLIP_ENDSmodule. Users can now also define its output format by setting the--clipkit_out_formatparameter (default:clipkit)
Dependencies
| Tool | Previous version | New version |
|---|---|---|
| ClipKIT | 2.4.0 | 2.4.1 |
| pyfastx | 2.2.0 | |
| hhsuite | 3.3.0 | |
| multiqc | 1.27 | 1.28 |
Deprecated
- #32 - Deprecated
CLIP_ENDSmodule and--clipping_toolparameter. The only option now isClipKIT, covering both previous modes, via setting--trim_ends_only
Initial release of nf-core/proteinfamilies, created with the nf-core template.
Added
- Amino acid sequence clustering (mmseqs)
- Multiple sequence alignment (famsa, mafft, clipkit)
- Hidden Markov Model generation (hmmer)
- Between families redundancy removal (hmmer)
- In-family sequence redundancy removal (mmseqs)
- Family updating (hmmer, seqkit, mmseqs, famsa, mafft, clipkit)
- Family statistics presentation (multiqc)
By @vagkaratzas and @mberacochea.