I performed taxonomic annotation of eukaryotic bins with both mmseqs taxonomy and metauk easy-predict + metaeuk taxtocontig with the same reference db (UniRef90). If I look at the output files:
_taxResult_tax_per_contig.tsv for metaeuk and _lca.tsv for mmseqs
the fifth column (retained fragments) show really different results, with higher numbers for mmseqs. Can be the following a good explanation?
"For mmseqs “fragments” are protein fragments that passsed the prefilter step generated during the translated search, while in MetaEuk, they refer to the MetaEuk predictions that are subsequently taxonomically labelled, not the enormous set of raw six-frame fragments. MetaEuk first extracts putative coding fragments, finds compatible exon sets, performs redundancy reduction, and produces predicted proteins; taxtocontig then assigns taxonomy to those predictions and aggregates them to the contig. That is why MetaEuk's number is much smaller. It has already filtered and consolidated the initial fragments into gene/protein predictions before taxtocontig does its taxonomy aggregation."
Thank you!
I performed taxonomic annotation of eukaryotic bins with both mmseqs taxonomy and metauk easy-predict + metaeuk taxtocontig with the same reference db (UniRef90). If I look at the output files:
_taxResult_tax_per_contig.tsv for metaeuk and _lca.tsv for mmseqs
the fifth column (retained fragments) show really different results, with higher numbers for mmseqs. Can be the following a good explanation?
"For mmseqs “fragments” are protein fragments that passsed the prefilter step generated during the translated search, while in MetaEuk, they refer to the MetaEuk predictions that are subsequently taxonomically labelled, not the enormous set of raw six-frame fragments. MetaEuk first extracts putative coding fragments, finds compatible exon sets, performs redundancy reduction, and produces predicted proteins; taxtocontig then assigns taxonomy to those predictions and aggregates them to the contig. That is why MetaEuk's number is much smaller. It has already filtered and consolidated the initial fragments into gene/protein predictions before taxtocontig does its taxonomy aggregation."
Thank you!