Readme
The Big Fantastic Virus Database (BFVD) is a repository of 351,242 protein structures predicted by applying ColabFold to the viral sequence representatives of the UniRef30 clusters.
BFVD holds a unique repertoire of protein structures, spanning major viral clades.
Kim R, Levy Karin E, Steinegger M. BFVD - a large repository of predicted viral protein structures Nucleic Acids Research doi: doi.org/10.1093/nar/gkae1119 (2024)
Updates
- 2024-09-04: First distribution of BFVD.
- 2024-11-01 (2023_02_v1): 175,454 Base-MSA & 175,788 Base+Logan-MSA
Of the 351,242 BFVD entries initially predicted with a base multiple sequence alignment (base-MSA), 175,788 lacked detectable homologs.
For these, we augmented the alignments using Logan’s large-scale assemblies, reinforcing nearly half of all BFVD entries.
- 2025-03-17 (2023_02_v2): 205,681 Base-MSA & 37,296 Base+Logan-MSA & 108,265 Base+Logan-MSA + 12-recycles
Of the 351,242 BFVD entries, 175,788 lacked identifiable homologs in their base MSAs. The remaining entries, which had sufficient homologs, were left unchanged.
For those insufficient homologs, we augmented the alignments using Logan-based data and performed 12-cycle predictions.
This process generated three versions for each affected entry: (1) Base-MSA, (2) Base+Logan-MSA, and (3) Base+Logan-MSA & 12-recycles.
Finally, we kept the best-scoring model (based on pLDDT) for each entry.
Availability
- BFVD is browsable with UniProt accessions through website
- BFVD is searchable through Foldseek webserver
- Scripts for BFVD analyses are available at Zenodo
- PDB files of BFVD are also available at Zenodo
Data description
1-bfvd.tar.gz: 351,242 predicted structures of BFVD.
2-bfvd.version: version file.
3-bfvd_foldcompdb.tar.gz: Compressed version of Foldseek database using Foldcomp.
Only 347,481 structures, none of which are discontinuous, were included.
4-bfvd_foldseekdb.tar.gz: Foldseek databse of 351,242 predicted structures of BFVD.
5-bfvd_metadata.tsv: General information of each model.
- UniRef100: UniRef100 identifier of the sequence
- model: File name of the predicted protein structure
- avg_pLDDT: Average pLDDT score of the predicted protein structure
- pTM: pTM score of the predicted protein structure
- splitted: Whether the protein sequence of UniRef100 entry was splitted into multiple models
We splitted the protein sequences if their length are above 1500. (0 = not splitted, 1 = splitted)
- version: Specifies the MSA/refinement pipeline used to produce the final BFVD structure. (BASE, BASE+LOGAN, or BASE+LOGAN+12CY)
6-msa.tar: MSAs for each BFVD entries
7-bfvd_taxid.tsv: BFVD entry and their taxonomic identifier.
- model: File name of the BFVD.
- taxId: Taxonomy identifier of the protein.The protein ID, the portion before the first underscore in model, was used to retrieve the taxonomy ID.
8-bfvd_taxID_rank_scientificname_lineage.tsv: BFVD entry and their taxonomic information.
- model: File name of the BFVD.
- taxId: Taxonomy identifier of the protein.The protein ID, the portion before the first underscore in model, was used to retrieve the taxonomy ID.
- rank: rank of the taxonomy.
- scientific name: scientific name of the corresponding taxonomy identifier.
- lineage: lineage of the taxonomy.
9-uniref30_2302_virus-rep_mem.tsv: UniRef30 virus clusters.
- repId: Cluster representatives used for structure prediction
- memId: Member corresponding to the representative