Readme
The Big Fantastic Virus Database (BFVD) is a repository of 5,776,417 viral protein
structures predicted with ColabFold-AlphaFold2. BFVD v3 covers 98.9% of the viral
sequences in UniProt 2025_03, fully covers 72.6% of viral reference proteomes, and
spans 72.7% of ICTV-recognized virus species, with 75.3% of entries predicted at
high confidence (pLDDT ≥ 70).
Kim R, Levy Karin E, Steinegger M. BFVD - a large repository of predicted viral protein structures Nucleic Acids Research doi: doi.org/10.1093/nar/gkae1119 (2024)
Updates
- 2024-09-04: First distribution of BFVD.
- 2024-11-01 (2023_02_v1): 175,454 Base-MSA & 175,788 Base+Logan-MSA
Of the 351,242 BFVD entries initially predicted with a base multiple sequence alignment (base-MSA), 175,788 lacked detectable homologs.
For these, we augmented the alignments using Logan's large-scale assemblies, reinforcing nearly half of all BFVD entries.
- 2025-03-17 (2023_02_v2): 205,681 Base-MSA & 37,296 Base+Logan-MSA & 108,265 Base+Logan-MSA + 12-recycles
Of the 351,242 BFVD entries, 175,788 lacked identifiable homologs in their base MSAs. The remaining entries, which had sufficient homologs, were left unchanged.
For those insufficient homologs, we augmented the alignments using Logan-based data and performed 12-cycle predictions.
This process generated three versions for each affected entry: (1) Base-MSA, (2) Base+Logan-MSA, and (3) Base+Logan-MSA & 12-recycles.
Finally, we kept the best-scoring model (based on pLDDT) for each entry.
- 2026-09-15 (2025_03_v3): 5,776,417 structures, a 16.4-fold increase over v2.
Predicted for individual UniProt 2025_03 viral sequences (≤2,000 residues)
rather than for UniRef30 cluster representatives.
Every entry was predicted from a Base+Logan50-MSA.
Availability
- BFVD is browsable with UniProt accessions through website
- BFVD is searchable through Foldseek webserver
- Scripts for BFVD analyses are available at Zenodo
- PDB files of BFVD are also available at Zenodo
Data description
Files below are those published for the current release (2025_03_v3).
Earlier releases are under archived/ and contain a different set of files.
1-bfvd_v3_pdbs.tar.zst: 5,776,417 predicted structures of BFVD, as PDB files.
2-bfvd.version: version file.
3-bfvd_foldseekdb.tar.gz: Foldseek database of 5,776,417 predicted structures of BFVD.
4-bfvd_v3_metadata.tsv.tar.gz: General information of each entry.
- accession: UniProt accession of the sequence
- protein_name: Protein name from UniProt
- length: Number of modelled residues
- plddt: Average pLDDT score of the predicted protein structure
- ptm: pTM score of the predicted protein structure (NA for ProteinTTT models)
- model: Prediction method used for the released structure (ColabFold-AF2 or ProteinTTT)
- basemsa: Number of sequences in the base MSA
- loganmsa: Number of sequences in the MSA after adding Logan50 homologs
- taxid: Taxonomy identifier of the protein
- taxname: Scientific name of the taxonomy identifier
- ictv_id: ICTV identifier mapped from the taxonomy identifier
- uniprot_host: Host organism retrieved by UniProt
- ictv_host_category: ICTV host category
- proteome_id: UniProt proteome(s) the entry belongs to, separated by ';'
Absent values are given as NA.
5-bfvd_v3_taxID_rank_scientificname_lineage.tsv.tar.gz: BFVD entry and their taxonomic information.
- model: File name of the BFVD.
- taxId: Taxonomy identifier of the protein.The protein ID, the portion before the first underscore in model, was used to retrieve the taxonomy ID.
- rank: rank of the taxonomy.
- scientific name: scientific name of the corresponding taxonomy identifier.
- lineage: lineage of the taxonomy.