All structures registered in 3decision are automatically classified against a curated set of protein families as part of the post-registration analysis pipeline. Family definitions are based on the protein family definitions used by MOE Project and PSILO and provided by Chemical Computing Group (CCG).
Each protein family is represented by a high-quality multiple sequence alignment containing carefully selected reference sequences and structures. From this alignment, a Hidden Markov Model (HMM) is generated to capture the characteristic sequence signature of the family.
Protein family classification is performed at the biomolecular sequence level rather than directly on individual structures. Biomolecular sequences in 3decision, including UniProt sequences and generic PRO_ sequences, are compared against the family HMMs using HMMER. A sequence is assigned to a family when its match score falls below the family-specific E-value cutoff.
When a structure is registered, its chains are mapped to biomolecular sequences. Protein family matches are then projected from the biomolecular sequence onto the corresponding residue range represented in the structure. As a result:
If a newly registered structure contains a sequence that is not yet represented in 3decision, a new PRO_ biomolecular sequence is created. This sequence is then searched against all supported family HMMs using HMMER to determine whether it belongs to one or more protein families and to assign any applicable family annotations.
The Family definitions also include biologically meaningful annotations that identify important regions within family members. Examples include:
These annotations are defined on the family's reference alignment and projected onto matching biomolecular sequences based on their alignment to the family HMM. Through residue-level mapping, the corresponding annotation ranges become available on structure chains wherever they overlap the resolved portion of the sequence.
The current integration supports a fixed set of 18 protein families, covering major drug discovery target classes and therapeutic modalities.
Once a structure has been classified, family assignments and associated annotations become available throughout 3decision.
The Protein Families section displays:
The Annotation Browser displays:
Family information can be visualised directly in the structure viewer:
Protein Family can be used as a search criterion when querying the database.
Because 3decision also supports the existing ChEMBL-based protein classification, family options are labelled according to their source:
Protein Family can be used as a filter to focus results on a specific target class or therapeutic modality.
Protein family classification and annotation for an existing structure are refreshed if the Sequence Mapping analysis is relaunched.Once the sequence mapping has been updated, 3decision automatically reassess any applicable protein family assignments and annotations.