Bioinformatics: Sequence Comparison and Scoring

Master bioinformatics sequence comparison and scoring. Learn about dot plots, scoring matrices (PAM, BLOSUM), gap penalties, and log-odds scores. Optimize your understanding!

Podcast

Bioinformatika: DNA detektivové0:00 / 18:58
0:001:00 remaining

Bioinformatics is a rapidly growing field that bridges biology and computer science, essential for analyzing vast biological datasets. Understanding sequence comparison and scoring is fundamental to this discipline, allowing researchers to identify similarities between DNA, RNA, or protein sequences. This process is crucial for predicting gene and protein functions, reconstructing evolutionary relationships, and identifying new biological sequences.

Unpacking Bioinformatics: Sequence Comparison and Scoring Essentials

Bioinformatics helps in analyzing large datasets generated by high-throughput studies, moving beyond single experiments. Key tasks include DNA sequencing, sequence assembly, genome annotation, gene and protein function prediction, and computational evolutionary biology. At its core, bioinformatics allows us to compare biological sequences like proteins and nucleic acids to understand functional and structural patterns, assess similarities, and perform evolutionary analyses.

Initial Sequence Comparison: The Dot Plot Method

One of the most basic methods for comparing sequences is the dot plot. This visual tool helps quickly identify regions of local alignment, repeats, insertions, and deletions. While simple and intuitive, dot plots do not provide robust statistical measures of similarity.

To create a dot plot, one sequence is written horizontally and the other vertically on a matrix. Dots are placed where characters match. Areas of similarity appear as diagonal lines.

  • Word-based method: Dot plots are a word-based method, where the 'word' size can be adjusted. For example, a word size of 1 means every single character match is dotted.
  • Thresholding: A threshold can be applied, such as requiring 4 out of 5 characters to match before a dot is placed.
  • Identifying features: Diagonal lines indicate matches, extra diagonals reveal repeats (like in human mucin), and solid boxes can signify low complexity regions or biased composition.

The Need for Numerical Scoring Matrices in Sequence Alignment

Since dot plots lack statistical measures, numerical methods are essential to quantify sequence similarity. Scoring matrices are empirical weighting schemes developed for both nucleotide and amino acid character comparisons. They account for factors beyond simple matches and mismatches, providing a statistically sound measure of relatedness.

Key Biological Factors in Protein Scoring Matrices:

  1. Conservation: Matrices account for absolute conservation (identical residues) and conservative substitutions (residues with similar physicochemical properties, like a Tyr for a Trp).
  2. Frequency: Rare residues are weighted differently from common ones. For example, a tryptophan (Trp) matching Trp scores high (e.g., 11 in BLOSUM62) because it's a rare amino acid.
  3. Evolution: Different matrices are designed for varying evolutionary rates or distances between sequences.

Understanding Log-Odds Scores

Scoring matrices utilize log-odds scores ($S_{i,j}$) to represent the ratio of observed versus random frequency of substitution. The formula is: $S_{i,j} = ext{log} rac{(q_{i,j})}{(p_i p_j)}$.

  • $p_i$ and $p_j$ are the probabilities of residues $i$ and $j$ occurring among all proteins.
  • $q_{i,j}$ represents how often amino acids $i$ and $j$ are empirically observed to align with each other.

A positive score indicates that two residues replace each other more often than by chance. A negative score means they replace each other less often than by chance.

  • Exact match: Always positive. For example, Trp matching Trp might score 11.
  • Conservative substitution: Observed more often than by chance, but weaker than an exact match. For example, a Tyr for a Trp might score 2.
  • Non-conservative substitution: Observed less often than by chance, resulting in a negative score. For instance, a Val for a Trp might score -3.

Nucleotide Scoring Matrices

The landscape for nucleotide scoring matrices is simpler than for proteins. They often focus on counting matches and mismatches, frequently assuming equal base frequency. Sometimes, purine-purine (A, G) versus purine-pyrimidine (A, T, C, G) substitutions may be scored differently. It's important to note that protein searches are generally more powerful for detecting distant evolutionary relationships than nucleotide searches.

ATGCSWRYKMBVHDN
A5-4-4-4-411-4-41-4-1-1-1-2
T-45-4-4-41-411-4-1-4-1-1-2
G-4-45-41-41-41-4-1-1-4-1-2
C-4-4-451-4-41-41-1-1-1-4-2
S-4-411-1-4-2-2-2-2-1-1-3-3-1
W11-4-4-4-1-2-2-2-2-3-3-1-1-1
R1-41-4-2-2-1-4-2-2-3-1-3-1-1
Y-41-41-2-2-4-1-2-2-1-3-1-3-1
K-411-4-2-2-2-2-1-4-1-3-3-1-1
M1-4-41-2-2-2-2-4-1-3-1-1-3-1
B-4-1-1-1-1-3-3-1-1-3-1-2-2-2-1
V-1-4-1-1-1-3-1-3-3-1-2-1-2-2-1
H-1-1-4-1-3-1-3-1-3-1-2-2-1-2-1
D-1-1-1-4-3-1-1-3-1-3-2-2-2-1-1
N-2-2-2-2-1-1-1-1-1-1-1-1-1-1-1

Gap Penalties: Essential for Accurate Alignments

When aligning sequences, gaps represent insertions or deletions (indels) that have occurred during evolution. To prevent biologically implausible scenarios, gaps must be limited. Gaps are considered evolutionarily expensive events and are therefore penalized in scoring algorithms.

There are two main types of gap penalties: linear and affine.

  • Linear gap penalties apply a constant penalty for each gap, regardless of its length.
  • Affine gap penalties are more sophisticated. They use the formula G + Ln, where G is the gap opening penalty, L is the gap extension penalty, and n is the length of the gap. Critically, G is always greater than L. This models a parsimonious view of evolution, where a single large gap (representing one evolutionary event) is more likely than multiple small gaps (requiring multiple events). This approach aims for the simplest explanation with the fewest evolutionary steps.

High gap opening and low gap extension penalties encourage fewer, larger gaps in alignments.

PAM Matrices: Point Accepted Mutation

PAM (Point Accepted Mutation) matrices, developed by Dayhoff et al. in 1978, were among the first widely used scoring matrices. They are based on observed substitution patterns in closely related proteins (sharing >85% similarity) aligned globally.

Key characteristics and assumptions of PAM matrices:

  • Empirical Basis: Derived from 1572 observed changes across 71 closely related protein groups.
  • PAM Unit: One PAM unit corresponds to approximately 1 amino acid change per 100 residues (1% divergence).
  • Extrapolation: Matrices for longer evolutionary distances (e.g., PAM160, PAM250) are extrapolated from the PAM1 matrix, making them predictive rather than purely empirical.
  • Assumptions: A key assumption is that amino acid replacement is independent of previous mutations. Also, all sites are assumed to be equally mutable, without accounting for conserved motifs.
  • Limitations: Based on a relatively small number of sequences available in 1978, with a bias towards small, globular proteins. They implicitly assume that evolutionary forces are constant over long periods.

BLOSUM Matrices: Blocks of Amino Acid Substitution

BLOSUM (Blocks Of Amino Acid Substitution) matrices, developed by Henikoff & Henikoff in 1992, offer an alternative approach. They are based on conserved, ungapped blocks of sequences found in the BLOCKS database, reflecting local alignments.

Key characteristics of BLOSUM matrices:

  • Empirical Basis: Directly calculated from substitution patterns in conserved local alignments across varying evolutionary distances, not extrapolated.
  • Dataset: Derived from over 2,000 blocks representing more than 500 groups of related proteins.
  • Clustering: Sequences with identity above a certain threshold are clustered, and their contribution is weighted to one. For example, BLOSUM62 represents sequences with no more than 62% identity.
  • Direct Calculation: BLOSUM matrices are more directly empirical for longer evolutionary distances compared to extrapolated PAM matrices.

Flashcards

1 / 23

What biological factors influence protein scoring matrices?

Conservation (absolute and conservative substitutions), residue frequency (rare vs common residues), and evolution (different matrices for different e

Tap to flip · Swipe to navigate

Selecting the Right Scoring Matrix

Choosing the appropriate scoring matrix is crucial for successful sequence comparison. The choice depends on the expected similarity between the sequences you are aligning. The "Similarity (%)" column indicates the range of sequence similarities each matrix is best designed to detect.

Guide for Selecting an Appropriate Scoring Matrix:

MatrixBest UseSimilarity (%)
PAM40Short alignments that are highly similar70–90
PAM160Detecting members of a protein family50–60
PAM250Longer alignments of more divergent sequences~30
BLOSUM90Short alignments that are highly similar70–90
BLOSUM80Detecting members of a protein family50–60
BLOSUM62Most effective in finding all potential similarities30–40
BLOSUM30Longer alignments of more divergent sequences<30

For instance, BLOSUM62 is often considered a good general-purpose matrix because it effectively finds similarities in the 30-40% range, a common level of divergence for functional homologs.

BLOSUM62 Example:

ARNDCQEGHILKMFPSTWYVBZX*
A4-1-2-20-1-10-2-1-1-1-1-2-110-3-20-2-10-4
R-150-2-310-20-3-22-1-3-2-1-1-3-2-3-10-1-4
N-2061-30001-3-30-2-3-210-4-2-330-1-4
D-2-216-302-1-1-3-4-1-3-3-10-1-4-3-341-1-4
C0-3-3-39-3-4-3-3-1-1-3-1-2-3-1-1-2-2-1-3-3-2-4
Q-1100-352-20-3-210-3-10-1-2-1-203-1-4
E-1002-425-20-3-31-2-3-10-1-3-2-214-1-4
G0-20-1-3-2-26-2-4-4-2-3-3-20-2-2-3-3-1-2-1-4
H-201-1-300-28-3-3-1-2-1-2-1-2-22-300-1-4
I-1-3-3-3-1-3-3-4-342-310-3-2-1-3-13-3-3-1-4
L-1-2-3-4-1-2-3-4-324-220-3-2-1-2-11-4-3-1-4
K-120-1-311-2-1-3-25-1-3-10-1-3-2-201-1-4
M-1-1-2-3-10-2-3-212-150-2-1-1-1-11-3-1-1-4
F-2-3-3-3-2-3-3-3-100-306-4-2-213-1-3-3-1-4
P-1-2-2-1-3-1-1-2-2-3-3-1-2-47-1-1-4-3-2-2-1-2-4
S1-110-1000-1-2-20-1-2-141-3-2-2000-4
T0-10-1-1-1-1-2-2-1-1-1-1-2-115-2-20-1-10-4
W-3-3-4-4-2-2-3-2-2-3-2-3-11-4-3-2112-3-4-3-2-4
Y-2-2-2-3-2-1-2-32-1-1-2-13-3-2-227-1-3-2-1-4
V0-3-3-3-1-2-2-3-331-21-1-2-20-3-14-3-2-1-4
B-2-134-301-10-3-40-3-3-20-1-4-3-341-1-4
Z-1001-334-20-3-31-1-3-10-1-3-2-214-1-4
X0-1-1-1-2-1-1-1-1-1-1-1-1-1-200-2-1-1-1-1-1-4
*-4-4-4-4-4-4-4-4-4-4-4-4-4-4-4-4-4-4-4-4-4-41

Frequently Asked Questions about Bioinformatics Sequence Comparison

What is the main purpose of sequence comparison in bioinformatics?

The main purpose of sequence comparison is to identify new sequences, understand functional and structural patterns within biological molecules, compare how similar sequences are, and perform evolutionary analysis. It helps in predicting the function of an unknown sequence by comparing it to sequences with known functions.

How do PAM and BLOSUM matrices differ in their construction?

PAM matrices (Point Accepted Mutation) are derived from global alignments of very closely related sequences and then extrapolated for longer evolutionary distances. In contrast, BLOSUM matrices (Blocks Of Amino Acid Substitution) are directly calculated from conserved, ungapped blocks of local alignments across various evolutionary distances, making them more empirically derived for divergent sequences. PAM relies on extrapolation, while BLOSUM uses direct observation from sequence families.

Why are gaps penalized in sequence alignments?

Gaps represent insertions or deletions, which are considered evolutionarily expensive events. Penalizing gaps helps to limit the number of gaps introduced into an alignment, preventing biologically implausible scenarios. It encourages finding alignments that reflect a more parsimonious evolutionary history, meaning one with the fewest necessary changes.

When should you use a PAM matrix versus a BLOSUM matrix?

Choose a PAM matrix for aligning highly similar sequences (e.g., PAM40 for 70-90% similarity) or for studies explicitly modeling evolutionary distance from a common ancestor. Opt for a BLOSUM matrix, particularly BLOSUM62, for more divergent sequences (e.g., 30-40% similarity) or for general-purpose searches where detecting more distant relationships is key, as they are often more robust for sequences with lower identity. BLOSUM matrices are generally preferred for database searches like BLAST due to their empirical derivation from diverse blocks.

What are log-odds scores in scoring matrices?

Log-odds scores are a mathematical representation in scoring matrices that quantify the likelihood of an amino acid substitution occurring by biological evolution versus by random chance. A positive log-odds score suggests the substitution is observed more often than expected randomly, indicating a conservative or accepted change. A negative score means the substitution is less common than random chance, suggesting it's less favorable biologically.

Sign up to access full content

Create a free account to unlock all study materials, take interactive tests, listen to podcasts and more.

Create free account

Related topics