Teórica 3
Búsquedas de secuencias en bases de datos
Algoritmos heurísticos de búsqueda de secuencias similares en bases de datos. BLAST. FASTA.
Estadística de las búsquedas y significancia de los hits y alineamientos identificados. E-values, distribución de valores extremos.
Material de lectura y consulta
- Bioinformatics (2020). Edited by Baxevanis, Bader, Wishart. 4th Edition, Wiley. Hay copia en el laboratorio de Bioinformática (IIB)
- The Statistics of Sequence Similarity Scores
Karlin S, Altschul SF. Methods for assessing the statistical significance of molecular sequence features by using general scoring schemes (1990). Proc Natl Acad Sci U S A. 87(6):2264-8. doi: 10.1073/pnas.87.6.2264. PMID: 2315319; PMCID: PMC53667.
Pearson WR, Lipman DJ. Improved tools for biological sequence comparison (1988). Proc Natl Acad Sci U S A. 85(8):2444-8. doi: 10.1073/pnas.85.8.2444. PMID: 3162770; PMCID: PMC280013.
- Selecting the Right Similarity-Scoring Matrix. Pearson, W. R. (2013) Curr. Prot. Bioinformatics Chapter 3: Unit 3.5
Having a BLAST with bioinformatics (and avoiding BLASTphemy). Pertsemlidis, A., Fondon, J.W. (2001) Genome Biol 2, reviews2002.1
- An Introduction to Similarity ("Homology") Searching. Pearson, W. R. (2013) Curr. Prot. Bioinformatics Chapter 3: Unit 3.1
- The FASTA package of programs (W. R. Pearson and D. J. Lipman (1988)
- Download BLAST Software and Databases (NCBI)
A Novel algorithm for identifying low-complexity regions in a protein sequence. Xuehui Li, Tamer Kahveci (2006) Bioinformatics 22: 2980--2987
- Software to detect low complexity regions