Toward the development of a gene index to the human genome: an assessment of the nature of high-throughput EST sequence data.-Reference-Cited by-同舟云学术

Toward the development of a gene index to the human genome: an assessment of the nature of high-throughput EST sequence data.

Published:1996-09 Issue:9 Volume:6 Page:829-845
ISSN:1088-9051
Container-title:Genome Research
language:en
Short-container-title:Genome Res.

Author:

Aaronson J S,Eckman B,Blevins R A,Borkowski J A,Myerson J,Imran S,Elliston K O

Abstract

A rigorous analysis of the Merck-sponsored EST data with respect to known gene sequences increases the utility of the data set and helps refine methods for building a gene index. A highly curated human transcript data base was used as a reference data set of known genes. A detailed analysis of EST sequences derived from known genes was performed to assess the accuracy of EST sequence annotation. The EST data was screened to remove low-quality and low-complexity sequences. A set of high-quality ESTs similar to the transcript data base was identified using BLAST; this subset of ESTs was compared with the set of known genes using the Smith-Waterman algorithm. Error rates of several types were assessed based on a flexible match criterion defining sequence identity. The rate of lane-tracking errors is very low, approximately 0.5%. Insert size data is accurate within approximately 20%. Reversed clone and internal priming error rates are approximately 5% and 2.5%, respectively, contributing to the incorrect identification of reads as 3' ends of genes. Follow-up investigation reveals that a significant number of clones, miscategorized as reversed, represent overlapping genes on the opposite strand of entries in the transcript data base. Relevance of these results to the creation of a high-quality index to the human genome capable of supporting diverse genomic investigations is discussed.

Publisher

Cold Spring Harbor Laboratory

Subject

Genetics(clinical),Genetics

Reference33 articles.

1. Aaronson, J.S. and K.O. Elliston. 1996. ftp://avery.merck.com/mgi/IndexReport.

2. Complementary DNA Sequencing: Expressed Sequence Tags and Human Genome Project

3. Initial assessment of human gene diversity and expression patterns based upon 83 million nucleotides of cDNA sequence.;Nature,1995

4. GenBank

Cited by 99 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Overview of Cancer Genomics, Organization, and Variations in the Human Genome;'Essentials of Cancer Genomic, Computational Approaches and Precision Medicine;2020

2. Template-switching artifacts resemble alternative polyadenylation;BMC Genomics;2019-11-08

3. Development of expressed sequenced tags (EST) to identify some pathogen resistance genes expressed in Gossypium arboreum;Gene Reports;2019-06

4. Understanding non-coding DNA regions in yeast;Biochemical Society Transactions;2013-11-20

5. Overview of DNA Microarrays: Types, Applications, and Their Future;Current Protocols in Molecular Biology;2013-01