1. Schuler, G. D., Epstein, J. A., Ohkawa, H. & Kans, J. A. Entrez: molecular biology database and retrieval system. Methods Enzymol. 266, 141–162 (1996)
2. Wilbur, W. J. & Yang, Y. An analysis of statistical term strength and its use in the indexing and retrieval of molecular biology texts. Comput. Biol. Med. 26, 209–222 (1996).Describes the vector-space model used by Entrez, the literature-search service maintained by the NCBI.
3. Renner, A. & Aszodi, A. High-throughput functional annotation of novel gene products using document clustering. Proc. Pacific Symp. Biocomp. 5, 54–68 (2000).
4. Shatkay, H., Edwards, S., Wilbur, W. J. & Boguski, M. Genes, themes, and microarrays. Proc. Int. Conf. Intell. Syst. Mol. Biol. 8, 317–327 (2000).
5. Manning, C. D. & Schutze, H. S. in Foundations of Statistical Natural Language Processing 85 (MIT press, Cambridge, Massachusetts, 1999).The indispensable reference for anyone who is interested in statistical natural language processing (NLP).