Please use this identifier to cite or link to this item: http://hdl.handle.net/10397/16406
Title: Attribute clustering for grouping, selection, and classification of gene expression data
Authors: Au, WH
Chan, KCC 
Wong, AKC
Yang, W
Keywords: Attribute clustering
Data mining
Gene expression classification
Gene selection
Microarray analysis
Issue Date: 2005
Publisher: ACM Special Interest Group
Source: IEEE/ACM transactions on computational biology and bioinformatics, 2005, v. 2, no. 2, p. 83-101 How to cite?
Journal: IEEE/ACM transactions on computational biology and bioinformatics 
Abstract: This paper presents an attribute clustering method which is able to group genes based on their interdependence so as to mine meaningful patterns from the gene expression data. It can be used for gene grouping, selection, and classification. The partitioning of a relational table into attribute subgroups allows a small number of attributes within or across the groups to be selected for analysis. By clustering attributes, the search dimension of a data mining algorithm is reduced. The reduction of search dimension is especially important to data mining in gene expression data because such data typically consist of a huge number of genes (attributes) and a small number of gene expression profiles (tuples). Most data mining algorithms are typically developed and optimized to scale to the number of tuples instead of the number of attributes. The situation becomes even worse when the number of attributes overwhelms the number of tuples, in which case, the likelihood of reporting patterns that are actually irrelevant due to chances becomes rather high. It is for the aforementioned reasons that gene grouping and selection are important preprocessing steps for many data mining algorithms to be effective when applied to gene expression data. This paper defines the problem of attribute clustering and introduces a methodology to solving it. Our proposed method groups interdependent attributes into clusters by optimizing a criterion function derived from an information measure that reflects the interdependence between attributes. By applying our algorithm to gene expression data, meaningful clusters of genes are discovered. The grouping of genes based on attribute interdependence within group helps to capture different aspects of gene association patterns in each group. Significant genes selected from each group then contain useful information for gene expression classification and identification. To evaluate the performance of the proposed approach, we applied it to two well-known gene expression data sets and compared our results with those obtained by other methods. Our experiments show that the proposed method is able to find the meaningful clusters of genes. By selecting a subset of genes which have high multiple-interdependence with others within clusters, significant classification information can be obtained. Thus, a small pool of selected genes can be used to build classifiers with very high classification rate. From the pool, gene expressions of different categories can be identified.
URI: http://hdl.handle.net/10397/16406
ISSN: 1545-5963
EISSN: 1557-9964
DOI: 10.1109/TCBB.2005.17
Appears in Collections:Journal/Magazine Article

Access
View full-text via PolyU eLinks SFX Query
Show full item record

SCOPUSTM   
Citations

112
Last Week
0
Last month
2
Citations as of Aug 12, 2017

WEB OF SCIENCETM
Citations

73
Last Week
0
Last month
0
Citations as of Aug 13, 2017

Page view(s)

43
Last Week
2
Last month
Checked on Aug 13, 2017

Google ScholarTM

Check

Altmetric



Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.