Please use this identifier to cite or link to this item: https://hdl.handle.net/20.500.11851/6634
Title: Employing Machine Learning Techniques for Data Enrichment: Increasing the number of samples for effective gene expression data analysis
Authors: Erdoğdu, Utku
Tan, Mehmet
Alhajj, Reda
Polat, Faruk
Demetrick, Douglas
Rokne, Jon
Keywords: gene expression data
sample generation
learning
genetic algorithm
probabilistic boolean network
Issue Date: 2011
Publisher: IEEE Computer Soc
Source: IEEE International Conference on Bioinformatics and Biomedicine (BIBM) -- NOV 12-15, 2011 -- Atlanta, GA
Series/Report no.: IEEE International Conference on Bioinformatics and Biomedicine-BIBM
Abstract: For certain domains, e. g. bioinformatics, producing more real samples is costly, error prone and time consuming. Therefore, there is a need for an intelligent automated process capable of substituting the real samples by artificial samples that carry the same characteristics as the real samples and hence could be used for running comprehensive testing of new methodologies. Motivated by this need, we describe a novel approach that integrates Probabilistic Boolean Network and genetic algorithm based techniques into a framework that uses some existing real samples as input and successfully produces new samples as output. The new samples will inspire the characteristics of the existing samples without duplicating them. This leads to diversity in the samples and hence a more rich set of samples to be used in testing. The developed framework incorporates two models (perspectives) for sample generation. We illustrate its applicability for producing new gene expression data samples; a high demanding area that has not received attention. The two perspectives employed in the process are based on models that are not closely related; the independence eliminates the bias of having the produced approach covering only certain characteristics of the domain and leading to samples skewed towards one direction. The produced results are very promising in showing the effectiveness, usefulness and applicability of the proposed multi-model framework.
URI: https://doi.org/10.1109/BIBM.2011.105
https://hdl.handle.net/20.500.11851/6634
ISBN: 978-0-7695-4574-5
ISSN: 2156-1125
2156-1133
Appears in Collections:Bilgisayar Mühendisliği Bölümü / Department of Computer Engineering
Scopus İndeksli Yayınlar Koleksiyonu / Scopus Indexed Publications Collection
WoS İndeksli Yayınlar Koleksiyonu / WoS Indexed Publications Collection

Show full item record

CORE Recommender

SCOPUSTM   
Citations

3
checked on Sep 23, 2022

WEB OF SCIENCETM
Citations

4
checked on Sep 24, 2022

Page view(s)

26
checked on Feb 6, 2023

Google ScholarTM

Check

Altmetric


Items in GCRIS Repository are protected by copyright, with all rights reserved, unless otherwise indicated.