CloudForest: A Scalable and Efficient Random Forest Implementation for Biological Data

Bressler, Ryan; Kreisberg, Richard B.; Bernard, Brady; Niederhuber, John E.; Vockley, Joseph G.; Shmulevich, Ilya; Knijnenburg, Theo A.

Published in

Public Library of Science, PLoS ONE, 12(10), p. e0144820, 2015

DOI: 10.1371/journal.pone.0144820

Tools

Export citation

Search in Google Scholar

CloudForest: A Scalable and Efficient Random Forest Implementation for Biological Data

Journal article published in 2015 by Ryan Bressler, Richard B. Kreisberg, Brady Bernard, John E. Niederhuber, Joseph G. Vockley, Ilya Shmulevich, Theo A. Knijnenburg

This paper is made freely available by the publisher.

Full text: Download

Preprint: archiving allowed

Upload

Postprint: archiving allowed

Upload

Published version: archiving allowed

Upload

Policy details

Data provided by

Abstract

Random Forest has become a standard data analysis tool in computational biology. However, extensions to existing implementations are often necessary to handle the complexity of biological datasets and their associated research questions. The growing size of these datasets requires high performance implementations. We describe CloudForest, a Random Forest package written in Go, which is particularly well suited for large, heterogeneous, genetic and biomedical datasets. CloudForest includes several extensions, such as dealing with unbalanced classes and missing values. Its flexible design enables users to easily implement additional extensions. CloudForest achieves fast running times by effective use of the CPU cache, optimizing for different classes of features and efficiently multi-threading. https://github.com/ilyalab/CloudForest.

Published in

Links

Tools

CloudForest: A Scalable and Efficient Random Forest Implementation for Biological Data

Abstract