Estimate the Performance of Cloudera Decision Support Queries
DOI:
https://doi.org/10.3991/ijoe.v18i01.27877Keywords:
Hadoop, Impala, Hive, Massive parallel processing, Big data, TPC-H, Graphics Data expo 2009Abstract
Hive and Impala queries are used to process a big amount of data. The overwriting amount of information requires an efficient data processing system. When we deal with a long-term batch query and analysis Hive will be more suitable for this query. Impala is the most powerful system suitable for real-time interactive Structured Query Language (SQL) query which are added a massive parallel processing to Hadoop distributed cluster. The data growth makes a problem with SQL Cluster because the execution processing time is increased. In this paper, a comparison is demonstrated between the performance time of Hive, Impala and SQL on two different data models with different queries chosen to test the performance. The results demonstrate that Impala outperforms Hive and SQL cluster when it comes to analyze data and processing tasks. Using two benchmark datasets, TPC-H and statistical computing, we compare the performance of Hive, Impala, and SQL clusters 2009 Statistical Graphics Data Expo.
Downloads
Published
2022-01-26
How to Cite
Allam, T. M. (2022). Estimate the Performance of Cloudera Decision Support Queries. International Journal of Online and Biomedical Engineering (iJOE), 18(01), pp. 127–138. https://doi.org/10.3991/ijoe.v18i01.27877
Issue
Section
Papers
License
Copyright (c) 2021 Tahani M. Allam
This work is licensed under a Creative Commons Attribution 4.0 International License.