Towards the Development of Large-Scale Data Warehouse Application Frameworks

Defining data warehouse requirements is widely recognized as one of the most important steps in the larger data warehouse system development process. This paper examines the potential risks and pitfalls within the data warehouse requirement collection and definition process. A real scenario of a large-scale data warehouse implementation is given, and details of this project, which ultimately failed due to inadequate requirement collection and definition process, are described. The presented case underscores and illustrates the impact of the requirement collection and definition process on the data warehouse implementation, while the case is analyzed within the context of the existing approaches, methodologies, and best practices for prevention and avoidance of typical data warehouse requirement errors and oversights.

Download Full-text

Data Warehousing Requirements Collection and Definition

International Journal of Business Intelligence Research ◽

10.4018/jbir.2010070105 ◽

2010 ◽

Vol 1 (3) ◽

pp. 66-76

Author(s):

Nenad Jukic ◽

Miguel Velasco

Keyword(s):

Data Warehouse ◽

Large Scale ◽

Development Process ◽

System Development ◽

Large Scale Data ◽

Typical Data ◽

Potential Risks ◽

Real Scenario ◽

The Impact ◽

Scale Data

Defining data warehouse requirements is widely recognized as one of the most important steps in the larger data warehouse system development process. This paper examines the potential risks and pitfalls within the data warehouse requirement collection and definition process. A real scenario of a large-scale data warehouse implementation is given, and details of this project, which ultimately failed due to inadequate requirement collection and definition process, are described. The presented case underscores and illustrates the impact of the requirement collection and definition process on the data warehouse implementation, while the case is analyzed within the context of the existing approaches, methodologies, and best practices for prevention and avoidance of typical data warehouse requirement errors and oversights.

Download Full-text

A Context-Based Performance Enhancement Algorithm for Columnar Storage in MapReduce with Hive

International Journal of Cloud Applications and Computing ◽

10.4018/ijcac.2013100104 ◽

2013 ◽

Vol 3 (4) ◽

pp. 38-50 ◽

Cited By ~ 1

Author(s):

Yashvardhan Sharma ◽

Saurabh Verma ◽

Sumit Kumar ◽

Shivam U.

Keyword(s):

Data Warehouse ◽

Large Scale ◽

Performance Enhancement ◽

High Reliability ◽

Data Management System ◽

Mapreduce Framework ◽

Large Scale Data ◽

Query Engine ◽

Commodity Clusters ◽

Scale Data

To achieve high reliability and scalability, most large-scale data warehouse systems have adopted the cluster-based architecture. In this context, MapReduce has emerged as a promising architecture for large scale data warehousing and data analytics on commodity clusters. The MapReduce framework offers several lucrative features such as high fault-tolerance, scalability and use of a variety of hardware from low to high range. But these benefits have resulted in substantial performance compromise. In this paper, we propose the design of a novel cluster-based data warehouse system, Daenyrys for data processing on Hadoop – an open source implementation of the MapReduce framework under the umbrella of Apache. Daenyrys is a data management system which has the capability to take decision about the optimum partitioning scheme for the Hadoop's distributed file system (DFS). The optimum partitioning scheme improves the performance of the complete framework. The choice of the optimum partitioning is query-context dependent. In Daenyrys, the columns are formed into optimized groups to provide the basis for the partitioning of tables vertically. Daenyrys has an algorithm that monitors the context of current queries and based on the observations, it re-partitions the DFS for better performance and resource utilization. In the proposed system, Hive, a MapReduce-based SQL-like query engine is supported above the DFS.

Download Full-text

HDW: A High Performance Large Scale Data Warehouse

2008 International Multi-symposiums on Computer and Computational Sciences ◽

10.1109/imsccs.2008.16 ◽

2008 ◽

Cited By ~ 3

Author(s):

Jinguo You ◽

Jianqing Xi ◽

Chuan Zhang ◽

Gengqi Guo

Keyword(s):

Data Warehouse ◽

High Performance ◽

Large Scale ◽

Large Scale Data ◽

Scale Data

Download Full-text

Large-Scale Data Learning Method for Anomaly Detection using Machine Learning for Monitoring Vibration in Vehicle Equipment

IEEJ Transactions on Industry Applications ◽

10.1541/ieejias.140.480 ◽

2020 ◽

Vol 140 (6) ◽

pp. 480-487

Author(s):

Minoru Kondo

Keyword(s):

Machine Learning ◽

Anomaly Detection ◽

Large Scale ◽

Learning Method ◽

Large Scale Data ◽

Scale Data

Download Full-text

Faculty Opinions recommendation of Comparative assessment of large-scale data sets of protein-protein interactions.

Faculty Opinions – Post-Publication Peer Review of the Biomedical Literature ◽

10.3410/f.1006598.82257 ◽

2002 ◽

Author(s):

Rob Russell

Keyword(s):

Protein Interactions ◽

Large Scale ◽

Comparative Assessment ◽

Data Sets ◽

Protein Protein Interactions ◽

Large Scale Data ◽

Scale Data ◽

Large Scale Data Sets

Download Full-text

ProGen:Provenance database generator for large-scale data set

Journal of Computer Applications ◽

10.3724/sp.j.1087.2008.02737 ◽

2009 ◽

Vol 28 (11) ◽

pp. 2737-2740

Author(s):

Xiao ZHANG ◽

Shan WANG ◽

Na LIAN

Keyword(s):

Large Scale ◽

Data Set ◽

Large Scale Data ◽

Scale Data

Download Full-text

Construction of integrated particle rendering environment for large scale data visualization

Impact ◽

10.21820/23987073.2018.11.9 ◽

2018 ◽

Vol 2018 (11) ◽

pp. 9-11

Author(s):

Koji Koyamada

Keyword(s):

Data Visualization ◽

Large Scale ◽

Large Scale Data ◽

Scale Data

Download Full-text

COMMUNITY-CURATED DATA RESOURCES AND LARGE-SCALE DATA-MODEL SYNTHESES: THE CHILDREN OF COHMAP

10.1130/abs/2016am-286533 ◽

2016 ◽

Author(s):

John W. Williams ◽

◽

Simon Goring ◽

Eric Grimm ◽

Jason McLachlan

Keyword(s):

Data Model ◽

Large Scale ◽

Large Scale Data ◽

Scale Data

Download Full-text