High Availability Ultra-Large Cloud-Oriented Storage Systems: A Review, Mind Mapping and Open Research Directions
Abstract
Cloud computing paradigm enables the delivery of computing, storage, and applications as services as subscription-oriented services to end users over the Internet. It reduces the need for highly expensive computing infrastructure, dedicated storage servers, and software applications to be procured, deployed, and managed as in house capabilities. The amount of data generated by enterprise and business applications has increased exponentially. Addressing and dealing with enormous datasets is a challenging and very time-consuming task that necessitates a high level of computing infrastructure to enable proper data processing and analysis. This study aims to compare and analyze the different aspects of ultra-large data storage systems in cloud computing. The description, features, and classification of cloud storage systems are discussed, along with certain cloud computing considerations. Relationships between big data and cloud computing are also explored, as are ultra-large storage systems and Hadoop technology. Furthermore, research concerns such as dynamicity, scalability, availability, data integrity, self-organizing, data quality, data diversity, privacy, and legal and regulatory issues have been investigated. The sole motivation of this study has been thoroughly explained with the help of a mind-mapping diagram of cloud-oriented data storage (CODS) elements. The seven major elements of CODS—the data model, data consistency, data replication, data scalability, data integration, data virtualization, and transaction—have been systematically discussed in this study. This study also highlights some open research issues related to ultra-large storage systems, which summarize sufficient research efforts.