Hive:数据仓库利器

1. 简介

Hive是一个基于Hadoop的开源数据仓库工具,可以用来存储、查询和分析大规模数据。Hive使用SQL-like的HiveQL语言来查询数据,并将其结果存储在Hadoop的文件系统中。

2. 基本概念

介绍 Hive 的核心概念,例如表、分区、桶、HQL 等。

2.1 架构

Design - Apache Hive - Apache Software Foundation

|----------------------|-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| 组成 | 详情 |
| UI | The user interfacefor users to submit queries and other operations to the system. As of 2011 the system had a command line interface and a web based GUI was being developed. |
| Driver | The component whichreceives the queries. This component implements the notion of session handles and provides execute and fetch APIs modeled on JDBC/ODBC interfaces. |
| Compiler | The component thatparses the query, does semantic analysis on the different query blocks and query expressions and eventually generates an execution plan with the help of the table and partition metadata looked up from the metastore. |
| Metastore | The component that stores all the structure information of the various tables and partitions in the warehouse including column and column type information, the serializers and deserializers necessary to read and write data and the corresponding HDFS files where the data is stored. |
| Execution Engine | The component which executes the execution plan created by the compiler. The plan is a DAG of stages. The execution engine manages the dependencies between these different stages of the plan and executes these stages on the appropriate system components. |

2.2 Data Model

|----------------|-----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| 类型 | 详情 |
| Tables | These are analogous to Tables in Relational Databases. Tables can be filtered, projected, joined and unioned. Additionally all the data of a table is stored in a directory in HDFS. Hive also supports the notion of external tables wherein a table can be created on prexisting files or directories in HDFS by providing the appropriate location to the table creation DDL. The rows in a table are organized into typed columns similar to Relational Databases. |
| Partitions | Each Table can have one or more partition keys which determine how the data is stored, for example a table T with a date partition column ds had files with data for a particular date stored in the <table location>/ds=<date> directory in HDFS. Partitions allow the system to prune data to be inspected based on query predicates, for example a query that is interested in rows from T that satisfy the predicate T.ds = '2008-09-01' would only have to look at files in <table location>/ds=2008-09-01/ directory in HDFS. |
| Buckets | **Data in each partition may in turn be divided into Buckets based on the hash of a column in the table.**Each bucket is stored as a file in the partition directory. Bucketing allows the system to efficiently evaluate queries that depend on a sample of data (these are queries that use the SAMPLE clause on the table). |

3. 实践应用

3.1 数仓建设

4. 性能优化

介绍如何优化 Hive 的性能

5. 常见问题解答

5.1 常用SQL

|--------|--------------------------------|
| 场景 | SQL |
| 连续n天登录 | sql SELECT * FROM test; |
| | |

6. 总结

总结 Hive 的关键知识点,并提供学习资源和进一步研究方向。

相关推荐
糖醋_诗酒1 天前
数据仓库 - 转转 - 一面凉经
数据仓库
醉颜凉1 天前
数据仓库实战:自动化数据质量检测全流程——精准提升数据准确性与完整性
大数据·数据仓库·自动化
RestCloud1 天前
信创数据集成实践:某央企ETL全链路国产化迁移架构拆解
数据仓库·架构·数据安全·etl·etlcloud·数据传输·国产化替代
SelectDB技术团队3 天前
雨润集团 统一实时数据仓库:Apache Doris / SelectDB 的技术能力与实践
数据仓库·实时数仓·apache doris·selectdb
Irene19913 天前
剑指大数据:企业级数据仓库项目实战(金融租赁版)读书笔记__大数据测试集群的服务器节点服务分配规划表__解读(附:技术选型清单)
大数据·数据仓库·金融
Htr_4 天前
Outcome 核心概念与实战应用指南
大数据·hadoop·apache
SelectDB技术团队4 天前
抖音集团 实时数据仓库:Apache Doris / SelectDB 的技术能力与实践
数据仓库·apache
小马过河R5 天前
数据仓库入门:是什么、怎么做、和数据库有什么区别?
大数据·数据库·数据仓库·架构·驾驭工程·fde
BYSJMG5 天前
计算机毕业设计选题推荐|【基于大数据的城市噪音数据可视化分析】Spark+K-Means+FP-Growth实战
大数据·hadoop·信息可视化·spark·课程设计
慧一居士7 天前
Apache StreamPark 功能和使用场景介绍、使用步骤详细示例
数据仓库