A Primer on HBase Filters

Motivation Databases need flexible query patterns. A KV store with only Get, Put, and Scan would frustrate users—real workloads are richer. For orders: “this user’s orders in the last three months” needs at least (1) filter by user and (2) filter by time range, combined with AND. Scanning the whole table on the client and filtering locally would crush the cluster. Server-side filters are required. You also see OR, e.g. orders for Alice or Bob in the last three months—AND vs OR, often mixed....

July 2, 2019 · Zheng Hu

Nebula Beijing Meetup Summary

Nebula held its second Beijing meetup in Jingdong Beichen. CEO Xiaomeng Ye led Geobase, Ant Financial’s graph database. Two technical directors—Heng Chen and Fenglin Hou (dutor)—drive storage (similar to TiKV) and the query engine (similar to TiDB), respectively. The architecture separates storage and compute. The compute layer is SQL-like, with Go syntax for edge hops and pipelines to simplify nested subqueries. Storage is key-value with static hash partitioning by key (no dynamic partitioning yet)....

June 29, 2019 · Zheng Hu

Further GC optimization for HBase3.x: Reading HFileBlock into offheap directly

In HBASE-21879, we redesigned the offheap read path: read the HFileBlock from HDFS to pooled offheap ByteBuffers directly, while before HBASE-21879 we just read the HFileBlock to heap which would still lead to high GC pressure. After few months of development and testing, all subtasks have been resovled now except the HBASE-21946 (It depends on HDFS-14483 and our HDFS teams are working on this, we expect the HDFS-14483 to be included in hadoop 2....

June 23, 2019 · Zheng Hu

From HBase Off-Heap to Netty Memory Management

HBase Off-Heap Today HBase is a widely used distributed NoSQL database. Many workloads—feeds, ads, and similar—demand high throughput and low latency. HBase 2.0 off-heaped the core read and write paths: allocations go to JVM off-heap memory, which is not GC-managed and must be freed explicitly. On the write path, request buffers are allocated off-heap until data is written to the WAL and memstore. The memstore’s ConcurrentSkipListMap holds references to cells, not cell bodies; actual data lives in MSLAB chunks for easier off-heap management....

February 23, 2019 · Zheng Hu

HBaseCon West 2018 Talk - HBase Practice at Xiaomi

HBaseConWest2018 was held on June 18 in San Jose, California, hosted by Hortonworks. Attending HBaseCon West in Silicon Valley each year has become routine for the Xiaomi HBase team—our community presence is well known (seven HBase Committers, two PMC members), and the company is willing to share a year-in-review of internal practice and community contributions. In 2018 we submitted the talk “HBase Practice at Xiaomi,” spent considerable effort preparing it, and rehearsed in English three times internally....

June 18, 2018 · Zheng Hu

Becoming an HBase Committer

On October 20, I accepted an invitation from the Apache HBase community and became an HBase Committer. At Xiaomi I maintain our internal HBase branch and production clusters, and the company already had six HBase Committers—including one PMC member—so becoming a Committer was a natural next step rather than something extraordinary. Compared with someone doing HBase at a company with no Committers, it takes considerably more time and effort. Some observations about the community:...

October 22, 2017 · Zheng Hu

HBase Region Balance in Practice

HBase is a distributed key-value database that supports automatic load balancing. With the balance switch (balance_switch) enabled, the HMaster process automatically selects regions according to a specified policy and assigns them to RegionServers with lower load. The official distribution currently supports two region-selection policies: DefaultLoadBalancer and StochasticLoadBalancer, both described in detail below. Because all HBase data (including HLog, meta, HStoreFile, and so on) is written to HDFS, region moves are very lightweight....

June 28, 2017 · Zheng Hu

HBaseCon West 2017 Session Notes

Notes on HBaseCon West 2017 presentations: 1. HBase at Xiaomi Presented jointly by Zhe Yang and Guanghao Zhang—both became HBase Committers in 2016 (Xiaomi has produced eight HBase Committers in total, including two PMC members, and has resolved hundreds of issues). Highlights included: Lessons from upgrading clusters from 0.94 to 0.98. Experience using G1GC for internal HBase deployments. Community contributions and improvements in 2016, including ordered replication log push, Scan optimizations, async client development, and related benchmark results....

June 28, 2017 · Zheng Hu

HBase HLog Replay Ordering Inconsistency

In an HBase master-slave replication cluster, as shown on the left in the figure below, Region-Server-X and Region-Server-Y are two RegionServers in the master cluster. Under normal conditions, writes to Region-A append logs to Hlog-X on Region-Server-X, and Region-Server-X asynchronously applies those HLog entries in batches to the slave cluster. If Region-Server-X then crashes, Region-A is taken over by Region-Server-Y. Subsequent writes to Region-A append logs to Hlog-Y, while Region-Server-Y starts a new thread to replay Hlog-X....

June 13, 2016 · Zheng Hu

TokuDB's Multi-Version Concurrency Control (MVCC)

This article covers transaction isolation in TokuDB. The source implementation is complex; for clarity, we focus on the most essential parts and omit minor details. Background In traditional relational databases (Oracle, MySQL, SQL Server, and others), transactions are central to both engineering and discussion. The core properties of a transaction are ACID. A (atomicity) means a transaction’s sub-operations have only two outcomes: all succeed on commit, or all are undone on rollback....

December 13, 2015 · Zheng Hu