rocksdb

History

Changyu Bi cc23b46da1 Support using ZDICT_finalizeDictionary to generate zstd dictionary (#9857 ) Summary: An untrained dictionary is currently simply the concatenation of several samples. The ZSTD API, ZDICT_finalizeDictionary(), can improve such a dictionary's effectiveness at low cost. This PR changes how dictionary is created by calling the ZSTD ZDICT_finalizeDictionary() API instead of creating raw content dictionary (when max_dict_buffer_bytes > 0), and pass in all buffered uncompressed data blocks as samples. Pull Request resolved: https://github.com/facebook/rocksdb/pull/9857 Test Plan: #### db_bench test for cpu/memory of compression+decompression and space saving on synthetic data: Set up: change the parameter [here](`fb9a167a55/tools/db_bench_tool.cc (L1766)`) to 16384 to make synthetic data more compressible. ``` # linked local ZSTD with version 1.5.2 # DEBUG_LEVEL=0 ROCKSDB_NO_FBCODE=1 ROCKSDB_DISABLE_ZSTD=1 EXTRA_CXXFLAGS="-DZSTD_STATIC_LINKING_ONLY -DZSTD -I/data/users/changyubi/install/include/" EXTRA_LDFLAGS="-L/data/users/changyubi/install/lib/ -l:libzstd.a" make -j32 db_bench dict_bytes=16384 train_bytes=1048576 echo "========== No Dictionary ==========" TEST_TMPDIR=/dev/shm ./db_bench -benchmarks=filluniquerandom,compact -num=10000000 -compression_type=zstd -compression_max_dict_bytes=0 -block_size=4096 -max_background_jobs=24 -memtablerep=vector -allow_concurrent_memtable_write=false -disable_wal=true -max_write_buffer_number=8 >/dev/null 2>&1 TEST_TMPDIR=/dev/shm /usr/bin/time ./db_bench -use_existing_db=true -benchmarks=compact -compression_type=zstd -compression_max_dict_bytes=0 -block_size=4096 2>&1 \| grep elapsed du -hc /dev/shm/dbbench/sst \| grep total echo "========== Raw Content Dictionary ==========" TEST_TMPDIR=/dev/shm ./db_bench_main -benchmarks=filluniquerandom,compact -num=10000000 -compression_type=zstd -compression_max_dict_bytes=$dict_bytes -block_size=4096 -max_background_jobs=24 -memtablerep=vector -allow_concurrent_memtable_write=false -disable_wal=true -max_write_buffer_number=8 >/dev/null 2>&1 TEST_TMPDIR=/dev/shm /usr/bin/time ./db_bench_main -use_existing_db=true -benchmarks=compact -compression_type=zstd -compression_max_dict_bytes=$dict_bytes -block_size=4096 2>&1 \| grep elapsed du -hc /dev/shm/dbbench/sst \| grep total echo "========== FinalizeDictionary ==========" TEST_TMPDIR=/dev/shm ./db_bench -benchmarks=filluniquerandom,compact -num=10000000 -compression_type=zstd -compression_max_dict_bytes=$dict_bytes -compression_zstd_max_train_bytes=$train_bytes -compression_use_zstd_dict_trainer=false -block_size=4096 -max_background_jobs=24 -memtablerep=vector -allow_concurrent_memtable_write=false -disable_wal=true -max_write_buffer_number=8 >/dev/null 2>&1 TEST_TMPDIR=/dev/shm /usr/bin/time ./db_bench -use_existing_db=true -benchmarks=compact -compression_type=zstd -compression_max_dict_bytes=$dict_bytes -compression_zstd_max_train_bytes=$train_bytes -compression_use_zstd_dict_trainer=false -block_size=4096 2>&1 \| grep elapsed du -hc /dev/shm/dbbench/sst \| grep total echo "========== TrainDictionary ==========" TEST_TMPDIR=/dev/shm ./db_bench -benchmarks=filluniquerandom,compact -num=10000000 -compression_type=zstd -compression_max_dict_bytes=$dict_bytes -compression_zstd_max_train_bytes=$train_bytes -block_size=4096 -max_background_jobs=24 -memtablerep=vector -allow_concurrent_memtable_write=false -disable_wal=true -max_write_buffer_number=8 >/dev/null 2>&1 TEST_TMPDIR=/dev/shm /usr/bin/time ./db_bench -use_existing_db=true -benchmarks=compact -compression_type=zstd -compression_max_dict_bytes=$dict_bytes -compression_zstd_max_train_bytes=$train_bytes -block_size=4096 2>&1 \| grep elapsed du -hc /dev/shm/dbbench/sst \| grep total # Result: TrainDictionary is much better on space saving, but FinalizeDictionary seems to use less memory. # before compression data size: 1.2GB dict_bytes=16384 max_dict_buffer_bytes = 1048576 space cpu/memory No Dictionary 468M 14.93user 1.00system 0:15.92elapsed 100%CPU (0avgtext+0avgdata 23904maxresident)k Raw Dictionary 251M 15.81user 0.80system 0:16.56elapsed 100%CPU (0avgtext+0avgdata 156808maxresident)k FinalizeDictionary 236M 11.93user 0.64system 0:12.56elapsed 100%CPU (0avgtext+0avgdata 89548maxresident)k TrainDictionary 84M 7.29user 0.45system 0:07.75elapsed 100%CPU (0avgtext+0avgdata 97288maxresident)k ``` #### Benchmark on 10 sample SST files for spacing saving and CPU time on compression: FinalizeDictionary is comparable to TrainDictionary in terms of space saving, and takes less time in compression. ``` dict_bytes=16384 train_bytes=1048576 for sst_file in `ls ../temp/myrock-sst/` do echo "******** $sst_file ********" echo "========== No Dictionary ==========" ./sst_dump --file="../temp/myrock-sst/$sst_file" --command=recompress --compression_level_from=6 --compression_level_to=6 --compression_types=kZSTD echo "========== Raw Content Dictionary ==========" ./sst_dump --file="../temp/myrock-sst/$sst_file" --command=recompress --compression_level_from=6 --compression_level_to=6 --compression_types=kZSTD --compression_max_dict_bytes=$dict_bytes echo "========== FinalizeDictionary ==========" ./sst_dump --file="../temp/myrock-sst/$sst_file" --command=recompress --compression_level_from=6 --compression_level_to=6 --compression_types=kZSTD --compression_max_dict_bytes=$dict_bytes --compression_zstd_max_train_bytes=$train_bytes --compression_use_zstd_finalize_dict echo "========== TrainDictionary ==========" ./sst_dump --file="../temp/myrock-sst/$sst_file" --command=recompress --compression_level_from=6 --compression_level_to=6 --compression_types=kZSTD --compression_max_dict_bytes=$dict_bytes --compression_zstd_max_train_bytes=$train_bytes done 010240.sst (Size/Time) 011029.sst 013184.sst 021552.sst 185054.sst 185137.sst 191666.sst 7560381.sst 7604174.sst 7635312.sst No Dictionary 28165569 / 2614419 32899411 / 2976832 32977848 / 3055542 31966329 / 2004590 33614351 / 1755877 33429029 / 1717042 33611933 / 1776936 33634045 / 2771417 33789721 / 2205414 33592194 / 388254 Raw Content Dictionary 28019950 / 2697961 33748665 / 3572422 33896373 / 3534701 26418431 / 2259658 28560825 / 1839168 28455030 / 1846039 28494319 / 1861349 32391599 / 3095649 33772142 / 2407843 33592230 / 474523 FinalizeDictionary 27896012 / 2650029 33763886 / 3719427 33904283 / 3552793 26008225 / 2198033 28111872 / 1869530 28014374 / 1789771 28047706 / 1848300 32296254 / 3204027 33698698 / 2381468 33592344 / 517433 TrainDictionary 28046089 / 2740037 33706480 / 3679019 33885741 / 3629351 25087123 / 2204558 27194353 / 1970207 27234229 / 1896811 27166710 / 1903119 32011041 / 3322315 32730692 / 2406146 33608631 / 570593 ``` #### Decompression/Read test: With FinalizeDictionary/TrainDictionary, some data structure used for decompression are in stored in dictionary, so they are expected to be faster in terms of decompression/reads. ``` dict_bytes=16384 train_bytes=1048576 echo "No Dictionary" TEST_TMPDIR=/dev/shm/ ./db_bench -benchmarks=filluniquerandom,compact -compression_type=zstd -compression_max_dict_bytes=0 > /dev/null 2>&1 TEST_TMPDIR=/dev/shm/ ./db_bench -use_existing_db=true -benchmarks=readrandom -cache_size=0 -compression_type=zstd -compression_max_dict_bytes=0 2>&1 \| grep MB/s echo "Raw Dictionary" TEST_TMPDIR=/dev/shm/ ./db_bench -benchmarks=filluniquerandom,compact -compression_type=zstd -compression_max_dict_bytes=$dict_bytes > /dev/null 2>&1 TEST_TMPDIR=/dev/shm/ ./db_bench -use_existing_db=true -benchmarks=readrandom -cache_size=0 -compression_type=zstd -compression_max_dict_bytes=$dict_bytes 2>&1 \| grep MB/s echo "FinalizeDict" TEST_TMPDIR=/dev/shm/ ./db_bench -benchmarks=filluniquerandom,compact -compression_type=zstd -compression_max_dict_bytes=$dict_bytes -compression_zstd_max_train_bytes=$train_bytes -compression_use_zstd_dict_trainer=false > /dev/null 2>&1 TEST_TMPDIR=/dev/shm/ ./db_bench -use_existing_db=true -benchmarks=readrandom -cache_size=0 -compression_type=zstd -compression_max_dict_bytes=$dict_bytes -compression_zstd_max_train_bytes=$train_bytes -compression_use_zstd_dict_trainer=false 2>&1 \| grep MB/s echo "Train Dictionary" TEST_TMPDIR=/dev/shm/ ./db_bench -benchmarks=filluniquerandom,compact -compression_type=zstd -compression_max_dict_bytes=$dict_bytes -compression_zstd_max_train_bytes=$train_bytes > /dev/null 2>&1 TEST_TMPDIR=/dev/shm/ ./db_bench -use_existing_db=true -benchmarks=readrandom -cache_size=0 -compression_type=zstd -compression_max_dict_bytes=$dict_bytes -compression_zstd_max_train_bytes=$train_bytes 2>&1 \| grep MB/s No Dictionary readrandom : 12.183 micros/op 82082 ops/sec 12.183 seconds 1000000 operations; 9.1 MB/s (1000000 of 1000000 found) Raw Dictionary readrandom : 12.314 micros/op 81205 ops/sec 12.314 seconds 1000000 operations; 9.0 MB/s (1000000 of 1000000 found) FinalizeDict readrandom : 9.787 micros/op 102180 ops/sec 9.787 seconds 1000000 operations; 11.3 MB/s (1000000 of 1000000 found) Train Dictionary readrandom : 9.698 micros/op 103108 ops/sec 9.699 seconds 1000000 operations; 11.4 MB/s (1000000 of 1000000 found) ``` Reviewed By: ajkr Differential Revision: D35720026 Pulled By: cbi42 fbshipit-source-id: 24d230fdff0fd28a1bb650658798f00dfcfb2a1f		2022-05-20 12:09:09 -07:00
..
adaptive	More refactoring ahead of footer & meta changes (#9240 )	2021-12-10 08:13:26 -08:00
block_based	Support using ZDICT_finalizeDictionary to generate zstd dictionary (#9857 )	2022-05-20 12:09:09 -07:00
cuckoo	Remove own ToString() (#9955 )	2022-05-06 13:03:58 -07:00
plain	Added GetFactoryCount/Names/Types to ObjectRegistry (#9358 )	2022-05-16 09:44:43 -07:00
block_fetcher.cc	Remove own ToString() (#9955 )	2022-05-06 13:03:58 -07:00
block_fetcher.h	More refactoring ahead of footer & meta changes (#9240 )	2021-12-10 08:13:26 -08:00
block_fetcher_test.cc	Make MemoryAllocator into a Customizable class (#8980 )	2021-12-17 04:20:47 -08:00
cleanable_test.cc	Eliminate unnecessary (slow) block cache Ref()ing in MultiGet (#9899 )	2022-04-26 21:59:24 -07:00
format.cc	Set Read rate limiter priority dynamically and pass it to FS (#9996 )	2022-05-18 19:41:44 -07:00
format.h	Optimize & clean up footer code (#9280 )	2021-12-13 17:43:07 -08:00
get_context.cc	Support readahead during compaction for blob files (#9187 )	2021-11-19 17:53:47 -08:00
get_context.h	Cleanup includes in dbformat.h (#8930 )	2021-09-29 04:04:40 -07:00
internal_iterator.h	Reuse internal auto readhead_size at each Level (expect L0) for Iterations (#9056 )	2021-11-10 16:20:04 -08:00
iter_heap.h	Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433 )	2020-02-20 12:09:57 -08:00
iterator.cc	Eliminate unnecessary (slow) block cache Ref()ing in MultiGet (#9899 )	2022-04-26 21:59:24 -07:00
iterator_wrapper.h	Reuse internal auto readhead_size at each Level (expect L0) for Iterations (#9056 )	2021-11-10 16:20:04 -08:00
merger_test.cc	Cleanup multiple implementations of VectorIterator (#8901 )	2021-10-06 07:48:31 -07:00
merging_iterator.cc	MergingIterator: rearrange fields to reduce paddings (#9024 )	2021-10-14 12:01:56 -07:00
merging_iterator.h	Cleanup includes in dbformat.h (#8930 )	2021-09-29 04:04:40 -07:00
meta_blocks.cc	Use std::numeric_limits<> (#9954 )	2022-05-05 13:08:21 -07:00
meta_blocks.h	Tests for filter compatibility (#9773 )	2022-04-06 15:54:40 -07:00
mock_table.cc	Add rate limiter priority to ReadOptions (#9424 )	2022-02-16 23:18:14 -08:00
mock_table.h	Fix some minor issues in the Customizable infrastructure (#8566 )	2021-08-19 10:10:47 -07:00
multiget_context.h	Multi file concurrency in MultiGet using coroutines and async IO (#9968 )	2022-05-19 15:36:27 -07:00
persistent_cache_helper.cc	New stable, fixed-length cache keys (#9126 )	2021-12-16 17:15:13 -08:00
persistent_cache_helper.h	Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433 )	2020-02-20 12:09:57 -08:00
persistent_cache_options.h	Use STATIC_AVOID_DESTRUCTION for static objects with non-trivial destructors (#9958 )	2022-05-17 09:39:22 -07:00
scoped_arena_iterator.h	Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433 )	2020-02-20 12:09:57 -08:00
sst_file_dumper.cc	Support using ZDICT_finalizeDictionary to generate zstd dictionary (#9857 )	2022-05-20 12:09:09 -07:00
sst_file_dumper.h	Support using ZDICT_finalizeDictionary to generate zstd dictionary (#9857 )	2022-05-20 12:09:09 -07:00
sst_file_reader.cc	Fast path for detecting unchanged prefix_extractor (#9407 )	2022-01-21 11:37:46 -08:00
sst_file_reader_test.cc	Use the comparator from the sst file table properties in sst_dump_tool (#9491 )	2022-02-08 12:15:35 -08:00
sst_file_writer.cc	`compression_per_level` should be used for flush and changeable (#9658 )	2022-03-07 18:06:19 -08:00
sst_file_writer_collectors.h	Remove own ToString() (#9955 )	2022-05-06 13:03:58 -07:00
table_builder.h	Fast path for detecting unchanged prefix_extractor (#9407 )	2022-01-21 11:37:46 -08:00
table_factory.cc	Restore Regex support for ObjectLibrary::Register, rename new APIs to allow old one to be deprecated in the future (#9362 )	2022-01-11 06:33:48 -08:00
table_properties.cc	Remove own ToString() (#9955 )	2022-05-06 13:03:58 -07:00
table_properties_internal.h	Improve / clean up meta block code & integrity (#9163 )	2021-11-18 11:43:44 -08:00
table_reader.h	Multi file concurrency in MultiGet using coroutines and async IO (#9968 )	2022-05-19 15:36:27 -07:00
table_reader_bench.cc	Fast path for detecting unchanged prefix_extractor (#9407 )	2022-01-21 11:37:46 -08:00
table_reader_caller.h	Fix and detect headers with missing dependencies (#8893 )	2021-09-10 10:00:26 -07:00
table_test.cc	Adjust public APIs to prefer 128-bit SST unique ID (#10009 )	2022-05-17 18:43:48 -07:00
two_level_iterator.cc	Clarify caching behavior for index and filter partitions (#9068 )	2021-10-27 17:23:04 -07:00
two_level_iterator.h	Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433 )	2020-02-20 12:09:57 -08:00
unique_id.cc	Track SST unique id in MANIFEST and verify (#9990 )	2022-05-19 11:04:21 -07:00
unique_id_impl.h	Track SST unique id in MANIFEST and verify (#9990 )	2022-05-19 11:04:21 -07:00