rocksdb

mirror of https://github.com/facebook/rocksdb.git synced 2024-11-29 18:33:58 +00:00

History

Peter Dillinger 239d17a19c Support optimize_filters_for_memory for Ribbon filter (#7774 ) Summary: Primarily this change refactors the optimize_filters_for_memory code for Bloom filters, based on malloc_usable_size, to also work for Ribbon filters. This change also replaces the somewhat slow but general BuiltinFilterBitsBuilder::ApproximateNumEntries with implementation-specific versions for Ribbon (new) and Legacy Bloom (based on a recently deleted version). The reason is to emphasize speed in ApproximateNumEntries rather than 100% accuracy. Justification: ApproximateNumEntries (formerly CalculateNumEntry) is only used by RocksDB for range-partitioned filters, called each time we start to construct one. (In theory, it should be possible to reuse the estimate, but the abstractions provided by FilterPolicy don't really make that workable.) But this is only used as a heuristic estimate for hitting a desired partitioned filter size because of alignment to data blocks, which have various numbers of unique keys or prefixes. The two factors lead us to prioritize reasonable speed over 100% accuracy. optimize_filters_for_memory adds extra complication, because precisely calculating num_entries for some allowed number of bytes depends on state with optimize_filters_for_memory enabled. And the allocator-agnostic implementation of optimize_filters_for_memory, using malloc_usable_size, means we would have to actually allocate memory, many times, just to precisely determine how many entries (keys) could be added and stay below some size budget, for the current state. (In a draft, I got this working, and then realized the balance of speed vs. accuracy was all wrong.) So related to that, I have made CalculateSpace, an internal-only API only used for testing, non-authoritative also if optimize_filters_for_memory is enabled. This simplifies some code. Pull Request resolved: https://github.com/facebook/rocksdb/pull/7774 Test Plan: unit test updated, and for FilterSize test, range of tested values is greatly expanded (still super fast) Also tested `db_bench -benchmarks=fillrandom,stats -bloom_bits=10 -num=1000000 -partition_index_and_filters -format_version=5 [-optimize_filters_for_memory] [-use_ribbon_filter]` with temporary debug output of generated filter sizes. Bloom+optimize_filters_for_memory: 1 Filter size: 197 (224 in memory) 134 Filter size: 3525 (3584 in memory) 107 Filter size: 4037 (4096 in memory) Total on disk: 904,506 Total in memory: 918,752 Ribbon+optimize_filters_for_memory: 1 Filter size: 3061 (3072 in memory) 110 Filter size: 3573 (3584 in memory) 58 Filter size: 4085 (4096 in memory) Total on disk: 633,021 (-30.0%) Total in memory: 634,880 (-30.9%) Bloom (no offm): 1 Filter size: 261 (320 in memory) 1 Filter size: 3333 (3584 in memory) 240 Filter size: 3717 (4096 in memory) Total on disk: 895,674 (-1% on disk vs. +offm; known tolerable overhead of offm) Total in memory: 986,944 (+7.4% vs. +offm) Ribbon (no offm): 1 Filter size: 2949 (3072 in memory) 1 Filter size: 3381 (3584 in memory) 167 Filter size: 3701 (4096 in memory) Total on disk: 624,397 (-30.3% vs. Bloom) Total in memory: 690,688 (-30.0% vs. Bloom) Note that optimize_filters_for_memory is even more effective for Ribbon filter than for cache-local Bloom, because it can close the unused memory gap even tighter than Bloom filter, because of 16 byte increments for Ribbon vs. 64 byte increments for Bloom. Reviewed By: jay-zhuang Differential Revision: D25592970 Pulled By: pdillinger fbshipit-source-id: 606fdaa025bb790d7e9c21601e8ea86e10541912		2020-12-18 14:31:03 -08:00
..
advisor	remediation of S205607	2020-07-17 17:20:49 -07:00
block_cache_analyzer	In ParseInternalKey(), include corrupt key info in Status (#7515 )	2020-10-28 10:12:58 -07:00
dump	Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433 )	2020-02-20 12:09:57 -08:00
rdb	Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433 )	2020-02-20 12:09:57 -08:00
analyze_txn_stress_test.sh	Add copyright headers per FB open-source checkup tool. (#5199 )	2019-04-18 10:55:01 -07:00
auto_sanity_test.sh	Add copyright headers per FB open-source checkup tool. (#5199 )	2019-04-18 10:55:01 -07:00
benchmark.sh	Fixed typo in benchmark.sh (#6434 )	2020-02-19 17:08:02 -08:00
benchmark_leveldb.sh	Add copyright headers per FB open-source checkup tool. (#5199 )	2019-04-18 10:55:01 -07:00
blob_dump.cc	Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433 )	2020-02-20 12:09:57 -08:00
check_all_python.py	Allow missing "unversioned" python, as in CentOS 8 (#6883 )	2020-05-29 11:29:23 -07:00
check_format_compatible.sh	add 6.15.fb to check_format_compatible.sh (#7738 )	2020-12-03 12:45:14 -08:00
CMakeLists.txt	Mark dependencies as PRIVATE and fix missing dependencies in tools. (#6790 )	2020-05-12 21:07:55 -07:00
db_bench.cc	Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433 )	2020-02-20 12:09:57 -08:00
db_bench_tool.cc	Support optimize_filters_for_memory for Ribbon filter (#7774 )	2020-12-18 14:31:03 -08:00
db_bench_tool_test.cc	Fix db_bench_tool_test. Fixes 7341 (#7344 )	2020-09-09 09:07:16 -07:00
db_crashtest.py	Inject the random write error to stress test (#7653 )	2020-12-17 11:52:28 -08:00
db_repl_stress.cc	More Makefile Cleanup (#7097 )	2020-07-09 14:35:17 -07:00
db_sanity_test.cc	Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433 )	2020-02-20 12:09:57 -08:00
dbench_monitor
Dockerfile
generate_random_db.sh	Add copyright headers per FB open-source checkup tool. (#5199 )	2019-04-18 10:55:01 -07:00
ingest_external_sst.sh	Add copyright headers per FB open-source checkup tool. (#5199 )	2019-04-18 10:55:01 -07:00
io_tracer_parser.cc	Add IO Tracer Parser (#7333 )	2020-09-23 15:50:26 -07:00
io_tracer_parser_test.cc	Add IO Tracer Parser (#7333 )	2020-09-23 15:50:26 -07:00
io_tracer_parser_tool.cc	Add IO Tracer Parser (#7333 )	2020-09-23 15:50:26 -07:00
io_tracer_parser_tool.h	Add IO Tracer Parser (#7333 )	2020-09-23 15:50:26 -07:00
ldb.cc	Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433 )	2020-02-20 12:09:57 -08:00
ldb_cmd.cc	Remove unused includes (#7604 )	2020-10-28 23:22:27 -07:00
ldb_cmd_impl.h	add `ldb unsafe_remove_sst_file` subcommand (#7335 )	2020-09-03 16:54:51 -07:00
ldb_cmd_test.cc	Store FileSystemPtr object that contains FileSystem ptr (#7180 )	2020-08-12 17:31:23 -07:00
ldb_test.py	Allow missing "unversioned" python, as in CentOS 8 (#6883 )	2020-05-29 11:29:23 -07:00
ldb_tool.cc	add `ldb unsafe_remove_sst_file` subcommand (#7335 )	2020-09-03 16:54:51 -07:00
pflag
reduce_levels_test.cc	Replace reinterpret_cast with static_cast_with_check (#7067 )	2020-07-02 19:25:41 -07:00
regression_test.sh	Add copyright headers per FB open-source checkup tool. (#5199 )	2019-04-18 10:55:01 -07:00
report_lite_binary_size.sh	Add copyright headers per FB open-source checkup tool. (#5199 )	2019-04-18 10:55:01 -07:00
rocksdb_dump_test.sh	Add copyright headers per FB open-source checkup tool. (#5199 )	2019-04-18 10:55:01 -07:00
run_flash_bench.sh	Add copyright headers per FB open-source checkup tool. (#5199 )	2019-04-18 10:55:01 -07:00
run_leveldb.sh	Add copyright headers per FB open-source checkup tool. (#5199 )	2019-04-18 10:55:01 -07:00
sample-dump.dmp
sst_dump.cc	Implement a new subcommand "identify" for sst_dump (#6943 )	2020-06-08 13:58:28 -07:00
sst_dump_test.cc	Fix compile error for old gcc-4.8 (#7358 )	2020-09-08 12:09:34 -07:00
sst_dump_tool.cc	In ParseInternalKey(), include corrupt key info in Status (#7515 )	2020-10-28 10:12:58 -07:00
trace_analyzer.cc	Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433 )	2020-02-20 12:09:57 -08:00
trace_analyzer_test.cc	Add trace_analyzer_test to ASSERT_STATUS_CHECKED list (#7480 )	2020-10-01 15:58:52 -07:00
trace_analyzer_tool.cc	Remove unused includes (#7604 )	2020-10-28 23:22:27 -07:00
trace_analyzer_tool.h	Add trace_analyzer_test to ASSERT_STATUS_CHECKED list (#7480 )	2020-10-01 15:58:52 -07:00
verify_random_db.sh	Add copyright headers per FB open-source checkup tool. (#5199 )	2019-04-18 10:55:01 -07:00
write_external_sst.sh	Add copyright headers per FB open-source checkup tool. (#5199 )	2019-04-18 10:55:01 -07:00
write_stress.cc	Replace namespace name "rocksdb" with ROCKSDB_NAMESPACE (#6433 )	2020-02-20 12:09:57 -08:00
write_stress_runner.py	Allow missing "unversioned" python, as in CentOS 8 (#6883 )	2020-05-29 11:29:23 -07:00