Skip to content

[python] Avoid Bloom false negatives for floating-point BTree keys - #10364

Open
TheR1sing3un wants to merge 1 commit into
apache:masterfrom
TheR1sing3un:fix/python-nan-bloom-lookups
Open

TheR1sing3un wants to merge 1 commit into
apache:masterfrom
TheR1sing3un:fix/python-nan-bloom-lookups

Conversation

@TheR1sing3un

Copy link
Copy Markdown
Member

Purpose

Python's floating-point BTree comparator treats different NaN payloads as equal, and likewise treats positive and negative zero as equal. The optional Bloom filter hashes serialized bytes, so equality/IN queries using an equivalent value with different bytes can incorrectly return no rows.

Route NaN and zero point queries through the existing comparator-based range lookup. This preserves current comparison semantics and supports already-written indexes without changing key encodings or the file format. Other point queries retain the Bloom fast path.

Tests

  • Cover FLOAT and DOUBLE NaN payloads/signs and both signed-zero directions, with Bloom enabled/disabled, across multiple data blocks, including multiple rows sharing a key and equality/IN queries.
  • btree_bloom_filter_test.py: 12 passed, including the existing assertion that missing ordinary keys read only the Bloom filter. Seven new cases fail against the unchanged base; one signed-zero case happens to be a Bloom false positive.
  • Existing btree_thread_safety_test.py and global_index_build_test.py: 47 passed.
  • Flake8 for changed files and git diff --check pass.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant