Repository navigation
[Meta] Research: what can we test with Hypothesis? #107862
Description
Activity
- addedtype-featureA feature request or enhancementA feature request or enhancementtestsTests in the Lib/test dirTests in the Lib/test dir
on Aug 11, 2023 Right now I've ported and modified several
binasciitests from https://github.com/Zac-HD/stdlib-property-tests/blob/master/tests/test_encode_decode.py :)I'd suggest looking at PyPy's property-based tests, since they're known to be useful for developers of a Python implementation.
I'd also reconsider (3). While complicated strategies are obviously more work to develop and use, in my experience they're also disproportionately likely to find bugs - precisely because code working with complicated data structures are even more difficult to test without Hypothesis, and so there are usually weird edge cases.
Examples of this effect include #84838, #86384, #89901, and #83134. The
hypothesmithsource-code strategies are at a pre-alpha / proof-of-concept stage and have more third-party dependencies so I'm not sure they'd be a good fit for CPython at this time, but you get the idea. I don't have capacity to implement it myself, but could supervise a volunteer or contractor to build a production-ready strategy if there's interest.Reacted by sobolevnPersonally, I don't think we should go down this path. Hypothesis has a certain amount of randomness to it. In general, our tests should be specifically designed to cover every path of interest.
When someone is designing new code, it is perfectly reasonable to use Hypothesis. However, for existing code, we "let's use Hypothesis" isn't a goal directed as a specific problem. One issue that we've had with people pursuing a "let's use Hypothesis" goal is that they are rarely working on modules that they fully understand, that the invariants they mentally invent aren't the actual invariants and reflect bugs in their own understanding rather than actual bugs in the code. For example, we got reports on
colorsysconversions not being exactly invertible; however, due to color gamut limitations they can't always be inverted (information is lost) and the conversion tables were dictated by published standards that were invertible. We also got false report on the random module methods by people who just let Hypothesis plug in extreme values without any thought of what the method was actually supposed to do in real use cases. There may have been one actual (but very minor) bug in the standard library found by Hypothesis, but everything else was just noise.It would be perfectly reasonable to use Hypothesis outside of our test suite and then report an actual bug if found. Otherwise, I think actually checking in the h-tests would just garbage-up our test suite, make it run less deterministically, and no explicitly write-out all the cases being covered.
Reacted by Blaise Pabonbrandonardenwalli commented
on Aug 15, 2023 on Aug 15, 2023 · Hidden as off-topicshow commentMore actionssobolevn commented
on Aug 15, 2023 on Aug 15, 2023 · Hidden as off-topicAuthorshow commentMore actionsbrandonardenwalli commented
on Aug 16, 2023 on Aug 16, 2023 · Hidden as off-topicshow commentMore actionsIt would be perfectly reasonable to use Hypothesis outside of our test suite and then report an actual bug if found. Otherwise, I think actually checking in the h-tests would just garbage-up our test suite, make it run less deterministically, and no explicitly write-out all the cases being covered.
But, we already have Hypothesis as a part of your test suite and CI. It already has some tests to it. Link:
cpython/.github/workflows/build.yml
Lines 361 to 467 in 5d936b6
test_hypothesis: name: "Hypothesis tests on Ubuntu" runs-on: ubuntu-20.04 timeout-minutes: 60 needs: check_source if: needs.check_source.outputs.run_tests == 'true' && needs.check_source.outputs.run_hypothesis == 'true' env: OPENSSL_VER: 1.1.1v PYTHONSTRICTEXTENSIONBUILD: 1 steps: - uses: actions/checkout@v3 - name: Register gcc problem matcher run: echo "::add-matcher::.github/problem-matchers/gcc.json" - name: Install Dependencies run: sudo ./.github/workflows/posix-deps-apt.sh - name: Configure OpenSSL env vars run: | echo "MULTISSL_DIR=${GITHUB_WORKSPACE}/multissl" >> $GITHUB_ENV echo "OPENSSL_DIR=${GITHUB_WORKSPACE}/multissl/openssl/${OPENSSL_VER}" >> $GITHUB_ENV echo "LD_LIBRARY_PATH=${GITHUB_WORKSPACE}/multissl/openssl/${OPENSSL_VER}/lib" >> $GITHUB_ENV - name: 'Restore OpenSSL build' id: cache-openssl uses: actions/cache@v3 with: path: ./multissl/openssl/${{ env.OPENSSL_VER }} key: ${{ runner.os }}-multissl-openssl-${{ env.OPENSSL_VER }} - name: Install OpenSSL if: steps.cache-openssl.outputs.cache-hit != 'true' run: python3 Tools/ssl/multissltests.py --steps=library --base-directory $MULTISSL_DIR --openssl $OPENSSL_VER --system Linux - name: Add ccache to PATH run: | echo "PATH=/usr/lib/ccache:$PATH" >> $GITHUB_ENV - name: Configure ccache action uses: hendrikmuhs/ccache-action@v1.2 - name: Setup directory envs for out-of-tree builds run: | echo "CPYTHON_RO_SRCDIR=$(realpath -m ${GITHUB_WORKSPACE}/../cpython-ro-srcdir)" >> $GITHUB_ENV echo "CPYTHON_BUILDDIR=$(realpath -m ${GITHUB_WORKSPACE}/../cpython-builddir)" >> $GITHUB_ENV - name: Create directories for read-only out-of-tree builds run: mkdir -p $CPYTHON_RO_SRCDIR $CPYTHON_BUILDDIR - name: Bind mount sources read-only run: sudo mount --bind -o ro $GITHUB_WORKSPACE $CPYTHON_RO_SRCDIR - name: Restore config.cache uses: actions/cache@v3 with: path: ${{ env.CPYTHON_BUILDDIR }}/config.cache key: ${{ github.job }}-${{ runner.os }}-${{ needs.check_source.outputs.config_hash }} - name: Configure CPython out-of-tree working-directory: ${{ env.CPYTHON_BUILDDIR }} run: | ../cpython-ro-srcdir/configure \ --config-cache \ --with-pydebug \ --with-openssl=$OPENSSL_DIR - name: Build CPython out-of-tree working-directory: ${{ env.CPYTHON_BUILDDIR }} run: make -j4 - name: Display build info working-directory: ${{ env.CPYTHON_BUILDDIR }} run: make pythoninfo - name: Remount sources writable for tests # some tests write to srcdir, lack of pyc files slows down testing run: sudo mount $CPYTHON_RO_SRCDIR -oremount,rw - name: Setup directory envs for out-of-tree builds run: | echo "CPYTHON_BUILDDIR=$(realpath -m ${GITHUB_WORKSPACE}/../cpython-builddir)" >> $GITHUB_ENV - name: "Create hypothesis venv" working-directory: ${{ env.CPYTHON_BUILDDIR }} run: | VENV_LOC=$(realpath -m .)/hypovenv VENV_PYTHON=$VENV_LOC/bin/python echo "HYPOVENV=${VENV_LOC}" >> $GITHUB_ENV echo "VENV_PYTHON=${VENV_PYTHON}" >> $GITHUB_ENV ./python -m venv $VENV_LOC && $VENV_PYTHON -m pip install -U hypothesis - name: 'Restore Hypothesis database' id: cache-hypothesis-database uses: actions/cache@v3 with: path: ./hypothesis key: hypothesis-database-${{ github.head_ref || github.run_id }} restore-keys: | - hypothesis-database- - name: "Run tests" working-directory: ${{ env.CPYTHON_BUILDDIR }} run: | # Most of the excluded tests are slow test suites with no property tests # # (GH-104097) test_sysconfig is skipped because it has tests that are # failing when executed from inside a virtual environment. ${{ env.VENV_PYTHON }} -m test \ -W \ -o \ -j4 \ -x test_asyncio \ -x test_multiprocessing_fork \ -x test_multiprocessing_forkserver \ -x test_multiprocessing_spawn \ -x test_concurrent_futures \ -x test_socket \ -x test_subprocess \ -x test_signal \ -x test_sysconfig - uses: actions/upload-artifact@v3 if: always() with: name: hypothesis-example-db path: .hypothesis/examples/
I propose adding more cases, where it makes sense.In general, our tests should be specifically designed to cover every path of interest.
I agree that our regular tests should cover all paths, but there are more to it. Path coverage is only as good as 100%. But, we are obviously limited by the number of data we can provide. We cannot come up with lots of data, our current test suites proove my point.
But, Hypothesis can. This is exactly what it is good at: proving lots of correctly structured data.
Hypothesis has a certain amount of randomness to it
There are different way on how we control the randomness. First, we use a database with examples:
Plus, Hypothesis itself controls how data is generated to be repeatable.cpython/Lib/test/support/hypothesis_helper.py
Lines 22 to 35 in 5d936b6
from hypothesis.database import ( GitHubArtifactDatabase, MultiplexedDatabase, ReadOnlyDatabase, ) hypothesis.settings.register_profile( "cpython-local-dev", database=MultiplexedDatabase( hypothesis.settings.default.database, ReadOnlyDatabase(GitHubArtifactDatabase("python", "cpython")), ), ) hypothesis.settings.load_profile("cpython-local-dev") One issue that we've had with people pursuing a "let's use Hypothesis" goal is that they are rarely working on modules that they fully understand, that the invariants they mentally invent aren't the actual invariants and reflect bugs in their own understanding rather than actual bugs in the code
This is a valid concern, I hope that collaboration among developers can solve this.
And this is exactly why I created this issue: to figure out what is worth testing and what not.There may have been one actual (but very minor) bug in the standard library found by Hypothesis, but everything else was just noise.
I don't think that this is actually correct. First, we cannot know what bugs were found by Hypothesis, because people might just post the simplest repro without naming the tool.
Second, we have these bugs that do mention Hypothesis:
Refs #86275
btw, I'm new to Hypothesis and the
test_zoneinfofrom @pganssle has lots of different use cases, the location is
https://github.com/pganssle/zoneinfo/blob/master/tests/test_zoneinfo_property.py
Edit: Thanks for all your kind warnings, I didn't actually try to use it. I just stared into the abyss and it stared back at me. It's fascinating in an academic sense. The way that some people read about https://en.wikipedia.org/wiki/Australian_funnel-web_spider
I do think we should be expanding hypothesis tests. There was no stipulation in the original acceptance that they would only be used for zoneinfo, the idea was to use property tests more widely.
We should at least start by migrating whichever tests in Zac's repo still make sense (probably most of them).
Since they don't run on all CI, they should always come with
@examples, in which case they are essentially parametrized tests (the stubs run the examples).Reacted by Blaise Pabon and Zhe ZhangHypothesis has a certain amount of randomness to it.
That's a good thing!
In general, our tests should be specifically designed to cover every path of interest.
I don't see a big problem here. You can always use "example" decorator to include all relevant cases.
However, for existing code, we "let's use Hypothesis" isn't a goal directed as a specific problem.
The goal might be a better testing of the existing codebase. I could bet, that issue like #111342 (just a typo) would be quickly discovered by hypothesis-based tests. Shouldn't sumprod() be commutative?
I think that hypothesis is excellent in testing invariants, like here:
cpython/Lib/test/test_float.py
Lines 1543 to 1562 in 046a4e3
def test_roundtrip(self): def roundtrip(x): return fromHex(toHex(x)) for x in [NAN, INF, self.MAX, self.MIN, self.MIN-self.TINY, self.TINY, 0.0]: self.identical(x, roundtrip(x)) self.identical(-x, roundtrip(-x)) # fromHex(toHex(x)) should exactly recover x, for any non-NaN float x. import random for i in range(10000): e = random.randrange(-1200, 1200) m = random.random() s = random.choice([1.0, -1.0]) try: x = s*ldexp(m, e) except OverflowError: pass else: self.identical(x, fromHex(toHex(x)))
This example, probably, will gain not too much (except readability) with using hypothesis. But sometimes we use just a fixed data set to test such cases:
cpython/Lib/test/test_capi/test_float.py
Lines 156 to 180 in 046a4e3
def test_pack_unpack_roundtrip(self): pack = _testcapi.float_pack unpack = _testcapi.float_unpack large = 2.0 ** 100 values = [1.0, 1.5, large, 1.0/7, math.pi] if HAVE_IEEE_754: values.extend((INF, NAN)) for value in values: for size in (2, 4, 8,): if size == 2 and value == large: # too large for 16-bit float continue rel_tol = EPSILON[size] for endian in (BIG_ENDIAN, LITTLE_ENDIAN): with self.subTest(value=value, size=size, endian=endian): data = pack(size, value, endian) value2 = unpack(data, endian) if math.isnan(value): self.assertTrue(math.isnan(value2), (value, value2)) elif size < 8: self.assertTrue(math.isclose(value2, value, rel_tol=rel_tol), (value, value2)) else: self.assertEqual(value2, value) invariants they mentally invent aren't the actual invariants and reflect bugs in their own understanding
Why it's the tool problem, not people who written bad tests? :-)
We now have both the code and CI job to run
hypothesistests.In this issue I invite everyone to propose ideas: what can we test with it?
Criteria:
hypothesishypothesiswill produce a lot of cases to tests, we cannot do any slow / network related things)What else?
Known applications
Property-based tests are great when there are certain patterns:
decodeandencodedumpsandloadsThere are also some hidden patterns as (example):
'a' in sthens.index('a')must return an integerExisting work
Good examples of exising stuff:
hypothesistests intest_zoneinfoby @pganssleLinked PRs
hypothesistests totest_binascii#107863