Repository navigation
Reproducible pyc: frozenset is not serialized in a deterministic order #81777
Description
Activity
See bpo-29708 meta issue and https://reproducible-builds.org/ for reproducible builds.
pyc files are not fully reproducible yet: frozenset items are not serialized in a deterministic order
One solution would be to modify marshal to sort frozenset items before serializing them. The issue is how to handle items which cannot be compared. Example:
>>> l=[float("nan"), b'bytes', 'unicode'] >>> l.sort() Traceback (most recent call last): File "<stdin>", line 1, in <module> TypeError: '<' not supported between instances of 'bytes' and 'float'
One workaround for types which cannot be compared is to use the type name in the key used to compare items:
>>> l.sort(key=lambda x: (type(x).__name__, x)) >>> l [b'bytes', nan, 'unicode']
Note: comparison between bytes and str raises a BytesWarning exception when using python3 -bb.
Second problem: how to handle exceptions when comparison raises an error anyway?
Another solution would be to use the PYTHONHASHSEED environment variable. For example, if SOURCE_DATE_EPOCH is set, PYTHONHASHSEED would be set to 0. This option is not my favorite because it disables a security fix against denial of service on dict and set:
https://python-security.readthedocs.io/vuln/hash-dos.html--
Previous discussions on reproducible frozenset:
- https://mail.python.org/pipermail/python-dev/2018-July/154604.html
- https://bugs.python.org/issue34093#msg321523
See also bpo-34093: "Reproducible pyc: FLAG_REF is not stable" and PEP-552 "Deterministic pycs".
- added3.9 (EOL)end of lifeend of lifeinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)
on Jul 15, 2019 bpo-34722 also talks about frozenset, nondeterministic order and sorting. Maybe this ticket and that one are for the same issue?
Normal sets have the same issue, see bpo-43850.
Would it be reasonable to make it so that sets are always created with the definition order? Looking at the set implementation, this seems perfectly possible.
Would it be reasonable to make it so that sets are
always created with the definition order?No, it would not. We would also have to maintain order across set operations such as intersection which which would become dramatically more expensive if they had to maintain order. For example intersecting a million element set with a ten element set always takes ten steps regardless of the order of arguments, but to maintain order of the left hand operand could take a hundred times more work.
No, it would not. We would also have to maintain order across set operations such as intersection which which would become dramatically more expensive if they had to maintain order. For example intersecting a million element set with a ten element set always takes ten steps regardless of the order of arguments, but to maintain order of the left hand operand could take a hundred times more work.
Can these operations happen during bytecode generation? I am fairly new to these internals so my understanding is not great. During bytecode generation is can code that performs such operations run?
s/hundred/hundred thousand/
s/is can/can/
Another idea, would it be possible to add a flag to turn on reproducibility, sacrificing performance? This flag could be set when generating bytecode, where the performance hit shouldn't be that relevant.
Another idea, would it be possible to add a flag to turn on reproducibility, sacrificing performance?
The flag is the SOURCE_DATE_EPOCH env var, no?
I would not expect SOURCE_DATE_EPOCH to sacrifice performance. During packaging, SOURCE_DATE_EPOCH is always set, and sometimes we need to perform expensive operations. We only need this behavior during cache generation, making the solution not optimal.
Backtracking a bit to your proposal for sorting the elements. Is it possible to have two different types with the same name? We need a unique identifier for each type.
After that, we need the type to allow sorting/comparing items, which AFAIK is not something we can guarantee.
We could certainly do the sorting where we are able to, and bail out if impossible, which I feel should handle the majority of cases. This is not optimal, but reasonable.Is there any way we could something like resetting the hash seed during cache generation?
Possible solution: add an ordered subtype of frozenset which would keep an array of items in the original order. The compiler only creates frozenset when optimizes "x in {1, 2}" or "for x in {1, 2}". It should now create an ordered frozenset from a list of constants (removing possible duplicates). The marshal module should save items in that order and restore ordered frozensets when load data. It should not increase memory consumption too much, because frozenset constants in code are rare and small.
What about normal sets? They also suffer from the same issue.
What about normal sets?
pyc files don't contain a regular set. So it is out of scope of this issue.
32 remaining items
- added 5 commits that reference this issue
on Apr 2, 2025 - added a commit that references this issue
on Apr 3, 2025 - added 4 commits that reference this issue
on Apr 3, 2025 - added 2 commits that reference this issue
on Jun 13, 2025
setandfrozensetmarshalling deterministic #27926set/frozensetmarshalling code #28068test_deterministic_setsto correctly handle different string hash algorithms #28147Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields: