Repository navigation
Decide how to format Misc/NEWS entries #66
Description
Activity
I agree with @larryhastings that towncrier's directory for categories is nice, and so we should try and go with it. Larry is also right, though, that we will need some dummy file in the directory so git will track them when we start a new
next/, but I think a simpleREADME.rstthat explains what the directory for will suffice (and it has the next side-effect of GitHub rendering theREADME.rstso it's easy to read through the web UI).As for the metadata for What's New, I would like @1st1 and anyone else who has had to fill in that doc to say whether having the "why" something should be added to What's New is useful, or just simply knowing what needs inclusion? (Category versus flag, IOW, since a flag is much easier to manage than a category.) Another approach is to have a "What's New" and "boring" label for PRs and have a status check that requires one of those two labels be present before one can merge a PR. (Once again, I would want @1st1 or someone to say if seeing a list of PRs or grepping through files would be easier.)
For the file name, I'm fine with a chronological sort, but I don't think a nonce is truly necessary. The most common case is going to be data + issue number, so I think the format is
YYYY-MM-DD.bpo-NNNN.rstwill keep the vast majority of clashes from occurring and give a good enough chronological sort at the day level (I'm not about to ask people to add the time since if we expect people to do this by hand on occasion then they will inevitably screw up the UTC conversion, especially after a daylight saving switch). We can officially make the format have an optional nonce for actual clashes, but that will probably be extremely rare. So maybe the format should ber"(?P<date>\d{4}-\d{2}-\d{2})\.bpo-(?P<issues>\d+(,\d+)*)(?:\.\w+)?\.rst"?I do like the idea from @ncoghlan of using reST comments as a way to store metadata. Question becomes what metadata do we want? 😉 If we go with directories for categories ala towncrier and we store the issue number(s) in the file name along with the date, then in the common case nothing is really necessary. We could store who to thank, but I'm assuming that for
Misc/ACKSwe are either going to continue to have people enter themselves and others manually and maybe supplement it with what Rust and Rails and extract the info from the git metadata (see #7 to discuss that topic). So that only leaves What's New and we need to first decide what we want for that and whether we want it with the news entry or if we want it on the PR.But since what metadata we want to store is technically optional once we decide on a format for it as long as it's not always required, all we really have to decide upfront is the directory, file name, and file format; we can add metadata later. Do people think we can figure these formatting details in a week so we can have a decision made on Friday, April 7? Then if we are still discussing possible metadata we can just add them in later.
re: Nikola format
It's really the same issue as with field-lists.- STORY: Users can specify an issue record start (like a context manager) and or end in order to concatenate documents for sphinx without unnecessary transformation.
RST :field-lists:
:versions: 2.7, 3.6.1 :tags: bugfix, security :pr: 123, :issue: 22 :cve: 2011-1015
:propertyname: value- RST "field lists" https://docutils.readthedocs.io/en/sphinx-docs/ref/rst/restructuredtext.html#field-lists
- what happens when concatenating RST documents with redundant field list properties?
- is there a collision on
:propertyname:?- how to define a subject (subject, propertyname, value)?
- from the filename
- from the RST heading
- this would work for merged files if there are headings for each feature
- [...]
- how to define a subject (subject, propertyname, value)?
- is there a collision on
- Misc/NEWS data is published with Sphinx:
- Nikola metadata is specified per-post with docutils directives:
Just an FYI, I just blocked Wes from the Python organization since he did not heed my warning. I would delete the messages he left but I'm using them as evidence as to why I put in the organization-level block.
Reacted by Zachary Ware, Mariatta, Gregory P. Smith, Brian Curtin, Berker Peksag, Alyssa Coghlan and Giampaolo RodolaWould PR contributors be expected to add the NEWS files in their PRs? As I understand it, one of the goals of this change is to minimize the amount of manual work core devs have to do to merge a PR.
But a PR probably won't be merged on the same day it's opened, so the date in the filename will be wrong if we use the date the PR was opened. I can't really think of a better solution than having the core dev edit the filename while merging, but maybe we can do better.
The exact date doesn't matter too much, as the main requirement is to have a robust sort key so regenerating NEWS isn't dependent on things like the directory listing order for the snippets directory.
Probably the simplest way to handle that is to say that formally the date field is just "the date the NEWS snippet was written", and we'd advise folks to leave writing it until relatively late in the PR process for complex patches.
Reacted by Jelle Zijlstra and Brett CannonYeah, the date field is there just to provide a unique filename and a reasonably accurate chronological ordering. Misc/NEWS has historically been roughly (but not perfectly) chronologically ordered, and I was trying to preserve that. But it doesn't have to be exact.
Reacted by Alyssa CoghlanSo what's left to decide? Are we happy with using directories and the proposed file name format? Are we assuming if we decide to add metadata we can do that later? Am I missing anything?
Checking I understand what the current plans actually are:
- Pending news entries will go into
Misc/News/next/{category}and then automatically get moved toMisc/News/{version}/{category}once they're included in a release - The snippet name format will be
YYYY-MM-DD.bpo-NNNN[.whatever].rst(where the square brackets indicate the human readable slug is optional) - If we later decide we need some additional metadata, it will use the
.. Name: Valueformat and be separated from the snippet text by a blank line
If I've understood correctly, then that's a definite +1 from me.
- Pending news entries will go into
Placing the version information in the file path will make it harder to deal with backports of patches, since any patch with a blurb input file in it will have to be modified to move the file into the right version directory for the branch.
If you cherry-pick the original checkin, you'll get the original path in the
nextdirectory, which has no version information.@ncoghlan Yep, you understood correctly. The only possible change is supporting multiple issue numbers in the file name, but I'm really not bothered by dropping that if we are going to still state that manually in the message itself.
One thing I just thought of: I guess we don't really need to care about line wrapping if we use textwrap or something in the final output? I bring this up because if we leave out the bullet point --
--- then people will be off by 2 characters potentially if we want to explicitly wrap at 72 or 80 characters. So does it make sense to say "wrap at 80 characters" knowing full well we will re-wrap on output to take the bullet point and indent into account? What do you think @larryhastings ?@brettcannon Is bpo # mandatory? As a What's New editor I'd love to have an unambiguous and reliable connection between the NEWS entry and the actual code changes.
I don't think we should impose the addition of What's New-specific metadata on the committers at this point.
@elprans Yes, the idea is to have the issue number be mandatory since anything that's worth having in What's New should have a relevant issue to track the discussion (and then indirectly track the PR/patch that made the change through the issue). Unless there was something else you had in mind?
23 remaining items
While writing out my language summit slides I realized that we should just require a nonce in the filename, even when written by hand to prevent any chance of collision. People can bash on their keyboard for all I care, just as long as there's something to make sure collisions just won't happen.
I agree, and I was going to write that in my directions for writing blurb files by hand. There'll be a slot for it in the filename and it should be non-empty.
FWIW I now generate the nonce as follows: compute the MD5 of the body text, convert to urlsafe base64, and use just the first six characters.
Oh, and, one persnickety pedantic reply I meant to make @brettcannon:
We're starting off with no metadata required;
Well, we're starting off with no metadata specified inside the file. There are four pieces of metadata: section, bpo, date, and nonce. But we're storing all those in the filename.
The modern version of "tidy" will actually be a single blurb file containing multiple entries, each delimited by
..lines, with this metadata stored as actual metadata (e.g... section: Library).@larryhastings Glad we agree on the nonce bit. 😄 And the calculated value sounds 👍
And yes, you're right about no metadata in the file; allergies have hit me hard and so my mental capacity isn't 💯 at the moment.
I've updated blurb on github ( https://github.com/larryhastings/blurb ). It's updated to this modern data format. Docs aren't updated yet, sorry.
Other changes:
- Instead of a zillion files, it maintains one big file per version. All metadata is preserved.
- When you
blurb releasea new version, it merges all thenextfiles into a single file for that version. - blurb now aggressively reflows text. If you
blurb splitthenblurb merge, you'll have a lot of text-reflowing diffs inMisc/NEWS.
One question: do we prefer "Issue #" or "bpo-"? Both are in use right now. blurb prints "Issue #" because it's old school, but I don't have a strong opinion.
@larryhastings we're trying to move to "bpo-" to namespace issue numbers.
Hello, what's the progress on that issue? (Sorry, too lazy to read this super long issue.) Maybe send a report to python-dev?
@Haypo or just attend the language summit tomorrow. 😉 (And if you can't make it, then just trust that @larryhastings and I are working on rolling this out 😄).
In Misc/NEWS a bunch of entries look like this::
- [Security] Issue #27278: Fix os.urandom() implementation using getrandom() on Linux. Truncate size to INT_MAX and loop until we collected enough random bytes, instead of casting a directly Py_ssize_t to int.Apart from screwing up my parser, these entries seem to think that tagging an entry with
[Security]is meaningful somehow. I'm not sure it is. Maybe?I'm going to have to detect the
[Security]part in order to parse properly. I can just throw it away, or I could do any of the following:- Add a metadata field that says ".. security: True"
- Add
*(Security)*to the end of the entry inMisc/NEWSwhen I generate it (blurb merge).
My inclination is to add the metadata field now, and we can decide on formatting stuff later.
@larryhastings adding at the end is fine by me. I think it's just a marker for Linux vendors and the like to especially pay attention to a specific news entry.
Aye, flagging things specifically as security issues serves as a "You should backport this even if none of your customers explicitly request it" marker for redistributors that standardise on a particular maintenance release as their baseline version. It also helps in keeping the list at http://python-security.readthedocs.io/ up to date.
However, the exact format doesn't matter, so long as its noted somewhere, and can be added after the fact if a bug fix is later determined to have security implications.
- For security entries, I suggest to move them to a new Security section.Reacted by Alyssa Coghlan and Brett Cannon
Can we close this? python/devguide PR #212 is blocking on it.
Yep, we can close it.
Now that we have chosen what NEWS tools to go with, I'm starting a fresh discussion on how exactly we want to format the entries.
@larryhastings said:
I haven't found the time to sit down and write this properly, so here's a quick note on this topic before Brett makes up his mind. sorry if it's a bit long / messy.
First, I don't have a strong opinion about what the input format to blurb should look like. If there was a consensus about "it should look like X", then I'd make it look like X.
We don't seem to have a consensus about what the input format should look like, because I don't think we've reached consensus about what metadata the tool needs. We need to figure that out first.
Obviously necessary:
I would also like:
3. some datestamp/nonce that ensures the news entries remain in some sort of stable order (I prefer chronological ordering, sadly git doesn't maintain timestamps)
I believe Brett is also asking for:
4. an optional "please consider for the next What's New document" flag
5. a suggested category for "What's New"
By the way, towncrier's approach of pre-created directories named for the categories (2.) is a nice idea. That ensures people don't misspell the category name. blurb could easily switch to that. the only downside I know of is that iiuc git doesn't track directories as first-class objects, so we'd have to have an extra empty file in each directory.
blurb current supports 1-3. it uses the filename for two bits of metadata (stable sorting order and category), and the contents of the file are simply the news entry. but it's kind of reached my comfort level regarding storing metadata in the filename.
if we only want to add 4, a simple "consider for what's new please", then okay I think we could live with sticking that in the filename. Like we add .wn just before the extension, for example.
if we want to add 5, then my inclination is to add a simple metadata blob to the contents of the file:
#is a line commentIf we do that, then my inclination is further to move all the metadata into that blob:
category=Librarynonce=20170513062235.ef4c88a1# what's new = Improved Modules(uncomment "what's new" to use it)
The "blurb" tool would make it easy to add these, but users could also create the file by hand using a web page form that formats the output for them. (Making the entry entirely by hand might be tricky, since the nonce should be in a standardized format. Maybe we could give them a short blob of Python they run to generate one?)
If all the metadata lives inside the file, then we don't care what the filename is, it just needs to be unique.
@ncoghlan said:
For metadata-in-the-file, I quite like the format that the Nikola static blog generator uses:
As an added bonus, when the file has the
.rstextension, my editor automatically grays out the metadata as line comments, and assuming we're planning to use ReST in the snippets for ease of Sphinx integration, that would also apply here.Using structured metadata like that would also open up future options for acknowledgements that aren't directly tracked in the git metadata - cases where we built on a patch written by someone else, or someone contributed API design ideas that someone else implemented, etc. At the moment we put that in the snippet body ("Initial patch by ..." and so forth), but a metadata field could more easily feed into ideas like auto-generating Misc/ACKS in addition to Misc/NEWS.
As far as a stable sort algorithm for display goes, we could then define that as:
datetime.utcnow().strftime("%Y-%m-%d %H:%M:%S UTC"))