diff --git a/AGENTS.md b/AGENTS.md index 08c79f45..49560d0a 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -271,7 +271,7 @@ The library has two layers: `nameparser/config/` (data) and `nameparser/parser.p **Sweep the forms a change can reach: comma shapes, then `name_order`.** No comma, a FULL name before the comma, and a ONE-WORD name before the comma are three paths, and the third is the miss — #429, #430 and #432 are all that path disagreeing with the full-name path on inputs the full-name path reads correctly. For orders the sweep already exists (`tests/v2/test_cases.py` runs every row under all three) but carries one assertion, R2's family partition, so it checks nothing a new change moves; coverage there has been vacuous before (PR #394's review found the suite passed with `name_order` discarded from grouping). Parse your change's names in each comma shape and each order, and read the ones you did not predict. -**Then sweep the spellings: a class reached by vocabulary is also reached by SHAPE.** A test that uses only the listed spelling of a word class walks only the vocabulary path, and the shape path is a separate branch that can be missing while every test passes. Roman numerals are the sharpest case: `ii`, `iii` and `iv` are suffix vocabulary; `i` and `v` are vocabulary too but initial-shaped, so `is_suffix_piece` vetoes them and only the numeral fork reads them; and `vi`, `vii`, `ix` and every longer one are in no list at all, read by the fork's shape test alone (checked 2026-10-04 against `Lexicon.default()`). So a change that handles `V` and `III` can still miss `VI` — #610 found exactly that, the maiden take's candidate loop admitting a shape-only numeral in one position where the peel reads it in two (`Jane Doe nee Smith VI Prof.` kept maiden `Smith VI` while `V` and `III` worked). The same split runs through the credentials (`MA` listed, `X.Y.Z.` by the dotted shape, `XYZ` by the caps shape, and `Ph. D.` merged by group into one flagged piece the peel's walk never holds) and the titles (`Prof.` period-marked and read at the end of a name, `Prof` bare and a name word there). When a change touches a word class, test one spelling from each path. +**Then sweep the spellings: a class reached by vocabulary is also reached by SHAPE.** A test that uses only the listed spelling of a word class walks only the vocabulary path, and the shape path is a separate branch that can be missing while every test passes. Roman numerals are the sharpest case: `ii`, `iii` and `iv` are suffix vocabulary; `i` and `v` are vocabulary too but initial-shaped, so `is_suffix_piece` vetoes them and only the numeral fork reads them; and `vi`, `vii`, `ix` and every longer one are in no list at all, read by the fork's shape test alone (checked 2026-10-04 against `Lexicon.default()`). So a change that handles `V` and `III` can still miss `VI` — #610 found exactly that, the maiden take's candidate loop admitting a shape-only numeral in one position where the peel reads it in two (`Jane Doe nee Smith VI Prof.` kept maiden `Smith VI` while `V` and `III` worked). The same split runs through the credentials (`MA` listed, `X.Y.Z.` by the dotted shape, `XYZ` by the caps shape, and `Ph. D.` merged by group into one flagged piece the peel's walk never holds -- #603 found #602's run-start predicate refusing it too, so `Smith, John Ph. D. Jones` kept middle 'Jones' where `Smith, John PhD Jones` read suffix 'PhD Jones') and the titles (`Prof.` period-marked and read at the end of a name, `Prof` bare and a name word there). When a change touches a word class, test one spelling from each path. ### Configuration layer (`nameparser/config/`) @@ -320,7 +320,7 @@ The 2.0 rewrite lands as underscore-private modules alongside the v1 code. These - **Method organization**, fixed section order in every class: fields + `__post_init__` validation → alternative constructors → dunders (construction/equality → protocol → operators) → properties → public methods by concern (access → editing → comparison → rendering delegates) → private helpers last, except a helper serving exactly one section may sit at that section's head. Sanctioned deviation, facade layer only: `HumanName` and the shim `Constants` organize by v1 concern groups (`# -- render defaults --`, `# -- config / parsing --`, `# -- fields --`, ..., dunders and pickle last) — the classes mirror v1's own surface and die in 3.0; the canonical order still binds every core type. - **Validation is eager and fail-loud**: every `raise` states the offending value, the expected form, and the fix. Exception taxonomy: wrong type — including wrong element type inside a collection, bare `str` where an iterable of strings is expected, or a `Mapping` where a plain iterable is expected — raises `TypeError`; well-typed but unacceptable values raise `ValueError`; failed enum lookups stay `ValueError` for any input (stdlib `EnumType` precedent). **When the message hands the reader code to paste, that code has to survive a type checker** — nameparser ships `py.typed`. #337's segmenterless warning offered `Policy(segment_scripts=())`, an `arg-type` error, because these fields are annotated with what they STORE rather than everything the constructor accepts. Prefer the `frozenset()` / `()` spellings in messages and docstrings, and pin the offered spelling in a test — the warning tests matched on `ja_segmenter` and never checked the actionable half of the message. **A warning emitted in `Parser.__post_init__` needs `parser_for` to re-emit it from its own frame** (the `catch_warnings(record=True)` block at its return): `__post_init__`'s `stacklevel` is sized for direct `Parser(...)` construction, and through `parser_for`'s extra frame the default one-line rendering attributes the warning to the library's own `return Parser(...)` — the exact call the message tells the user to change becomes invisible. No single stacklevel serves both entry points; a new construction warning gets the re-emission for free, but a new CONSTRUCTION SITE for `Parser` inside this package needs its own re-emission or its callers get library-attributed warnings (#337 review). - **Guard, hint, and emit for the WHOLE family, and parametrize the test over it**: a check added to one member of a set belongs on all of it, and the test must sweep the family, not one example. This session shipped `_reject_str_and_mapping` on `Policy` but not `PolicyPatch`, the bytes decode hint on three of five config entry points, and a regex-sync roster missing four of its copies — each a separate follow-up bug that a `{class} × {field} × {bad-value}` parametrization would have caught and a per-example test hid. When you find you're guarding member N, grep for the other members first. -- **Ambiguities are emitted at the DECISION site**: an `Ambiguity` records a fork the parse had to call, not a token that sits in an ambiguous vocabulary. Emit where the branch is taken — the trailing-suffix peel in `_assign`, the delimiter escape's follow-up in `classify` — never by scanning for a `vocab:*-ambiguous` tag. The same tagged token is a genuine fork in one position and unremarkable in another (`do` mid-name in "Joao da Silva do Amaral de Souza" chooses nothing). **A branch that runs but changes nothing is not a decision either** -- the prefix chain's `merge(k, j)` executes even when `j == k + 1`, folding a piece into itself, and keying on "the code got here" reported a fork for all ambiguous particles on "Do Van Jr." (`Dr.` when that was written, before #367 made a plain title transparent and put the shape out of the loop's reach entirely), where the particle stayed a lone leading name piece — the GIVEN name under the default order, the family name under `FAMILY_FIRST` — and `_assign` reported the same token again. Check that the branch actually claimed something (`j > k + 1`) before recording. Structure often settles the question before it arises, which is why `PARTICLE_OR_GIVEN` is not emitted on the `FAMILY_COMMA` path's WHOLLY-FAMILY read -- the comma fixed which piece is the family -- and `SUFFIX_OR_NAME` is not emitted for "Ma, Jack". A tail segment is the same case from the other side: assign reads it wholly as suffixes, a maiden marker there included since #601 (`Jane Doe, PhD, Jr nee van Ma` reads suffix `PhD, Jr nee van Ma`, the marker an ordinary word, rules.md#M2), so group's two chain emitters are handed no report list there, after either comma — `John Smith, Jr., Freiherr von Richthofen` still chains `von` and reports only `comma-structure` (rules.md#C2, 2026-09-28). Read that scope narrowly: the comma settles nothing about a particle trailing the given name, so P6's attachment in `post_rules` decides that fork on the same path and reports it (#405), in the kind naming the reading it OVERRODE, which is the reading assign made and not the word's vocabulary: `SUFFIX_OR_NAME` where assign had read the run as a post-nominal (`vd`, `mc`), else `PARTICLE_OR_GIVEN` where the run holds an ambiguous particle (`van`, and `do`, which is in the suffix vocabulary too but in its AMBIGUOUS half, so no credential reading was overridden), else silence. The decision site also has the token index and the detail text in hand, which the tag scan would have to reconstruct. **If a fork's two branches are taken in DIFFERENT stages, every one of them needs the emitter** -- `PARTICLE_OR_GIVEN` is decided in `_assign` when the ambiguous particle stays a lone leading piece, in `_group` when something shifts it off the name's leading piece and the prefix chain claims it, and in `post_rules` when P6's attachment takes a trailing particle into the family after a comma, so all three report; for two years only the first did. What can still do the shifting is narrow, and #367 is why: a plain title no longer can (`Dr. Van Johnson` reads as `Van Johnson` does and reports from `_assign`), so the `_group` emitter needs a word that is BOTH a title and a particle — measured, `TITLES ∩ particles_ambiguous` is `{freiherr, st}` in the default vocabulary (`do` left TITLES in #296's audit; decisions.md's Excluded block records the before and after), plus any overlap a caller's config creates — standing ahead of the chained particle as the LEADING NAME word. Titles may precede it, so `Dr. St van Johnson` reaches the emitter and `St van Johnson` does too; a given name may not, so `Jan Freiherr von Richthofen` does not reach it while `Freiherr von Richthofen` and `Dr. Freiherr von Richthofen` do. Two shapes that look like they should reach it and do NOT, both measured by stepping `STAGES` and watching where `ambiguities` grows: `Dr. Do van Johnson` and `Do St Johnson` report from `assign`, not `group`, because `do` is no longer a title and so stays the leading name piece assign reports on — a both-vocabulary word CHAINED (`Jan St Johnson`) reports nothing at all. When checking whether that emitter is dead, a both-vocabulary word in the leading name position is the thing to look for, and the answer is that it is not dead. The stage-ownership map in `tests/v2/pipeline/test_state.py` must list `ambiguities` for each such stage, and it passes vacuously until a case row exercises the path, so add the row too. Report BOTH directions of a two-way fork — "John Smith MA" (read as a suffix) and "Jack MA" (read as the family name) are equally guesses. Every kind needs a trigger in `tests/v2/test_contracts.py::_AMBIGUITY_TRIGGERS` (an explicit `None`, strict-xfail, while reserved), and case-table rows pin expected kinds exactly, so a new emitter shows up in both immediately. **Pin the decision, not the vocabulary**: the only titled-particle test used an UNAMBIGUOUS particle, so it walked the right code path and proved nothing about the branch under test -- two criticals passed 1539 tests. A row contrasting the two readings ("John Smith V" against "John Smith B") is what makes an emitter's absence meaningful. +- **Ambiguities are emitted at the DECISION site**: an `Ambiguity` records a fork the parse had to call, not a token that sits in an ambiguous vocabulary. Emit where the branch is taken — the trailing-suffix peel in `_assign`, the delimiter escape's follow-up in `classify` — never by scanning for a `vocab:*-ambiguous` tag. The same tagged token is a genuine fork in one position and unremarkable in another (`do` mid-name in "Joao da Silva do Amaral de Souza" chooses nothing). **A branch that runs but changes nothing is not a decision either** -- the prefix chain's `merge(k, j)` executes even when `j == k + 1`, folding a piece into itself, and keying on "the code got here" reported a fork for all ambiguous particles on "Do Van Jr." (`Dr.` when that was written, before #367 made a plain title transparent and put the shape out of the loop's reach entirely), where the particle stayed a lone leading name piece — the GIVEN name under the default order, the family name under `FAMILY_FIRST` — and `_assign` reported the same token again. Check that the branch actually claimed something (`j > k + 1`) before recording. Structure often settles the question before it arises, which is why `PARTICLE_OR_GIVEN` is not emitted on the `FAMILY_COMMA` path's WHOLLY-FAMILY read -- the comma fixed which piece is the family -- and `SUFFIX_OR_NAME` is not emitted for "Ma, Jack". A tail segment is the same case from the other side: assign reads it as suffixes, a maiden marker there included since #601, and since #603 a title word past the second comma as a title (rules.md#C2, so `Freiherr` below reads title) (`Jane Doe, PhD, Jr nee van Ma` reads suffix `PhD, Jr nee van Ma`, the marker an ordinary word, rules.md#M2), so group's two chain emitters are handed no report list there, after either comma — `John Smith, Jr., Freiherr von Richthofen` still chains `von` and reports only `comma-structure` (rules.md#C2, 2026-09-28). Read that scope narrowly: the comma settles nothing about a particle trailing the given name, so P6's attachment in `post_rules` decides that fork on the same path and reports it (#405), in the kind naming the reading it OVERRODE, which is the reading assign made and not the word's vocabulary: `SUFFIX_OR_NAME` where assign had read the run as a post-nominal (`vd`, `mc`), else `PARTICLE_OR_GIVEN` where the run holds an ambiguous particle (`van`, and `do`, which is in the suffix vocabulary too but in its AMBIGUOUS half, so no credential reading was overridden), else silence. The decision site also has the token index and the detail text in hand, which the tag scan would have to reconstruct. **If a fork's two branches are taken in DIFFERENT stages, every one of them needs the emitter** -- `PARTICLE_OR_GIVEN` is decided in `_assign` when the ambiguous particle stays a lone leading piece, in `_group` when something shifts it off the name's leading piece and the prefix chain claims it, and in `post_rules` when P6's attachment takes a trailing particle into the family after a comma, so all three report; for two years only the first did. What can still do the shifting is narrow, and #367 is why: a plain title no longer can (`Dr. Van Johnson` reads as `Van Johnson` does and reports from `_assign`), so the `_group` emitter needs a word that is BOTH a title and a particle — measured, `TITLES ∩ particles_ambiguous` is `{freiherr, st}` in the default vocabulary (`do` left TITLES in #296's audit; decisions.md's Excluded block records the before and after), plus any overlap a caller's config creates — standing ahead of the chained particle as the LEADING NAME word. Titles may precede it, so `Dr. St van Johnson` reaches the emitter and `St van Johnson` does too; a given name may not, so `Jan Freiherr von Richthofen` does not reach it while `Freiherr von Richthofen` and `Dr. Freiherr von Richthofen` do. Two shapes that look like they should reach it and do NOT, both measured by stepping `STAGES` and watching where `ambiguities` grows: `Dr. Do van Johnson` and `Do St Johnson` report from `assign`, not `group`, because `do` is no longer a title and so stays the leading name piece assign reports on — a both-vocabulary word CHAINED (`Jan St Johnson`) reports nothing at all. When checking whether that emitter is dead, a both-vocabulary word in the leading name position is the thing to look for, and the answer is that it is not dead. The stage-ownership map in `tests/v2/pipeline/test_state.py` must list `ambiguities` for each such stage, and it passes vacuously until a case row exercises the path, so add the row too. Report BOTH directions of a two-way fork — "John Smith MA" (read as a suffix) and "Jack MA" (read as the family name) are equally guesses. Every kind needs a trigger in `tests/v2/test_contracts.py::_AMBIGUITY_TRIGGERS` (an explicit `None`, strict-xfail, while reserved), and case-table rows pin expected kinds exactly, so a new emitter shows up in both immediately. **Pin the decision, not the vocabulary**: the only titled-particle test used an UNAMBIGUOUS particle, so it walked the right code path and proved nothing about the branch under test -- two criticals passed 1539 tests. A row contrasting the two readings ("John Smith V" against "John Smith B") is what makes an emitter's absence meaningful. - **A kind is worth adding only if a reader would hesitate too**: the test is not "does the code take a branch" but whether a person reading that input would genuinely be unsure. "Smith, John V" reads as a middle initial to anyone -- the comma settles it -- so reporting it would be noise that teaches callers to ignore the field, which costs more than the missing report. Reachability of the second branch is necessary, not sufficient. Prefer leaving a fork silent and documenting the omission over emitting on input nobody finds ambiguous. - **Parser owns config-dependent conveniences**: `Parser.matches`/`Parser.capitalized`/`Parser.revise` exist because the `ParsedName` equivalents fall back to DEFAULT config for str/omitted arguments (documented loudly in both docstrings). `revise` harvests tokens from a full sub-parse of each replacement value (tags kept minus `FOLDED_TAG`, roles forced, the R1 entry pass `suffix_entries` re-run over the forced state so a suffix value's entries follow its own commas, ambiguities discarded); the merge tail is shared with `replace()` via `ParsedName._with_field_tokens`. `Parser.capitalized` delegates through `name.capitalized(self.lexicon)` specifically so `_parser` never imports `_render` — keep it that way. - **Per-word vocabulary fields warn on multi-word entries** (`_normset`/`_normpairs` via `_warn_dead_entry`, UserWarning, never a raise — see the given_name_titles Gotcha for why raising is wrong). `given_name_titles` is the one multi-word-matched field and is exempt; `_edit` passes `warn=False` (add() warns once via the new instance's `__post_init__`; remove() stores nothing). The default vocabulary and every locale pack must stay warning-free (`test_default_lexicon_builds_warning_free`, `test_pack_vocabulary_entries_are_single_words`). @@ -419,7 +419,7 @@ Don't use the bare `python3 -m doctest .rst` CLI (no `optionflags`) to che **Prefix-join uses value-based `list.index()`** in `join_on_conjunctions` — fragile when a token value repeats (e.g. a trailing title that's also a suffix acronym, or two `van`s); constrain such lookups to start at `i + 1`. See #100. -**Title vs suffix is positional for BARE words, and the leading period-abbreviation rule overrides even that** — a word matching `TITLES` at the front of a name becomes `title`; the same word matching `SUFFIX_ACRONYMS`/`SUFFIX_WORDS` at the end becomes `suffix` (never both, regardless of the word's real-world meaning). The `TITLES`/suffix overlap was audited in #296 (2026-08-23): the pure postnominals (`jr`, `junior`, `phd`, `do`, `se`) left `TITLES`, the v1-residue `dr`/`sra` left the suffix sets, and the twelve words still in both (`md`, `ms`, `sa`, `sr`, `lt`, `ra`, `vc`, and the ranks `cpl`, `cpo`, `cpt`, `csm`, `sgm`) are deliberate duals that position decides. External test sources (old issue gists, etc.) sometimes assert `suffix` for a leading professional abbreviation like `RA`/`PD`/`Dipl.-Ing.` — that's the source data being wrong, not a parser bug. Verify position before "fixing" it. Two qualifications the older "purely positional" wording papered over, the first measured 2026-08-01 and the second answered 2026-09-08 (#316): a PERIOD-marked leading word is claimed by the shape rule before any vocabulary is read (`"Esq. Smith"` → `title`, though `esq` is suffix-only), and a PERIOD-marked trailing word is claimed by VOCABULARY — first the suffix sets, which the peel reads before anything else (`"John Smith Esq."` → `suffix`), then the titles (`"John Smith Prof."` → `title='Prof.'`, `rules.md#H5`) — while an unlisted abbreviation there stays a name part (`"John Smith Xyz."` → `family='Xyz.'`) unless it is written as two or more period-separated chunks, which since 2.4 IS a trailing shape rule (`Policy.unlisted_dotted_suffixes`, default on: `"John Smith X.Y.Z."` → `suffix='X.Y.Z.'`, read by the same words-to-spare count a bare ambiguous acronym takes, `rules.md#S3`). The SINGLE trailing period is still not one, deliberately — it is the abbreviation shape any word can wear. The bare word is where "positional" still holds whole, and the reason it must: `TITLES` holds words that are in no suffix set — 746 of them, measured on this tree with `L = Parser().lexicon; len(L.titles - L.suffix_acronyms - L.suffix_words)`, a figure that grows with the vocabulary and never shrinks the argument — and many are ordinary surnames (`king`, `bishop`, `prince`, `pope`, `judge`, `sheriff`, `baron`, `master`, ...), so a vocabulary-first rule over BARE trailing words would read `"Mary Jane King"` as `title='King'`, `family='Jane'`. The period is what separates the safe case from that one, and it separates it by being a WRITING convention rather than by making the collision go away: `"Mary Jane King"` keeps `family='King'` while `"Mary Jane King."` reads `title='King.'`, `family='Jane'`, a cost accepted under the input-is-a-name premise rather than one the rule prevents (decisions.md#P5's trailing-position bullet, measured 2026-09-09). Since #316 that sentence describes the shipped rule rather than an aspiration. +**Title vs suffix is positional for BARE words, and the leading period-abbreviation rule overrides even that** — a word matching `TITLES` at the front of a name becomes `title`; the same word matching `SUFFIX_ACRONYMS`/`SUFFIX_WORDS` at the end becomes `suffix` (never both, regardless of the word's real-world meaning). The `TITLES`/suffix overlap was audited in #296 (2026-08-23): the pure postnominals (`jr`, `junior`, `phd`, `do`, `se`) left `TITLES`, the v1-residue `dr`/`sra` left the suffix sets, and the twelve words still in both (`md`, `ms`, `sa`, `sr`, `lt`, `ra`, `vc`, and the ranks `cpl`, `cpo`, `cpt`, `csm`, `sgm`) are deliberate duals that position decides. External test sources (old issue gists, etc.) sometimes assert `suffix` for a leading professional abbreviation like `RA`/`PD`/`Dipl.-Ing.` — that's the source data being wrong, not a parser bug. Verify position before "fixing" it. Two qualifications the older "purely positional" wording papered over, the first measured 2026-08-01 and the second answered 2026-09-08 (#316): a PERIOD-marked leading word is claimed by the shape rule before any vocabulary is read (`"Esq. Smith"` → `title`, though `esq` is suffix-only), and a PERIOD-marked trailing word is claimed by VOCABULARY — first the suffix sets, which the peel reads before anything else (`"John Smith Esq."` → `suffix`), then the titles (`"John Smith Prof."` → `title='Prof.'`, `rules.md#H5`) — while an unlisted abbreviation there stays a name part (`"John Smith Xyz."` → `family='Xyz.'`) unless it is written as two or more period-separated chunks, which since 2.4 IS a trailing shape rule (`Policy.unlisted_dotted_suffixes`, default on: `"John Smith X.Y.Z."` → `suffix='X.Y.Z.'`, read by the same words-to-spare count a bare ambiguous acronym takes, `rules.md#S3`). The SINGLE trailing period is still not one, deliberately — it is the abbreviation shape any word can wear. The bare word is where "positional" still holds whole -- outside a part past the second comma, which holds no name word, so C2 reads a bare title word there as a title since #603 (`Eric H. Holder, Jr., Attorney General`) -- and the reason it must: `TITLES` holds words that are in no suffix set — 746 of them, measured on this tree with `L = Parser().lexicon; len(L.titles - L.suffix_acronyms - L.suffix_words)`, a figure that grows with the vocabulary and never shrinks the argument — and many are ordinary surnames (`king`, `bishop`, `prince`, `pope`, `judge`, `sheriff`, `baron`, `master`, ...), so a vocabulary-first rule over BARE trailing words would read `"Mary Jane King"` as `title='King'`, `family='Jane'`. The period is what separates the safe case from that one, and it separates it by being a WRITING convention rather than by making the collision go away: `"Mary Jane King"` keeps `family='King'` while `"Mary Jane King."` reads `title='King.'`, `family='Jane'`, a cost accepted under the input-is-a-name premise rather than one the rule prevents (decisions.md#P5's trailing-position bullet, measured 2026-09-09). Since #316 that sentence describes the shipped rule rather than an aspiration. ### Tests (`tests/`) diff --git a/docs/customize.rst b/docs/customize.rst index 3c108ad7..c8caf804 100644 --- a/docs/customize.rst +++ b/docs/customize.rst @@ -1021,19 +1021,26 @@ Suffixes not separated by commas ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ ``extra_suffix_delimiters`` handles sources that separate post-nominals -with something other than a comma. The default reading of such a name -is bad enough to be the reason you'd go looking: +with something other than a comma. Undeclared, the separator is a word. +Since 2.4 such a name still parses when a credential opens the part +after the comma, because that makes the whole part the suffix, but the +separator stays in the suffix as written. Declared, it splits the +suffix into groups the way a comma does: .. doctest:: >>> name = parse("Jane Smith, RN - CRNA") >>> name.given, name.family, name.suffix - ('RN', 'Jane Smith', 'CRNA') + ('Jane', 'Smith', 'RN - CRNA') >>> policy = Policy(extra_suffix_delimiters=frozenset({" - "})) >>> name = Parser(policy=policy).parse("Jane Smith, RN - CRNA") >>> name.given, name.family, name.suffix ('Jane', 'Smith', 'RN, CRNA') +Through 2.3 the undeclared reading was given ``RN``, family ``Jane +Smith`` and suffix ``CRNA``, so the delimiter was the only way to get +the name right. + .. _unlisted-credentials: Credentials the vocabulary doesn't list diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 5b3301cd..38f0aa38 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -958,6 +958,15 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): MEASURED 2026-10-03, branch against 061f02da. POPULATION: each head in {`Smith, John,`, `Smith, John, Jr.,`, `John Smith,`, `John Smith`} followed by every 2-, 3- and 4-word product of {PhD, MD, Puig, i, y, -, Jr., Soler, Mr., née, and} that contains `-`, 19,972 texts, plus the 1,505 distinct names of `tools/differential/corpus*.jsonl` at 061f02da. Under `extra_suffix_delimiters=(" - ",)` 2,624 of the generated parses move (suffix alone 2,384, suffix and maiden 240 -- every one of the 240 a maiden name lost to a marker beside a core: 100 with `née -`, 108 with `- née`, 32 with both); under `(" / ",)`, with every `-` word rewritten to `/`, the same 2,624 (2,384 and 240), the move being a property of the core's position rather than its spelling; and under `(" and ",)` over the texts as written, where `and` is now the core and the dash a word, 2,528 (2,468 and 60). No other field and no ambiguity report moves under any of the three. In the corpus, one name moves under each of two policies: `Smith, John, PhD née Puig - i Soler` under ` - ` (above), and `Doe, Jane, and Jr.` under ` and `, suffix 'and Jr.' to 'Jr.', a leading core dropped as any lone core is. COMMA-TWIN AGREEMENT over the generated texts under ` - `, comparing the seven role fields of each against the text with ` - ` replaced by `, ` under the default policy and keeping only texts whose core is neither first nor last after the head nor beside another core: `Smith, John,` 752 of 2,100 disagreed before and 0 after, `Smith, John, Jr.,` 752 of 2,100 and 0; `John Smith,` 1,956 of 2,100 both before and after, the head where the typed comma moves the name's structure. (A first draft counted 973 of 3,310 for the `Jr.,` head, its filter reading `Jr.,` as part of the body; the review of this change caught it.) RECOMPUTE: parse each text under each policy at both trees (pin `nameparser.__file__` on each side, AGENTS.md's two-tree gotcha), compare the seven role fields and the sorted ambiguity kinds, and count the texts whose record differs; for the twin, compare the role fields alone. SUPERSEDES, under M2: the 2026-09-26 #538 entry (its stepping, its invariant and its Not-repaired list, which this closes); the 2026-09-22 #397 follow-up's bound argument (a core can no longer stand between a marker and the clause's first word, the segment being cut at it first); and the #418 reorder entry's closing repair, which screened a tail's cores out of the marker pass so that `Smith, John, PhD née - Jones` read maiden 'Jones' -- it reads suffix 'PhD née, Jones' now, as above. Entries there that describe the walk stepping over a tail's cores (the #424 and #533 bullets) describe machinery this change deletes; they are dated and left as written. Not changed: a core OUTSIDE a tail segment is still a word (C1's Accepted entry), and a core inside a token (`RN/CRNA` under `/`) still reads by `splits_into_suffixes` and keeps the token whole. +- 2026-10-04 (Derek), #603 — A CREDENTIAL OPENING THE PART AFTER THE COMMA MAKES IT THE POSTNOMINAL PART, WHATEVER FOLLOWS; AND A TITLE WORD PAST THE SECOND COMMA IS A TITLE. One word that is neither suffix nor title vocabulary used to flip the whole comma reading: `John Smith, PhD Jones` read family 'John Smith', given 'PhD', where `John Smith, MD FACS` and `Eric H. Holder, Jr. Attorney General` were already read as a name and its postnominals. Now the first suffix word of the part, with only titles in front of it, opens the part when it starts S2's run (`_pieces.starts_a_credential_run`, #602's own predicate, so the two runs cannot disagree about what starts one), and every word after it reads as that run reads: a title word as a title, any other as a suffix. The mechanism is the one #296 built for a part holding no name word, extended rather than added beside: `segment` still decides FAMILY_COMMA, and `segment_suffix_reading` returns a reading instead of None, so assign's no-name path reads the part before the comma positionally where it holds two name words. Every word the opened part takes that would otherwise have made it the given part is reported (`suffix-or-name`), as #602's run reports what it absorbs. + THE ONE-WORD PATH, DECIDED (Derek): the whole part is suffixes. `Doe, PhD Jones` reads family 'Doe', suffix 'PhD Jones', no given name, as `Smith, PhD` already reads. DECLINED: reading the credential as a suffix and the next name word as the given name (`Smith, PhD Jones` → given 'Jones'), proposed by analogy with a title at the head of the given part. Asked whether that is a real format, it is not -- nobody writes `Family, Credential Given` on purpose -- and the one realistic input it served, a misplaced generation (`Smith, Jr. John` for `Smith Jr., John`), is garbage in, garbage out. One reading for both comma paths was the deciding ground: the one-word path disagreeing with the full-name path is the shape #429, #430 and #432 each found. Accepted cost: `Smith, Jr. John` loses 'John' to the suffix, where 2.3 read given 'John' with 'Jr.' misfiled as a title. + A TITLE/SUFFIX DUAL OPENS NOTHING (Derek). A word in both vocabularies (`md`, `ms`, `sr`, the ranks `lt`, `cpl` and the rest) is the listing form's title in front of a given name, `Smith, Ms Jane`, and a credential behind it opens nothing either (`Smith, Ms PhD Jones` keeps given 'PhD', as `Smith, MD PhD Ma` keeps given 'Ma' under #544). DECLINED after building it: a dual opening the part behind two or more words before the comma. Derek first chose it for `John Smith, MD Jones`; the reviews found what the count cannot see -- a surname of two words (`García Márquez, Ms Gabriela`, `Smith Jones, Sr. Maria`), a title in front of a one-word surname (`Dr. Smith, Ms Jane`, the count taking the title as a word), and spaced initials (`García Márquez, Ms G. J.`) each lost the given name -- and with the word explained (a dual is a word in both lists, decided by position), Derek chose that duals never open. So `John Smith, MD Jones` and the `MD - PhD - FACS` delimiter rows keep their 2.3 reading. An ambiguous-class member opens nothing either (`John Smith, Ma Jones` keeps given 'Ma'), nor does a suffix word S2's run start refuses (`John Smith, vd Jones`). + STRICT MODE: the veto refuses an initial-shaped word the suffix READING, not the run. `Smith, PSM I.` under `lenient_comma_suffixes=False` reads suffix 'PSM I.' (2.3: given 'PSM', suffix 'I.'), reporting the absorbed 'I.', which matches #602's no-comma run (`John Smith PhD V.` reports; its given-part run does not, a disagreement #602 left and this does not touch). + THE SPELLING GAP THIS FOUND, FIXED HERE. `starts_a_credential_run` refused any piece longer than one token, so the split credential `Ph. D.` group merges started neither #602's run nor this one while `PhD` started both: `Smith, John Ph. D. Jones` read middle 'Jones' where `Smith, John PhD Jones` read suffix 'PhD Jones'. The merged piece now starts the run, the "suffix" piece tag being the Ph./D. merge's alone; `trailing_candidates` had already admitted it, so its pre-check and the predicate now agree. The no-comma spelling `John Smith Ph. D. Jones` still reads family 'Jones': the trailing peel's walk never holds the merged piece, which is a separate path and is left (AGENTS.md's spelling sweep). + C2, WIDENED (Derek): a word of the title vocabulary in a part past the second comma that is not also suffix vocabulary is a title -- the part stands behind a comma, the line H5 draws for reading a title from the end of a name -- and a piece holding a suffix word stays a suffix, a dual and a connective join that took one (`PhD - and MD`) alike. `Eric H. Holder, Jr., Attorney General` reads title 'Attorney General', and a part of title and suffix words is recognized rather than flagged. The flag's test is asked word by word against the title vocabulary, a period-joined word through `period_joined_vocab` as classify tags it (`John Smith, Jr., Lt.Gov.`, which the code review found still flagged), so `Secretary of State` reads its title but keeps `comma-structure`: knowing that group's title chain takes the connective would be `segment` modelling `group`, the sign #611 is about. Every release through 2.3.0 read these as suffixes, 1.4.0 included (`Eric H. Holder, Jr., Attorney General` suffix 'Jr., Attorney General', `Smith, John, Prof.` suffix 'Prof.', measured with the released wheels 2026-10-04); a first draft of this entry said 1.4.0 read them as titles, inferring it from the 1.4.0 run reporting no unexplained diff, when the reason was the broad `fix(comma-family)` rule there claiming them, which #603's two rules now precede by declaration. Accepted consequence: a noble rank there reads as a title, `John Smith, Jr., Freiherr von Richthofen` → title 'Freiherr', suffix 'Jr., von Richthofen'. + WHAT MOVES, measured 2026-10-04 against master at 97d36e0d, after the dual decision above. The differential corpora at 97d36e0d (1,519 name and order entries over 1,515 distinct names): four names, every one intended -- the dash-delimited `Steven Hardman, RN - CRNA` now reads given 'Steven', family 'Hardman', suffix 'RN - CRNA', and `John, Smith, Dr.`, `Andrew Perkins, Jr., Col. (Ret)` and `1 & 2, 3 4 5, Mr.` read their tail title. A generated grid: every name of `tools/differential/corpus*.jsonl` at 97d36e0d, bare-string and object lines alike, plus 120,000 one-to-seven-word texts drawn with `random.Random(601)` from the 54 words {Jane, Doe, Smith, John, nee, née, geb., z domu, PhD, MA, Ma, MD, Jr., Jr, III, V, VI, i, y, de, de la, van, vd, mc, do, Do, DO, abd, abdul, Dr., Prof., King., Le, Attorney, General, Jones, `,`, ba, R.A.I., (nee, Smith), Ó, binti, Ph., D., Chief, Justice, Ms, Sr., RN, -, G.J., MJ, Col.}, plus each of the heads {Jane Doe, Doe,, Smith, John, John Smith, Dr., Mai Le, Jane van der Berg} followed by every three-word product of {nee Smith, PhD, MA, Prof., de, Jr., Jones, VI, do, i Soler, nee}; 98,630 texts, each parsed under the default policy, `name_order` family-first and a ` - ` delimiter, the seven role fields, the ambiguity kinds and details and `initials()` compared against 97d36e0d. 588 texts move: 581 have a part after a comma whose first non-title word is a suffix word that is not also a title, or a title word past the second comma, or both; the other seven were read by hand -- a leading period-shape title (`geb.`) in front of the opening credential, and #602's given-part run started by `Ph. D.` (`, abdul Ph. D. Do binti Smith`). Frames per parse through `Parser().parse`, py3.11: the reference band and plain names unchanged (`Dr. Juan Q. Xavier de la Vega III` 449, `Smith, John` 179, `John Smith, Jr.` 210, `John Smith, Dr.` 261, `Smith, Ms Jane` 256); a part of suffix words behind one word gets cheaper, the positional count now leaving in C for a single piece (`Smith, Jr.` 163 → 156, `Smith, MD PhD` 260 → 253); and the opened part pays for what it reads (`John Smith, PhD Jones` 338 → 362). +- 2026-10-04 #611 — C1, S2 AND P6 GREW BY PATCHING, AND EACH SHOWS THE DECISIVE SIGN; THE RETHINKS ARE FILED, NOT DONE. The question was docs/design/AGENTS.md's, asked of the three longest rules after M2 (#601): does a guard model what another stage would do? C1: `segment` predicts group's particle chain (the #562 pair test) and assign's capitals reading (the skip #562 broke), mirrors classify's tags in its unit counts, and decides the structure in two stages, assign re-deciding #296's positional read with a second count -- the copy #603 extended. S2: the words-to-spare count runs over pieces group merges, so the chain re-asks the peel and rolls back (`John van Mc`), P5 compares two views, group emits S2's report, and the peel runs up to five times a parse. P6: assign holds back exactly the particle tail post_rules will attach (`particle_tail(..., sticky_from)`). Decided: #603 shipped as a targeted change inside today's machinery; the rethinks are #613 (C1, read the comma once after classify and bind, P6 folded in) and #614 (S2, read the trailing run once over units and bind), C1 first because S2's unit count should be shared with it, neither for 2.4. + ### T1 — separators, not joiners - 2026-07 (v2 core, PR #288) — v1's squash_emoji/squash_bidi REMOVED the character and joined its neighbors. (v1.3.0 had no bidi handling at all: squash_bidi entered late v1 via #266, 2026-07-07, on the emoji precedent's shape.) ('A😀B' → 'AB'); v2 makes an ignorable character a separator ('A😀B' → 'A', 'B'). The unavoidable consequence of every part being an exact positioned piece of the input: with no rewriting stage, nothing can splice two half-words together. diff --git a/docs/design/mechanisms.md b/docs/design/mechanisms.md index 729cf491..fd0dad72 100644 --- a/docs/design/mechanisms.md +++ b/docs/design/mechanisms.md @@ -200,6 +200,7 @@ Problem shape. A test pins an ordering, a sort, a dedup or a partition, and its - Run all the gates, not the ones you remember: ruff runs before mypy and pytest in CI, and each has caught what the others passed. - Purge __pycache__ between same-length source mutations; stale bytecode makes a changed file measure as unchanged. - After NARROWING a rule, check the receiver: the names a narrowed rule sheds land on a neighbour, and nothing guarantees the neighbour's prose describes what it inherited — #375 fixed an over-claiming rule and relocated the bug onto its neighbour. Ask also what the old behavior was CONCEALING: #379's attachment removed the input a test used to build an all-particle middle name (#402), and #400's reserve fix exposed the dual-membership count shape #397 names -- twice in one session a fix's real yield was a defect it stopped hiding. +- A gate run with no UNEXPLAINED line for a name is not evidence that the baseline read it the same: a broad rule may have claimed the diff. #603's first draft wrote that 1.4.0 read a title past the second comma as a title because no 1.4.0 name moved unexplained, when the 1.4.0 ledger's `fix(comma-family)` rule, a regex over every Latin comma name, was claiming each one in the classified summary; the released wheel read suffix 'Jr., Attorney General'. Read the classified summary for the name, and ask the wheel (`PYTHONSAFEPATH=1 uv run --isolated --no-project --with nameparser==X`, from outside the worktree) before writing what a release did. - A ledger rule EXPLAINS a diff; nothing checks that its own sentence still DESCRIBES it, so a rule can go on covering a name it has stopped being true of and the gate stays green. Measured on #533: a `(?i)` anchor written for `John née Jones Smith Ma` — the clause keeping the credential — also reached the capitals spelling, which that change made read the opposite way, the clause giving the credential up. The rule still matched the name and still covered its fields, so `unexplained: 0` was silent about it. What found it was ATTRIBUTION rather than the gate: measuring every corpus name the change touches against the PARENT commit, the EXPLAINED ones as well as the unexplained, and asking of each which change actually moves it. Do that before writing a ledger rule's prose, not only when a name goes unexplained — the failure mode is a rule that reads as authoritative and argues for the reading it lost. Second instance, #397/#461 (2026-09-20), and it is the same method finding the OPPOSITE shape: `fix(initials-per-word)` had been silently ABSORBING eleven 1.4.0-parity breaks that belong to the new connective rule, so the gate was green and the attribution was wrong about which change owned them. A rule can over-cover as well as mis-describe, and neither is a diff the gate can report. - A skip is indistinguishable from "correctly declined": pytest turns an empty parametrize into a skip, and a filter that widens its own skip set cannot fail. After changing any selection shape, verify the guard still REACHES the code it watches — assert the selected set is non-empty, or force-a-decision on its size. - A differential corpus cannot evidence behavior keyed to OUT-of-vocabulary shapes: it holds only names someone wrote down, and an unrecognized word is by definition outside the vocabulary — a green run over the corpus proves nothing about such a rule. diff --git a/docs/design/rules.md b/docs/design/rules.md index 8a0033d6..82b76f88 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1666,8 +1666,9 @@ C1. Rationale: a credential run after the comma means the name is in With a comma present, the name reads as trailing suffixes when the part after the first comma is entirely suffix words and more than one word precedes the comma; otherwise it reads as the - listing form, the part before the comma being the family name. - Only the part after the first comma decides. + listing form, the part before the comma being the family name + unless the part after it fixes no family boundary (below). Only + the part after the first comma decides. Wherever this rule counts the words before the comma, a particle run and the one name word it attaches to are one word, the reach P1's fold counts rather than the whole of P2's chain: the listing @@ -1798,7 +1799,22 @@ C1. Rationale: a credential run after the comma means the name is in only, fixes no family boundary, so a part before the comma with more than one name word keeps its positional read, order and all. A name word in the part after the comma makes it the - given name, with titles before it and suffixes after. + given name, with titles before it and suffixes after, unless a + credential opens the part. A credential opening the part makes it + the postnominal part however it goes on: where the first suffix + word of the part, with only titles in front of it, is one that + starts S2's run, which S2 bounds (the split 'Ph. D.' among them), + every word after it reads as that run reads, a title word as a + title and any other word as a suffix, under strict mode too, + whose veto refuses a word the suffix reading and not the run + ('Smith, PSM I.'). The comma has then fixed nothing, so a part + before it holding two or more name words keeps its positional + read, and a part holding one is the family ('John Smith, PhD + Jones', 'Doe, PhD Jones'). A word of both the title and the + suffix vocabulary opens nothing, and nor does a credential behind + one: in front of a given name it is the listing form's title + ('Smith, Ms Jane'), and no count of the words before the comma + tells a surname of two words from a given name and a family name. A one-character suffix word — the only kind a reader could take for an initial — is read by what stands before it. Behind another suffix it is describing that suffix and continues the @@ -1815,7 +1831,8 @@ C1. Rationale: a credential run after the comma means the name is in credential run, and a letter in it continues that run up to the further comma. Longer suffix words are not in question either way, and the strict knob above still vetoes the initial-shaped - ones, so the run ends at them there. + ones, so the run ends at them there unless a credential opened + the part, above. A delimiter the policy declares parts a trailing suffix part as a comma would: the words on each side of it read exactly as the parts of the same text written with a comma in its place, so no @@ -1836,7 +1853,7 @@ C1. Rationale: a credential run after the comma means the name is in "Smith, Sr." → suffix="Sr." "Smith, PSM I" → suffix="PSM I" "Smith, PSM I." → suffix="PSM I." - "Smith, PSM I." strict-comma-suffixes → suffix="I." + "Smith, PSM I." strict-comma-suffixes → suffix="PSM I." "Smith, John V." → middle="V." "Smith, John PhD I." → suffix="PhD I." "Smith, John V" → suffix="V" · boundary @@ -1845,6 +1862,13 @@ C1. Rationale: a credential run after the comma means the name is in "Doe, MA PhD" → ambiguities=("suffix-or-name",) "Smith, Dr." → title="Dr." "Smith, Dr. Jr." → suffix="Jr." + "John Smith, PhD Jones" → given="John" + "John Smith, PhD Jones" → suffix="PhD Jones" + "John Smith, PhD Jones" → ambiguities=("suffix-or-name",) + "Doe, PhD Jones" → suffix="PhD Jones" + "John Smith, Ph. D. Jones" → suffix="Ph. D. Jones" + "John Smith, MD Jones" → given="Jones" · boundary + "Smith, Ms Jane" → title="Ms" · boundary "John Smith, Mr." → given="John" "John Smith, Mr." → family="Smith" "John Smith, Mr. Jr." → given="John" @@ -1952,8 +1976,13 @@ C1. Rationale: a credential run after the comma means the name is in "Smith, John, MD - née Jones Smith" extra_suffix_delimiters-dash → suffix="MD, née Jones Smith" Accepted: a delimiter core the policy names (T1) is a word here, not structure — v1 applied the delimiter to the suffix-comma - form alone, and that limitation is kept as parity: "Smith, RN - - CRNA" reads given "RN" under the policy as without it. + form alone, and that limitation is kept as parity: the core in + 'Smith, RN - CRNA' is a word under the policy as without it. + Since #603 it is a word of the run the credential opens, so the + part reads suffix "RN - CRNA" where 2.0 through 2.3, like v1, + read given "RN". + "Smith, RN - CRNA" → suffix="RN - CRNA" + "Steven Hardman, RN - CRNA" → given="Steven" Accepted: the further-comma qualifier carries no example line of its own. It discriminates PAIRS and spans both branches, so exemplifying it means a with-comma partner for each — every one @@ -1971,11 +2000,21 @@ C1. Rationale: a credential run after the comma means the name is in C2. Rationale: text beyond the recognized comma parts should be taken in without silent guessing. - Parts beyond the second are consumed as suffixes either way; a - non-empty extra part that is not entirely suffix words is - flagged as a structural ambiguity rather than rejected — parsing - never fails on content. An empty part between doubled commas is - consumed silently. + Parts beyond the second are consumed as suffixes either way, + except that a word of the title vocabulary there that is not also + suffix vocabulary is a title: a part past the second comma holds + no name word, so the surnames H5 keeps a bare trailing title word + from misreading are not there to protect, and a piece holding a + suffix word, a title/suffix dual among them, stays the postnominal + the slot makes it (#603). A non-empty extra part that + is not entirely suffix words is flagged as a structural ambiguity + rather than rejected — parsing never fails on content — unless + its words are title words and suffix words, one title word at + least, a word whose period-joined chunks read as a title + ('Lt.Gov.') counting as one. That is asked word by word, so a part + a connective joins into one title ('Secretary of State') keeps + the flag. An empty + part between doubled commas is consumed silently. A part the parse reads as a credential run by some route other than the suffix vocabulary is recognized and is not flagged: a run of ambiguous acronyms whose written case leans credential @@ -1998,6 +2037,12 @@ C2. Rationale: text beyond the recognized comma parts should be name. What such a part reports is its own — C1's flip and this rule's flag. "John Smith, MD, Bart" → suffix="MD, Bart" + "Eric H. Holder, Jr., Attorney General" → title="Attorney General" + "Eric H. Holder, Jr., Attorney General" → ambiguities=() + "Eric H. Holder, Jr., Secretary of State" → ambiguities=("comma-structure",) · boundary + "Smith, John, Prof." → title="Prof." + "John Smith, Jr., Lt.Gov." → ambiguities=() + "John Smith, MD, Ms" → suffix="MD, Ms" · boundary "John Smith, MD,, Jr." → suffix="MD, Jr." · boundary "John Smith, MD, R.A.I." → suffix="MD, R.A.I." "John Smith, MD, R.A.I." → ambiguities=() @@ -2016,7 +2061,7 @@ C2. Rationale: text beyond the recognized comma parts should be reports on outside a tail: the_chain_reports_the_acronym_it_takes and titled_particle_chain_survives_a_title_that_is_also_a_particle. - history: decisions.md#C1, decisions.md#S2 · interacts: C1, P2, S2, S3 · implemented: nameparser/_pipeline/_segment.py, nameparser/_pipeline/_group.py + history: decisions.md#C1, decisions.md#S2 · interacts: C1, H5, P2, S2, S3 · implemented: nameparser/_pipeline/_segment.py, nameparser/_pipeline/_assign.py, nameparser/_pipeline/_group.py ## Name order (O) diff --git a/docs/release_log.rst b/docs/release_log.rst index 405b4c3f..3e011650 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -32,10 +32,14 @@ Release Log - **Change where a maiden marker takes a maiden name: only behind a surname, and only up to the credentials and titles the name ends with.** ``HumanName("Jane Doe nee Smith")``, ``Mai Le née Nguyen`` and ``Doe nee Smith, Jane`` give maiden ``Smith`` and ``Nguyen`` as before, but a marker counts only in the name before any comma or in the surname part before a family comma, and only behind a name word. Anywhere else it is an ordinary word, as in 1.4.0: ``Doe, Jane nee Smith`` gives middle ``nee Smith``, ``Smith, John, PhD née Jones`` gives suffix ``PhD née Jones``, and ``Dr. nee Smith`` gives first ``nee``, last ``Smith``, where 2.0 through 2.3 gave maiden ``Smith``, ``Jones`` and ``Smith`` -- the last with no name at all. A credential in front of the marker ends the name (see the credential-run entry below), so ``Jane Doe PhD nee Smith`` gives suffix ``PhD nee Smith`` where 2.0 through 2.3 gave suffix ``PhD``, maiden ``Smith``. A particle or a lone word is a surname here: ``Jane van nee Smith`` gives last ``van``, maiden ``Smith``, and ``Jane Smith née V`` gives last ``Smith``, maiden ``V``, as 2.0 and 2.1 read it, where 2.2 and 2.3 gave last ``née``, suffix ``V``. The words the marker takes now end where the run of post-nominals and titles the name would end with if the clause were not written begins, and the words it gives up read as they would there: ``Jane Doe nee Smith MA`` gives maiden ``Smith``, suffix ``MA``, where 2.0 through 2.3 gave maiden ``Smith MA`` and said nothing, as do the one-case ``JANE DOE NEE SMITH MA`` and ``jane doe nee smith ma``; ``Jane Doe nee Smith MA PhD`` gives suffix ``MA PhD`` where 2.3.0 gave maiden ``Smith MA``, suffix ``PhD``, so the two orders of the credentials now agree; ``Jane Doe nee Smith DO DO`` gives suffix ``DO DO``; ``Jane Doe nee Smith Prof.``, ``Mary Smith née Jones Prof.`` and ``Jane van der Berg nee Smith Prof.`` give title ``Prof.``; ``Jane Doe nee Smith King.`` gives title ``King.``; ``Jane Doe nee Smith MA Prof.`` and ``Jane Doe nee Smith Prof. MA`` both give title ``Prof.``, suffix ``MA``; and ``Jane Doe nee Smith V Prof.`` gives suffix ``V`` -- each where 2.3.0 kept every word in the maiden name. A credential the clause gives up or keeps at that boundary is reported as ``suffix-or-name``. The writing still decides, as it does at the end of a name with no clause: ``Jane Doe nee Smith Ma`` keeps maiden ``Smith Ma``, and ``Jane Doe nee Yo-Yo Ma`` keeps a two-word birth surname whole. The first word after the marker is always taken, the marker having announced a name: ``Jane Doe nee MA`` and ``Jane Doe nee King.`` keep it, and ``Jane Doe nee Prof. Dr.`` gives maiden ``Prof.``, title ``Dr.``, where 2.3.0 gave maiden ``Prof. Dr.``. Before a family comma the clause keeps every word, as before: ``Doe nee Smith Prof., Jane`` keeps maiden ``Smith Prof.``. Delimiters settle the question outright: ``Jane Doe (nee Smith MA)`` keeps the whole span, and brackets around a clause ending in a period are dropped as in 2.3.0, so ``Jane Doe (nee Smith Prof.)`` gives title ``Prof.``. The name #548 reported, ``Dr. nee Smith PhD Prof.``, gives title ``Dr. Prof.``, first ``nee``, last ``Smith``, suffix ``PhD``, where 2.3.0 gave last ``PhD``, maiden ``Smith``. ``John Smith nee Jones R.A.I.`` gives suffix ``R.A.I.``, as 2.3.0 read it. The rule replaces the clause rules this cycle's #533 and #535 had added, and the readings those changes gave here are superseded. See the ``M2`` entry of ``docs/design/decisions.md`` (closes #601) - - **Fix a comma part read wholly as suffixes reporting a particle in it as chained onto a name.** ``parse("John Smith, Jr., Freiherr von Richthofen").ambiguities`` names ``comma-structure`` alone, where 2.0 through 2.3 also named ``particle-or-given`` for ``von`` -- a word the same parse had put in the suffix, so the report described a reading it never made. No field moves. A part the parser consumes as suffixes, after a suffix comma or past the second comma, reports what the part is and nothing about its words as names, and the particle chain's credential-acronym report this release adds (above) keeps the same bound: ``John Smith, PhD Do Ma`` reports the comma's decision once and not again for ``Ma``. See the 2026-09-28 bullet of the ``C1`` entry in ``docs/design/decisions.md`` + - **Fix a comma part read wholly as suffixes reporting a particle in it as chained onto a name.** ``parse("John Smith, Jr., Freiherr von Richthofen").ambiguities`` names ``comma-structure`` alone, where 2.0 through 2.3 also named ``particle-or-given`` for ``von`` -- a word the same parse had put in the suffix, so the report described a reading it never made. This fix moves no field; the title change below (#603) moves ``Freiherr`` to the title. A part the parser consumes as suffixes, after a suffix comma or past the second comma, reports what the part is and nothing about its words as names, and the particle chain's credential-acronym report this release adds (above) keeps the same bound: ``John Smith, PhD Do Ma`` reports the comma's decision once and not again for ``Ma``. See the 2026-09-28 bullet of the ``C1`` entry in ``docs/design/decisions.md`` - **A credential after the name core starts a suffix run to the end of its part.** ``HumanName("John Smith PhD Jones")`` gives suffix ``PhD Jones`` and reports ``suffix-or-name`` for the name word it took, where 2.3.0 gave middle ``Smith PhD``, last ``Jones``; ``Smith, John PhD Jones`` gives suffix ``PhD Jones`` where 2.3.0 gave middle ``Jones``, suffix ``PhD``. A title in the run is a title, so ``Eric H. Holder Jr. Attorney General`` gives title ``Attorney General``, suffix ``Jr.``, as the comma spelling ``Eric H. Holder, Jr. Attorney General`` already read, where 2.3.0 gave middle ``H. Holder Jr. Attorney``, last ``General``. Only an unambiguous credential or generational word starts the run, and only behind the name core -- two name words with no comma, or the given part after one. A credential that is also a name does not start one, nor does a surname particle or a single letter, so ``John Smith MA Jones`` and ``Mohamed Ali Abd Allah`` read as before, and a bare title word starts nothing: ``Mary Jane King Smith`` keeps middle ``Jane King``. See the ``S2`` entry of ``docs/design/decisions.md`` (closes #602) + - **Fix a comma part that opens with a credential being read as a given name.** ``HumanName("John Smith, PhD Jones")`` gives first ``John``, last ``Smith``, suffix ``PhD Jones`` and reports ``suffix-or-name`` for the name word it took, where 2.3.0 gave first ``PhD``, middle ``Jones``, last ``John Smith`` (and 1.4.0 title ``PhD``, first ``Jones``). With one word before the comma the whole part is suffixes: ``Doe, PhD Jones`` gives last ``Doe``, suffix ``PhD Jones``. A title word in the part stays a title. A word that is a title as well as a credential (``MD``, ``Ms``) opens nothing, so ``Smith, Ms Jane`` keeps title ``Ms``, first ``Jane``. An undeclared delimiter becomes one more word of the part: ``Steven Hardman, RN - CRNA`` gives first ``Steven``, last ``Hardman``, suffix ``RN - CRNA``, where 1.4.0 through 2.3.0 gave first ``RN``, middle ``-``, last ``Steven Hardman``. The split spelling now counts like the joined one, so ``Smith, John Ph. D. Jones`` gives suffix ``Ph. D. Jones``, where 2.3.0 gave middle ``Jones``, suffix ``Ph. D.``. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #603) + + - **Change a title word in a part after a second comma to read as a title.** ``HumanName("Eric H. Holder, Jr., Attorney General")`` gives title ``Attorney General``, suffix ``Jr.``, where 1.4.0 through 2.3.0 gave suffix ``Jr., Attorney General``, and the part no longer reports ``comma-structure``. A word that is also a suffix stays a suffix (``John Smith, MD, Ms``). A title joined by a connective is read as a title, but its part keeps the report (``Secretary of State``). (#603) + - **Add the Catalan and Polish surname link.** ``parse("Josep Carod i Rovira")`` gives family ``Carod i Rovira``, where every release from 1.4.0 through 2.3.0 gave middle ``Carod i`` with family ``Rovira``; ``Josep Lluis Carod i Rovira`` gives middle ``Lluis`` with that same family; and ``Carod i Rovira, Josep`` gives it too, where they read family ``Carod Rovira`` and took the link into ``suffix`` as a generation marker. ``i`` is connective vocabulary now, the way ``y`` already was, and a connective counts as a name word wherever the three-word carve-out counts them -- whatever else the vocabulary says the word is, which matters here because ``i`` is also the roman numeral. A connective that is also generational vocabulary joins only where a name word stands on each side of it, so ``John Quincy Smith i`` keeps suffix ``i``, ``Josep Lluis Carod i III`` keeps suffix ``i III``, and the two-word ``Carod i`` keeps its generation reading. Written wholly in one case the letter reads as an initial and says so: ``JOSEP CAROD I ROVIRA`` and ``josep carod i rovira`` keep the fields they had and gain a ``conjunction-or-initial`` report, which a one-case name gains wherever a bare ``i`` or ``I`` stands among the name's own words -- a letter inside a maiden clause is read by the clause's rules and stays silent, as ``e`` already was -- and in an all-lower name that reading can move a field, each such name now reading as its all-caps twin already did (``parse("john smith i jr")`` gives middle ``smith``, family ``i`` and suffix ``jr`` where it gave family ``smith`` and suffix ``i jr``). Case repair follows the reading: a lower-case ``i`` the parse read as the generation is still title-cased by ``capitalize(force=True)`` (``Carod i`` gives ``Carod I``, as every release did), while one standing among the name words keeps its lower case as ``y`` always has (``Carod i Rovira`` gives ``Carod i Rovira``, where ``Carod I Rovira`` was the pre-2.4 answer). A link inside a maiden clause stays in the birth name, which no release read that way: ``HumanName("Jane Doe nee Puig i Soler")`` gives maiden ``Puig i Soler`` with last ``Doe``, where 2.0 through 2.3 ended the birth name at the link and gave maiden ``Puig`` with middle ``Doe i``, last ``Soler`` -- and the same words would have joined into last ``Doe i Soler`` under the change above, carrying a word of the birth name into the current surname. ``Jane Doe née Kowalska i Nowak`` moves with it, as does the all-lower ``jane doe nee puig i soler``; the ``y`` spelling always read this way and is untouched. The link still has to be joining: ``Jane Doe nee Puig i`` keeps maiden ``Puig`` with suffix ``i``, and ``Jane Doe nee Puig i III`` suffix ``i III``. A caller with Catalan or Polish data removes the entry from ``conjunctions_ambiguous`` and gets the join in the one-case names too; a caller who wants none of this removes ``i`` from ``conjunctions``, which restores every prior FIELD and every prior report, with two readings it does not restore and cannot: a letter the two vocabularies disagree about being an initial reads as one here and as the generation there, and case repair leaves a connective the parse placed among the NAME words in lower case where the off switch title-cases it -- ``parse("Dr. John i Smith").capitalized(force=True)`` keeps ``i`` where the off switch gives ``Dr. John I Smith``, and ``Carod i Rovira`` and ``Josep i Rovira`` are the same shape. Those two are the whole of what the switch does not undo, and ``tests/v2/test_properties.py`` states them as its invariants' only exemptions. A delimiter the caller declares through ``Policy(extra_suffix_delimiters=...)`` parts a trailing suffix part as a comma does (see the suffix-delimiter entry below), so no link joins across one: under ``(" - ",)``, ``Smith, John, PhD - i Soler`` keeps suffix ``PhD, i Soler``, as 2.3.0 read it. The default policy declares no such delimiter. See the ``P3`` and ``M2`` entries of ``docs/design/decisions.md`` (closes #397, closes #538) - **Fix a connective contributing no initial even where it is joining nothing.** ``parse("Juan de y").initials()`` gives ``J. y.``, where every release gave ``J.`` while ``family_base`` said ``y`` -- two views of one parse disagreeing about one token. A connective contributes nothing where it is JOINING, and initials like any other name word where its part holds nothing else for it to join. One rule for all three groups, so ``John and Jane Smith`` gives ``J. J. S.`` where 2.0 through 2.3 gave ``J. a. J. S.`` and 1.4.0 the run-together ``J a J. S.``, ``Duke of Edinburgh`` gives ``D. E.`` where 2.0 through 2.3 gave ``D. o. E.`` and 1.4.0 ``D o E.``, and ``John & Jane`` gives ``J. J.``. The question is asked of the whole part and never of a word count, so ``Jon Dough and`` has base ``Dough and`` and keeps ``J. D.``, and ``Juan Velasquez y Garcia`` keeps ``J. V. G.``. ``HumanName.initials()`` moves with the core -- over the differential corpora the two surfaces move on the same names and give the same values, reading one mark. Two names come back into 1.4.0 parity rather than away from it: ``JUAN Y GARCIA`` and ``محمد و علي`` both give the answer 1.4.0 gave. Parsing got cheaper by the same change -- the marks come off one pass instead of two, six fewer Python frames per name on 3.11. Two limits carried over from the 2.4 facade fix above: case repair still keeps such a connective lower-case, so ``initials()`` and ``capitalize()`` disagree about it on purpose, and a name restored from a pickle or a copy, or built from keyword fields, carries no tags and takes the older reading. See the ``R3`` entry of ``docs/design/decisions.md`` (closes #461) diff --git a/docs/usage.rst b/docs/usage.rst index 4e159afa..f43a1cd2 100644 --- a/docs/usage.rst +++ b/docs/usage.rst @@ -40,12 +40,17 @@ Understood by default Every piece of each is optional: 1. ``Title Given "Nickname" Middle Middle Family Suffix`` -2. ``Family [Suffix], Title Given (Nickname) Middle Middle[,] Suffix [, Suffix]`` -3. ``Title Given Middle Family [Suffix], Suffix [, Suffix]`` +2. ``Family [Suffix], Title Given (Nickname) Middle Middle[,] Suffix [, Suffix or Title]`` +3. ``Title Given Middle Family [Suffix], Suffix or Title [, Suffix or Title]`` + +A title written with a period (``Prof.``) is also read at the end of a +name in forms 1 and 3 and at the end of the given name in form 2; +written bare there (``Prof``), it is a name word, because many title +words are also surnames. The last two differ in what the comma is doing. In form 2 it separates the family name from the rest, so the family name comes first; in form -3 it only sets off suffixes, and the name before it is still +3 it only sets off suffixes and titles, and the name before it is still given-then-family: .. doctest:: @@ -57,6 +62,18 @@ given-then-family: >>> parse("John Doe, Jr.").family # form 3 'Doe' +A part after the comma in form 3, or after a further comma in either +form, holds suffixes and titles, a word there that is a title rather +than a suffix being read as the title (after a further comma, since +2.4): + +.. doctest:: + + >>> parse("Eric H. Holder, Attorney General").title + 'Attorney General' + >>> parse("Eric H. Holder, Jr., Attorney General").title + 'Attorney General' + Family-first forms you declare ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ diff --git a/nameparser/_pipeline/_assign.py b/nameparser/_pipeline/_assign.py index 64d60eef..f6eb7731 100644 --- a/nameparser/_pipeline/_assign.py +++ b/nameparser/_pipeline/_assign.py @@ -30,31 +30,39 @@ nonempty nickname beside exactly one piece in total puts that piece in FAMILY. FAMILY_COMMA: segment 0 wholly FAMILY (v1 parity) UNLESS segment 1 -holds no name word (titles and suffixes only), which fixed no family -boundary -- there segment 0 takes the NO_COMMA positional read instead, -order and all ('John Smith, Dr.', 'John Smith, Mr. Jr.'); segment -1 is wholly SUFFIX when it is nothing but suffix pieces ('Smith, Jr.', -'Smith, Ph. D. Jr.' -- the credential run C1 describes, in the listing -form), else gets leading titles, then the same trailing title run +holds no name word (titles and suffixes only) or a credential opens +it (#603), either of which fixed no family boundary -- there segment 0 +takes the NO_COMMA positional read instead, order and all ('John +Smith, Dr.', 'John Smith, Mr. Jr.', 'John Smith, PhD Jones'), given +two name words to read; segment 1 is wholly SUFFIX when it is nothing +but suffix pieces ('Smith, Jr.', 'Smith, Ph. D. Jr.' -- the credential +run C1 describes, in the listing form), titles and suffixes as #602's +run reads them when a credential opens it ('Smith, PhD Jones'), else +gets leading titles, then the same trailing title run over the pieces this segment does not read as suffixes ('Smith, John Prof.'), then given, then middles with those suffix-reading pieces to -suffix -- one predicate for both; segments 2+ are suffixes (lenient -- -segment already flagged non-suffixy ones COMMA_STRUCTURE). -SUFFIX_COMMA: segment 0 as NO_COMMA; segments 1+ wholly SUFFIX. +suffix -- one predicate for both; segments 2+ are suffixes but for +their title words (#603), lenient -- segment already flagged the ones +that are neither COMMA_STRUCTURE. +SUFFIX_COMMA: segment 0 as NO_COMMA; segment 1 wholly SUFFIX, and +segments 2+ as after a family comma. Emits PARTICLE_OR_GIVEN when the leading name piece is a lone particles_ambiguous token with more pieces following ("Van Johnson", and since #367 "Dr. Van Johnson" too, a title no longer displacing the particle out of that position) -- whatever role name_order assigns. -Emits SUFFIX_OR_NAME at SIX sites: the trailing roman numeral, each +Emits SUFFIX_OR_NAME at these sites: the trailing roman numeral, each ambiguous acronym the trailing peel had to resolve, the bare-suffix carve-out where an input that is nothing but post-nominal vocabulary gets its first word made into the name (H4's suffix half, #491), -- since #289 -- the FAMILY-COMMA path's own read of the first post-comma piece, -- since #531 -- the class member ENDING that path's given part, which the first-piece emitter could never reach, -and -- since #544 -- each member of a post-comma part read wholly as +-- since #544 -- each member of a post-comma part read wholly as credentials that was read as one only because a credential in front -of it anchors it. +of it anchors it, -- since #602 -- each name word the credential run +absorbs, in the no-comma name and in the given part after a family +comma, and -- since #603 -- each name word in a post-comma part a +credential opens. Further emitters of the same kind live in `_segment.py`, `_group.py` and `_post_rules.py`; they are not assign's and are not counted here. And @@ -574,10 +582,11 @@ def assign(state: ParseState) -> ParseState: # Mr.' has two pieces and one name, and read positionally lost # its family (the code review). anchored_picks: list[int] = [] + absorbed: list[int] = [] reading = segment_suffix_reading( state.pieces[1], state.piece_tags[1], tokens, state.policy.lenient_comma_suffixes, state.one_case, - anchored_picks) + anchored_picks, absorbed) # rules.md#C1's exception, scoped to the ambiguous credential # class: this is the first report of the comma's OWN decision # (listing or credential run), where the writing left the @@ -996,6 +1005,14 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: f"the comma is also an ordinary name word; read " f"as a credential", (i2,))) + # rules.md#C1: "A credential opening the part makes it + # the postnominal part however it goes on" (#603), and a + # word in it that would have made it the given part is + # reported, as #602's run reports the name words it + # takes + for k in absorbed: + if has_name_content(pieces[k], tokens): + ambiguities.append(_absorbed(pieces[k], tokens)) else: n = _peel_leading_titles(pieces, ptags, tokens) # rules.md#H5: "the title is TRANSPARENT to the suffix @@ -1044,14 +1061,17 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: for k in titled_idx: _set_roles(tokens, pieces[k], Role.TITLE) # v1 walk order: the first non-title piece is ALWAYS the - # given, before any suffix check -- 'Hardman, RN - CRNA' - # keeps first='RN'. The one deliberate 2.0 deviation, - # classified fix(comma-family) -- a last piece that is - # unambiguously suffix-shaped is a suffix, where v1 made - # it the given ('Andrews, M.D.', 'Smith, Dr. Jr.') -- is - # the no-name read above now: a segment whose only - # non-title piece is a suffix piece holds no name word, - # so the walk here never meets the case. + # given, before any suffix check -- 'Smith, V. Jones' keeps + # first='V.'. Two 2.x deviations are the segment reading + # above now (`reading is not None`), so the walk here never + # meets either: a last piece + # that is unambiguously suffix-shaped is a suffix, where v1 + # made it the given ('Andrews, M.D.', 'Smith, Dr. Jr.'; + # classified fix(comma-family)), a segment whose only + # non-title piece is a suffix piece holding no name word; + # and a credential opening the segment makes it the + # postnominal part whatever follows ('Hardman, RN - CRNA' + # kept first='RN' until #603). if n < len(pieces): _set_roles(tokens, pieces[n], Role.GIVEN) # The chain's floor leaves a name piece standing, so the @@ -1072,8 +1092,10 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: sticky_from = len(pieces) for m in range(n + 1, len(pieces)): tags = tokens[pieces[m][0]].tags - if (m not in titled_idx and "vocab:suffix" in tags - and tags.isdisjoint(_NOT_A_RUN_START) + if (m not in titled_idx + and ("suffix" in ptags[m] + or ("vocab:suffix" in tags + and tags.isdisjoint(_NOT_A_RUN_START))) and starts_a_credential_run(pieces[m], ptags[m], tokens)): sticky_from = m @@ -1212,9 +1234,14 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: # (`_vocab.surname_unit_facts`), over EVERY token, so a suffix # word still stops a particle ('van Jr. Berg, Mr.' is two name # words); a unit counts when a non-suffix piece holds a token - # of it. + # of it. One piece is one name word at most, and that is a + # DECISION, not only a saving: a connective join makes one piece + # of several units ('Vega y Lopez'), which the unit count alone + # would read positionally, losing the family to the given name + # behind a part a credential opens ('Vega y Lopez, PhD Jones'). + # Settled in C before any of that is built ('Smith, Jr.'). positional = False - if reading is not None: + if reading is not None and len(fam_pieces) > 1: idx = [i for piece in fam_pieces for i in piece] named = {i for k, piece in enumerate(fam_pieces) if not is_suffix_piece(piece, fam_tags[k], tokens) @@ -1252,10 +1279,28 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: else: _set_roles(tokens, piece, Role.FAMILY) tail = 2 - # segments past the structure's name segments are wholly suffixes + # rules.md#C2: segments past the structure's name segments are + # consumed as suffixes, except "a word of the title vocabulary there + # that is not also suffix vocabulary is a title" (#603) -- a part + # past the second comma holds no name word, so the surnames H5 + # keeps a bare trailing title word from misreading are not there to + # protect ('Eric H. Holder, Jr., Attorney General'). A + # piece holding a suffix word stays the postnominal the slot makes + # it, a dual alone as after a family comma (C1) and a dual group's + # connective join took into a title ('PhD - and MD') alike. for seg_idx in range(tail, len(state.segments)): - for piece in state.pieces[seg_idx]: - _set_roles(tokens, piece, Role.SUFFIX) + for piece, ptags_ in zip(state.pieces[seg_idx], + state.piece_tags[seg_idx]): + # the tag tests inline ahead of the predicate, so a tail of + # suffix words pays no frame for the question + _set_roles(tokens, piece, + Role.TITLE + if (("title" in ptags_ + or "vocab:title" in tokens[piece[0]].tags) + and is_title_piece(piece, ptags_, tokens) + and not any("vocab:suffix" in tokens[i].tags + for i in piece)) + else Role.SUFFIX) return copy_with(state, tokens=tuple(tokens), order=order, ambiguities=tuple(ambiguities)) diff --git a/nameparser/_pipeline/_group.py b/nameparser/_pipeline/_group.py index 4e8b6c1d..6bbe990c 100644 --- a/nameparser/_pipeline/_group.py +++ b/nameparser/_pipeline/_group.py @@ -1282,9 +1282,10 @@ def group(state: ParseState) -> ParseState: # Suppressed after a family comma for the same reason _assign # suppresses it there: the family name is already fixed, so # there is no fork left to report. Suppressed in a tail segment - # as well, after either comma: assign reads that segment - # wholly as suffixes, a marker standing in it being an ordinary - # word (rules.md#M2, #601), so a + # as well, after either comma: assign reads that segment as + # suffixes and its title words as titles (#603), never as a + # name, a marker standing in it being an ordinary word + # (rules.md#M2, #601), so a # chain report there -- a particle # chained onto a name piece, or an acronym taken into the name # -- names a reading the parse never takes. rules.md#C2: "a part diff --git a/nameparser/_pipeline/_pieces.py b/nameparser/_pipeline/_pieces.py index ac75196d..42be7fd1 100644 --- a/nameparser/_pipeline/_pieces.py +++ b/nameparser/_pipeline/_pieces.py @@ -277,7 +277,23 @@ def starts_a_credential_run(piece: Sequence[int], ptags: Set[str], than a numeral, and carries no `initial` tag), a connective, a word that is also bound given-name or particle vocabulary (`abd` heads `Abd Allah`; `vd`, `mc`), and the non-Latin honorific words. + + The split credential group merges ('Ph. D.', the one piece carrying + the "suffix" piece tag) starts it as its one-token spelling 'PhD' + does: a run that started for one spelling of a word and not the + other is the vocabulary-against-shape split AGENTS.md's spelling + sweep names (#603 found it, 'Smith, John Ph. D. Jones' keeping + middle 'Jones' where 'Smith, John PhD Jones' reads suffix + 'PhD Jones'). The no-comma spelling is not reached: `run_start` + asks only of the pieces `peel_walk` keeps, and the walk drops the + merged piece, so 'John Smith Ph. D. Jones' still reads family + 'Jones'. """ + if "suffix" in ptags: + # the Ph./D. pair alone: a connective join keeps the tag on the + # wider piece ('Ph. D. and Mary'), which starts nothing, as + # 'PhD and Mary' does not + return len(piece) == 2 if len(piece) != 1 or not is_suffix_piece(piece, ptags, tokens): return False tok = tokens[piece[0]] @@ -489,6 +505,7 @@ def segment_suffix_reading(pieces: Sequence[Sequence[int]], lenient: bool, one_case: bool | None, anchored: list[int] | None = None, + absorbed: list[int] | None = None, ) -> tuple[bool, ...] | None: """How each piece of a no-name segment reads: True a suffix, False a title. None when the segment holds a name word and so is not a @@ -543,8 +560,27 @@ class by SHAPE takes the count instead, which is decided at the initial there being no shape anyone writes (#430). A title resets that: what follows a bare title is not continuing a credential. + A part OPENED by a credential is the postnominal part however it + goes on (#603, rules.md#C1): where an unambiguous suffix word that + starts #602's run (`starts_a_credential_run`) is the first suffix + word of the part, with only titles in front of it (a title/suffix + dual among them ends the chance: a credential behind one opens + nothing), every later + word is read as #602's run reads the given part -- a title word as + a title, anything else as a suffix -- so 'John Smith, PhD Jones' + and 'Smith, PhD Jones' read suffix 'PhD Jones'. A word of the + title vocabulary as well opens nothing (#603, Derek): in front of + a given name it is the listing form's title ('Smith, Ms Jane'), + and no count of the words before the comma can tell a two-word + surname from a given name and a family name, so none decides it. + `absorbed`, when passed, receives the index + of every piece the opened part takes that would otherwise have + made this None -- the caller reports them, as #602's run reports + the name words it takes. + None covers both ways a segment can fail to be a run: a name word - anywhere in it, and no pieces at all ('Doe,, Jr.', which holds no + anywhere in it (in a part no credential opened), and no pieces at + all ('Doe,, Jr.', which holds no title to read by). `lenient` is Policy.lenient_comma_suffixes, and only the numeral @@ -584,6 +620,14 @@ class by SHAPE takes the count instead, which is decided at the # MD Ma' keeps its anchored 'Ma' only by this walk's reading). leading = True dual_led = False + # #603: the first suffix word of the leading run, which may open + # the part unless a dual stood ahead of it; asked whether it does + # only once a piece reaches the title and name branches below, so a + # part of nothing but suffix words ('Smith, Jr.', 'Smith, PhD MA') + # never pays for the question + opener: tuple[Sequence[int], Set[str]] | None = None + asked = False + opened = False for piece, tags in zip(pieces, ptags): # the verdict just recorded IS "stands behind a suffix" -- keeping # a separate flag meant maintaining that equality by hand at three @@ -595,6 +639,8 @@ class by SHAPE takes the count instead, which is decided at the and "vocab:title" in tokens[piece[0]].tags): dual_led = True else: + if leading and not dual_led: + opener = (piece, tags) leading = False out.append(True) continue @@ -621,6 +667,19 @@ class by SHAPE takes the count instead, which is decided at the and _numeral_behind_the_initial_veto(piece, tokens)): leading = False out.append(True) + continue + if not asked and opener is not None: + asked = True + opened = starts_a_credential_run(opener[0], opener[1], tokens) + if opened: + # #602's run: a title word reads as a title, anything else + # as a suffix -- vocabulary only, as in the given part + if is_title_piece(piece, tags, tokens): + out.append(False) + else: + if absorbed is not None: + absorbed.append(len(out)) + out.append(True) elif is_leading_title(piece, tags, tokens): out.append(False) else: diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index fb6cc763..494af780 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -5,7 +5,8 @@ Produces: segments (runs of main-token indices; interior segments may be EMPTY -- doubled commas keep their structural position), structure, one_case where the comma form asked for it, COMMA_STRUCTURE -ambiguities for unrecognized extra segments, and SUFFIX_OR_NAME where +ambiguities for unrecognized extra segments (a part of title and +suffix words is recognized, #603), and SUFFIX_OR_NAME where the comma FLIPPED the structure for a member of the ambiguous credential class, or for a run of suffix words holding one (#544). The flip and nothing else: where the structure did @@ -18,7 +19,10 @@ strict token test; Policy.extra_suffix_delimiters gives v1 suffix_delimiter parity, a delimiter-core token being transparent); Lexicon.maiden_markers DIRECTLY, for the own-words span the lazy case -gate takes (_pieces.own_words); Lexicon title and suffix vocabulary +gate takes (_pieces.own_words); Lexicon.titles DIRECTLY too, for +C2's test that a tail part is titles and suffixes (#603), with +_vocab.period_joined_vocab for a period-joined title; Lexicon +title and suffix vocabulary plus Policy.lenient_comma_suffixes again through _vocab. name_word_count, which counts NAME words for the class's own comma rule; and, since 2.4, Policy.unlisted_dotted_suffixes through @@ -55,6 +59,7 @@ ) from nameparser._pipeline._vocab import ( ambiguous_class_candidate, ambiguous_class_member, ambiguous_lean, + period_joined_vocab, caps_shape_candidate, is_one_case, is_paired_initials, claimed_as_non_name, is_single_letter_numeral, is_wholly_suffix, written_as_a_name, @@ -215,6 +220,19 @@ def class_run(seg: tuple[int, ...]) -> bool: state.lexicon) for i in seg) + def titles_and_suffixes(seg: tuple[int, ...]) -> bool: + # rules.md#C2: "a word of the title vocabulary there that is not + # also suffix vocabulary is a title" -- at least one title + # word, and every other word a suffix word + # period_joined_vocab for 'Lt.Gov.', one token that classify + # (not yet run) tags a title by this same call, and that assign + # then reads as one + titles = state.lexicon.titles + rest = tuple(i for i in seg + if _normalize(t := state.tokens[i].text) not in titles + and period_joined_vocab(t, state.lexicon) != "title") + return len(rest) < len(seg) and (not rest or suffixy(rest)) + # rules.md#C1: "the name reads as trailing suffixes when the part # after the first comma is entirely suffix words and more than one # word precedes the comma; otherwise it reads as the listing form" @@ -549,9 +567,10 @@ def class_run(seg: tuple[int, ...]) -> bool: groups[1])) # rules.md#C2: "a non-empty extra part that is not entirely suffix # words is flagged as a structural ambiguity rather than rejected" - # -- parts[2:] are consumed as suffixes unconditionally either - # way, so a non-suffix tail segment gets the COMMA_STRUCTURE - # flag, not a structure veto. The lean reaches this reading too: + # -- parts[2:] are consumed as suffixes either way, their title + # words as titles (#603), so a tail segment that is neither all + # suffix words nor title words beside suffix words gets the + # COMMA_STRUCTURE flag, not a structure veto. The lean reaches this reading too: # a tail of leaning credentials is a credential run, which is the # one place this design quiets a report rather than adding one. for seg in groups[2:]: @@ -569,8 +588,15 @@ def class_run(seg: tuple[int, ...]) -> bool: # further out: it is the only one of the three that walks the # by-shape class, and it is asked only of a run BOTH suffix # readings have already declined. + # #603: a part assign reads as titles and suffixes is + # recognized too -- word by word, off the title vocabulary + # (classify has not run yet), so a part holding a connective ('Secretary of State') keeps the + # flag: knowing that group's title chain takes it would be a + # model of group, the shape mechanisms.md's + # READ-WITHOUT-THEN-BIND exists to remove. Last, as the + # rarest reading. if (seg and not suffixy(seg) and not suffixy(seg, case_class()) - and not class_run(seg)): + and not class_run(seg) and not titles_and_suffixes(seg)): texts_joined = " ".join(texts(seg)) ambiguities.append(PendingAmbiguity( AmbiguityKind.COMMA_STRUCTURE, diff --git a/tests/test_suffixes.py b/tests/test_suffixes.py index f97f3e83..91eb240f 100644 --- a/tests/test_suffixes.py +++ b/tests/test_suffixes.py @@ -301,12 +301,17 @@ def test_suffix_delimiter_constants_level(self) -> None: self.m(hn.suffix, "RN, CRNA", hn) def test_suffix_delimiter_none_by_default_known_limitation(self) -> None: - # Without suffix_delimiter set, " - " between suffixes breaks parsing. - # This test documents the known limitation — do not "fix" it. + # Without suffix_delimiter set, " - " is a word, not a separator. + # Through 2.3 (and in v1) that broke parsing -- first 'RN', last + # 'Steven Hardman', suffix 'CRNA' -- and this test pinned the + # limitation. Since #603 a credential opening the part after the + # comma makes it the suffix part (rules.md#C1), so the undeclared + # delimiter is one more word of that part and the name parses; + # the name of the test keeps the history. hn = HumanName("Steven Hardman, RN - CRNA") - self.m(hn.first, "RN", hn) - self.m(hn.last, "Steven Hardman", hn) - self.m(hn.suffix, "CRNA", hn) + self.m(hn.first, "Steven", hn) + self.m(hn.last, "Hardman", hn) + self.m(hn.suffix, "RN - CRNA", hn) def test_suffix_delimiter_trailing_delimiter_ignored(self) -> None: # Trailing delimiter must not defeat suffix detection. Using a diff --git a/tests/v2/cases.py b/tests/v2/cases.py index e8b384f2..b83811e7 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -2102,16 +2102,19 @@ def _check_cjk_shape_purity(self) -> None: "untitled 'St John Smith' into one given name"), Case("a_part_past_the_second_reports_no_particle_fork", "John Smith, Jr., Freiherr von Richthofen", - {"given": "John", "family": "Smith", - "suffix": "Jr., Freiherr von Richthofen"}, + {"given": "John", "family": "Smith", "title": "Freiherr", + "suffix": "Jr., von Richthofen"}, ambiguities=("comma-structure",), notes="the row above as a part past the second comma, which " - "is consumed wholly as suffixes (C2): 'von' still " - "chains in group, but it ends a suffix, so the " - "particle-or-given report that 'von' was chained onto " - "a name piece named a reading the parse never makes. " - "Every 2.x release through 2.3.0 carried it; the " - "comma-structure flag is the part's own report"), + "is consumed as suffixes but for its title words (C2): " + "'von' still chains in group, but it ends a suffix, so " + "the particle-or-given report that 'von' was chained " + "onto a name piece named a reading the parse never " + "makes. Every 2.x release through 2.3.0 carried it; the " + "comma-structure flag is the part's own report. Since " + "#603 'Freiherr', a title word, reads as the title it " + "is there; until then the part read wholly as suffix " + "'Jr., Freiherr von Richthofen'"), Case("titled_ambiguous_particle_no_op_chain", "St Van Jr.", {"title": "St", "family": "Van", "suffix": "Jr."}, notes="the piece after the particle is a suffix, so the chain " @@ -6122,11 +6125,101 @@ def _check_cjk_shape_purity(self) -> None: notes="v1 #144: the trailing piece of a two-part comma name " "takes the lenient suffix test"), Case("family_comma_first_piece_is_given", "Steven Hardman, RN - CRNA", - {"given": "RN", "middle": "-", "family": "Steven Hardman", - "suffix": "CRNA"}, - notes="v1 walk order: the first post-comma piece is the given " - "before any suffix check (the delimiter is UNSET here " - "-- v1's documented limitation, kept)"), + {"given": "Steven", "family": "Hardman", "suffix": "RN - CRNA"}, + classification="fix(#603)", + notes="a credential opening the part after the comma makes it " + "the postnominal part (rules.md#C1, #603), so the " + "pre-comma name keeps its positional read and the " + "undeclared delimiter is one more word of the run. v1 and " + "2.0 through 2.3 read v1's walk order here, the first " + "post-comma piece the given before any suffix check -- " + "given 'RN', middle '-', family 'Steven Hardman', the " + "documented limitation of an UNSET delimiter -- which the " + "row id still names"), + Case("a_credential_opening_the_comma_part_needs_an_unambiguous_one", + "John Smith, Ma Jones", + {"given": "Ma", "middle": "Jones", "family": "John Smith"}, + ambiguities=("suffix-or-name",), + notes="the boundary of #603's opened part (rules.md#C1): only " + "a word that starts S2's run opens it, and a member of " + "the ambiguous credential class does not, so the part " + "keeps the listing form's given name and the first " + "post-comma piece reports the fork it read (#289)"), + Case("a_dual_opens_no_part_behind_a_two_word_surname", + "García Márquez, Ms Gabriela", + {"title": "Ms", "given": "Gabriela", "family": "García Márquez"}, + notes="#603's boundary for a title/suffix dual (rules.md#C1): it " + "opens nothing, being the listing form's title in front " + "of a given name, and no count before the comma tells a " + "two-word surname from a given name and a family name. A " + "first draft opened the part behind two words and read " + "given 'García', suffix 'Ms Gabriela' here; Derek chose " + "that duals never open (2026-10-04)"), + Case("the_split_credential_starts_the_given_part_run", + "Smith, John Ph. D. Jones", + {"given": "John", "family": "Smith", "suffix": "Ph. D. Jones"}, + ambiguities=("suffix-or-name",), + classification="fix(#603)", + notes="#602's run in the given part after a family comma, " + "started by the split credential group merges: " + "starts_a_credential_run refused any piece longer than " + "one token, so 'Ph. D.' started no run where 'PhD' did, " + "and this read middle 'Jones', suffix 'Ph. D.' (2.3.0 " + "too). Found by #603, whose opened part asks the same " + "predicate (AGENTS.md's spelling sweep)"), + Case("a_split_credential_joined_to_a_name_opens_nothing", + "John Smith, Ph. D. and Mary Jones", + {"given": "Ph. D. and Mary", "middle": "Jones", + "family": "John Smith"}, + notes="the two spellings agree (rules.md#C1, #603): group's " + "connective join keeps the merged credential's 'suffix' " + "piece tag on the wider piece, and starts_a_credential_run " + "accepted any piece so tagged, so this read suffix " + "'Ph. D. and Mary Jones' with 'Mary' taken in silently " + "while 'John Smith, PhD and Mary Jones' reads as here; the " + "PR review found it, and only the bare Ph./D. pair starts " + "a run now"), + Case("a_one_piece_name_before_an_opened_part_is_the_family", + "Vega y Lopez, PhD Jones", + {"family": "Vega y Lopez", "suffix": "PhD Jones"}, + ambiguities=("suffix-or-name",), + classification="fix(#603)", + notes="one piece before the comma is one name word, even where " + "a connective joined three units into it; without " + "assign's len(fam_pieces) > 1 exit the unit count reads " + "it positionally, given 'Vega y Lopez' and no family. " + "2.3.0 read given 'PhD', middle 'Jones'"), + Case("only_the_first_suffix_word_can_open_the_comma_part", + "Smith, MA PhD Jones", + {"given": "MA", "family": "Smith", "suffix": "PhD Jones"}, + ambiguities=("suffix-or-name", "suffix-or-name"), + notes="#603's opener is the part's FIRST suffix word with only " + "titles in front of it (rules.md#C1): 'MA', a member of " + "the ambiguous class, opens nothing and ends the leading " + "run, so the 'PhD' behind it opens nothing either and the " + "part is the walk's. Letting a later credential open it " + "reads suffix 'MA PhD Jones'"), + Case("a_misplaced_generation_opens_the_comma_part", "Smith, Jr. John", + {"family": "Smith", "suffix": "Jr. John"}, + ambiguities=("suffix-or-name",), + classification="fix(#603)", + notes="decisions.md#C1's accepted cost: a generational word " + "starts S2's run, so it opens the part and 'John' is " + "absorbed, one word before the comma reading the whole " + "part as suffixes (Derek). 2.3.0 read title 'Jr.', given " + "'John'"), + Case("a_part_of_titles_behind_a_credential_was_already_postnominal", + "Eric H. Holder, Jr. Attorney General", + {"title": "Attorney General", "given": "Eric", "middle": "H.", + "family": "Holder", "suffix": "Jr."}, + notes="the shape #603 generalizes: a part after the comma " + "holding no name word fixed no family boundary, so the " + "name before it keeps its positional read and the part " + "reads as its titles and suffixes. Unmoved by #603, which " + "extends that reading to a part a credential opens " + "whatever stands behind it ('Eric H. Holder, Jr. Chief " + "Justice', 'justice' being no title word, reads suffix " + "'Jr. Justice', title 'Chief')"), Case("family_comma_lone_suffix_piece", "Andrews, M.D.", {"family": "Andrews", "suffix": "M.D."}, classification="fix(comma-family)", @@ -8701,31 +8794,23 @@ def _check_cjk_shape_purity(self) -> None: "there is no abbreviation, so the numeral is the " "generation it looks like. This row is what makes " "#432's fix a period test rather than a numeral test"), - Case("family_comma_title_resets_the_credential_run", "Smith, PSM Dr. I", - {"title": "Dr.", "given": "PSM", "family": "Smith", - "suffix": "I"}, - classification="fix(#316)", + Case("family_comma_title_resets_the_credential_run", "Smith, MD Dr. I", + {"title": "MD Dr.", "given": "I", "family": "Smith"}, notes="THE RESET. A title ends the run: what follows a bare " "title is not continuing a credential, so the numeral " "behind it does not join and the segment is no run at " - "all. Removing that one line leaves the whole suite " - "green while this becomes title 'Dr.' + suffix 'PSM " - "I' -- the reset fires across the suite and until this " - "row no input observed it, which is the " - "inert-measurement shape. Since #316 the word it " - "resets ON is a title here rather than a middle name: " - "'I' is what this segment reads as its suffix, so " - "'Dr.' is the trailing piece and the walk takes it. " - "Transparency does NOT reach this row, and that is the " - "reset itself: 'Smith, PSM I' reads suffix 'PSM I' " - "with no given name at all (measured 2026-09-09), " - "because with no title between them the numeral " - "CONTINUES the credential run. Removing the title " - "removes the reset, so the shorter spelling is a " - "different reading and not this one minus a word. " - "1.4.0 read suffix 'Dr., I' -- 'dr' was still " - "postnominal vocabulary before #296's audit, so the " - "row's old parity claim had outlived it"), + "all, leaving the walk to read the dual and the title as " + "the given part's titles and 'I' as its given name. " + "Without the reset the numeral continues the 'MD' in " + "front of the title and the part reads title 'Dr.', " + "suffix 'MD I'. Spelled with the dual 'MD' since #603: " + "'Smith, PSM Dr. I', this row until then, is a part the " + "credential 'PSM' OPENS, which reads as #602's run reads " + "whatever stands behind it (rules.md#C1) -- title 'Dr.', " + "suffix 'PSM I' -- so the reset no longer decides it. A " + "dual opens nothing, and the reset is what is left " + "deciding this spelling. 1.4.0 read the " + "old spelling suffix 'Dr., I'"), Case("family_comma_run_numeral_after_a_split_credential", "Smith, Ph. D. I", {"family": "Smith", "suffix": "Ph. D. I"}, @@ -8754,13 +8839,23 @@ def _check_cjk_shape_purity(self) -> None: "about, and no longer decides the separator"), Case("family_comma_strict_keeps_the_initial_veto", "Smith, PSM I.", - {"given": "PSM", "family": "Smith", "suffix": "I."}, + {"family": "Smith", "suffix": "PSM I."}, + ambiguities=("suffix-or-name",), policy=Policy(lenient_comma_suffixes=False), - notes="C1's strict knob still vetoes initial-shaped words, so " - "the run ends at the numeral where lenient continues " - "through it. #430's first draft read no policy at all " - "and silently overrode the one knob a caller sets to " - "prevent exactly this; nothing in the suite saw it"), + classification="fix(#603)", + notes="C1's strict knob still vetoes initial-shaped words as " + "SUFFIX words, so the numeral does not continue the run " + "as a suffix word (#430's first draft read no policy at " + "all and silently overrode the knob; " + "test_strict_ends_the_run_at_the_initial_shaped_numeral " + "pins that reading in test_pieces.py). Since #603 the " + "credential 'PSM' opens the part, and the run it opens " + "takes 'I.' in as a word, as #602's run takes it in the " + "given part under strict ('Smith, John PhD V.'), and " + "reports it, a word strict would otherwise read as a " + "name, as the no-comma run reports 'John Smith PhD V.'. " + "Until then this row read given 'PSM', family 'Smith', " + "suffix 'I.'"), Case("family_comma_run_with_a_name_is_not_a_run", "Smith, John Jr.", {"given": "John", "family": "Smith", "suffix": "Jr."}, notes="the non-flip: a name word in the run makes it the " @@ -9250,17 +9345,16 @@ def _check_cjk_shape_purity(self) -> None: "measured 2026-09-09 in test_pieces.py)"), Case("family_comma_then_a_lone_suffix_word_segment", "Smith, John, Prof.", - {"given": "John", "family": "Smith", "suffix": "Prof."}, - ambiguities=("comma-structure",), - classification="parity", + {"given": "John", "family": "Smith", "title": "Prof."}, + classification="fix(#603)", notes="the trailing slot is a segment away: a second comma " "makes the last part its own segment, which the tail " - "consumes as a suffix before any trailing walk reads a " - "piece -- so 'prof' leaving the suffix vocabulary " - "(#296) does not reach this shape and 'Smith, John, " - "Prof.' still reads suffix 'Prof.' where 'Smith, John " - "Prof.' reads title. Unmoved by this bundle and by " - "1.4.0 alike (measured 2026-09-09)"), + "reads before any trailing walk reads a piece. Since " + "#603 a title word there is a title (rules.md#C2), so " + "'Smith, John, Prof.' reads title 'Prof.' as 'Smith, " + "John Prof.' does, and the part is recognized rather " + "than flagged. Until then, and in 1.4.0, it read suffix " + "'Prof.' with a comma-structure flag"), # -- #271: script-scoped order + segmentation (amendment 2026-07-27) Case("ko_unspaced_default", "김민준", diff --git a/tests/v2/pipeline/test_pieces.py b/tests/v2/pipeline/test_pieces.py index e643b07f..42d3728b 100644 --- a/tests/v2/pipeline/test_pieces.py +++ b/tests/v2/pipeline/test_pieces.py @@ -129,12 +129,23 @@ def test_strict_ends_the_run_at_the_initial_shaped_numeral() -> None: Lenient continues the credential run through a one-character suffix word; strict vetoes initial-shaped words, so the run is no - run at all and the segment falls to the walk. + run at all and the segment falls to the walk. Asked behind a + title/suffix dual, which opens no part (#603). Behind a credential that does + open it ('PSM'), strict refuses the numeral the suffix reading but + the opened part takes it in all the same, and hands it to the + caller to report. """ - state = _through_group("Smith, PSM I") + state = _through_group("Smith, MD I") args = (state.pieces[1], state.piece_tags[1], list(state.tokens)) assert segment_suffix_reading(*args, True, state.one_case) == (True, True) assert segment_suffix_reading(*args, False, state.one_case) is None + state = _through_group("Smith, PSM I") + args = (state.pieces[1], state.piece_tags[1], list(state.tokens)) + for lenient, taken in ((True, []), (False, [1])): + absorbed: list[int] = [] + assert segment_suffix_reading(*args, lenient, state.one_case, + None, absorbed) == (True, True) + assert absorbed == taken def _comma_part_reading(text: str) -> tuple[tuple[bool, ...] | None, diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index 2db0ac9e..5626aa7d 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -2270,6 +2270,26 @@ class _LatinCopy(NamedTuple): #: question someone answers in writing, not something to skip past. _NOT_A_VOCABULARY_COPY = frozenset({ frozenset({"^", " "}), # the honorific rule's leading anchor + # #603's two rules (2026-10-04), one alternative per corpus name: + # what selects them is the SHAPE -- a credential opening the part + # after the comma, a title word past the second -- and no wordlist. + frozenset({"Doe, PhD Jones", "John Smith, PhD Jones", + "John Smith, Ph\\. D\\. Jones", "Smith, RN - CRNA", + "Steven Hardman, RN - CRNA"}), + # ... and the title rule, without 'John, Smith, Dr.' where the + # 2.0.0 and 2.1.0 ledgers' #296 rule already claims that name + frozenset({"1 & 2, 3 4 5, Mr\\.", + "Andrew Perkins, Jr\\., Col\\. \\x28Ret\\x29", + "Eric H\\. Holder, Jr\\., Attorney General", + "Eric H\\. Holder, Jr\\., Secretary of State", + "John Smith, Jr\\., Lt\\.Gov\\.", + "Smith, John, Prof\\."}), + frozenset({"1 & 2, 3 4 5, Mr\\.", + "Andrew Perkins, Jr\\., Col\\. \\x28Ret\\x29", + "Eric H\\. Holder, Jr\\., Attorney General", + "Eric H\\. Holder, Jr\\., Secretary of State", + "John Smith, Jr\\., Lt\\.Gov\\.", + "John, Smith, Dr\\.", "Smith, John, Prof\\."}), # The #533 review's two literal-anchored rules, one alternative # per corpus name. Lists of names, not copies of any wordlist: # what selects them is the SHAPE the clause's decline turns on -- @@ -3820,6 +3840,7 @@ def _claim(rule: dict) -> _Claim: # 2026-10-04, #601/#602: 443 -> 438; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. # 2026-10-04, #601/#602 review: 438 -> 439, rules.md#S2's new # example 'Smith, John PhD de Jr.'. Reach, verified by name. + # 2026-10-04, #603: 441 -> 452; the rules.md#C1 and #C2 examples #603 added entered the corpus. "fix(comma-family) lone post-comma piece routes to suffix/title, not first": # 2026-10-01, #564: 423 -> 430, 'John Smith, XYZ', 'Smith, # XYZ', 'García Márquez, MJ', 'García Márquez, MJ PhD', @@ -3843,7 +3864,7 @@ def _claim(rule: dict) -> _Claim: # rules.md#H1 comma example. Reach. # 2026-10-04, merging #606 in as well: 441 over the merged # corpus. Reach. - _Claim(441, ('given', 'suffix', 'title'), "a3028e60b1e7", None), + _Claim(452, ('given', 'suffix', 'title'), "256da88fcab5", None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3952,6 +3973,7 @@ def _claim(rule: dict) -> _Claim: # 2026-10-04, #601/#602: 443 -> 438; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. # 2026-10-04, #601/#602 review: 438 -> 439, rules.md#S2's new # example 'Smith, John PhD de Jr.'. Reach, verified by name. + # 2026-10-04, #603: 441 -> 452; the rules.md#C1 and #C2 examples #603 added entered the corpus. "fix(comma-precomma-family) pre-comma run reads as family, not given": # 2026-10-01, #564: 423 -> 430, 'John Smith, XYZ', 'Smith, # XYZ', 'García Márquez, MJ', 'García Márquez, MJ PhD', @@ -3969,7 +3991,7 @@ def _claim(rule: dict) -> _Claim: # 2026-10-04, #606: 443 -> 444, 'Greve, Anna'. Reach. # 2026-10-04, merging #606 in as well: 441 over the merged # corpus. Reach. - _Claim(441, ('family', 'given'), "a3028e60b1e7", None), + _Claim(452, ('family', 'given'), "256da88fcab5", None), # 2026-10-01, #575: new, 4; 'De La Cruz, Ed', 'Freiherr von # Berg, Ed', 'Van Buren, Ed', 'de la Cruz, Ma'. "fix(#575) a particle surname before a comma is one name word": @@ -4243,8 +4265,9 @@ def _claim(rule: dict) -> _Claim: # 2026-10-03, #549: 105 -> 106, 'Smith, John, PhD - and MD'. # Reach. # 2026-10-04, #601/#602: 106 -> 107; 'Jane and née Jones', a new rules.md#M2 example, is in reach -- a marker behind a connective is a word in the run since #601. + # 2026-10-04, #603: 107 -> 108; the rules.md#C1 and #C2 examples #603 added entered the corpus. "fix(initials-per-word) a connective run initials each word (facade, since 2.0.0)": - _Claim(107, ('_initials',), "9a8658d181fb", ('DEFAULT',)), + _Claim(108, ('_initials',), "b398136feeef", ('DEFAULT',)), # 2026-09-19, #533: 41 -> 43. Two new corpus names opening # with a bound-given word, 'Berg, abdul MA' and 'Berg, abdul # nee Jones MA' -- the P5 pair this change added to record @@ -4278,8 +4301,9 @@ def _claim(rule: dict) -> _Claim: _Claim(123, ('_initials',), "d69d11083dba", ('DEFAULT',)), # 2026-09-23, #459: 18 -> 19, 'john smith ph. d.', rules.md#R4's # two-token line. Reach, verified name by name. + # 2026-10-04, #603: 19 -> 20; the rules.md#C1 and #C2 examples #603 added entered the corpus. "fix(initials-per-word) the Ph. D. merge initials each word (facade, since 2.0.0)": - _Claim(19, ('_initials',), "adfdec5a9e13", ('DEFAULT',)), + _Claim(20, ('_initials',), "d89231d36185", ('DEFAULT',)), # The 2.3 title-run bundle's five rules, last in every # ledger. All five are anchored on NAMES, so the reach IS the # mover list: 2 names for the run keying, 1 for the esq drop, @@ -4579,6 +4603,12 @@ def _claim(rule: dict) -> _Claim: # Accepted example for a surname-borne title. "fix(#606) salutations and titles in other languages": _Claim(1, ('given', 'title'), "0c78dbac8a6b", None), + # 2026-10-04, #603: new, 5; rules.md#C1's opened-part examples and the delimiter rows it moves. + "fix(#603) a credential opening the part after the comma makes it the postnominal part": + _Claim(5, ('family', 'given', 'middle', 'suffix', 'title'), "a4ebde5b8cd7", None), + # 2026-10-04, #603: new, 7; rules.md#C2's tail-title examples and the radar names it moves. + "fix(#603) a title word in a part past the second comma is a title": + _Claim(7, ('suffix', 'title'), "8cbad6f13c6b", None), }, "expected_since_2.0.0.toml": { # The ph removal (#459/#521): one literal name, the cases.py @@ -4910,8 +4940,11 @@ def _claim(rule: dict) -> _Claim: _Claim(1, ('family', 'suffix', 'title'), "01bf2bd3f895", None), "fix(#296) do is a name, so it no longer stops the leading-particle scan as a title": _Claim(1, ('family', 'given'), "faa2c70fc49e", None), + # 2026-10-04, #603: roles lost `_ambiguities`; a title word past + # the second comma is a title now, so 'John, Smith, Dr.' raises + # no comma-structure report and the rule's fields narrowed. "fix(#296) dr is not postnominal vocabulary, so 'John Smith, Dr.' keeps its split and its title": - _Claim(2, ('_ambiguities', 'suffix', 'title'), "34d3d96adb65", ('DEFAULT', 'FAMILY_FIRST')), + _Claim(2, ('suffix', 'title'), "34d3d96adb65", ('DEFAULT', 'FAMILY_FIRST')), # 2026-09-18: 1 -> 2. One new corpus name, 'Jack MA.', the # all-caps spelling #289's case contrast now reads as a # credential; the rule's own 'Jack Ma.' is untouched. @@ -5322,6 +5355,12 @@ def _claim(rule: dict) -> _Claim: # Accepted example for a surname-borne title. "fix(#606) salutations and titles in other languages": _Claim(1, ('given', 'title'), "0c78dbac8a6b", None), + # 2026-10-04, #603: new, 5; rules.md#C1's opened-part examples and the delimiter rows it moves. + "fix(#603) a credential opening the part after the comma makes it the postnominal part": + _Claim(5, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "a4ebde5b8cd7", None), + # 2026-10-04, #603: new, 6; rules.md#C2's tail-title examples and the radar names it moves. + "fix(#603) a title word in a part past the second comma is a title": + _Claim(6, ('_ambiguities', 'suffix', 'title'), "eeb11557d7f0", None), }, # The 2.3 cycle's first rule, and a facade-only render fix: every # role is identical, so `_initials` alone. Reach and digest as in @@ -5788,6 +5827,12 @@ def _claim(rule: dict) -> _Claim: # Accepted example for a surname-borne title. "fix(#606) salutations and titles in other languages": _Claim(1, ('given', 'title'), "0c78dbac8a6b", None), + # 2026-10-04, #603: new, 5; rules.md#C1's opened-part examples and the delimiter rows it moves. + "fix(#603) a credential opening the part after the comma makes it the postnominal part": + _Claim(5, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "a4ebde5b8cd7", None), + # 2026-10-04, #603: new, 7; rules.md#C2's tail-title examples and the radar names it moves. + "fix(#603) a title word in a part past the second comma is a title": + _Claim(7, ('_ambiguities', 'suffix', 'title'), "8cbad6f13c6b", None), }, "expected_since_2.1.0.toml": { # The ph removal (#459/#521): one literal name, the cases.py @@ -6072,8 +6117,11 @@ def _claim(rule: dict) -> _Claim: _Claim(1, ('suffix', 'title'), "01bf2bd3f895", None), "fix(#296) do is a name, so it no longer stops the leading-particle scan as a title": _Claim(1, ('family', 'given'), "faa2c70fc49e", None), + # 2026-10-04, #603: roles lost `_ambiguities`; a title word past + # the second comma is a title now, so 'John, Smith, Dr.' raises + # no comma-structure report and the rule's fields narrowed. "fix(#296) dr is not postnominal vocabulary, so 'John Smith, Dr.' keeps its split and its title": - _Claim(2, ('_ambiguities', 'suffix', 'title'), "34d3d96adb65", ('DEFAULT', 'FAMILY_FIRST')), + _Claim(2, ('suffix', 'title'), "34d3d96adb65", ('DEFAULT', 'FAMILY_FIRST')), # 2026-09-18: 1 -> 2. One new corpus name, 'Jack MA.', the # all-caps spelling #289's case contrast now reads as a # credential; the rule's own 'Jack Ma.' is untouched. @@ -6479,6 +6527,12 @@ def _claim(rule: dict) -> _Claim: # Accepted example for a surname-borne title. "fix(#606) salutations and titles in other languages": _Claim(1, ('given', 'title'), "0c78dbac8a6b", None), + # 2026-10-04, #603: new, 5; rules.md#C1's opened-part examples and the delimiter rows it moves. + "fix(#603) a credential opening the part after the comma makes it the postnominal part": + _Claim(5, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "a4ebde5b8cd7", None), + # 2026-10-04, #603: new, 6; rules.md#C2's tail-title examples and the radar names it moves. + "fix(#603) a title word in a part past the second comma is a title": + _Claim(6, ('_ambiguities', 'suffix', 'title'), "eeb11557d7f0", None), }, "expected_since_2.3.0.toml": { # The ph removal (#459/#521): one literal name, the cases.py @@ -6798,6 +6852,12 @@ def _claim(rule: dict) -> _Claim: # Accepted example for a surname-borne title. "fix(#606) salutations and titles in other languages": _Claim(1, ('given', 'title'), "0c78dbac8a6b", None), + # 2026-10-04, #603: new, 5; rules.md#C1's opened-part examples and the delimiter rows it moves. + "fix(#603) a credential opening the part after the comma makes it the postnominal part": + _Claim(5, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "a4ebde5b8cd7", None), + # 2026-10-04, #603: new, 7; rules.md#C2's tail-title examples and the radar names it moves. + "fix(#603) a title word in a part past the second comma is a title": + _Claim(7, ('_ambiguities', 'suffix', 'title'), "8cbad6f13c6b", None), }, } @@ -8558,6 +8618,13 @@ def test_a_rule_reaching_no_corpus_name_says_why_it_is_kept() -> None: "fix(comma-family) lone post-comma piece routes to suffix/title, not first", 4), ("fix(#575) a particle surname before a comma is one name word", "fix(comma-precomma-family) pre-comma run reads as family, not given", 4), + # 2026-10-04, #603: both declared on the #603 opened-part rule, + # over its five names, ahead of the two comma rules for #575's + # reason -- the move is a whole part's, not one piece's. + ("fix(#603) a credential opening the part after the comma makes it the postnominal part", + "fix(comma-family) lone post-comma piece routes to suffix/title, not first", 5), + ("fix(#603) a credential opening the part after the comma makes it the postnominal part", + "fix(comma-precomma-family) pre-comma run reads as family, not given", 5), ("fix(#296) a lone post-comma credential is a suffix", "fix(suffix-routing) a two-token name ending in the suffix word jr keeps it in `suffix`", 2), ("fix(#400/#274) bound-given join and maiden consumption in one name", diff --git a/tests/v2/test_order_correspondence.py b/tests/v2/test_order_correspondence.py index 5fa1f781..6fa9a406 100644 --- a/tests/v2/test_order_correspondence.py +++ b/tests/v2/test_order_correspondence.py @@ -3,7 +3,7 @@ Form 4 (Title Family Given Middle Middle [Particle] [, Suffix] under FAMILY_FIRST) is form 2 (Family [Suffix], Title Given (Nickname) -Middle Middle[,] Suffix [, Suffix]) with the comma removed and the +Middle Middle[,] Suffix [, Suffix or Title]) with the comma removed and the family inline -- the two notations quoted verbatim from tools/differential/shapes.py. The pairs are GENERATED and asserted equal, which consults no vocabulary and no rule -- so it catches a diff --git a/tests/v2/test_properties.py b/tests/v2/test_properties.py index 8eceadfe..a681783e 100644 --- a/tests/v2/test_properties.py +++ b/tests/v2/test_properties.py @@ -3386,7 +3386,11 @@ def test_a_one_letter_particle_with_its_period_reads_as_any_initial() -> None: #: The recorded negative control: how many grid texts disagree with the -#: veto removed from ONE site, measured 2026-10-04. Every site moves +#: veto removed from ONE site, measured 2026-10-04 (classify 412 and +#: the facade 121 until #603 the same day, whose opened comma part +#: reads some texts the same with the veto off or on, the credential +#: opening the part either way: nine of classify's, 'DO Ó., ...' +#: shapes, and one of the facade's). Every site moves #: some, so none is decoration. _render's is measured without the #: particle case masks, as for a caller's one-letter particle that has #: none: with the shipped 'ó' mask the repair takes the mask before @@ -3397,10 +3401,10 @@ def test_a_one_letter_particle_with_its_period_reads_as_any_initial() -> None: #: changed nothing and was left out (decisions.md#P7 says how that was #: measured). _P7_SITE_EFFECT = { - "nameparser._pipeline._classify": 412, + "nameparser._pipeline._classify": 403, "nameparser._pipeline._vocab": 21, "nameparser._render": 338, - "nameparser._facade": 121, + "nameparser._facade": 120, } diff --git a/tools/differential/compare.py b/tools/differential/compare.py index 8e40fa80..45d1c926 100644 --- a/tools/differential/compare.py +++ b/tools/differential/compare.py @@ -2134,7 +2134,12 @@ class _ShapeMismatch(NamedTuple): #: by necessity, not a pytest-speed roster check, and it is not here. _WATCHED_DIFFS: dict[str, dict[str, tuple[str, ...]]] = { "expected_since_1.4.0.toml": { - "1 & 2, 3 4 5, Mr.": ("_initials",), + # 2026-10-04 (#603): the shape moved from `_initials` to + # {suffix, title}. A title word past the second comma is a title + # now (rules.md#C2), so 'Mr.' leaves the suffix; 1.4.0 read it + # as suffix 'Mr.' here, and the initials-only difference this + # row watched is now inside a role diff. + "1 & 2, 3 4 5, Mr.": ("suffix", "title"), "Anh do": ("_initials",), "Anna Müller (geb. Schmidt)": ("maiden", "nickname"), "Anna Müller geb. Schmidt": ("family", "maiden", "middle"), @@ -2200,7 +2205,13 @@ class _ShapeMismatch(NamedTuple): # credential again by its capitals, as this baseline read it # by vocabulary, so only the comma's report differs. "John Smith, RAI": ("_ambiguities",), - "John, Smith, Dr.": ("_ambiguities",), + # 2026-10-04 (#603): the shape moved from `_ambiguities` to + # {suffix, title}. A title word past the second comma is a + # title now (rules.md#C2), so 'Dr.' reads title and the part is + # recognized: the comma-structure report this row watched is + # gone, and the role diff is the one the baseline's suffix 'Dr.' + # gives. + "John, Smith, Dr.": ("suffix", "title"), "Jong van der": ("_initials",), "Jong, van der": ("_initials",), "Jose E. Maria Santos": ("_initials",), @@ -2249,7 +2260,13 @@ class _ShapeMismatch(NamedTuple): # credential again by its capitals, as this baseline read it # by vocabulary, so only the comma's report differs. "John Smith, RAI": ("_ambiguities",), - "John, Smith, Dr.": ("_ambiguities",), + # 2026-10-04 (#603): the shape moved from `_ambiguities` to + # {suffix, title}. A title word past the second comma is a + # title now (rules.md#C2), so 'Dr.' reads title and the part is + # recognized: the comma-structure report this row watched is + # gone, and the role diff is the one the baseline's suffix 'Dr.' + # gives. + "John, Smith, Dr.": ("suffix", "title"), "Jong van der": ("_initials",), "Jong, van der": ("_initials",), "Jose E. Maria Santos": ("_initials",), diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index 987e9276..02afcb85 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -61,6 +61,7 @@ "Doe, John van DO" "Doe, John van Ma" "Doe, MA PhD" +"Doe, PhD Jones" "Dr Jr" "Dr King Jr" "Dr." @@ -80,6 +81,8 @@ "Dr. nee Smith PhD Prof." "Duke of Edinburgh" "Eric H. Holder Jr. Attorney General" +"Eric H. Holder, Jr., Attorney General" +"Eric H. Holder, Jr., Secretary of State" "Esq. Smith" "Freiherr von Berg MA" "Freiherr von Berg, Ed" @@ -211,18 +214,23 @@ "John Smith, Jones" "John Smith, Jr do" "John Smith, Jr vd" +"John Smith, Jr., Lt.Gov." "John Smith, LEED AP" "John Smith, MA" +"John Smith, MD Jones" "John Smith, MD, Bart" "John Smith, MD, Ma" +"John Smith, MD, Ms" "John Smith, MD, R.A.I." "John Smith, MD, XYZ" "John Smith, MD,, Jr." "John Smith, Ma" "John Smith, Mr." "John Smith, Mr. Jr." +"John Smith, Ph. D. Jones" "John Smith, PhD" "John Smith, PhD DO DO" +"John Smith, PhD Jones" "John Smith, PhD MEng" "John Smith, PhD X.Y." "John Smith, PhD XYZ" @@ -366,6 +374,7 @@ "Smith, John, PhD - and MD" "Smith, John, PhD - i Soler" "Smith, John, PhD née Jones" +"Smith, John, Prof." "Smith, John, Puig - i Soler" "Smith, John, and" "Smith, Jr." @@ -374,6 +383,7 @@ "Smith, MD PhD Ma" "Smith, Ma" "Smith, Major. John" +"Smith, Ms Jane" "Smith, Ms Ma" "Smith, Ms." "Smith, Ms. Jane" @@ -383,11 +393,13 @@ "Smith, PhD" "Smith, PhD MEng" "Smith, PhD Ma" +"Smith, RN - CRNA" "Smith, Sr." "Smith, XYZ" "Smith, de Mesnil Jean" "Smith. John" "Steven Hardman, MD, DO, DDS" +"Steven Hardman, RN - CRNA" "The Right Hon. the President of the Queen's Bench Division" "Van Buren, Ed" "Van Johnson" diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index 697a52af..d492c7e7 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -785,6 +785,55 @@ surname is one name: that is rules.md#C1's count (#575), which this rule names. """ +[[change]] +# #603's two rules stand ahead of the comma-family catch-all below +# on purpose: each names its corpus names one by one, the narrower +# claim, so the classified summary credits #603 rather than that +# broad rule, which admits the same diffs on any Latin comma name. +# rules.md#C1: "A credential opening the part makes it the postnominal +# part however it goes on" (#603, 2026-10-04): where the first suffix +# word after the comma starts S2's run, every word behind it reads as +# that run reads, so a full name before the comma keeps its positional +# read ('John Smith, PhD Jones') and a one-word name is the family +# ('Doe, PhD Jones'). The undeclared dash of the delimiter rows is one +# more word of the run; a title/suffix dual opens nothing. +issue = "fix(#603) a credential opening the part after the comma makes it the postnominal part" +name_regex = "^(?:Doe, PhD Jones|John Smith, Ph\\. D\\. Jones|John Smith, PhD Jones|Smith, RN - CRNA|Steven Hardman, RN - CRNA)$" +fields = ["family", "given", "middle", "suffix", "title"] + +[[change.precedes_narrower]] +issue = "fix(comma-family) lone post-comma piece routes to suffix/title, not first" +why = """ +That rule describes ONE post-comma piece leaving `first` for `suffix` +or `title`. These names move a whole part: a credential opening it +takes every word behind it into the suffix (rules.md#C1, #603), name +words included, and a full name before the comma keeps its positional +read, so `family` and `middle` move with it. One-word names whose diff +fits that rule's fields ('Doe, PhD Jones', 'Smith, RN - CRNA') are +still this rule's: the cause is the opened part, not the lone piece. +""" + +[[change.precedes_narrower]] +issue = "fix(comma-precomma-family) pre-comma run reads as family, not given" +why = """ +That rule describes the pre-comma run reading as the family. Here the +pre-comma run does the opposite where it holds two name words -- the +comma has fixed nothing, so it keeps its positional read ('John Smith, +PhD Jones' gives given 'John', family 'Smith') -- and the move is the +opened part's (rules.md#C1, #603), not that rule's. +""" + +[[change]] +# rules.md#C2: "a word of the title vocabulary there that is not also +# suffix vocabulary is a title" (#603, 2026-10-04): a part past the +# second comma reads its title words as titles, and a part of title and +# suffix words is recognized rather than flagged ('Eric H. Holder, Jr., +# Attorney General'). 1.4.0, like 2.0 through 2.3, read them as +# suffixes. +issue = "fix(#603) a title word in a part past the second comma is a title" +name_regex = "^(?:1 & 2, 3 4 5, Mr\\.|Andrew Perkins, Jr\\., Col\\. \\x28Ret\\x29|Eric H\\. Holder, Jr\\., Attorney General|Eric H\\. Holder, Jr\\., Secretary of State|John Smith, Jr\\., Lt\\.Gov\\.|John, Smith, Dr\\.|Smith, John, Prof\\.)$" +fields = ["suffix", "title"] + [[change]] issue = "fix(comma-family) lone post-comma piece routes to suffix/title, not first" # 'Smith, Dr.' / 'Andrews, M.D.': v1 put the lone strict-suffix-or-title diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 6c8a0864..723af852 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -1058,10 +1058,14 @@ issue = "fix(#296) dr is not postnominal vocabulary, so 'John Smith, Dr.' keeps # sets the post-comma 'Dr.' is a title only, so the comma is followed # by nothing but titles and the pre-comma name keeps its split (the # no-name-word repair): suffix 'Dr.' -> title 'Dr.' for the two-part -# spelling. The three-part spelling keeps suffix 'Dr.' (a third part -# is suffix by position) and moves only by gaining the +# spelling. The three-part spelling read suffix 'Dr.' (a third part +# was suffix by position) and moved only by gaining the # COMMA_STRUCTURE report C2 gives a third part that is not suffix -# words. +# words, until #603 (2026-10-04) made a title word there a title: +# rules.md#C2, "a word of the title vocabulary there that is not also +# suffix vocabulary is a title". It now moves suffix 'Dr.' -> title +# 'Dr.' and no report, the same {title, suffix} as the two-part +# spelling, so `_ambiguities` left this rule's fields that day. # # The two-part spelling is ALSO compared under FAMILY_FIRST, as shape # 4 of corpus_shapes.jsonl: the declared order applies to the @@ -1075,7 +1079,7 @@ issue = "fix(#296) dr is not postnominal vocabulary, so 'John Smith, Dr.' keeps # this string under it, so a diff arriving there would be one nobody # has looked at. name_regex = "(?i)^john,?\\s+smith,\\s*dr\\.?$" -fields = ["title", "suffix", "_ambiguities"] +fields = ["title", "suffix"] orders = ["DEFAULT", "FAMILY_FIRST"] [[change]] @@ -3972,3 +3976,26 @@ issue = "fix(#606) salutations and titles in other languages" # The given name becomes the title: 'Greve'. name_regex = "^Greve Anna$" fields = ["title", "given"] + +[[change]] +# rules.md#C1: "A credential opening the part makes it the postnominal +# part however it goes on" (#603, 2026-10-04): where the first suffix +# word after the comma starts S2's run, every word behind it reads as +# that run reads, so a full name before the comma keeps its positional +# read ('John Smith, PhD Jones') and a one-word name is the family +# ('Doe, PhD Jones'). The undeclared dash of the delimiter rows is one +# more word of the run; a title/suffix dual opens nothing. +issue = "fix(#603) a credential opening the part after the comma makes it the postnominal part" +name_regex = "^(?:Doe, PhD Jones|John Smith, Ph\\. D\\. Jones|John Smith, PhD Jones|Smith, RN - CRNA|Steven Hardman, RN - CRNA)$" +fields = ["_ambiguities", "family", "given", "middle", "suffix", "title"] + +[[change]] +# rules.md#C2: "a word of the title vocabulary there that is not also +# suffix vocabulary is a title" (#603, 2026-10-04): a part past the +# second comma reads its title words as titles, and a part of title and +# suffix words is recognized rather than flagged ('Eric H. Holder, Jr., +# Attorney General'). 1.4.0, like 2.0 through 2.3, read them as +# suffixes. +issue = "fix(#603) a title word in a part past the second comma is a title" +name_regex = "^(?:1 & 2, 3 4 5, Mr\\.|Andrew Perkins, Jr\\., Col\\. \\x28Ret\\x29|Eric H\\. Holder, Jr\\., Attorney General|Eric H\\. Holder, Jr\\., Secretary of State|John Smith, Jr\\., Lt\\.Gov\\.|Smith, John, Prof\\.)$" +fields = ["_ambiguities", "suffix", "title"] diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index 989df2c9..eafdcce3 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -724,10 +724,14 @@ issue = "fix(#296) dr is not postnominal vocabulary, so 'John Smith, Dr.' keeps # sets the post-comma 'Dr.' is a title only, so the comma is followed # by nothing but titles and the pre-comma name keeps its split (the # no-name-word repair): suffix 'Dr.' -> title 'Dr.' for the two-part -# spelling. The three-part spelling keeps suffix 'Dr.' (a third part -# is suffix by position) and moves only by gaining the +# spelling. The three-part spelling read suffix 'Dr.' (a third part +# was suffix by position) and moved only by gaining the # COMMA_STRUCTURE report C2 gives a third part that is not suffix -# words. +# words, until #603 (2026-10-04) made a title word there a title: +# rules.md#C2, "a word of the title vocabulary there that is not also +# suffix vocabulary is a title". It now moves suffix 'Dr.' -> title +# 'Dr.' and no report, the same {title, suffix} as the two-part +# spelling, so `_ambiguities` left this rule's fields that day. # # The two-part spelling is ALSO compared under FAMILY_FIRST, as shape # 4 of corpus_shapes.jsonl: the declared order applies to the @@ -741,7 +745,7 @@ issue = "fix(#296) dr is not postnominal vocabulary, so 'John Smith, Dr.' keeps # this string under it, so a diff arriving there would be one nobody # has looked at. name_regex = "(?i)^john,?\\s+smith,\\s*dr\\.?$" -fields = ["title", "suffix", "_ambiguities"] +fields = ["title", "suffix"] orders = ["DEFAULT", "FAMILY_FIRST"] [[change]] @@ -3917,3 +3921,26 @@ issue = "fix(#606) salutations and titles in other languages" # The given name becomes the title: 'Greve'. name_regex = "^Greve Anna$" fields = ["title", "given"] + +[[change]] +# rules.md#C1: "A credential opening the part makes it the postnominal +# part however it goes on" (#603, 2026-10-04): where the first suffix +# word after the comma starts S2's run, every word behind it reads as +# that run reads, so a full name before the comma keeps its positional +# read ('John Smith, PhD Jones') and a one-word name is the family +# ('Doe, PhD Jones'). The undeclared dash of the delimiter rows is one +# more word of the run; a title/suffix dual opens nothing. +issue = "fix(#603) a credential opening the part after the comma makes it the postnominal part" +name_regex = "^(?:Doe, PhD Jones|John Smith, Ph\\. D\\. Jones|John Smith, PhD Jones|Smith, RN - CRNA|Steven Hardman, RN - CRNA)$" +fields = ["_ambiguities", "family", "given", "middle", "suffix", "title"] + +[[change]] +# rules.md#C2: "a word of the title vocabulary there that is not also +# suffix vocabulary is a title" (#603, 2026-10-04): a part past the +# second comma reads its title words as titles, and a part of title and +# suffix words is recognized rather than flagged ('Eric H. Holder, Jr., +# Attorney General'). 1.4.0, like 2.0 through 2.3, read them as +# suffixes. +issue = "fix(#603) a title word in a part past the second comma is a title" +name_regex = "^(?:1 & 2, 3 4 5, Mr\\.|Andrew Perkins, Jr\\., Col\\. \\x28Ret\\x29|Eric H\\. Holder, Jr\\., Attorney General|Eric H\\. Holder, Jr\\., Secretary of State|John Smith, Jr\\., Lt\\.Gov\\.|Smith, John, Prof\\.)$" +fields = ["_ambiguities", "suffix", "title"] diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index 4f14e7c3..1d69a140 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -2397,3 +2397,26 @@ issue = "fix(#606) salutations and titles in other languages" # The given name becomes the title: 'Greve'. name_regex = "^Greve Anna$" fields = ["title", "given"] + +[[change]] +# rules.md#C1: "A credential opening the part makes it the postnominal +# part however it goes on" (#603, 2026-10-04): where the first suffix +# word after the comma starts S2's run, every word behind it reads as +# that run reads, so a full name before the comma keeps its positional +# read ('John Smith, PhD Jones') and a one-word name is the family +# ('Doe, PhD Jones'). The undeclared dash of the delimiter rows is one +# more word of the run; a title/suffix dual opens nothing. +issue = "fix(#603) a credential opening the part after the comma makes it the postnominal part" +name_regex = "^(?:Doe, PhD Jones|John Smith, Ph\\. D\\. Jones|John Smith, PhD Jones|Smith, RN - CRNA|Steven Hardman, RN - CRNA)$" +fields = ["_ambiguities", "family", "given", "middle", "suffix"] + +[[change]] +# rules.md#C2: "a word of the title vocabulary there that is not also +# suffix vocabulary is a title" (#603, 2026-10-04): a part past the +# second comma reads its title words as titles, and a part of title and +# suffix words is recognized rather than flagged ('Eric H. Holder, Jr., +# Attorney General'). 1.4.0, like 2.0 through 2.3, read them as +# suffixes. +issue = "fix(#603) a title word in a part past the second comma is a title" +name_regex = "^(?:1 & 2, 3 4 5, Mr\\.|Andrew Perkins, Jr\\., Col\\. \\x28Ret\\x29|Eric H\\. Holder, Jr\\., Attorney General|Eric H\\. Holder, Jr\\., Secretary of State|John Smith, Jr\\., Lt\\.Gov\\.|John, Smith, Dr\\.|Smith, John, Prof\\.)$" +fields = ["_ambiguities", "suffix", "title"] diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index 3435d2ef..0ab0f16f 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -1683,3 +1683,26 @@ issue = "fix(#606) salutations and titles in other languages" # The given name becomes the title: 'Greve'. name_regex = "^Greve Anna$" fields = ["title", "given"] + +[[change]] +# rules.md#C1: "A credential opening the part makes it the postnominal +# part however it goes on" (#603, 2026-10-04): where the first suffix +# word after the comma starts S2's run, every word behind it reads as +# that run reads, so a full name before the comma keeps its positional +# read ('John Smith, PhD Jones') and a one-word name is the family +# ('Doe, PhD Jones'). The undeclared dash of the delimiter rows is one +# more word of the run; a title/suffix dual opens nothing. +issue = "fix(#603) a credential opening the part after the comma makes it the postnominal part" +name_regex = "^(?:Doe, PhD Jones|John Smith, Ph\\. D\\. Jones|John Smith, PhD Jones|Smith, RN - CRNA|Steven Hardman, RN - CRNA)$" +fields = ["_ambiguities", "family", "given", "middle", "suffix"] + +[[change]] +# rules.md#C2: "a word of the title vocabulary there that is not also +# suffix vocabulary is a title" (#603, 2026-10-04): a part past the +# second comma reads its title words as titles, and a part of title and +# suffix words is recognized rather than flagged ('Eric H. Holder, Jr., +# Attorney General'). 1.4.0, like 2.0 through 2.3, read them as +# suffixes. +issue = "fix(#603) a title word in a part past the second comma is a title" +name_regex = "^(?:1 & 2, 3 4 5, Mr\\.|Andrew Perkins, Jr\\., Col\\. \\x28Ret\\x29|Eric H\\. Holder, Jr\\., Attorney General|Eric H\\. Holder, Jr\\., Secretary of State|John Smith, Jr\\., Lt\\.Gov\\.|John, Smith, Dr\\.|Smith, John, Prof\\.)$" +fields = ["_ambiguities", "suffix", "title"] diff --git a/tools/differential/shapes.py b/tools/differential/shapes.py index 568890ef..b0d175f6 100644 --- a/tools/differential/shapes.py +++ b/tools/differential/shapes.py @@ -77,10 +77,11 @@ class Shape(NamedTuple): "1.4.0"), 2: Shape(None, "Family [Suffix], Title Given (Nickname) Middle Middle[,] " - "Suffix [, Suffix]", + "Suffix [, Suffix or Title]", "1.4.0"), 3: Shape(None, - "Title Given Middle Family [Suffix], Suffix [, Suffix]", + "Title Given Middle Family [Suffix], Suffix or Title " + "[, Suffix or Title]", "1.4.0"), 4: Shape("FAMILY_FIRST", "Title Family Given Middle Middle [Particle] [, Suffix]",