Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 3 additions & 3 deletions AGENTS.md

Large diffs are not rendered by default.

13 changes: 10 additions & 3 deletions docs/customize.rst
Original file line number Diff line number Diff line change
Expand Up @@ -1021,19 +1021,26 @@ Suffixes not separated by commas
~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~

``extra_suffix_delimiters`` handles sources that separate post-nominals
with something other than a comma. The default reading of such a name
is bad enough to be the reason you'd go looking:
with something other than a comma. Undeclared, the separator is a word.
Since 2.4 such a name still parses when a credential opens the part
after the comma, because that makes the whole part the suffix, but the
separator stays in the suffix as written. Declared, it splits the
suffix into groups the way a comma does:

.. doctest::

>>> name = parse("Jane Smith, RN - CRNA")
>>> name.given, name.family, name.suffix
('RN', 'Jane Smith', 'CRNA')
('Jane', 'Smith', 'RN - CRNA')
>>> policy = Policy(extra_suffix_delimiters=frozenset({" - "}))
>>> name = Parser(policy=policy).parse("Jane Smith, RN - CRNA")
>>> name.given, name.family, name.suffix
('Jane', 'Smith', 'RN, CRNA')

Through 2.3 the undeclared reading was given ``RN``, family ``Jane
Smith`` and suffix ``CRNA``, so the delimiter was the only way to get
the name right.

.. _unlisted-credentials:

Credentials the vocabulary doesn't list
Expand Down
9 changes: 9 additions & 0 deletions docs/design/decisions.md
Original file line number Diff line number Diff line change
Expand Up @@ -958,6 +958,15 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py):
MEASURED 2026-10-03, branch against 061f02da. POPULATION: each head in {`Smith, John,`, `Smith, John, Jr.,`, `John Smith,`, `John Smith`} followed by every 2-, 3- and 4-word product of {PhD, MD, Puig, i, y, -, Jr., Soler, Mr., née, and} that contains `-`, 19,972 texts, plus the 1,505 distinct names of `tools/differential/corpus*.jsonl` at 061f02da. Under `extra_suffix_delimiters=(" - ",)` 2,624 of the generated parses move (suffix alone 2,384, suffix and maiden 240 -- every one of the 240 a maiden name lost to a marker beside a core: 100 with `née -`, 108 with `- née`, 32 with both); under `(" / ",)`, with every `-` word rewritten to `/`, the same 2,624 (2,384 and 240), the move being a property of the core's position rather than its spelling; and under `(" and ",)` over the texts as written, where `and` is now the core and the dash a word, 2,528 (2,468 and 60). No other field and no ambiguity report moves under any of the three. In the corpus, one name moves under each of two policies: `Smith, John, PhD née Puig - i Soler` under ` - ` (above), and `Doe, Jane, and Jr.` under ` and `, suffix 'and Jr.' to 'Jr.', a leading core dropped as any lone core is. COMMA-TWIN AGREEMENT over the generated texts under ` - `, comparing the seven role fields of each against the text with ` - ` replaced by `, ` under the default policy and keeping only texts whose core is neither first nor last after the head nor beside another core: `Smith, John,` 752 of 2,100 disagreed before and 0 after, `Smith, John, Jr.,` 752 of 2,100 and 0; `John Smith,` 1,956 of 2,100 both before and after, the head where the typed comma moves the name's structure. (A first draft counted 973 of 3,310 for the `Jr.,` head, its filter reading `Jr.,` as part of the body; the review of this change caught it.) RECOMPUTE: parse each text under each policy at both trees (pin `nameparser.__file__` on each side, AGENTS.md's two-tree gotcha), compare the seven role fields and the sorted ambiguity kinds, and count the texts whose record differs; for the twin, compare the role fields alone.
SUPERSEDES, under M2: the 2026-09-26 #538 entry (its stepping, its invariant and its Not-repaired list, which this closes); the 2026-09-22 #397 follow-up's bound argument (a core can no longer stand between a marker and the clause's first word, the segment being cut at it first); and the #418 reorder entry's closing repair, which screened a tail's cores out of the marker pass so that `Smith, John, PhD née - Jones` read maiden 'Jones' -- it reads suffix 'PhD née, Jones' now, as above. Entries there that describe the walk stepping over a tail's cores (the #424 and #533 bullets) describe machinery this change deletes; they are dated and left as written. Not changed: a core OUTSIDE a tail segment is still a word (C1's Accepted entry), and a core inside a token (`RN/CRNA` under `/`) still reads by `splits_into_suffixes` and keeps the token whole.

- 2026-10-04 (Derek), #603 — A CREDENTIAL OPENING THE PART AFTER THE COMMA MAKES IT THE POSTNOMINAL PART, WHATEVER FOLLOWS; AND A TITLE WORD PAST THE SECOND COMMA IS A TITLE. One word that is neither suffix nor title vocabulary used to flip the whole comma reading: `John Smith, PhD Jones` read family 'John Smith', given 'PhD', where `John Smith, MD FACS` and `Eric H. Holder, Jr. Attorney General` were already read as a name and its postnominals. Now the first suffix word of the part, with only titles in front of it, opens the part when it starts S2's run (`_pieces.starts_a_credential_run`, #602's own predicate, so the two runs cannot disagree about what starts one), and every word after it reads as that run reads: a title word as a title, any other as a suffix. The mechanism is the one #296 built for a part holding no name word, extended rather than added beside: `segment` still decides FAMILY_COMMA, and `segment_suffix_reading` returns a reading instead of None, so assign's no-name path reads the part before the comma positionally where it holds two name words. Every word the opened part takes that would otherwise have made it the given part is reported (`suffix-or-name`), as #602's run reports what it absorbs.
THE ONE-WORD PATH, DECIDED (Derek): the whole part is suffixes. `Doe, PhD Jones` reads family 'Doe', suffix 'PhD Jones', no given name, as `Smith, PhD` already reads. DECLINED: reading the credential as a suffix and the next name word as the given name (`Smith, PhD Jones` → given 'Jones'), proposed by analogy with a title at the head of the given part. Asked whether that is a real format, it is not -- nobody writes `Family, Credential Given` on purpose -- and the one realistic input it served, a misplaced generation (`Smith, Jr. John` for `Smith Jr., John`), is garbage in, garbage out. One reading for both comma paths was the deciding ground: the one-word path disagreeing with the full-name path is the shape #429, #430 and #432 each found. Accepted cost: `Smith, Jr. John` loses 'John' to the suffix, where 2.3 read given 'John' with 'Jr.' misfiled as a title.
A TITLE/SUFFIX DUAL OPENS NOTHING (Derek). A word in both vocabularies (`md`, `ms`, `sr`, the ranks `lt`, `cpl` and the rest) is the listing form's title in front of a given name, `Smith, Ms Jane`, and a credential behind it opens nothing either (`Smith, Ms PhD Jones` keeps given 'PhD', as `Smith, MD PhD Ma` keeps given 'Ma' under #544). DECLINED after building it: a dual opening the part behind two or more words before the comma. Derek first chose it for `John Smith, MD Jones`; the reviews found what the count cannot see -- a surname of two words (`García Márquez, Ms Gabriela`, `Smith Jones, Sr. Maria`), a title in front of a one-word surname (`Dr. Smith, Ms Jane`, the count taking the title as a word), and spaced initials (`García Márquez, Ms G. J.`) each lost the given name -- and with the word explained (a dual is a word in both lists, decided by position), Derek chose that duals never open. So `John Smith, MD Jones` and the `MD - PhD - FACS` delimiter rows keep their 2.3 reading. An ambiguous-class member opens nothing either (`John Smith, Ma Jones` keeps given 'Ma'), nor does a suffix word S2's run start refuses (`John Smith, vd Jones`).
STRICT MODE: the veto refuses an initial-shaped word the suffix READING, not the run. `Smith, PSM I.` under `lenient_comma_suffixes=False` reads suffix 'PSM I.' (2.3: given 'PSM', suffix 'I.'), reporting the absorbed 'I.', which matches #602's no-comma run (`John Smith PhD V.` reports; its given-part run does not, a disagreement #602 left and this does not touch).
THE SPELLING GAP THIS FOUND, FIXED HERE. `starts_a_credential_run` refused any piece longer than one token, so the split credential `Ph. D.` group merges started neither #602's run nor this one while `PhD` started both: `Smith, John Ph. D. Jones` read middle 'Jones' where `Smith, John PhD Jones` read suffix 'PhD Jones'. The merged piece now starts the run, the "suffix" piece tag being the Ph./D. merge's alone; `trailing_candidates` had already admitted it, so its pre-check and the predicate now agree. The no-comma spelling `John Smith Ph. D. Jones` still reads family 'Jones': the trailing peel's walk never holds the merged piece, which is a separate path and is left (AGENTS.md's spelling sweep).
C2, WIDENED (Derek): a word of the title vocabulary in a part past the second comma that is not also suffix vocabulary is a title -- the part stands behind a comma, the line H5 draws for reading a title from the end of a name -- and a piece holding a suffix word stays a suffix, a dual and a connective join that took one (`PhD - and MD`) alike. `Eric H. Holder, Jr., Attorney General` reads title 'Attorney General', and a part of title and suffix words is recognized rather than flagged. The flag's test is asked word by word against the title vocabulary, a period-joined word through `period_joined_vocab` as classify tags it (`John Smith, Jr., Lt.Gov.`, which the code review found still flagged), so `Secretary of State` reads its title but keeps `comma-structure`: knowing that group's title chain takes the connective would be `segment` modelling `group`, the sign #611 is about. Every release through 2.3.0 read these as suffixes, 1.4.0 included (`Eric H. Holder, Jr., Attorney General` suffix 'Jr., Attorney General', `Smith, John, Prof.` suffix 'Prof.', measured with the released wheels 2026-10-04); a first draft of this entry said 1.4.0 read them as titles, inferring it from the 1.4.0 run reporting no unexplained diff, when the reason was the broad `fix(comma-family)` rule there claiming them, which #603's two rules now precede by declaration. Accepted consequence: a noble rank there reads as a title, `John Smith, Jr., Freiherr von Richthofen` → title 'Freiherr', suffix 'Jr., von Richthofen'.
WHAT MOVES, measured 2026-10-04 against master at 97d36e0d, after the dual decision above. The differential corpora at 97d36e0d (1,519 name and order entries over 1,515 distinct names): four names, every one intended -- the dash-delimited `Steven Hardman, RN - CRNA` now reads given 'Steven', family 'Hardman', suffix 'RN - CRNA', and `John, Smith, Dr.`, `Andrew Perkins, Jr., Col. (Ret)` and `1 & 2, 3 4 5, Mr.` read their tail title. A generated grid: every name of `tools/differential/corpus*.jsonl` at 97d36e0d, bare-string and object lines alike, plus 120,000 one-to-seven-word texts drawn with `random.Random(601)` from the 54 words {Jane, Doe, Smith, John, nee, née, geb., z domu, PhD, MA, Ma, MD, Jr., Jr, III, V, VI, i, y, de, de la, van, vd, mc, do, Do, DO, abd, abdul, Dr., Prof., King., Le, Attorney, General, Jones, `,`, ba, R.A.I., (nee, Smith), Ó, binti, Ph., D., Chief, Justice, Ms, Sr., RN, -, G.J., MJ, Col.}, plus each of the heads {Jane Doe, Doe,, Smith, John, John Smith, Dr., Mai Le, Jane van der Berg} followed by every three-word product of {nee Smith, PhD, MA, Prof., de, Jr., Jones, VI, do, i Soler, nee}; 98,630 texts, each parsed under the default policy, `name_order` family-first and a ` - ` delimiter, the seven role fields, the ambiguity kinds and details and `initials()` compared against 97d36e0d. 588 texts move: 581 have a part after a comma whose first non-title word is a suffix word that is not also a title, or a title word past the second comma, or both; the other seven were read by hand -- a leading period-shape title (`geb.`) in front of the opening credential, and #602's given-part run started by `Ph. D.` (`, abdul Ph. D. Do binti Smith`). Frames per parse through `Parser().parse`, py3.11: the reference band and plain names unchanged (`Dr. Juan Q. Xavier de la Vega III` 449, `Smith, John` 179, `John Smith, Jr.` 210, `John Smith, Dr.` 261, `Smith, Ms Jane` 256); a part of suffix words behind one word gets cheaper, the positional count now leaving in C for a single piece (`Smith, Jr.` 163 → 156, `Smith, MD PhD` 260 → 253); and the opened part pays for what it reads (`John Smith, PhD Jones` 338 → 362).
- 2026-10-04 #611 — C1, S2 AND P6 GREW BY PATCHING, AND EACH SHOWS THE DECISIVE SIGN; THE RETHINKS ARE FILED, NOT DONE. The question was docs/design/AGENTS.md's, asked of the three longest rules after M2 (#601): does a guard model what another stage would do? C1: `segment` predicts group's particle chain (the #562 pair test) and assign's capitals reading (the skip #562 broke), mirrors classify's tags in its unit counts, and decides the structure in two stages, assign re-deciding #296's positional read with a second count -- the copy #603 extended. S2: the words-to-spare count runs over pieces group merges, so the chain re-asks the peel and rolls back (`John van Mc`), P5 compares two views, group emits S2's report, and the peel runs up to five times a parse. P6: assign holds back exactly the particle tail post_rules will attach (`particle_tail(..., sticky_from)`). Decided: #603 shipped as a targeted change inside today's machinery; the rethinks are #613 (C1, read the comma once after classify and bind, P6 folded in) and #614 (S2, read the trailing run once over units and bind), C1 first because S2's unit count should be shared with it, neither for 2.4.

### T1 — separators, not joiners

- 2026-07 (v2 core, PR #288) — v1's squash_emoji/squash_bidi REMOVED the character and joined its neighbors. (v1.3.0 had no bidi handling at all: squash_bidi entered late v1 via #266, 2026-07-07, on the emoji precedent's shape.) ('A😀B' → 'AB'); v2 makes an ignorable character a separator ('A😀B' → 'A', 'B'). The unavoidable consequence of every part being an exact positioned piece of the input: with no rewriting stage, nothing can splice two half-words together.
Expand Down
1 change: 1 addition & 0 deletions docs/design/mechanisms.md
Original file line number Diff line number Diff line change
Expand Up @@ -200,6 +200,7 @@ Problem shape. A test pins an ordering, a sort, a dedup or a partition, and its
- Run all the gates, not the ones you remember: ruff runs before mypy and pytest in CI, and each has caught what the others passed.
- Purge __pycache__ between same-length source mutations; stale bytecode makes a changed file measure as unchanged.
- After NARROWING a rule, check the receiver: the names a narrowed rule sheds land on a neighbour, and nothing guarantees the neighbour's prose describes what it inherited — #375 fixed an over-claiming rule and relocated the bug onto its neighbour. Ask also what the old behavior was CONCEALING: #379's attachment removed the input a test used to build an all-particle middle name (#402), and #400's reserve fix exposed the dual-membership count shape #397 names -- twice in one session a fix's real yield was a defect it stopped hiding.
- A gate run with no UNEXPLAINED line for a name is not evidence that the baseline read it the same: a broad rule may have claimed the diff. #603's first draft wrote that 1.4.0 read a title past the second comma as a title because no 1.4.0 name moved unexplained, when the 1.4.0 ledger's `fix(comma-family)` rule, a regex over every Latin comma name, was claiming each one in the classified summary; the released wheel read suffix 'Jr., Attorney General'. Read the classified summary for the name, and ask the wheel (`PYTHONSAFEPATH=1 uv run --isolated --no-project --with nameparser==X`, from outside the worktree) before writing what a release did.
- A ledger rule EXPLAINS a diff; nothing checks that its own sentence still DESCRIBES it, so a rule can go on covering a name it has stopped being true of and the gate stays green. Measured on #533: a `(?i)` anchor written for `John née Jones Smith Ma` — the clause keeping the credential — also reached the capitals spelling, which that change made read the opposite way, the clause giving the credential up. The rule still matched the name and still covered its fields, so `unexplained: 0` was silent about it. What found it was ATTRIBUTION rather than the gate: measuring every corpus name the change touches against the PARENT commit, the EXPLAINED ones as well as the unexplained, and asking of each which change actually moves it. Do that before writing a ledger rule's prose, not only when a name goes unexplained — the failure mode is a rule that reads as authoritative and argues for the reading it lost. Second instance, #397/#461 (2026-09-20), and it is the same method finding the OPPOSITE shape: `fix(initials-per-word)` had been silently ABSORBING eleven 1.4.0-parity breaks that belong to the new connective rule, so the gate was green and the attribution was wrong about which change owned them. A rule can over-cover as well as mis-describe, and neither is a diff the gate can report.
- A skip is indistinguishable from "correctly declined": pytest turns an empty parametrize into a skip, and a filter that widens its own skip set cannot fail. After changing any selection shape, verify the guard still REACHES the code it watches — assert the selected set is non-empty, or force-a-decision on its size.
- A differential corpus cannot evidence behavior keyed to OUT-of-vocabulary shapes: it holds only names someone wrote down, and an unrecognized word is by definition outside the vocabulary — a green run over the corpus proves nothing about such a rule.
Expand Down
Loading
Loading