Skip to content

Should a comma turn off the glued-honorific peel? ("田中さん, 太郎" keeps さん in the family name) #312

Description

@derek73

Scope corrected 2026-08-01. The example below was wrong and the fix it proposes does not work — see the scope comment. The issue now covers where the peel's SITE is, not only which gates it inherits. The reasoning in this body still stands; the measurements in this section did not.

#308 added a peel: a listed CJK honorific glued to the end of a name token is split off and routed to suffix. It runs inside script_segment, and therefore inherits that stage's two structural opt-outs — the family comma and the 间隔号. Neither gate was argued for the peel specifically; both came with the placement.

That produces spelling disagreements of exactly the kind #308 set out to remove — though not the ones this issue originally named (corrected 2026-08-01):

parse("김 민준씨")    # family 김, given 민준, suffix 씨
parse("김, 민준씨")   # given '민준씨', family 김        ← honorific absorbed

parse("田中 太郎さん")  # family 田中, given 太郎, suffix さん
parse("田中, 太郎さん") # given '太郎さん', family 田中     ← honorific absorbed

The honorific is glued to the GIVEN name in both, which under a family comma lives in segments[1] — where the peel never looks.

#308 treated this shape as a defect worth fixing elsewhere — ko_honorific_glued_given_trailing_suffix exists precisely so 김민준씨 Jr. and Dr 김민준씨, Jr. agree — and then left the identical disagreement standing under a family comma (ja_honorific_glued_family_comma). Two rows in the same table are pinned to opposite conclusions about the same phenomenon.

Why the gates may not be the peel's to inherit

Both gates answer where does a name divide into surname + given. AGENTS.md states the comma doctrine as "script-conditional behavior is ignored where a comma already decides the family", and the interpunct gate as "a divided name is a transcription — its pieces are syllable groups".

The peel does not ask that question. It asks a position-independent one: does this token end in a word that can never end a name? A comma elsewhere in the string does not change the answer, and the vetting behind GLUED_HONORIFICS is likewise position-independent. The peel is also explicitly not gated on segment_scripts for exactly this reason — the vocabulary carries its own license rather than borrowing the script's.

Measured, if the peel simply moves above both gates

田中さん, 太郎      family 田中さん, given 太郎     →  family 田中, given 太郎, suffix さん
田中さん, PhD       title PhD, family 田中さん      →  title PhD, family 田中, suffix さん
威廉·莎士比亚さん   given 威廉, family 莎士比亚さん →  given 威廉, family 莎士比亚, suffix さん
김, 민준씨          given 민준씨, family 김        →  unchanged

Four tests fail: the two stage tests pinning the gates, ja_honorific_glued_family_comma, and its facade twin.

The part that makes this design work, not a two-line move

Look at the last row. Under FAMILY_COMMA the name spans two segments:

"김, 민준씨"   structure FAMILY_COMMA   segments ((0,), (1,))   tokens ['김', '민준씨']

The peel scans segments[0] only. So moving the call fixes the pre-comma side and leaves the post-comma side glued — trading today's clean rule ("a family comma turns the peel off entirely") for a new asymmetry inside the comma case. Doing it properly means deciding what the peel's site is when the name spans segments, and for 김, 민준씨 the honorific is attached to the given name — a different question from the one the peel answers today.

Options

  1. Leave it. Document the comma boundary as intended and reword the two rows so they stop disagreeing. Cheapest; keeps one simple rule ("a comma turns off everything script-conditional") at the cost of the spelling disagreement.
  2. Move the peel above the gates and decide the multi-segment site question. Removes the disagreement for the pre-comma case; needs a rule for the post-comma one.
  3. Lift the peel into its own stage between segment and script_segment. Makes the ordering structural rather than a comment and ends gate inheritance by construction. Costs: _PEELED_TAG becomes cross-module, _split/_longest_entry need a shared home, and "the one stage licensed to change the token count" becomes two.

Timing

2.1.0 is unreleased. If this lands before the release, no user ever sees the inconsistency; if it doesn't, option 1's documentation is needed so the boundary is at least stated rather than accidental.

Found by an altitude review on #311.

Activity

  1. added this to the v2.1 milestone on Jul 31, 2026
  2. derek73 commented on Aug 1, 2026

    @derek73
    OwnerAuthor

    Decision: option 2, in a shape the original three options did not name. Recording the reasoning while it is fresh.

    The choice does not decide the hard part

    Both option 2 and option 3 still have to answer the same question: what the peel's site is when the name spans segments. Under FAMILY_COMMA, 김, 민준씨 is segments == ((0,), (1,)) with the honorific attached to the given name, across the comma. A separate stage does not dodge that — it still has to pick a rule. So the choice is about structure, not about the answer.

    What option 3 costs, beyond the mechanics listed above

    Two of the "costs" in the issue body are really one invariant and one contract:

    • It makes the token-count exemption plural. tests/v2/pipeline/test_state.py pins a per-stage ownership map, and script_segment is deliberately absent from the token-count assert — "the one contract this stage is exempt from." That single exemption is what keeps the anti-Strange parsing of name w lastname prefix and title before and after #100 span discipline enforceable: exactly one stage may change the token count, so exactly one place needs index-remap scrutiny. Option 3 makes it two.
    • _PEELED_TAG becomes inter-stage protocol. Today the peel emits it and the segmenter's neighbour test reads it, same module — which is why its comment says it needs no home in _types, unlike FOLDED_TAG. Split them and it inherits FOLDED_TAG's full obligation: a home in _types and a strip on the way out (_parser.py strips FOLDED_TAG from public tokens). Option 3 would remove a same-stage coupling by adding a cross-stage one.

    The shape chosen

    Keep one stage, restructure it into symmetric siblings: _peel_honorific_tail(state) and _split_surname_site(state), each owning its own gates, with script_segment reduced to universal preconditions plus two calls:

    if original.isascii():        return    # universal preconditions
    if not segments:              return
    state = _peel_honorific_tail(state)     # vocabulary-licensed
    if structure is FAMILY_COMMA: return    # segmentation-only gates
    if interpunct_offsets:        return
    ... surname / segmenter half
    

    This gets most of option 3's legibility with none of the invariant or tag cost, and it makes the distinction the issue is about visible in the code rather than only in prose. It also blunts option 2's stated weakness: with the peel above the gate block, a new gate naturally lands in the block below it, so re-capturing the peel takes an insertion against the visual grain rather than an ordinary edit.

    What would have changed the answer, and the evidence that arrived

    A stage earns itself when it has more than one inhabitant. The question was whether more vocabulary-conditional token surgery is coming.

    Speculating on the candidates, they are mostly leading peels, and a leading peel is harder to vet than a trailing one — the leading position is where surnames live, and the argument that clears 양 in #308 is exactly that a surname LEADS so the trailing-only gate never sees it. The Chinese familiar prefixes 老/小 (老王, 小王 — among the commonest informal address forms in Mandarin) fail that test outright: 小 is a common given-name character, and peeling it would cut 小明 — this repo's own worked example — in half.

    But #317 (Thai) is a real counterexample: นาย/นาง/นางสาว are closed-class address terms glued to the given name, plausibly clean to vet, and they need a head peel routed to title. So a second inhabitant is more likely than the 老/小 analysis alone suggested — and the sibling shape accommodates a _peel_honorific_head as a third sibling without prejudging the stage question.

    If a head peel ever ships, that is the moment the two-inhabitant argument for option 3 becomes real and informed, and the refactor is no harder then than now.

    Still to decide when this is implemented

    The multi-segment site question, which neither option answers: under FAMILY_COMMA the name spans segments[0] and segments[1], and 김, 민준씨 puts the honorific on the given-name side. Moving the peel above the gates without answering this fixes the pre-comma case and leaves the post-comma one glued — trading today's clean rule ("a family comma turns the peel off entirely") for a new asymmetry inside the comma case.

  3. derek73 commented on Aug 1, 2026

    @derek73
    OwnerAuthor

    Scope corrected, and widened

    Re-measuring before implementing found that this issue's headline example was wrong, and that the fix it proposed does not work. Both are corrected in the body; the detail is here.

    The example was backwards

    田中さん 太郎 and 田中さん, 太郎 agree — neither peels. The peel site is the last NON-POST-NOMINAL token, which is 太郎 in the spaced form, so nothing is there to peel. That example was written before #311's scan-back landed and never re-measured.

    Measured, the real disagreements:

    AGREE   田中さん 太郎     family 田中さん, given 太郎
            田中さん, 太郎    family 田中さん, given 太郎
    
    DIFFER  김 민준씨         family 김, given 민준, suffix 씨
            김, 민준씨        given '민준씨', family 김
    
    DIFFER  田中 太郎さん      family 田中, given 太郎, suffix さん
            田中, 太郎さん     given '太郎さん', family 田中
    
    DIFFER  田中さん PhD      family 田中, suffix 'さん, PhD'
            田中さん, PhD     title PhD, family 田中さん
    
    DIFFER  威廉·莎士比亚 さん  family 莎士比亚, suffix さん
            威廉·莎士比亚さん   family 莎士比亚さん
    

    Two of the four are the honorific glued to the given name, which under a family comma lives in segments[1] — where the peel never looks. One is the 间隔号, which has the identical problem for the identical reason.

    Option 2 as filed does not work

    Moving the peel above the gates and changing nothing else:

    • fixes 田中さん, PhD ✓
    • does nothing for 김, 민준씨 or 田中, 太郎さん — segments[0] is just 김/田中 ✗
    • makes 田中さん, 太郎 peel, introducing a disagreement where one did not exist ✗

    One of four fixed, two missed, one broken.

    What the issue now covers

    The site, not only the gates. Two changes that only work together:

    1. Site: the last non-post-nominal token across ALL segments, not just segments[0]. Flattening segments rather than the raw token stream preserves the extracted-nickname exclusion that ko_honorific_glued_given_nickname pins.
    2. Placement: after the universal preconditions (ASCII-only original, empty segments), before the FAMILY_COMMA and 间隔号 gates.

    Prototyped and measured: every glued/spaced pair above then agrees on the honorific, 田中さん, 太郎 stays unchanged, and three tests fail, all by design — test_interpunct_divided_name_never_peels, ja_honorific_glued_family_comma, and its facade twin. test_family_comma_skips_the_peel keeps passing but becomes misnamed: the peel is no longer skipped because of the comma, it simply finds no site.

    Structure stays as decided above — siblings under one stage, not a new stage.

    Net: after this the peel consults no comma structure and no dot, only its own vocabulary and the two universal preconditions. That is the claim #308 made for it and did not deliver.

    The ASCII bail is deliberately still in front of the peel; its known wart (a caller-configured Latin tail fires only when some other token in the name is non-ASCII) stays documented and out of scope here.

  4. added a commit that references this issue on Aug 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions