From 808b15bf01caf148d9079fb0533ceb16de1d7818 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sun, 4 Oct 2026 12:45:21 -0700 Subject: [PATCH 1/3] Share P6's particle tail and the maiden take's candidates as one predicate each (#610) `_pieces.particle_tail` is P6's walk, read off roles, so post_rules' attachment and assign's given-part credential run (#602) ask one walk instead of two that agreed only through the run's roles; it moved none of the 95,813 readings of the equivalence grid, and family-comma parses pay fewer frames ('Smith, John' 183 -> 180). `_pieces.trailing_candidates` is the maiden take's choice of clause words, written beside the peel whose admissions it must cover, and a sweep test pins that it is a superset of what `tail_reading` takes over a 37,044-parse grid. The test found the hand copy wrong twice: a shape-only numeral ('VI', in no wordlist) is the walk's last piece when a period title stands behind it or the merged 'Ph. D.' does, so 'Jane Doe nee Smith VI Prof.' kept maiden 'Smith VI' where 'John Smith VI Prof.' reads suffix 'VI'. Both now end the clause at the numeral, as rules.md#M2 says; the recorded negative control is 288 misses for the loop #601 shipped. Ten readings moved against master, all this class. Two case rows, an M2 example, the ledgers classifying that example, and a decisions.md#M2 entry; mechanisms.md#ONE-PREDICATE-PER-QUESTION names both new instances. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 1 + docs/design/mechanisms.md | 2 +- docs/design/rules.md | 5 +- nameparser/_pipeline/_assign.py | 42 ++---- nameparser/_pipeline/_group.py | 50 +------ nameparser/_pipeline/_pieces.py | 134 +++++++++++++++++-- nameparser/_pipeline/_post_rules.py | 51 ++----- nameparser/_pipeline/_state.py | 7 + tests/v2/cases.py | 21 +++ tests/v2/pipeline/test_pieces.py | 59 +++++++- tests/v2/test_ledger_guards.py | 26 +++- tools/differential/corpus_rules.jsonl | 1 + tools/differential/expected_since_1.4.0.toml | 5 +- tools/differential/expected_since_2.0.0.toml | 5 +- tools/differential/expected_since_2.1.0.toml | 5 +- tools/differential/expected_since_2.2.0.toml | 5 +- tools/differential/expected_since_2.3.0.toml | 5 +- 17 files changed, 287 insertions(+), 137 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index c12f5b14..89193681 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -244,6 +244,7 @@ the fullwidth-colon marker (旧姓:佐藤 arrives as one word; the head-peel q - 2026-09-27 (Derek), #544 — S2'S COMPANY CLAUSE DOES NOT REACH ACROSS A CLAUSE, recorded as an Accepted boundary rather than repaired. `Jane Doe Jr. nee Smith Ma` keeps maiden 'Smith Ma' and reports it, the clause's own words standing between 'Jr.' and 'Ma', where the clause-less `Jane Doe Jr. Ma` reads suffix 'Jr. Ma'. tests/v2/test_properties.py's clause-agreement walk pins the pairs that differ for exactly this reason as its `anchored_head` class — 810 of its 20,412 pairs, recorded 2026-09-27, 0 with the anchor off. Inside the clause the company reads as it does anywhere: `Jane Doe nee Smith PhD MEng` ends the clause at 'PhD' and reads suffix 'PhD MEng', maiden 'Smith'. - 2026-10-04 (Derek), #601 — A MARKER COUNTS ONLY BEHIND A SURNAME, AND IT TAKES THE WORDS AFTER IT UP TO THE TRAILING RUN THE NAME WOULD END WITH IF THE CLAUSE WERE NOT WRITTEN. This SUPERSEDES the release machinery of the #533 and #535 entries above (the release question, the per-stop checks, the anchored-head boundary of the #544 entry) and the given-part and suffix-comma reach the 2026-07-03 rule had; the #399, #420, #434 and delimiter entries stand. The trigger was #548, `Dr. nee Smith PhD Prof.`, where 2.3.0 read family 'PhD', maiden 'Smith': a clause whose head was nothing but a title, and a release check asking whether each word it gave up could still read as a post-nominal. #548 was closed into this issue and #602 (S2's run) because both halves of its answer were rules rather than repairs. THE HEAD RULE (Derek): a marker is one only in the name before any comma or the surname part before a family comma, behind at least one name word past the leading title run, and not straight behind an unambiguous suffix word or a connective. Anywhere else it is an ordinary word — in the given part after a family comma (`Doe, Jane nee Smith` reads middle 'nee Smith', 1.4.0's reading) and in a part after a suffix comma (`Smith, John, PhD née Jones` reads suffix 'PhD née Jones', also 1.4.0's). Derek's framing was that a maiden marker always follows a family name, and that a marker following something the parser has decided is a title, suffix, particle or connective is not a marker. THE PARTICLE HALF WAS DROPPED BY MEASUREMENT: prototype variant 1 refused a marker behind particle vocabulary and so declined `Anh Do née Tran` and `Mai Le née Nguyen`, Do, Le and Van being particles that are also the commonest Vietnamese surnames; a particle counts as the surname (`Jane van nee Smith` reads family 'van', maiden 'Smith'). THE BOUND (Derek): the clause takes its words up to the run of post-nominals and titles that the trailing rules read at the end of the name with the clause removed, and the words it gives up take the roles that reading gives them — the TRAILING rule's reading of the clause-free name, not the full parse of it, so a head the full parse would read differently (H1's lone title, a P5 join) cannot change where the clause ends. The first word after the marker is always taken, the marker having announced a name — a lone roman numeral included, so `Jane Smith née V` reads maiden 'V' where the clause-free `Jane Smith V` reads suffix 'V', accepted rather than given a numeral exception (rules.md#M2's Accepted block) — unless it is an unambiguous suffix word, where the marker declines (`Jane Smith nee PhD` reads family 'nee', suffix 'PhD', as before). Before a family comma the clause keeps the 2026-07-03 stop at the first suffix word, since no trailing rule reads that part's end. VARIANTS, prototyped 2026-10-03 behind an environment switch on one tree (the numbers are in #601's comments): v2/v3 grew the clause until the existing release check accepted it and reached 0 invariant violations but inherited the check's caution about particles (`Anh Do geb. de la Vega PhD Prof.` kept 'PhD Prof.' in the maiden name); v4 read the run once over the name as written without binding it, and is the recorded negative control in tests/v2/test_properties.py (350 of the clause grid's 4,482 parses fail under it, measured 2026-10-04); v5/v6 bound that reading as roles but still counted the clause's words among the words to spare, so `John nee Smith ba` read 'ba' as a credential where `John ba` reads a name; v7, the clause-free reading bound as roles, was adopted. Against #535's grids v7 left 0 violations of the M2 invariant where the tree then had 15,586 (grid A, 327,936 texts) and 8,713 (grid B, 173,057), measured 2026-10-03 with an invariant check written to #535's recipe. WHAT THE PROTOTYPE MISSED, found implementing it: the numeral fork. A trailing roman numeral is a suffix only by its shape against the word before it, and v7 read the clause's last word without that test, so `John née Jones Smith VI` kept 'VI' in the maiden name; the take now asks `is_trailing_numeral_suffix` of the last piece, and cases.py's `a_numeral_the_clause_free_name_reads_ends_the_clause` pins it. And the head's title run is measured over the words before the marker, a lone title excepted: measured over the whole segment, `Lord Chancellor née Jones` counted 'Chancellor' as a title and refused the marker. TRIAGE, 2026-10-04: the prototype's 68 failing case rows (40 facade failures among them) were re-pinned as `fix(#601)` with their group named in the note, bar the head-rule bug above; tests/v2/test_parser.py's corpus-append invariant became two-sided (appending ` née X` moves no other field, OR the take records why it declined), with a named set for the suffixes that themselves hold a marker. DEGENERATE INPUT, NO READING DESIGNED FOR (Derek, 2026-10-04: garbage in, garbage out). Two names that read worse by eye moved, and they are recorded here so the move is not mistaken for a regression, not because either reading is wanted: `Berg, Jane van der nee Smith DO` reads middle 'van der nee Smith', family 'Berg' where 2.3.0 read family 'van der Berg', maiden 'Smith DO' — the marker in a given part being a word, the particles no longer trail the given name (P6); and `Jane Doe, PhD née Smith` reads given 'PhD', middle 'née Smith', family 'Jane Doe' where 2.3.0 read given 'Jane', family 'Doe', suffix 'PhD', maiden 'Smith', which 2.3.0 reached only through the clause taking the marker (the comma part opening with a credential is #603's shape). Measured 2026-10-04 against 2.3.0. COST: the take is one function and no release helper survives; the #533 review's comma-head grids could no longer reach a take at all and were replaced by counting heads (tests/v2/test_properties.py's `_maiden_clause_grid`, 4,482 parses). +- 2026-10-04 #610 — THE MAIDEN TAKE'S CANDIDATE WORDS ARE ONE FUNCTION BESIDE THE PEEL, AND SHARING IT FIXED A GAP THE HAND COPY HAD. `_maiden_take` chose the clause words to read its clause-free view over with a filter of its own (`_clause_tail_word` plus a numeral clause), a hand copy of the peel's and the chain's admissions whose contract — a superset of what `tail_reading` takes — nothing stated or tested. It is now `_pieces.trailing_candidates`, written beside `peel_trailing`, and tests/v2/pipeline/test_pieces.py pins the superset over a generated grid (37,044 parses through `group`). The grid found the copy wrong twice. The numeral fork reads a shape-only roman numeral (`VI`, in no wordlist) as the WALK's last piece, and two things leave it last without standing at the end of the name: a period-marked title behind it, which the chain takes first, and the merged `Ph. D.`, which group flags and the walk never holds. The copy admitted the numeral only as the clause's last word, so `Jane Doe nee Smith VI Prof.` kept maiden 'Smith VI' where `John Smith VI Prof.` reads suffix 'VI', and `Jane Doe nee Smith VI Ph. D.` the same way. Both now end the clause at the numeral, as rules.md#M2's statement already said they should. Recorded negative control: the loop as #601 shipped it misses 288 texts of the grid, every one this shape; a first draft of the fix covered the title half and missed the `Ph. D.` half, 111 texts. Against master 8a8459a8 over the 95,813-text equivalence grid the change moved 10 readings, all of this class. The same issue lifted P6's particle-tail walk into `_pieces.particle_tail`, read off roles, so post_rules' attachment and assign's given-part credential run (#602) ask one walk; that half moved none of the 95,813 readings. ### N1 — the default delimiter pairs diff --git a/docs/design/mechanisms.md b/docs/design/mechanisms.md index 218be198..40a50821 100644 --- a/docs/design/mechanisms.md +++ b/docs/design/mechanisms.md @@ -63,7 +63,7 @@ Problem shape. "Which stage does X?" — asked before attributing behavior in pr ## ONE-PREDICATE-PER-QUESTION — one predicate answers it, and every other site calls that -Problem shape. Two stages need the same answer about the same input, and the one that does not own the decision is about to test for it. Contract statement. Where two sites ask the same question, exactly one predicate answers it and every other site calls that one — never a condition written to match it. The predicate belongs to the QUESTION, not to whichever stage decides: it may sit in a leaf both stages import, and for the leading-title test it must, since the deciding stage is assign and group cannot import assign. How it works. A hand-written mirror agrees with its original only until one of them moves, and the drift is invisible in both directions: each site keeps passing its own tests while they disagree about an input neither covers. Five instances, every one found as a defect before it was found as a pattern — #319 lifted the wholly-suffix predicate into the vocabulary layer "so the comma decision and the honorific peel's segment test cannot drift apart"; #401/#421 lifted the trailing-numeral fork out of assign so the bound-given reserve stopped carrying a copy, its hand-written mirror having been falsified in review more than once — the lesson recorded there being that what must be mirrored is assign's WALK, not merely its condition; #425 replaced that reserve's hand re-derivation of the trailing peel with one function over the view the join would leave; #424 moved assign's leading-title test down because group's own `title()` does not see H2's unlisted abbreviations, so `Xyz. van Johnson` chained where `Dr. van Johnson` did not; #429 moved the no-name-segment test down because group asked by segment INDEX where assign asks by CONTENT. The destination follows the LAYER, not the topic: a predicate over token text goes to `_vocab`, one over pieces and tags to `_pieces`. Both are leaves the stages sit on. The piece layer got its own module only in #439 — until then those predicates collected in `_group`, not because grouping owned them but because `_assign` imports `_group` and cannot be imported back, so group was the one place both stages could reach; five had accumulated across four PRs before the module existed. Stage order is this mechanism's limit, and it forecloses the alternative: where the reader comes AFTER the decider, record the answer on the state instead — `ParseState.order` is that shape, "Recorded rather than recomputed downstream, because the two can differ" — which is unavailable whenever the EARLIER stage is the one asking. (The concrete assign→group import that forced the `_group` collection is gone since #439; what remains is the ordering it was a symptom of, and tests/v2/test_layering.py is where the leaf's contract is now written down.) The cost is a second evaluation of the same predicate, measured for #429 at 1.2–2.2% of a family-comma parse and 0% of every other; recording that number was the right answer there over plumbing a state field the two sites would not otherwise share. Lives in. nameparser/_pipeline/_vocab.py over text (is_wholly_suffix; is_trailing_numeral_suffix — the #401/#421 instance, whose only caller since #439 is the shared peel rather than a stage; and maiden_marker_run, the #434 instance and the clearest two-stage case, called by classify over token texts and by extract over a clause's whitespace words, with group reading the tags classify recorded because it runs later; and delimiter_cores, the #436/#437 instance, read by group where a tail segment DROPS a configured delimiter core and by post_rules where the suffix view's entry boundary asks whether a dropped token was one, with a third reader inside this same module, is_wholly_suffix, where a configured core counts as suffix-shaped; and in_initialless_script, the #322/#323 instance and the only one here that is a REPERTOIRE test rather than a vocabulary one — the script half of the #320 initial veto, read by is_initial one function away and by _pieces.is_leading_title, so "a script with no initials has no period abbreviations either" is one predicate over _policy._NO_INITIALS rather than a second reading of that table; it lost its leading underscore when the second caller arrived; and caps_shape_candidate, the #516 instance and the newest, called from the sites that each needed the identical question answered — classify's own tag emission and _segment.py's two comma tests (the all-caps run and, since #564, the #544 run test's by-shape member), with its unit tests — where the usual reason for keeping such copies apart (a shared call costing every default-policy parse a frame it cannot use) does not hold: every caller asks first, in C, what the predicate would decline anyway — the setting and position, then alphabetic capitals and no listed suffix word — so only an unlisted all-caps word reaches the call, at the default as under any setting (decisions.md#S2)) and nameparser/_pipeline/_pieces.py over pieces: is_suffix_piece, leading_titles and peel_walk are called by both stages, while is_leading_title, is_title_piece and trailing_start are called by group alone (measured 2026-09-06 by call site: `is_leading_title` has no caller in `_assign.py`, which reads `leading_titles` instead — a first draft of this clause listed it among the shared ones) — `trailing_start` being the one to know, since it answers where the trailing run begins and is what P2's chain stops at, and M2's walk wherever no trailing rule reads the clause (elsewhere, since #535, the walk stops where `tail_reading` says) — and segment_suffix_reading by assign alone since #436/#437, that last one being #430's instance, where THREE readers shared one answer until the render join, group's third, was replaced by a rule over the commas the writer typed (decisions.md#C1, 2026-09-06); it stays where it is, one call site being no reason to move a predicate that two sites will contest again. `trailing_titles` was that last shape for one day (2026-09-08, the #316/#489 bundle, rules.md#H5), and since the /simplify round of 2026-09-09 the SHARED predicate is `tail_reading` instead — the peel-and-chain fixed point that answers where the name pieces end (decisions.md#H5). Assign calls it at its main walk and group's bound-given reserve calls it twice, once per view the join compares, because that reserve reads the name words assign will leave and this walk is half of what leaves them (rules.md#P5; counting a trailing title word among them joined 'Prof. abdul rahman Prof.' where 'Prof. abdul rahman' does not). Since #535 group's maiden walk calls it as well, over the clause and over the view its take would leave, wherever a trailing rule reads the clause (rules.md#M2), for the same reason: the walk's stops must end the clause where assign's reading of the name will begin. `peel_trailing` and `trailing_titles` are what that fixed point is BUILT from, and neither is a two-stage question any longer: `peel_trailing` has one caller outside `_pieces.py`, the maiden walk in `_group.py`, which asks the peel itself because it needs ONE half of the answer at a time -- the numeral's over the pieces as written and again over the view its take would leave (#424), the acronym's beside it (#533) -- where `trailing_start` and `tail_reading`, the two callers in the leaf, fold both halves into one index; since #535 it asks the bare peel only where no trailing rule reads the clause, and reads `tail_reading` everywhere else, so that a trailing title does not hide the numeral or credential in front of it; that walk is a reader of the peel and not a second spelling of it, the question being asked of a different name each time. `trailing_titles` has exactly one caller, assign's family-comma segment-1 walk, which reads the chain without the re-peel, and `_group.py` does not import it. The tail reading is in the leaf rather than inline because each assign site had been given a cheap frame-free gate written to match the walk's own first condition, which is a second implementation of the question and was removed in review; what the leaf costs is one frame per entry point, measured, and the walk's own first test is a compiled regex rather than a call, so an ordinary name pays a match and stops. The reserve's two calls cost the reference name nothing — it never enters that branch, having no bound given word — and the parse and facade frame counts did not move (measured 2026-09-09). Re-measured 2026-09-09 by an AST call-site census over `_pipeline/*.py` — every call node whose callee is one of these names, keyed by module and enclosing function, which is what caught the census claiming a share for `peel_trailing` that the round had just taken away — the rest of it holds unchanged: is_suffix_piece, leading_titles, peel_walk and now tail_reading shared, is_leading_title, is_title_piece and trailing_start group-only — assign still reads `leading_titles` and never `is_leading_title`, which is what keeps H2's shape inference out of the trailing slot. And nameparser/_pipeline/_post_rules.py over a state: suffix_entries, the #511 instance, the R1 entry pass as a function, the one instance living in a stage rather than in a leaf — it is a pass over a whole ParseState and no leaf takes one, and AGENTS.md names it as the exception — run by post_rules last in the stage (through its in-place worker) and by Parser.revise over a sub-parse whose roles it has forced, so a suffix value handed to revise() derives its entries by the rule a whole name uses rather than by a second reading of the value's commas (decisions.md#C1, 2026-09-06 #511). tests/v2/test_layering.py holds each module's contract, and a piece predicate growing a dependency on a STAGE shows up there as a widened entry. Reach for it when. You are about to write a condition that mirrors, matches or "does what X does" — or you find a comment saying one does. Grep for the other site's predicate and call it instead. +Problem shape. Two stages need the same answer about the same input, and the one that does not own the decision is about to test for it. Contract statement. Where two sites ask the same question, exactly one predicate answers it and every other site calls that one — never a condition written to match it. The predicate belongs to the QUESTION, not to whichever stage decides: it may sit in a leaf both stages import, and for the leading-title test it must, since the deciding stage is assign and group cannot import assign. How it works. A hand-written mirror agrees with its original only until one of them moves, and the drift is invisible in both directions: each site keeps passing its own tests while they disagree about an input neither covers. Five instances, every one found as a defect before it was found as a pattern — #319 lifted the wholly-suffix predicate into the vocabulary layer "so the comma decision and the honorific peel's segment test cannot drift apart"; #401/#421 lifted the trailing-numeral fork out of assign so the bound-given reserve stopped carrying a copy, its hand-written mirror having been falsified in review more than once — the lesson recorded there being that what must be mirrored is assign's WALK, not merely its condition; #425 replaced that reserve's hand re-derivation of the trailing peel with one function over the view the join would leave; #424 moved assign's leading-title test down because group's own `title()` does not see H2's unlisted abbreviations, so `Xyz. van Johnson` chained where `Dr. van Johnson` did not; #429 moved the no-name-segment test down because group asked by segment INDEX where assign asks by CONTENT. The destination follows the LAYER, not the topic: a predicate over token text goes to `_vocab`, one over pieces and tags to `_pieces`. Both are leaves the stages sit on. The piece layer got its own module only in #439 — until then those predicates collected in `_group`, not because grouping owned them but because `_assign` imports `_group` and cannot be imported back, so group was the one place both stages could reach; five had accumulated across four PRs before the module existed. Stage order is this mechanism's limit, and it forecloses the alternative: where the reader comes AFTER the decider, record the answer on the state instead — `ParseState.order` is that shape, "Recorded rather than recomputed downstream, because the two can differ" — which is unavailable whenever the EARLIER stage is the one asking. (The concrete assign→group import that forced the `_group` collection is gone since #439; what remains is the ordering it was a symptom of, and tests/v2/test_layering.py is where the leaf's contract is now written down.) The cost is a second evaluation of the same predicate, measured for #429 at 1.2–2.2% of a family-comma parse and 0% of every other; recording that number was the right answer there over plumbing a state field the two sites would not otherwise share. Lives in. nameparser/_pipeline/_vocab.py over text (is_wholly_suffix; is_trailing_numeral_suffix — the #401/#421 instance, whose only caller since #439 is the shared peel rather than a stage; and maiden_marker_run, the #434 instance and the clearest two-stage case, called by classify over token texts and by extract over a clause's whitespace words, with group reading the tags classify recorded because it runs later; and delimiter_cores, the #436/#437 instance, read by group where a tail segment DROPS a configured delimiter core and by post_rules where the suffix view's entry boundary asks whether a dropped token was one, with a third reader inside this same module, is_wholly_suffix, where a configured core counts as suffix-shaped; and in_initialless_script, the #322/#323 instance and the only one here that is a REPERTOIRE test rather than a vocabulary one — the script half of the #320 initial veto, read by is_initial one function away and by _pieces.is_leading_title, so "a script with no initials has no period abbreviations either" is one predicate over _policy._NO_INITIALS rather than a second reading of that table; it lost its leading underscore when the second caller arrived; and caps_shape_candidate, the #516 instance and the newest, called from the sites that each needed the identical question answered — classify's own tag emission and _segment.py's two comma tests (the all-caps run and, since #564, the #544 run test's by-shape member), with its unit tests — where the usual reason for keeping such copies apart (a shared call costing every default-policy parse a frame it cannot use) does not hold: every caller asks first, in C, what the predicate would decline anyway — the setting and position, then alphabetic capitals and no listed suffix word — so only an unlisted all-caps word reaches the call, at the default as under any setting (decisions.md#S2)) and nameparser/_pipeline/_pieces.py over pieces: is_suffix_piece, leading_titles and peel_walk are called by both stages, while is_leading_title, is_title_piece and trailing_start are called by group alone (measured 2026-09-06 by call site: `is_leading_title` has no caller in `_assign.py`, which reads `leading_titles` instead — a first draft of this clause listed it among the shared ones) — `trailing_start` being the one to know, since it answers where the trailing run begins and is what P2's chain stops at, and M2's walk wherever no trailing rule reads the clause (elsewhere, since #535, the walk stops where `tail_reading` says) — and segment_suffix_reading by assign alone since #436/#437, that last one being #430's instance, where THREE readers shared one answer until the render join, group's third, was replaced by a rule over the commas the writer typed (decisions.md#C1, 2026-09-06); it stays where it is, one call site being no reason to move a predicate that two sites will contest again. `trailing_titles` was that last shape for one day (2026-09-08, the #316/#489 bundle, rules.md#H5), and since the /simplify round of 2026-09-09 the SHARED predicate is `tail_reading` instead — the peel-and-chain fixed point that answers where the name pieces end (decisions.md#H5). Assign calls it at its main walk and group's bound-given reserve calls it twice, once per view the join compares, because that reserve reads the name words assign will leave and this walk is half of what leaves them (rules.md#P5; counting a trailing title word among them joined 'Prof. abdul rahman Prof.' where 'Prof. abdul rahman' does not). Since #535 group's maiden walk calls it as well, over the clause and over the view its take would leave, wherever a trailing rule reads the clause (rules.md#M2), for the same reason: the walk's stops must end the clause where assign's reading of the name will begin. `peel_trailing` and `trailing_titles` are what that fixed point is BUILT from, and neither is a two-stage question any longer: `peel_trailing` has one caller outside `_pieces.py`, the maiden walk in `_group.py`, which asks the peel itself because it needs ONE half of the answer at a time -- the numeral's over the pieces as written and again over the view its take would leave (#424), the acronym's beside it (#533) -- where `trailing_start` and `tail_reading`, the two callers in the leaf, fold both halves into one index; since #535 it asks the bare peel only where no trailing rule reads the clause, and reads `tail_reading` everywhere else, so that a trailing title does not hide the numeral or credential in front of it; that walk is a reader of the peel and not a second spelling of it, the question being asked of a different name each time. `trailing_titles` has exactly one caller, assign's family-comma segment-1 walk, which reads the chain without the re-peel, and `_group.py` does not import it. The tail reading is in the leaf rather than inline because each assign site had been given a cheap frame-free gate written to match the walk's own first condition, which is a second implementation of the question and was removed in review; what the leaf costs is one frame per entry point, measured, and the walk's own first test is a compiled regex rather than a call, so an ordinary name pays a match and stops. The reserve's two calls cost the reference name nothing — it never enters that branch, having no bound given word — and the parse and facade frame counts did not move (measured 2026-09-09). Re-measured 2026-09-09 by an AST call-site census over `_pipeline/*.py` — every call node whose callee is one of these names, keyed by module and enclosing function, which is what caught the census claiming a share for `peel_trailing` that the round had just taken away — the rest of it holds unchanged: is_suffix_piece, leading_titles, peel_walk and now tail_reading shared, is_leading_title, is_title_piece and trailing_start group-only — assign still reads `leading_titles` and never `is_leading_title`, which is what keeps H2's shape inference out of the trailing slot. And nameparser/_pipeline/_post_rules.py over a state: suffix_entries, the #511 instance, the R1 entry pass as a function, the one instance living in a stage rather than in a leaf — it is a pass over a whole ParseState and no leaf takes one, and AGENTS.md names it as the exception — run by post_rules last in the stage (through its in-place worker) and by Parser.revise over a sub-parse whose roles it has forced, so a suffix value handed to revise() derives its entries by the rule a whole name uses rather than by a second reading of the value's commas (decisions.md#C1, 2026-09-06 #511). Two more since #610 (2026-10-04), both in `_pieces.py`: `particle_tail`, P6's walk, asked by post_rules' attachment and by assign's given-part credential run, which had been two walks agreeing only through the run's roles; and `trailing_candidates`, the maiden take's choice of clause words, written beside the peel whose admissions it must cover, with a sweep test against `tail_reading` that found two shapes the hand copy had missed (decisions.md#M2). tests/v2/test_layering.py holds each module's contract, and a piece predicate growing a dependency on a STAGE shows up there as a widened entry. Reach for it when. You are about to write a condition that mirrors, matches or "does what X does" — or you find a comment saying one does. Grep for the other site's predicate and call it instead. ## RENDER-HONORS-THE-PARSE — the parse decides it, the views honor it diff --git a/docs/design/rules.md b/docs/design/rules.md index a53f5b42..8a0033d6 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -963,7 +963,7 @@ P6. Rationale: a particle ending the name has nothing to link negative-control sweep pinning the disagreeing set the precedence bullet above names. A change that breaks one side of that pair should expect that test, not this file, to say so first. - history: decisions.md#P6 · interacts: A1, C1, P1, S2, P5, M2 · implemented: nameparser/_pipeline/_assign.py, nameparser/_pipeline/_post_rules.py + history: decisions.md#P6 · interacts: A1, C1, P1, S2, P5, M2 · implemented: nameparser/_pipeline/_assign.py, nameparser/_pipeline/_pieces.py, nameparser/_pipeline/_post_rules.py P7. Rationale: a one-letter particle is spelled with the same letter as an initial, and the period is what tells them apart. An @@ -1533,6 +1533,7 @@ M2. Rationale: a maiden marker announces that what follows it is the "Jane Doe nee Smith MA Prof." → suffix="MA" "Jane Doe nee Smith Prof. MA" → maiden="Smith" "Jane Doe nee Smith V Prof." → suffix="V" + "Jane Doe nee Smith VI Prof." → suffix="VI" "Jane Doe nee King." → maiden="King." · boundary "Jane Doe nee Prof. Dr." → maiden="Prof." · boundary "Jane van der Berg nee Smith Prof." → title="Prof." @@ -1582,7 +1583,7 @@ M2. Rationale: a maiden marker announces that what follows it is the Accepted: a marker behind a credential is an ordinary word, and the credential's run (S2) takes it in with everything after it. "Jane Doe Jr. nee Smith" → suffix="Jr. nee Smith" · boundary - history: decisions.md#M2 · interacts: P2, P3, P5, P6, R1, R2, M1, S1, S2, H1, H5, C1 · implemented: nameparser/_pipeline/_group.py + history: decisions.md#M2 · interacts: P2, P3, P5, P6, R1, R2, M1, S1, S2, H1, H5, C1 · implemented: nameparser/_pipeline/_group.py, nameparser/_pipeline/_pieces.py M3. Rationale: an enclosure says nothing about whether it means maiden, but a recognized marker word inside it does — the clause diff --git a/nameparser/_pipeline/_assign.py b/nameparser/_pipeline/_assign.py index beaea73b..64d60eef 100644 --- a/nameparser/_pipeline/_assign.py +++ b/nameparser/_pipeline/_assign.py @@ -77,7 +77,7 @@ from nameparser._pipeline._pieces import ( anchor_in_reach, credential_at_the_given_slot, given_slot_anchors, _NOT_A_RUN_START, has_name_content, is_lone_never_given_particle, - is_suffix_piece, is_title_piece, is_wholly_particle, leading_titles, + is_suffix_piece, is_title_piece, leading_titles, particle_tail, listed_lean, peel_walk, segment_suffix_reading, starts_a_credential_run, tail_reading, trailing_titles, ) @@ -1080,35 +1080,21 @@ def reads_as_a_suffix(m: int, titled: frozenset[int]) -> bool: break # The particle tail P6 will attach is not the run's to take # (rules.md#P6: "a particle ending the name attaches to that - # family name", looking past the post-nominals behind it): - # the wholly-particle pieces ending the part, behind any run - # words, found as P6 finds them. They are left to the walk - # below and reach P6 with the role they had before #602. - # Absorbing them reported a suffix reading P6 then overrode, - # and P6 reported the override as a declined post-nominal - # ('Smith, John PhD de', 'Smith, John PhD de Jr.'). A - # particle P6 will NOT attach -- one with a credential - # behind it and another particle past that, 'Smith, John - # PhD de PhD van' -- stays in the run, as does a lone member - # of the ambiguous credential class (`do`): read as the - # credential, it is the word P6's #531 exception keeps out - # of the attachment, so the run and P6 agree on it. - # `pieces[p6_lo:p6_hi]` is that tail: back past the run - # words behind it, then over the particles, stopping at a - # class member + # family name"): `pieces[p6_lo:p6_hi]`, found by the walk P6 + # itself runs (#610), the run's words still holding no role + # here. They are left to the walk below and reach P6 with + # the role they had before #602. Absorbing them reported a + # suffix reading P6 then overrode, and P6 reported the + # override as a declined post-nominal ('Smith, John PhD de', + # 'Smith, John PhD de Jr.'). A particle P6 will NOT attach -- + # one with a credential behind it and another particle past + # that, 'Smith, John PhD de PhD van' -- stays in the run, as + # does a lone member of the ambiguous credential class + # (`do`), which the run reads as the credential and P6's + # #531 stop keeps out of the attachment. p6_lo = p6_hi = len(pieces) if sticky_from < len(pieces): - while (p6_hi > sticky_from - and not is_wholly_particle(pieces[p6_hi - 1], - tokens)): - p6_hi -= 1 - p6_lo = p6_hi - while (p6_lo > sticky_from - and is_wholly_particle(pieces[p6_lo - 1], tokens) - and not (len(pieces[p6_lo - 1]) == 1 - and not tokens[pieces[p6_lo - 1][0]].tags - .isdisjoint(_AMBIGUOUS_CREDENTIAL_TAGS))): - p6_lo -= 1 + p6_lo, p6_hi = particle_tail(pieces, tokens, sticky_from) for m in range(n + 1, len(pieces)): if m in titled_idx: continue diff --git a/nameparser/_pipeline/_group.py b/nameparser/_pipeline/_group.py index bdaa01b6..4e8b6c1d 100644 --- a/nameparser/_pipeline/_group.py +++ b/nameparser/_pipeline/_group.py @@ -48,8 +48,7 @@ from nameparser._lexicon import _run_addresses_by_given from nameparser._pipeline._pieces import ( is_leading_title, is_suffix_piece, is_title_piece, - is_trailing_title_word, starts_a_credential_run, - leading_titles, peel_walk, tail_reading, + leading_titles, peel_walk, tail_reading, trailing_candidates, trailing_start, trailing_start_past_titles, ) from nameparser._pipeline._state import ( @@ -58,7 +57,7 @@ ) from nameparser._pipeline._vocab import D, PH from nameparser._pipeline._vocab import ( - delimiter_cores, is_trailing_numeral_suffix, + delimiter_cores, ) from nameparser._types import AmbiguityKind, Role @@ -226,33 +225,6 @@ def _marker_run_pieces(pieces: Sequence[Sequence[int]], tokens[pieces[k][0]].tags for k in range(m + 1, len(pieces))) -#: the tags `_clause_tail_word` admits a lone word on -_CLAUSE_TAIL_TAGS = _AMBIGUOUS_CREDENTIAL_TAGS | {"vocab:suffix"} - - -# rules.md#M2: "It takes them up to the trailing run of post-nominals -# and titles that the end of the name reads as if the clause were not -# written" (#601). A word that may belong to that run -- the view -# decides which of them it actually takes. -def _clause_tail_word(piece: Sequence[int], ptags: Set[str], - tokens: Sequence[WorkToken]) -> bool: - """Suffix vocabulary, a title word H5's chain takes from the end - (period-marked: a BARE title word ending a name is a name word, and - pulled into the view it would only inflate the count of words to - spare), or a member of the ambiguous class: a word the end of the - clause-free name might read as a post-nominal or a title. Being one - is no answer -- `_maiden_take` reads the view to find out which of - them the trailing rule actually takes. A bare title inside #602's - credential run reaches the view another way, through the run's own - start.""" - if ("suffix" in ptags or is_suffix_piece(piece, ptags, tokens) - or is_trailing_title_word(piece, ptags, tokens)): - return True - return (len(piece) == 1 - and not tokens[piece[0]].tags.isdisjoint(_CLAUSE_TAIL_TAGS)) - - - def _maiden_take(pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], tokens: Sequence[WorkToken], @@ -324,23 +296,7 @@ def _maiden_take(pieces: Sequence[Sequence[int]], for j in range(lo, end)] if site is not ClauseSite.TRAILING: assert_never(site) - # the last piece is a candidate on the numeral fork's own SHAPE test - # too, which asks no vocabulary ('VI' is in no suffix list, and - # 'John Smith VI' reads it as the suffix all the same) - c = len(pieces) - while c - 1 > lo and ( - _clause_tail_word(pieces[c - 1], ptags[c - 1], tokens) - or (c == len(pieces) and len(pieces[c - 1]) == 1 - and is_trailing_numeral_suffix( - tokens[pieces[c - 1][0]].text, - tokens[pieces[c - 2][0]].text))): - c -= 1 - # the tag test is the predicate's own necessary half, inline so a - # name word in the clause pays no frame - c = next((j for j in range(lo + 1, c) - if ("suffix" in ptags[j] - or "vocab:suffix" in tokens[pieces[j][0]].tags) - and starts_a_credential_run(pieces[j], ptags[j], tokens)), c) + c = trailing_candidates(lo, pieces, ptags, tokens) end = len(pieces) released: list[tuple[int, Role]] = [] # nothing behind the clause that a trailing rule could take: the diff --git a/nameparser/_pipeline/_pieces.py b/nameparser/_pipeline/_pieces.py index e8f166f4..ac75196d 100644 --- a/nameparser/_pipeline/_pieces.py +++ b/nameparser/_pipeline/_pieces.py @@ -57,7 +57,8 @@ from typing import NamedTuple from nameparser._pipeline._state import ( - AMBIGUOUS_ACRONYM_TAG, SHAPE_ACRONYM_TAG, WorkToken, + AMBIGUOUS_ACRONYM_TAG, NAME_ROLES, SHAPE_ACRONYM_TAG, SUFFIX_OR_UNREAD, + WorkToken, ) from nameparser._pipeline._vocab import ( _PERIOD_ABBREV, Lean, ambiguous_lean, in_initialless_script, @@ -941,13 +942,66 @@ def run_start(rest: Sequence[int], names: int, return names -def is_wholly_particle(piece: Sequence[int], - tokens: Sequence[WorkToken]) -> bool: - """Whether every token of a piece is particle vocabulary -- the - unit rules.md#P6 attaches after a family comma.""" - if len(piece) == 1: - return "particle" in tokens[piece[0]].tags - return all("particle" in tokens[i].tags for i in piece) +# rules.md#P6: "a particle ending the name attaches to that family +# name" -- WHICH particles, asked by the attachment in post_rules and by +# assign's given-part credential run (#602), which leaves them to it +# (#610: the two had been two walks agreeing only by the run's roles) +def particle_tail(seg: Sequence[Sequence[int]], + tokens: Sequence[WorkToken], + floor: int = 0) -> tuple[int, int]: + """`seg[lo:hi]`, the run of wholly-particle pieces P6 attaches: back + from the end past pieces that hold no name word and are not + themselves particles -- a post-nominal is written BEHIND the + particle in this listing -- then back over particle pieces, + stopping at a lone member of the ambiguous credential class read as + the credential (#531: the capitals or a degree made it one, so it + is not the tussenvoegsel). Neither walk passes `floor`. + + The class-member stop keys on the VOCABULARY tag and not on the + suffix role alone, which is what gives P6's attachment precedence + over S2: `vd` and `mc` are unambiguous suffix vocabulary, also + particles, also suffix-roled, and the tag is what keeps them inside + the run, while the role alone would have stood the attachment down + for #531's member too ('Doe, John DO' read family 'DO Doe' with the + condition absent, verified when #531 landed). A trailing piece that + IS particle vocabulary ends the first walk rather than being looked + past, since `vd` arrives suffix-roled and is the run. + + Read off ROLES, which is what lets both callers ask it: post_rules + after assign has placed every word, and assign before it places the + words of a credential run, whose roles are then still None -- no + name role, and a class member not yet read as anything, the run + being what will read it as the credential. `lo == hi` where there + is no such run. + + The single-token tests are inline (the common piece is one token), + so the walk pays no frame per piece; it runs on every family-comma + parse.""" + hi = len(seg) + while hi > floor: + piece = seg[hi - 1] + if (all("particle" in tokens[i].tags for i in piece) + if len(piece) > 1 else + "particle" in tokens[piece[0]].tags): + break + if (any(tokens[i].role in NAME_ROLES for i in piece) + if len(piece) > 1 else + tokens[piece[0]].role in NAME_ROLES): + break + hi -= 1 + lo = hi + while lo > floor: + piece = seg[lo - 1] + if len(piece) == 1: + tok = tokens[piece[0]] + if "particle" not in tok.tags or ( + AMBIGUOUS_ACRONYM_TAG in tok.tags + and tok.role in SUFFIX_OR_UNREAD): + break + elif not all("particle" in tokens[i].tags for i in piece): + break + lo -= 1 + return lo, hi def has_name_content(piece: Sequence[int], @@ -1071,6 +1125,70 @@ def is_trailing_title_word(piece: Sequence[int], ptags: Set[str], and is_title_piece(piece, ptags, tokens)) +#: lone-word tags `trailing_candidates` admits on: the ambiguous class, +#: listed or by shape, and any suffix word -- an initial-shaped one +#: included, which `is_suffix_piece` vetoes and the numeral fork reads +_TRAILING_WORD_TAGS = frozenset({AMBIGUOUS_ACRONYM_TAG, SHAPE_ACRONYM_TAG, + "vocab:suffix"}) + + +# rules.md#M2: "It takes them up to the trailing run of post-nominals +# and titles that the end of the name reads as if the clause were not +# written" -- which words of a clause might be in that run, written +# beside the peel and the chain whose admissions it has to cover (#610) +def trailing_candidates(lo: int, pieces: Sequence[Sequence[int]], + ptags: Sequence[Set[str]], + tokens: Sequence[WorkToken]) -> int: + """The first of the pieces after `lo` that `tail_reading` might read + into the trailing run: `pieces[c:]`, a SUPERSET of what it takes, + for the maiden take to build its clause-free view over and let + `tail_reading` decide. `pieces[lo]` itself is never a candidate. + + Back from the end over every piece the peel or the chain admits on + its own terms -- a suffix piece, a member of the ambiguous class, + an initial-shaped suffix word, a period-marked title word -- and + over a roman numeral by the numeral fork's SHAPE test (which asks + no vocabulary: 'VI' is in no list), where nothing stands behind it + but such title words and group-flagged suffix pieces: the chain + takes the titles first and the walk never holds the flagged pieces + ('Ph. D.'), so either way the numeral is the walk's last piece and + the fork reads it. Then back to the first + credential that starts #602's run, whose words the run takes + whatever they are. A BARE title word is no candidate: H5's chain + does not take one, and in the view it would only inflate the count + of words to spare. + + The superset is pinned by tests/v2/pipeline/test_pieces.py over a + generated grid against `tail_reading` itself, so a new admission + in the peel or the chain fails there until this covers it.""" + c = len(pieces) + titles_behind = True + while c - 1 > lo: + k = c - 1 + piece = pieces[k] + title = is_trailing_title_word(piece, ptags[k], tokens) + if not (title or "suffix" in ptags[k] + or is_suffix_piece(piece, ptags[k], tokens) + or (len(piece) == 1 and ( + not tokens[piece[0]].tags.isdisjoint(_TRAILING_WORD_TAGS) + or (titles_behind and is_trailing_numeral_suffix( + tokens[piece[0]].text, + tokens[pieces[k - 1][0]].text))))): + break + # a group-flagged suffix piece is outside the peel's walk + # altogether (`peel_walk` drops it), so it is behind the + # numeral the way a chained title is + titles_behind = titles_behind and (title or "suffix" in ptags[k]) + c = k + # the tag test is the predicate's own necessary half, inline so a + # name word pays no frame + return next((j for j in range(lo + 1, c) + if ("suffix" in ptags[j] + or "vocab:suffix" in tokens[pieces[j][0]].tags) + and starts_a_credential_run(pieces[j], ptags[j], tokens)), + c) + + # rules.md#H5: "the title is TRANSPARENT to the suffix reading: where # two or more name words stand, what stands once the chain is taken # reads exactly as it would read written without the title, plus the diff --git a/nameparser/_pipeline/_post_rules.py b/nameparser/_pipeline/_post_rules.py index a1334d75..4d33444c 100644 --- a/nameparser/_pipeline/_post_rules.py +++ b/nameparser/_pipeline/_post_rules.py @@ -25,9 +25,12 @@ from nameparser._lexicon import _run_addresses_by_given from nameparser._pipeline._assign import _name_positions -from nameparser._pipeline._pieces import is_lone_never_given_particle +from nameparser._pipeline._pieces import ( + is_lone_never_given_particle, particle_tail, +) from nameparser._pipeline._state import ( - AMBIGUOUS_ACRONYM_TAG, ParseState, PendingAmbiguity, Structure, + NAME_ROLES, + ParseState, PendingAmbiguity, Structure, WorkToken, _NEVER_FLIPPED, comma_bucket, copy_with, ) from nameparser._pipeline._vocab import delimiter_cores, unit_ends @@ -51,8 +54,6 @@ r"^(оглу|оглы|оғлу|ўғли|угли|кызы|гызы|қызы|қизи|улы|ұлы|уулу)$", re.I) -_NAME_ROLES = (Role.GIVEN, Role.MIDDLE, Role.FAMILY) - #: The roles that are transparent to a run of post-nominals (R1's #: entry pass below). These three roles render into fields other than #: the name and the suffix, so a run of @@ -242,7 +243,7 @@ def _leading_name_piece(state: ParseState, if seg >= len(state.pieces): return () for piece in state.pieces[seg]: - if any(tokens[i].role in _NAME_ROLES for i in piece): + if any(tokens[i].role in NAME_ROLES for i in piece): return piece return () @@ -706,40 +707,14 @@ def post_rules(state: ParseState) -> ParseState: # the family view reads the tag and renders these before the base. if state.structure is Structure.FAMILY_COMMA and len(state.pieces) > 1: seg = state.pieces[1] - # A post-nominal sits BEHIND the tussenvoegsel in this listing - # ("Berg, Jan van Jr."), so the run is found by walking past a - # trailing piece that holds no name -- but only one that is not - # itself particle vocabulary, since `vd` arrives suffix-roled - # and IS the run. Without this the same name parsed two ways on - # whether a comma preceded the credential. - end = len(seg) - while (end - and not any(tokens[i].role in _NAME_ROLES - for i in seg[end - 1]) - and not all("particle" in tokens[i].tags - for i in seg[end - 1])): - end -= 1 - k = end - # #531: a class member the given-part slot read as a - # credential is NOT part of the run. P6 keys on vocabulary - # rather than role by design, which is what gives its - # attachment precedence over S2 -- so assign's suffix role - # alone does not stand it down, verified by running #531's - # assign half with this condition absent ('Doe, John DO' read - # family 'DO Doe', the suffix role silently overridden). - # Narrowed to AMBIGUOUS_ACRONYM_TAG rather than to the - # suffix role: `vd` and `mc` are unambiguous suffix - # vocabulary, also particles, also suffix-roled, and the tag - # is what keeps them inside the run. - while k and all("particle" in tokens[i].tags - for i in seg[k - 1]) \ - and not (len(seg[k - 1]) == 1 - and tokens[seg[k - 1][0]].role is Role.SUFFIX - and AMBIGUOUS_ACRONYM_TAG - in tokens[seg[k - 1][0]].tags): - k -= 1 + # `seg[k:end]` is the run, found by the walk assign's given-part + # credential run asks too (#610): past a trailing post-nominal, + # since one sits BEHIND the tussenvoegsel in this listing ("Berg, + # Jan van Jr."), and short of a class member read as the + # credential (#531). `particle_tail` carries both reasons. + k, end = particle_tail(seg, tokens) # GIVEN alone, which is what P6 says ("provided at least one - # given word remains"). Not `_NAME_ROLES`: P1's fold runs + # given word remains"). Not `NAME_ROLES`: P1's fold runs # earlier in this function and retags all of segment 1 to # FAMILY, so a test for "some name word remains" passes on # family text P1 just produced, and the rule then hoists the diff --git a/nameparser/_pipeline/_state.py b/nameparser/_pipeline/_state.py index 2b48585d..de173197 100644 --- a/nameparser/_pipeline/_state.py +++ b/nameparser/_pipeline/_state.py @@ -96,6 +96,13 @@ class WorkToken: _AMBIGUOUS_CREDENTIAL_TAGS = frozenset( {AMBIGUOUS_ACRONYM_TAG, SHAPE_ACRONYM_TAG}) +#: The roles that make a piece a name word (rules.md#P6's walk), and +#: the roles of a word not yet read as one: a suffix, or nothing yet +#: -- what `_pieces.particle_tail` reads, asked by post_rules after +#: assign and by assign before its credential run places its words. +NAME_ROLES = (Role.GIVEN, Role.MIDDLE, Role.FAMILY) +SUFFIX_OR_UNREAD = (Role.SUFFIX, None) + class Structure(Enum): """segment's comma-structure decision.""" diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 057a05e8..e8b384f2 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -1038,6 +1038,27 @@ def _check_cjk_shape_purity(self) -> None: "read the suffix. #601's prototype tested vocabulary only " "and kept 'VI' in the maiden name; a unit test under a " "reduced lexicon caught it"), + Case("a_shape_numeral_behind_which_a_title_stands_ends_the_clause", + "Jane Doe nee Smith VI Prof.", + {"title": "Prof.", "given": "Jane", "family": "Doe", + "suffix": "VI", "maiden": "Smith"}, + ambiguities=("suffix-or-name",), + notes="#610: the chain takes 'Prof.' first and leaves 'VI' the " + "walk's last piece, so the numeral fork reads it, as in " + "'John Smith VI Prof.'. The candidate loop #601 shipped " + "admitted a shape numeral only as the clause's last word " + "and kept maiden 'Smith VI'; sharing the candidates with " + "the peel (`_pieces.trailing_candidates`) closed it"), + Case("a_shape_numeral_before_a_merged_credential_ends_the_clause", + "Jane Doe nee Smith VI Ph. D.", + {"given": "Jane", "family": "Doe", "suffix": "VI Ph. D.", + "maiden": "Smith"}, + ambiguities=("suffix-or-name",), + notes="#610: the merged 'Ph. D.' is group-flagged and outside the " + "peel's walk, so 'VI' is the walk's last piece, as in 'Jane " + "Doe VI Ph. D.'. A first draft of the #610 fix covered the " + "title half only; the sweep test over `trailing_candidates` " + "found this one"), Case("a_bound_given_word_is_no_run", "Mohamed Ali Abd Allah", {"given": "Mohamed", "middle": "Ali Abd", "family": "Allah"}, notes="'abd' is in the credential list (ABD) and heads 'Abd " diff --git a/tests/v2/pipeline/test_pieces.py b/tests/v2/pipeline/test_pieces.py index 8e8eb247..6c092b36 100644 --- a/tests/v2/pipeline/test_pieces.py +++ b/tests/v2/pipeline/test_pieces.py @@ -22,7 +22,7 @@ credential_anchors, credential_at_the_given_slot, is_leading_title, leading_titles, own_words, peel_trailing, peel_walk, segment_suffix_reading, - trailing_titles, + tail_reading, trailing_candidates, trailing_titles, ) from nameparser._pipeline._segment import segment from nameparser._pipeline._state import ( @@ -886,3 +886,60 @@ def test_the_given_part_after_a_family_comma_reads_the_run_too() -> None: def test_a_title_inside_the_given_parts_run_is_a_title() -> None: name = parse("Holder, Eric Jr. Attorney General") assert (name.title, name.suffix) == ("Attorney General", "Jr.") + + +# #610: `trailing_candidates` is the maiden take's choice of which +# clause words to read the clause-free view over, and its contract is +# a SUPERSET of what `tail_reading` takes. The grid reads plain names, +# where the candidates' cut must fall at or before every piece the tail +# reading takes (the first name piece aside: in a clause that is the +# word after the marker, which the take always keeps). +_CANDIDATE_HEADS = ("Jane Doe", "J.", "Jane van der Berg", "Mai Le", + "abdul Berg", "John") +_CANDIDATE_WORDS = ("Smith", "VI", "V", "III", "MA", "Ma", "PhD", "Jr.", + "Prof.", "Dr.", "King.", "do", "DO", "de", "X.Y.Z.", + "Ph. D.", "i", "Jones") + + +def _candidate_misses() -> list[str]: + misses = [] + for head in _CANDIDATE_HEADS: + for n in (1, 2, 3): + for tail in itertools.product(_CANDIDATE_WORDS, repeat=n): + text = f"{head} {' '.join(tail)}" + state = _through_group(text) + pieces, ptags = state.pieces[0], state.piece_tags[0] + tokens = state.tokens + lead = leading_titles(pieces, ptags, tokens) + rest = peel_walk(lead, ptags) + if not rest: + continue + kept, chained, peel = tail_reading( + rest, pieces, ptags, tokens, state.one_case) + taken = (set(kept[peel.names:]) | set(chained)) - {rest[0]} + c = trailing_candidates(rest[0], pieces, ptags, tokens) + if any(j < c for j in taken): + misses.append(text) + return misses + + +def test_the_trailing_candidates_cover_what_the_tail_reading_takes() -> None: + """M2's take ends the clause at the trailing run the clause-free + name reads, which it can only find if every word that run might + hold is in its view. 6 heads x 18 words to a + tail of three, 37,044 parses through `group`, about 2.4s on py3.11 + (measured 2026-10-04). + + RECORDED NEGATIVE CONTROL, measured 2026-10-04 on this grid: the + candidate loop as #601 shipped it (master 8a8459a8), which admitted + a roman numeral by its shape only as the clause's LAST word, misses + 288 texts -- every one a numeral in no wordlist ('VI') with a + period-marked title or the merged 'Ph. D.' behind it. The chain + takes the title first, and the walk never holds the flagged + 'Ph. D.', so either way the numeral is the walk's last piece and + the fork reads it: 'Jane Doe nee Smith VI Prof.' kept maiden 'Smith + VI' where 'John Smith VI Prof.' reads suffix 'VI'. A first draft of + the fix covered the title and still missed the 'Ph. D.' half, 111 + texts, which this grid is what found. + """ + assert _candidate_misses() == [] diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index 694258df..2db0ac9e 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -3078,7 +3078,7 @@ class _LatinCopy(NamedTuple): frozenset({'Jane Doe nee Puig i Soler', 'Smith, John, PhD née Puig - i Soler', 'Smith, John, PhD née Puig Mr\\. - i Soler'}), - frozenset({'Jane Doe nee Smith DO DO', + frozenset({'Jane Doe nee Smith DO DO', 'Jane Doe nee Smith VI Prof\\.', 'Jane van der Berg nee Prof\\. King\\. MA', 'Jane van der Berg nee Smith Prof\\.', 'John nee Prof\\. ba MA'}), frozenset({'Jane Doe nee Smith MA JD', 'Jane Doe nee Smith MA PhD', @@ -3648,8 +3648,10 @@ def _claim(rule: dict) -> _Claim: # Smith', rules.md#C1's Accepted example. Reach -- it carries a # marker. # 2026-10-04, #601/#602: 109 -> 98; the rules.md#M2 examples #601 retired left the corpus, and the names #601 moved left this rule's regex. + # 2026-10-04, #610: +1, 'Jane Doe nee Smith VI Prof.', rules.md#M2's + # numeral example behind a title. Reach. "fix(#274) maiden markers consumed": - _Claim(98, ('family', 'maiden', 'middle'), "7b41ba0f0b38", None), + _Claim(99, ('family', 'maiden', 'middle'), "b9259620e118", None), # 2026-09-19, #533: 5 -> 6, the same one new corpus name # '田中 太郎 旧姓 佐藤 MA' as the CJK rule above. "fix(cjk-maiden-marker) maiden marker consumed, compounding with the CJK order flip": @@ -4559,8 +4561,10 @@ def _claim(rule: dict) -> _Claim: _Claim(6, ('family', 'middle', 'suffix', 'title'), "007582298dfc", None), "fix(#601/#602) a marker behind a title, a suffix word or a connective is an ordinary word": _Claim(4, ('family', 'middle', 'suffix', 'title'), "456e8a63ed75", None), + # 2026-10-04, #610: +1, 'Jane Doe nee Smith VI Prof.', rules.md#M2's + # numeral example behind a title. Reach. "fix(#274/#601) the clause ends at the clause-free name's trailing run, which the take consumes": - _Claim(4, ('family', 'given', 'maiden', 'middle', 'suffix', 'title'), "d15d58da6be1", None), + _Claim(5, ('family', 'given', 'maiden', 'middle', 'suffix', 'title'), "7d2d5efa8eb5", None), "fix(#274/#601) the first word after the marker is the maiden name": _Claim(1, ('family', 'maiden', 'middle', 'suffix'), "aaf53040b071", None), "fix(#602) a credential after the name core starts a run to the end of its part": @@ -5300,8 +5304,10 @@ def _claim(rule: dict) -> _Claim: _Claim(2, ('maiden', 'suffix'), "59dfc0c40e36", None), "fix(#601/#602) a marker behind a title, a suffix word or a connective is an ordinary word": _Claim(5, ('_ambiguities', 'family', 'given', 'maiden', 'suffix', 'title'), "1c9432e08a21", None), + # 2026-10-04, #610: +1, 'Jane Doe nee Smith VI Prof.', rules.md#M2's + # numeral example behind a title. Reach. "fix(#601) the clause ends at the clause-free name's trailing run, which the take consumes": - _Claim(4, ('_ambiguities', 'family', 'given', 'maiden', 'suffix', 'title'), "d15d58da6be1", None), + _Claim(5, ('_ambiguities', 'family', 'given', 'maiden', 'suffix', 'title'), "7d2d5efa8eb5", None), "fix(#602) a credential after the name core starts a run to the end of its part": _Claim(5, ('_ambiguities', 'family', 'middle', 'suffix', 'title'), "3129cd9609b9", None), "fix(#601/#602) a credential in the clause starts a run the take consumes": @@ -5764,8 +5770,10 @@ def _claim(rule: dict) -> _Claim: _Claim(2, ('maiden', 'suffix'), "59dfc0c40e36", None), "fix(#601/#602) a marker behind a title, a suffix word or a connective is an ordinary word": _Claim(5, ('_ambiguities', 'family', 'given', 'maiden', 'suffix', 'title'), "1c9432e08a21", None), + # 2026-10-04, #610: +1, 'Jane Doe nee Smith VI Prof.', rules.md#M2's + # numeral example behind a title. Reach. "fix(#601) the clause ends at the clause-free name's trailing run, which the take consumes": - _Claim(4, ('_ambiguities', 'maiden', 'suffix', 'title'), "d15d58da6be1", None), + _Claim(5, ('_ambiguities', 'maiden', 'suffix', 'title'), "7d2d5efa8eb5", None), "fix(#601) the first word after the marker is the maiden name": _Claim(1, ('_ambiguities', 'family', 'maiden', 'middle', 'suffix'), "aaf53040b071", None), "fix(#602) a credential after the name core starts a run to the end of its part": @@ -6453,8 +6461,10 @@ def _claim(rule: dict) -> _Claim: _Claim(2, ('maiden', 'suffix'), "59dfc0c40e36", None), "fix(#601/#602) a marker behind a title, a suffix word or a connective is an ordinary word": _Claim(5, ('_ambiguities', 'family', 'given', 'maiden', 'suffix', 'title'), "1c9432e08a21", None), + # 2026-10-04, #610: +1, 'Jane Doe nee Smith VI Prof.', rules.md#M2's + # numeral example behind a title. Reach. "fix(#601) the clause ends at the clause-free name's trailing run, which the take consumes": - _Claim(4, ('_ambiguities', 'family', 'given', 'maiden', 'suffix', 'title'), "d15d58da6be1", None), + _Claim(5, ('_ambiguities', 'family', 'given', 'maiden', 'suffix', 'title'), "7d2d5efa8eb5", None), "fix(#602) a credential after the name core starts a run to the end of its part": _Claim(5, ('_ambiguities', 'family', 'middle', 'suffix', 'title'), "3129cd9609b9", None), "fix(#601/#602) a credential in the clause starts a run the take consumes": @@ -6770,8 +6780,10 @@ def _claim(rule: dict) -> _Claim: _Claim(2, ('maiden', 'suffix'), "59dfc0c40e36", None), "fix(#601/#602) a marker behind a title, a suffix word or a connective is an ordinary word": _Claim(5, ('_ambiguities', 'family', 'given', 'maiden', 'suffix'), "1c9432e08a21", None), + # 2026-10-04, #610: +1, 'Jane Doe nee Smith VI Prof.', rules.md#M2's + # numeral example behind a title. Reach. "fix(#601) the clause ends at the clause-free name's trailing run, which the take consumes": - _Claim(4, ('_ambiguities', 'maiden', 'suffix', 'title'), "d15d58da6be1", None), + _Claim(5, ('_ambiguities', 'maiden', 'suffix', 'title'), "7d2d5efa8eb5", None), "fix(#601) the first word after the marker is the maiden name": _Claim(1, ('_ambiguities', 'family', 'maiden', 'middle', 'suffix'), "aaf53040b071", None), "fix(#602) a credential after the name core starts a run to the end of its part": diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index 6bc24104..987e9276 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -146,6 +146,7 @@ "Jane Doe nee Smith Prof." "Jane Doe nee Smith Prof. MA" "Jane Doe nee Smith V Prof." +"Jane Doe nee Smith VI Prof." "Jane Doe nee Smith X.Y.Z." "Jane Doe, MS LAc" "Jane Smith (Nee)" diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index 203308b3..697a52af 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -4738,6 +4738,9 @@ name_regex = "^(?:Dr\\. nee Smith PhD Prof\\.|Jane Doe Jr\\. nee Smith|Jane Doe fields = ["family", "middle", "suffix", "title"] [[change]] +# 'Jane Doe nee Smith VI Prof.': rules.md#M2's numeral example +# behind a title, which the take reaches since #610 shared its +# candidate words with the peel (2026-10-04). # rules.md#M2: "It takes them up to the trailing run of post-nominals and # titles that the end of the name reads as if the clause were not # written" (#601, 2026-10-04), and "The run so found ends the name: its @@ -4747,7 +4750,7 @@ fields = ["family", "middle", "suffix", "title"] # Against 1.4.0, which had no maiden reading, `maiden` fills # as well (#274). issue = "fix(#274/#601) the clause ends at the clause-free name's trailing run, which the take consumes" -name_regex = "^(?:Jane Doe nee Smith DO DO|Jane van der Berg nee Prof\\. King\\. MA|Jane van der Berg nee Smith Prof\\.|John nee Prof\\. ba MA)$" +name_regex = "^(?:Jane Doe nee Smith DO DO|Jane Doe nee Smith VI Prof\\.|Jane van der Berg nee Prof\\. King\\. MA|Jane van der Berg nee Smith Prof\\.|John nee Prof\\. ba MA)$" fields = ["family", "given", "maiden", "middle", "suffix", "title"] [[change]] diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 2ccc6d6b..6c8a0864 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -3877,6 +3877,9 @@ name_regex = "^(?:Dr\\. nee Smith PhD Prof\\.|Jane Doe Jr\\. nee Smith|Jane Doe fields = ["_ambiguities", "family", "given", "maiden", "suffix", "title"] [[change]] +# 'Jane Doe nee Smith VI Prof.': rules.md#M2's numeral example +# behind a title, which the take reaches since #610 shared its +# candidate words with the peel (2026-10-04). # rules.md#M2: "It takes them up to the trailing run of post-nominals and # titles that the end of the name reads as if the clause were not # written" (#601, 2026-10-04), and "The run so found ends the name: its @@ -3884,7 +3887,7 @@ fields = ["_ambiguities", "family", "given", "maiden", "suffix", "title"] # reaches into it". The old walk's release check kept these words in the # clause where a join it modelled might take them; the run is bound now. issue = "fix(#601) the clause ends at the clause-free name's trailing run, which the take consumes" -name_regex = "^(?:Jane Doe nee Smith DO DO|Jane van der Berg nee Prof\\. King\\. MA|Jane van der Berg nee Smith Prof\\.|John nee Prof\\. ba MA)$" +name_regex = "^(?:Jane Doe nee Smith DO DO|Jane Doe nee Smith VI Prof\\.|Jane van der Berg nee Prof\\. King\\. MA|Jane van der Berg nee Smith Prof\\.|John nee Prof\\. ba MA)$" fields = ["_ambiguities", "family", "given", "maiden", "suffix", "title"] [[change]] diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index 3fe80b3b..989df2c9 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -3822,6 +3822,9 @@ name_regex = "^(?:Dr\\. nee Smith PhD Prof\\.|Jane Doe Jr\\. nee Smith|Jane Doe fields = ["_ambiguities", "family", "given", "maiden", "suffix", "title"] [[change]] +# 'Jane Doe nee Smith VI Prof.': rules.md#M2's numeral example +# behind a title, which the take reaches since #610 shared its +# candidate words with the peel (2026-10-04). # rules.md#M2: "It takes them up to the trailing run of post-nominals and # titles that the end of the name reads as if the clause were not # written" (#601, 2026-10-04), and "The run so found ends the name: its @@ -3829,7 +3832,7 @@ fields = ["_ambiguities", "family", "given", "maiden", "suffix", "title"] # reaches into it". The old walk's release check kept these words in the # clause where a join it modelled might take them; the run is bound now. issue = "fix(#601) the clause ends at the clause-free name's trailing run, which the take consumes" -name_regex = "^(?:Jane Doe nee Smith DO DO|Jane van der Berg nee Prof\\. King\\. MA|Jane van der Berg nee Smith Prof\\.|John nee Prof\\. ba MA)$" +name_regex = "^(?:Jane Doe nee Smith DO DO|Jane Doe nee Smith VI Prof\\.|Jane van der Berg nee Prof\\. King\\. MA|Jane van der Berg nee Smith Prof\\.|John nee Prof\\. ba MA)$" fields = ["_ambiguities", "family", "given", "maiden", "suffix", "title"] [[change]] diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index 1577ce26..4f14e7c3 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -2315,6 +2315,9 @@ name_regex = "^(?:Dr\\. nee Smith PhD Prof\\.|Jane Doe Jr\\. nee Smith|Jane Doe fields = ["_ambiguities", "family", "given", "maiden", "suffix", "title"] [[change]] +# 'Jane Doe nee Smith VI Prof.': rules.md#M2's numeral example +# behind a title, which the take reaches since #610 shared its +# candidate words with the peel (2026-10-04). # rules.md#M2: "It takes them up to the trailing run of post-nominals and # titles that the end of the name reads as if the clause were not # written" (#601, 2026-10-04), and "The run so found ends the name: its @@ -2322,7 +2325,7 @@ fields = ["_ambiguities", "family", "given", "maiden", "suffix", "title"] # reaches into it". The old walk's release check kept these words in the # clause where a join it modelled might take them; the run is bound now. issue = "fix(#601) the clause ends at the clause-free name's trailing run, which the take consumes" -name_regex = "^(?:Jane Doe nee Smith DO DO|Jane van der Berg nee Prof\\. King\\. MA|Jane van der Berg nee Smith Prof\\.|John nee Prof\\. ba MA)$" +name_regex = "^(?:Jane Doe nee Smith DO DO|Jane Doe nee Smith VI Prof\\.|Jane van der Berg nee Prof\\. King\\. MA|Jane van der Berg nee Smith Prof\\.|John nee Prof\\. ba MA)$" fields = ["_ambiguities", "maiden", "suffix", "title"] [[change]] diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index 82746a49..3435d2ef 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -1601,6 +1601,9 @@ name_regex = "^(?:Dr\\. nee Smith PhD Prof\\.|Jane Doe Jr\\. nee Smith|Jane Doe fields = ["_ambiguities", "family", "given", "maiden", "suffix"] [[change]] +# 'Jane Doe nee Smith VI Prof.': rules.md#M2's numeral example +# behind a title, which the take reaches since #610 shared its +# candidate words with the peel (2026-10-04). # rules.md#M2: "It takes them up to the trailing run of post-nominals and # titles that the end of the name reads as if the clause were not # written" (#601, 2026-10-04), and "The run so found ends the name: its @@ -1608,7 +1611,7 @@ fields = ["_ambiguities", "family", "given", "maiden", "suffix"] # reaches into it". The old walk's release check kept these words in the # clause where a join it modelled might take them; the run is bound now. issue = "fix(#601) the clause ends at the clause-free name's trailing run, which the take consumes" -name_regex = "^(?:Jane Doe nee Smith DO DO|Jane van der Berg nee Prof\\. King\\. MA|Jane van der Berg nee Smith Prof\\.|John nee Prof\\. ba MA)$" +name_regex = "^(?:Jane Doe nee Smith DO DO|Jane Doe nee Smith VI Prof\\.|Jane van der Berg nee Prof\\. King\\. MA|Jane van der Berg nee Smith Prof\\.|John nee Prof\\. ba MA)$" fields = ["_ambiguities", "maiden", "suffix", "title"] [[change]] From c7f2b58e1a6348381cd7435075c6d875a607abbd Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sun, 4 Oct 2026 12:45:21 -0700 Subject: [PATCH 2/3] AGENTS.md: record lessons where the next session meets them; sweep spellings Two Workflow-level additions from #610. A rule to write down, in the same PR, the check that would have caught a defect, and where each kind of lesson goes (the sweep paragraphs, Gotchas, docs/design/AGENTS.md's axes, mechanisms.md). And a third sweep beside comma shapes and `name_order`: a word class reached by vocabulary is also reached by shape, with roman numerals as the case -- `ii`-`iv` listed, `i`/`v` listed but initial-shaped, `vi` and up in no list -- which is how #610's 'VI' gap hid behind working 'V' and 'III'. Co-Authored-By: Claude Opus 5.5 --- AGENTS.md | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/AGENTS.md b/AGENTS.md index f283e3fe..08c79f45 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -12,6 +12,8 @@ After a review round, **review the fix commit too** — second-round passes on # **Check whether the PR description still needs updating before merging.** A review that moves behavior or retracts a measurement leaves it stale, and a stale body reads as authoritative. +**When a defect is found that an earlier check would have caught, write the check down where the next person will meet it, in the same PR.** A review finding, a failing sweep or a regression that arrived because nobody thought to test a shape is a lesson about the repository, not only about the change, and it is lost when the PR merges unless it lands somewhere a later session reads before making the same mistake. Ask, before merging: what should I have checked, and where would I have looked? Then put it there — a missed input shape in the sweep paragraphs under Architecture (comma shapes, then `name_order`, then spellings, each earned this way); a trap in a module's behavior under Gotchas; a design-review miss in docs/design/AGENTS.md's axes; a measurement or test-design trap in mechanisms.md's Verification shapes or Field notes; a reusable design move in mechanisms.md proper. Name the PR or issue that earned it and the input that showed it, so the rule can be checked against its case. #610 is the shape: its sweep test found that the maiden take missed a shape-only numeral, and the spelling sweep below records the class (`VI` beside `V` and `III`). This is how the file accumulates expertise rather than history — prefer amending an existing paragraph to adding a new one, and delete a lesson the code has since made impossible. + **A scripted multi-edit that asserts as it goes discards everything when a late pattern misses.** Assert every pattern before applying any, then verify the edit landed — three times in one session a docs batch failed its last assertion, wrote nothing, and reported success, invisible because prose edits fail no test. ## Rules documentation (docs/design/) @@ -269,6 +271,8 @@ The library has two layers: `nameparser/config/` (data) and `nameparser/parser.p **Sweep the forms a change can reach: comma shapes, then `name_order`.** No comma, a FULL name before the comma, and a ONE-WORD name before the comma are three paths, and the third is the miss — #429, #430 and #432 are all that path disagreeing with the full-name path on inputs the full-name path reads correctly. For orders the sweep already exists (`tests/v2/test_cases.py` runs every row under all three) but carries one assertion, R2's family partition, so it checks nothing a new change moves; coverage there has been vacuous before (PR #394's review found the suite passed with `name_order` discarded from grouping). Parse your change's names in each comma shape and each order, and read the ones you did not predict. +**Then sweep the spellings: a class reached by vocabulary is also reached by SHAPE.** A test that uses only the listed spelling of a word class walks only the vocabulary path, and the shape path is a separate branch that can be missing while every test passes. Roman numerals are the sharpest case: `ii`, `iii` and `iv` are suffix vocabulary; `i` and `v` are vocabulary too but initial-shaped, so `is_suffix_piece` vetoes them and only the numeral fork reads them; and `vi`, `vii`, `ix` and every longer one are in no list at all, read by the fork's shape test alone (checked 2026-10-04 against `Lexicon.default()`). So a change that handles `V` and `III` can still miss `VI` — #610 found exactly that, the maiden take's candidate loop admitting a shape-only numeral in one position where the peel reads it in two (`Jane Doe nee Smith VI Prof.` kept maiden `Smith VI` while `V` and `III` worked). The same split runs through the credentials (`MA` listed, `X.Y.Z.` by the dotted shape, `XYZ` by the caps shape, and `Ph. D.` merged by group into one flagged piece the peel's walk never holds) and the titles (`Prof.` period-marked and read at the end of a name, `Prof` bare and a name word there). When a change touches a word class, test one spelling from each path. + ### Configuration layer (`nameparser/config/`) Most modules define a `frozenset` of known name pieces; `capitalization.py` and `regexes.py` define dicts. The SETS are frozen since 2.2 (#293): there is no `.add()`/`.remove()` on any of them, so a default word list is changed by configuring an object — a private `Constants` for `HumanName`, a `Lexicon` for the 2.0 API — never by editing the constant. A union of a `frozenset` with a set literal is still a `frozenset` (`TITLES`, `PARTICLES`), so the derived sets are frozen too. `CONSTANTS`/`Constants` still hand out mutable `SetManager`s; the freeze is on the module set constants they copy from. **Neither dict was frozen, and `CAPITALIZATION_EXCEPTIONS` is not covered by anything else either.** `REGEXES` is a compiled-pattern table rather than vocabulary and was never in #293's scope; `CAPITALIZATION_EXCEPTIONS` is vocabulary-shaped and is a decided, in-scope exemption. So the split-default hazard the freeze closes is still live for it, measured on 2.2 and re-measured 2026-09-23 with the key below once `'PhD'` became the shipped value: `CAPITALIZATION_EXCEPTIONS['dphil'] = 'DPhil'` reaches a freshly built `Constants` and neither the cached `Lexicon.default()` nor the shared `CONSTANTS`. Same advice — configure the object (`constants.capitalization_exceptions[...]`, or `dataclasses.replace(lexicon, capitalization_exceptions=...)`). `tests/v2/test_contracts.py::test_every_vocabulary_constant_is_frozen` names both dicts as explicit exemptions rather than letting its `isinstance` filter drop them. From 2db7ae0b6d9e79b67de2cd4927ba90443345ac57 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sun, 4 Oct 2026 12:56:28 -0700 Subject: [PATCH 3/3] Check #610's candidate superset at the take, on unjoined pieces Review found the sweep test read plain names through all of `group`, so the pieces it handed to `trailing_candidates` already carried the joins the maiden take runs before ('de VI Prof.' as one piece where the take sees three). The oracle now wraps `_group._maiden_take` and checks every take of a 37,044-clause grid on the pieces it was actually handed: the clause-free view, `tail_reading` over it, and every piece it takes at or after the cut. 0 misses; the loop #601 shipped misses 432 there, every one the shape-numeral class, which is the recorded negative control. decisions.md#M2 relabels the first grid's figures with their population, and mechanisms.md's Field notes record the trap. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 2 +- docs/design/mechanisms.md | 1 + tests/v2/pipeline/test_pieces.py | 93 +++++++++++++++++++++----------- 3 files changed, 65 insertions(+), 31 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 89193681..5b3301cd 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -244,7 +244,7 @@ the fullwidth-colon marker (旧姓:佐藤 arrives as one word; the head-peel q - 2026-09-27 (Derek), #544 — S2'S COMPANY CLAUSE DOES NOT REACH ACROSS A CLAUSE, recorded as an Accepted boundary rather than repaired. `Jane Doe Jr. nee Smith Ma` keeps maiden 'Smith Ma' and reports it, the clause's own words standing between 'Jr.' and 'Ma', where the clause-less `Jane Doe Jr. Ma` reads suffix 'Jr. Ma'. tests/v2/test_properties.py's clause-agreement walk pins the pairs that differ for exactly this reason as its `anchored_head` class — 810 of its 20,412 pairs, recorded 2026-09-27, 0 with the anchor off. Inside the clause the company reads as it does anywhere: `Jane Doe nee Smith PhD MEng` ends the clause at 'PhD' and reads suffix 'PhD MEng', maiden 'Smith'. - 2026-10-04 (Derek), #601 — A MARKER COUNTS ONLY BEHIND A SURNAME, AND IT TAKES THE WORDS AFTER IT UP TO THE TRAILING RUN THE NAME WOULD END WITH IF THE CLAUSE WERE NOT WRITTEN. This SUPERSEDES the release machinery of the #533 and #535 entries above (the release question, the per-stop checks, the anchored-head boundary of the #544 entry) and the given-part and suffix-comma reach the 2026-07-03 rule had; the #399, #420, #434 and delimiter entries stand. The trigger was #548, `Dr. nee Smith PhD Prof.`, where 2.3.0 read family 'PhD', maiden 'Smith': a clause whose head was nothing but a title, and a release check asking whether each word it gave up could still read as a post-nominal. #548 was closed into this issue and #602 (S2's run) because both halves of its answer were rules rather than repairs. THE HEAD RULE (Derek): a marker is one only in the name before any comma or the surname part before a family comma, behind at least one name word past the leading title run, and not straight behind an unambiguous suffix word or a connective. Anywhere else it is an ordinary word — in the given part after a family comma (`Doe, Jane nee Smith` reads middle 'nee Smith', 1.4.0's reading) and in a part after a suffix comma (`Smith, John, PhD née Jones` reads suffix 'PhD née Jones', also 1.4.0's). Derek's framing was that a maiden marker always follows a family name, and that a marker following something the parser has decided is a title, suffix, particle or connective is not a marker. THE PARTICLE HALF WAS DROPPED BY MEASUREMENT: prototype variant 1 refused a marker behind particle vocabulary and so declined `Anh Do née Tran` and `Mai Le née Nguyen`, Do, Le and Van being particles that are also the commonest Vietnamese surnames; a particle counts as the surname (`Jane van nee Smith` reads family 'van', maiden 'Smith'). THE BOUND (Derek): the clause takes its words up to the run of post-nominals and titles that the trailing rules read at the end of the name with the clause removed, and the words it gives up take the roles that reading gives them — the TRAILING rule's reading of the clause-free name, not the full parse of it, so a head the full parse would read differently (H1's lone title, a P5 join) cannot change where the clause ends. The first word after the marker is always taken, the marker having announced a name — a lone roman numeral included, so `Jane Smith née V` reads maiden 'V' where the clause-free `Jane Smith V` reads suffix 'V', accepted rather than given a numeral exception (rules.md#M2's Accepted block) — unless it is an unambiguous suffix word, where the marker declines (`Jane Smith nee PhD` reads family 'nee', suffix 'PhD', as before). Before a family comma the clause keeps the 2026-07-03 stop at the first suffix word, since no trailing rule reads that part's end. VARIANTS, prototyped 2026-10-03 behind an environment switch on one tree (the numbers are in #601's comments): v2/v3 grew the clause until the existing release check accepted it and reached 0 invariant violations but inherited the check's caution about particles (`Anh Do geb. de la Vega PhD Prof.` kept 'PhD Prof.' in the maiden name); v4 read the run once over the name as written without binding it, and is the recorded negative control in tests/v2/test_properties.py (350 of the clause grid's 4,482 parses fail under it, measured 2026-10-04); v5/v6 bound that reading as roles but still counted the clause's words among the words to spare, so `John nee Smith ba` read 'ba' as a credential where `John ba` reads a name; v7, the clause-free reading bound as roles, was adopted. Against #535's grids v7 left 0 violations of the M2 invariant where the tree then had 15,586 (grid A, 327,936 texts) and 8,713 (grid B, 173,057), measured 2026-10-03 with an invariant check written to #535's recipe. WHAT THE PROTOTYPE MISSED, found implementing it: the numeral fork. A trailing roman numeral is a suffix only by its shape against the word before it, and v7 read the clause's last word without that test, so `John née Jones Smith VI` kept 'VI' in the maiden name; the take now asks `is_trailing_numeral_suffix` of the last piece, and cases.py's `a_numeral_the_clause_free_name_reads_ends_the_clause` pins it. And the head's title run is measured over the words before the marker, a lone title excepted: measured over the whole segment, `Lord Chancellor née Jones` counted 'Chancellor' as a title and refused the marker. TRIAGE, 2026-10-04: the prototype's 68 failing case rows (40 facade failures among them) were re-pinned as `fix(#601)` with their group named in the note, bar the head-rule bug above; tests/v2/test_parser.py's corpus-append invariant became two-sided (appending ` née X` moves no other field, OR the take records why it declined), with a named set for the suffixes that themselves hold a marker. DEGENERATE INPUT, NO READING DESIGNED FOR (Derek, 2026-10-04: garbage in, garbage out). Two names that read worse by eye moved, and they are recorded here so the move is not mistaken for a regression, not because either reading is wanted: `Berg, Jane van der nee Smith DO` reads middle 'van der nee Smith', family 'Berg' where 2.3.0 read family 'van der Berg', maiden 'Smith DO' — the marker in a given part being a word, the particles no longer trail the given name (P6); and `Jane Doe, PhD née Smith` reads given 'PhD', middle 'née Smith', family 'Jane Doe' where 2.3.0 read given 'Jane', family 'Doe', suffix 'PhD', maiden 'Smith', which 2.3.0 reached only through the clause taking the marker (the comma part opening with a credential is #603's shape). Measured 2026-10-04 against 2.3.0. COST: the take is one function and no release helper survives; the #533 review's comma-head grids could no longer reach a take at all and were replaced by counting heads (tests/v2/test_properties.py's `_maiden_clause_grid`, 4,482 parses). -- 2026-10-04 #610 — THE MAIDEN TAKE'S CANDIDATE WORDS ARE ONE FUNCTION BESIDE THE PEEL, AND SHARING IT FIXED A GAP THE HAND COPY HAD. `_maiden_take` chose the clause words to read its clause-free view over with a filter of its own (`_clause_tail_word` plus a numeral clause), a hand copy of the peel's and the chain's admissions whose contract — a superset of what `tail_reading` takes — nothing stated or tested. It is now `_pieces.trailing_candidates`, written beside `peel_trailing`, and tests/v2/pipeline/test_pieces.py pins the superset over a generated grid (37,044 parses through `group`). The grid found the copy wrong twice. The numeral fork reads a shape-only roman numeral (`VI`, in no wordlist) as the WALK's last piece, and two things leave it last without standing at the end of the name: a period-marked title behind it, which the chain takes first, and the merged `Ph. D.`, which group flags and the walk never holds. The copy admitted the numeral only as the clause's last word, so `Jane Doe nee Smith VI Prof.` kept maiden 'Smith VI' where `John Smith VI Prof.` reads suffix 'VI', and `Jane Doe nee Smith VI Ph. D.` the same way. Both now end the clause at the numeral, as rules.md#M2's statement already said they should. Recorded negative control: the loop as #601 shipped it misses 288 texts of the grid, every one this shape; a first draft of the fix covered the title half and missed the `Ph. D.` half, 111 texts. Against master 8a8459a8 over the 95,813-text equivalence grid the change moved 10 readings, all of this class. The same issue lifted P6's particle-tail walk into `_pieces.particle_tail`, read off roles, so post_rules' attachment and assign's given-part credential run (#602) ask one walk; that half moved none of the 95,813 readings. +- 2026-10-04 #610 — THE MAIDEN TAKE'S CANDIDATE WORDS ARE ONE FUNCTION BESIDE THE PEEL, AND SHARING IT FIXED A GAP THE HAND COPY HAD. `_maiden_take` chose the clause words to read its clause-free view over with a filter of its own (`_clause_tail_word` plus a numeral clause), a hand copy of the peel's and the chain's admissions whose contract — a superset of what `tail_reading` takes — nothing stated or tested. It is now `_pieces.trailing_candidates`, written beside `peel_trailing`, and tests/v2/pipeline/test_pieces.py pins the superset at the take itself, on the unjoined pieces the take is handed, over a generated grid (37,044 clauses through `group`). The grid found the copy wrong twice. The numeral fork reads a shape-only roman numeral (`VI`, in no wordlist) as the WALK's last piece, and two things leave it last without standing at the end of the name: a period-marked title behind it, which the chain takes first, and the merged `Ph. D.`, which group flags and the walk never holds. The copy admitted the numeral only as the clause's last word, so `Jane Doe nee Smith VI Prof.` kept maiden 'Smith VI' where `John Smith VI Prof.` reads suffix 'VI', and `Jane Doe nee Smith VI Ph. D.` the same way. Both now end the clause at the numeral, as rules.md#M2's statement already said they should. Recorded negative control: the loop as #601 shipped it misses 432 of that grid's clauses, every one this shape. The test's first version read plain names through all of `group`, on pieces the joins had already merged, which review caught: there the old loop missed 288 names, and a first draft of the fix, covering the title half only, still missed 111 — the `Ph. D.` half, which that grid is what found. Against master 8a8459a8 over the 95,813-text equivalence grid the change moved 10 readings, all of this class. The same issue lifted P6's particle-tail walk into `_pieces.particle_tail`, read off roles, so post_rules' attachment and assign's given-part credential run (#602) ask one walk; that half moved none of the 95,813 readings. ### N1 — the default delimiter pairs diff --git a/docs/design/mechanisms.md b/docs/design/mechanisms.md index 40a50821..729cf491 100644 --- a/docs/design/mechanisms.md +++ b/docs/design/mechanisms.md @@ -186,6 +186,7 @@ Problem shape. A test pins an ordering, a sort, a dedup or a partition, and its ### Field notes — the traps themselves +- Test a function INSIDE a stage on the inputs its caller hands it, not on the stage's finished output. #610's first sweep test built pieces by running `group` whole and handed them to the maiden take's candidate function, but the take runs before any join, so the test read `de VI Prof.` as one piece where production hands it three: a guard checking a structure production never passes in, green for the wrong reason. Review moved the oracle to the take itself by wrapping it (tests/v2/pipeline/test_pieces.py), and recorded the old loop's misses there as the negative control (432), which is what shows the moved harness can still see a difference. - Enumerate the rules that BUILD a structure; do not recall them. #395's unit walk was written three times in one PR — P2's chain missing, then the conjunction and bound-given branches absorbing one token where the particle branch absorbed a unit, then the suffix stop — and each miss came from listing the joining rules from memory instead of reading them out of rules.md. Both later misses reproduced the very defect the first fix had just removed, mirrored. - Prefer monkeypatching a module-level helper to re-exec'ing a module: a caller looks a module global up at CALL time, so replacing `_units` or `_name_positions` reaches it and none of the traps below apply. Re-exec is only needed for logic INLINE in a function, and "the thing I want to mutate is not a module-level helper" is a hint that it could be one. - Re-exec'ing a module redefines the classes that module DEFINES, not the ones it imports — the exec re-runs its own import statements, which resolve through sys.modules. So `is` comparisons break selectively: re-exec'ing `_group.py` yields a second `BoundJoin` (defined there) while `Structure` (imported) stays identical, and `bound_join is not BoundJoin.DISABLED` is then always true, joining a family comma's own segment and reporting 3,150 phantom movers. Rebind what the module defines (`mod.__dict__["BoundJoin"] = real.BoundJoin`) and assert the identity rather than assuming it either way. diff --git a/tests/v2/pipeline/test_pieces.py b/tests/v2/pipeline/test_pieces.py index 6c092b36..e643b07f 100644 --- a/tests/v2/pipeline/test_pieces.py +++ b/tests/v2/pipeline/test_pieces.py @@ -7,7 +7,7 @@ """ import dataclasses import itertools -from collections.abc import Sequence, Set +from collections.abc import Callable, Sequence, Set import pytest @@ -16,6 +16,7 @@ from nameparser._pipeline import STAGES from nameparser._pipeline._assign import assign from nameparser._pipeline._classify import classify +from nameparser._pipeline import _group from nameparser._pipeline._group import group from nameparser._pipeline._pieces import ( _anchors, _numeral_behind_the_initial_veto, anchor_in_reach, @@ -26,7 +27,7 @@ ) from nameparser._pipeline._segment import segment from nameparser._pipeline._state import ( - AMBIGUOUS_ACRONYM_TAG, ParseState, WorkToken, + AMBIGUOUS_ACRONYM_TAG, ParseState, PendingAmbiguity, WorkToken, ) from nameparser._pipeline._tokenize import tokenize from nameparser._pipeline._vocab import is_one_case, is_title_shaped, tag_marker_runs @@ -890,10 +891,12 @@ def test_a_title_inside_the_given_parts_run_is_a_title() -> None: # #610: `trailing_candidates` is the maiden take's choice of which # clause words to read the clause-free view over, and its contract is -# a SUPERSET of what `tail_reading` takes. The grid reads plain names, -# where the candidates' cut must fall at or before every piece the tail -# reading takes (the first name piece aside: in a clause that is the -# word after the marker, which the take always keeps). +# a SUPERSET of what `tail_reading` takes from that view. Checked AT +# THE TAKE, on the pieces the take is handed -- before any join, which +# is where it runs -- by wrapping `_group._maiden_take` and the +# candidates call inside it: the clause-free name is the head plus every +# clause word after the first (which the take always keeps), and every +# piece the tail reading takes from it must fall at or after the cut. _CANDIDATE_HEADS = ("Jane Doe", "J.", "Jane van der Berg", "Mai Le", "abdul Berg", "John") _CANDIDATE_WORDS = ("Smith", "VI", "V", "III", "MA", "Ma", "PhD", "Jr.", @@ -901,45 +904,75 @@ def test_a_title_inside_the_given_parts_run_is_a_title() -> None: "Ph. D.", "i", "Jones") -def _candidate_misses() -> list[str]: - misses = [] +def _candidate_misses(monkeypatch: pytest.MonkeyPatch, + candidates: Callable[..., int] = trailing_candidates, + ) -> list[str]: + misses: list[str] = [] + seen: list[tuple[int, int]] = [] + real_take = _group._maiden_take + + def cut(lo: int, pieces: Sequence[Sequence[int]], + ptags: Sequence[Set[str]], + tokens: Sequence[WorkToken]) -> int: + c = candidates(lo, pieces, ptags, tokens) + seen.append((lo, c)) + return c + + def take(pieces: Sequence[Sequence[int]], ptags: Sequence[Set[str]], + tokens: Sequence[WorkToken], one_case: bool | None, + site: _group.ClauseSite, + ambiguities: list[PendingAmbiguity], + ) -> _group.MaidenIndices | None: + seen.clear() + answer = real_take(pieces, ptags, tokens, one_case, site, + ambiguities) + for lo, c in seen: + m = next(v for v in range(1, len(pieces)) + if _group._is_maiden_marker_piece(pieces[v], tokens)) + left = list(range(m)) + list(range(lo + 1, len(pieces))) + view = [pieces[q] for q in left] + vtags = [ptags[q] for q in left] + rest = peel_walk(leading_titles(view, vtags, tokens), vtags) + kept, chained, peel = tail_reading(rest, view, vtags, tokens, + one_case) + taken = {left[q] for q in (*kept[peel.names:], *chained) + if q >= m} + if any(j < c for j in taken): + misses.append(" ".join(tokens[i].text for p in pieces + for i in p)) + return answer + + monkeypatch.setattr(_group, "trailing_candidates", cut) + monkeypatch.setattr(_group, "_maiden_take", take) for head in _CANDIDATE_HEADS: for n in (1, 2, 3): for tail in itertools.product(_CANDIDATE_WORDS, repeat=n): - text = f"{head} {' '.join(tail)}" - state = _through_group(text) - pieces, ptags = state.pieces[0], state.piece_tags[0] - tokens = state.tokens - lead = leading_titles(pieces, ptags, tokens) - rest = peel_walk(lead, ptags) - if not rest: - continue - kept, chained, peel = tail_reading( - rest, pieces, ptags, tokens, state.one_case) - taken = (set(kept[peel.names:]) | set(chained)) - {rest[0]} - c = trailing_candidates(rest[0], pieces, ptags, tokens) - if any(j < c for j in taken): - misses.append(text) + _through_group(f"{head} nee Smith {' '.join(tail)}") return misses -def test_the_trailing_candidates_cover_what_the_tail_reading_takes() -> None: +def test_the_trailing_candidates_cover_what_the_tail_reading_takes( + monkeypatch: pytest.MonkeyPatch) -> None: """M2's take ends the clause at the trailing run the clause-free name reads, which it can only find if every word that run might - hold is in its view. 6 heads x 18 words to a - tail of three, 37,044 parses through `group`, about 2.4s on py3.11 + hold is in its view. 6 heads x 18 words to a tail of three after + 'nee Smith', 37,044 parses through `group`, about 2.6s on py3.11 (measured 2026-10-04). RECORDED NEGATIVE CONTROL, measured 2026-10-04 on this grid: the candidate loop as #601 shipped it (master 8a8459a8), which admitted a roman numeral by its shape only as the clause's LAST word, misses - 288 texts -- every one a numeral in no wordlist ('VI') with a - period-marked title or the merged 'Ph. D.' behind it. The chain + 432 texts -- every one a numeral in no wordlist ('VI') with + a period-marked title or the merged 'Ph. D.' behind it. The chain takes the title first, and the walk never holds the flagged 'Ph. D.', so either way the numeral is the walk's last piece and the fork reads it: 'Jane Doe nee Smith VI Prof.' kept maiden 'Smith VI' where 'John Smith VI Prof.' reads suffix 'VI'. A first draft of - the fix covered the title and still missed the 'Ph. D.' half, 111 - texts, which this grid is what found. + the fix covered the title half and missed the 'Ph. D.' half. + + A first version of this test read plain names through all of + `group`, so its pieces had the joins applied that the take runs + before ('de VI Prof.' arrived as one piece); review moved the + oracle to the take itself. """ - assert _candidate_misses() == [] + assert _candidate_misses(monkeypatch) == []