Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,6 +12,8 @@ After a review round, **review the fix commit too** — second-round passes on #

**Check whether the PR description still needs updating before merging.** A review that moves behavior or retracts a measurement leaves it stale, and a stale body reads as authoritative.

**When a defect is found that an earlier check would have caught, write the check down where the next person will meet it, in the same PR.** A review finding, a failing sweep or a regression that arrived because nobody thought to test a shape is a lesson about the repository, not only about the change, and it is lost when the PR merges unless it lands somewhere a later session reads before making the same mistake. Ask, before merging: what should I have checked, and where would I have looked? Then put it there — a missed input shape in the sweep paragraphs under Architecture (comma shapes, then `name_order`, then spellings, each earned this way); a trap in a module's behavior under Gotchas; a design-review miss in docs/design/AGENTS.md's axes; a measurement or test-design trap in mechanisms.md's Verification shapes or Field notes; a reusable design move in mechanisms.md proper. Name the PR or issue that earned it and the input that showed it, so the rule can be checked against its case. #610 is the shape: its sweep test found that the maiden take missed a shape-only numeral, and the spelling sweep below records the class (`VI` beside `V` and `III`). This is how the file accumulates expertise rather than history — prefer amending an existing paragraph to adding a new one, and delete a lesson the code has since made impossible.

**A scripted multi-edit that asserts as it goes discards everything when a late pattern misses.** Assert every pattern before applying any, then verify the edit landed — three times in one session a docs batch failed its last assertion, wrote nothing, and reported success, invisible because prose edits fail no test.

## Rules documentation (docs/design/)
Expand Down Expand Up @@ -269,6 +271,8 @@ The library has two layers: `nameparser/config/` (data) and `nameparser/parser.p

**Sweep the forms a change can reach: comma shapes, then `name_order`.** No comma, a FULL name before the comma, and a ONE-WORD name before the comma are three paths, and the third is the miss — #429, #430 and #432 are all that path disagreeing with the full-name path on inputs the full-name path reads correctly. For orders the sweep already exists (`tests/v2/test_cases.py` runs every row under all three) but carries one assertion, R2's family partition, so it checks nothing a new change moves; coverage there has been vacuous before (PR #394's review found the suite passed with `name_order` discarded from grouping). Parse your change's names in each comma shape and each order, and read the ones you did not predict.

**Then sweep the spellings: a class reached by vocabulary is also reached by SHAPE.** A test that uses only the listed spelling of a word class walks only the vocabulary path, and the shape path is a separate branch that can be missing while every test passes. Roman numerals are the sharpest case: `ii`, `iii` and `iv` are suffix vocabulary; `i` and `v` are vocabulary too but initial-shaped, so `is_suffix_piece` vetoes them and only the numeral fork reads them; and `vi`, `vii`, `ix` and every longer one are in no list at all, read by the fork's shape test alone (checked 2026-10-04 against `Lexicon.default()`). So a change that handles `V` and `III` can still miss `VI` — #610 found exactly that, the maiden take's candidate loop admitting a shape-only numeral in one position where the peel reads it in two (`Jane Doe nee Smith VI Prof.` kept maiden `Smith VI` while `V` and `III` worked). The same split runs through the credentials (`MA` listed, `X.Y.Z.` by the dotted shape, `XYZ` by the caps shape, and `Ph. D.` merged by group into one flagged piece the peel's walk never holds) and the titles (`Prof.` period-marked and read at the end of a name, `Prof` bare and a name word there). When a change touches a word class, test one spelling from each path.

### Configuration layer (`nameparser/config/`)

Most modules define a `frozenset` of known name pieces; `capitalization.py` and `regexes.py` define dicts. The SETS are frozen since 2.2 (#293): there is no `.add()`/`.remove()` on any of them, so a default word list is changed by configuring an object — a private `Constants` for `HumanName`, a `Lexicon` for the 2.0 API — never by editing the constant. A union of a `frozenset` with a set literal is still a `frozenset` (`TITLES`, `PARTICLES`), so the derived sets are frozen too. `CONSTANTS`/`Constants` still hand out mutable `SetManager`s; the freeze is on the module set constants they copy from. **Neither dict was frozen, and `CAPITALIZATION_EXCEPTIONS` is not covered by anything else either.** `REGEXES` is a compiled-pattern table rather than vocabulary and was never in #293's scope; `CAPITALIZATION_EXCEPTIONS` is vocabulary-shaped and is a decided, in-scope exemption. So the split-default hazard the freeze closes is still live for it, measured on 2.2 and re-measured 2026-09-23 with the key below once `'PhD'` became the shipped value: `CAPITALIZATION_EXCEPTIONS['dphil'] = 'DPhil'` reaches a freshly built `Constants` and neither the cached `Lexicon.default()` nor the shared `CONSTANTS`. Same advice — configure the object (`constants.capitalization_exceptions[...]`, or `dataclasses.replace(lexicon, capitalization_exceptions=...)`). `tests/v2/test_contracts.py::test_every_vocabulary_constant_is_frozen` names both dicts as explicit exemptions rather than letting its `isinstance` filter drop them.
Expand Down
1 change: 1 addition & 0 deletions docs/design/decisions.md
Original file line number Diff line number Diff line change
Expand Up @@ -244,6 +244,7 @@ the fullwidth-colon marker (旧姓:佐藤 arrives as one word; the head-peel q
- 2026-09-27 (Derek), #544 — S2'S COMPANY CLAUSE DOES NOT REACH ACROSS A CLAUSE, recorded as an Accepted boundary rather than repaired. `Jane Doe Jr. nee Smith Ma` keeps maiden 'Smith Ma' and reports it, the clause's own words standing between 'Jr.' and 'Ma', where the clause-less `Jane Doe Jr. Ma` reads suffix 'Jr. Ma'. tests/v2/test_properties.py's clause-agreement walk pins the pairs that differ for exactly this reason as its `anchored_head` class — 810 of its 20,412 pairs, recorded 2026-09-27, 0 with the anchor off. Inside the clause the company reads as it does anywhere: `Jane Doe nee Smith PhD MEng` ends the clause at 'PhD' and reads suffix 'PhD MEng', maiden 'Smith'.

- 2026-10-04 (Derek), #601 — A MARKER COUNTS ONLY BEHIND A SURNAME, AND IT TAKES THE WORDS AFTER IT UP TO THE TRAILING RUN THE NAME WOULD END WITH IF THE CLAUSE WERE NOT WRITTEN. This SUPERSEDES the release machinery of the #533 and #535 entries above (the release question, the per-stop checks, the anchored-head boundary of the #544 entry) and the given-part and suffix-comma reach the 2026-07-03 rule had; the #399, #420, #434 and delimiter entries stand. The trigger was #548, `Dr. nee Smith PhD Prof.`, where 2.3.0 read family 'PhD', maiden 'Smith': a clause whose head was nothing but a title, and a release check asking whether each word it gave up could still read as a post-nominal. #548 was closed into this issue and #602 (S2's run) because both halves of its answer were rules rather than repairs. THE HEAD RULE (Derek): a marker is one only in the name before any comma or the surname part before a family comma, behind at least one name word past the leading title run, and not straight behind an unambiguous suffix word or a connective. Anywhere else it is an ordinary word — in the given part after a family comma (`Doe, Jane nee Smith` reads middle 'nee Smith', 1.4.0's reading) and in a part after a suffix comma (`Smith, John, PhD née Jones` reads suffix 'PhD née Jones', also 1.4.0's). Derek's framing was that a maiden marker always follows a family name, and that a marker following something the parser has decided is a title, suffix, particle or connective is not a marker. THE PARTICLE HALF WAS DROPPED BY MEASUREMENT: prototype variant 1 refused a marker behind particle vocabulary and so declined `Anh Do née Tran` and `Mai Le née Nguyen`, Do, Le and Van being particles that are also the commonest Vietnamese surnames; a particle counts as the surname (`Jane van nee Smith` reads family 'van', maiden 'Smith'). THE BOUND (Derek): the clause takes its words up to the run of post-nominals and titles that the trailing rules read at the end of the name with the clause removed, and the words it gives up take the roles that reading gives them — the TRAILING rule's reading of the clause-free name, not the full parse of it, so a head the full parse would read differently (H1's lone title, a P5 join) cannot change where the clause ends. The first word after the marker is always taken, the marker having announced a name — a lone roman numeral included, so `Jane Smith née V` reads maiden 'V' where the clause-free `Jane Smith V` reads suffix 'V', accepted rather than given a numeral exception (rules.md#M2's Accepted block) — unless it is an unambiguous suffix word, where the marker declines (`Jane Smith nee PhD` reads family 'nee', suffix 'PhD', as before). Before a family comma the clause keeps the 2026-07-03 stop at the first suffix word, since no trailing rule reads that part's end. VARIANTS, prototyped 2026-10-03 behind an environment switch on one tree (the numbers are in #601's comments): v2/v3 grew the clause until the existing release check accepted it and reached 0 invariant violations but inherited the check's caution about particles (`Anh Do geb. de la Vega PhD Prof.` kept 'PhD Prof.' in the maiden name); v4 read the run once over the name as written without binding it, and is the recorded negative control in tests/v2/test_properties.py (350 of the clause grid's 4,482 parses fail under it, measured 2026-10-04); v5/v6 bound that reading as roles but still counted the clause's words among the words to spare, so `John nee Smith ba` read 'ba' as a credential where `John ba` reads a name; v7, the clause-free reading bound as roles, was adopted. Against #535's grids v7 left 0 violations of the M2 invariant where the tree then had 15,586 (grid A, 327,936 texts) and 8,713 (grid B, 173,057), measured 2026-10-03 with an invariant check written to #535's recipe. WHAT THE PROTOTYPE MISSED, found implementing it: the numeral fork. A trailing roman numeral is a suffix only by its shape against the word before it, and v7 read the clause's last word without that test, so `John née Jones Smith VI` kept 'VI' in the maiden name; the take now asks `is_trailing_numeral_suffix` of the last piece, and cases.py's `a_numeral_the_clause_free_name_reads_ends_the_clause` pins it. And the head's title run is measured over the words before the marker, a lone title excepted: measured over the whole segment, `Lord Chancellor née Jones` counted 'Chancellor' as a title and refused the marker. TRIAGE, 2026-10-04: the prototype's 68 failing case rows (40 facade failures among them) were re-pinned as `fix(#601)` with their group named in the note, bar the head-rule bug above; tests/v2/test_parser.py's corpus-append invariant became two-sided (appending ` née X` moves no other field, OR the take records why it declined), with a named set for the suffixes that themselves hold a marker. DEGENERATE INPUT, NO READING DESIGNED FOR (Derek, 2026-10-04: garbage in, garbage out). Two names that read worse by eye moved, and they are recorded here so the move is not mistaken for a regression, not because either reading is wanted: `Berg, Jane van der nee Smith DO` reads middle 'van der nee Smith', family 'Berg' where 2.3.0 read family 'van der Berg', maiden 'Smith DO' — the marker in a given part being a word, the particles no longer trail the given name (P6); and `Jane Doe, PhD née Smith` reads given 'PhD', middle 'née Smith', family 'Jane Doe' where 2.3.0 read given 'Jane', family 'Doe', suffix 'PhD', maiden 'Smith', which 2.3.0 reached only through the clause taking the marker (the comma part opening with a credential is #603's shape). Measured 2026-10-04 against 2.3.0. COST: the take is one function and no release helper survives; the #533 review's comma-head grids could no longer reach a take at all and were replaced by counting heads (tests/v2/test_properties.py's `_maiden_clause_grid`, 4,482 parses).
- 2026-10-04 #610 — THE MAIDEN TAKE'S CANDIDATE WORDS ARE ONE FUNCTION BESIDE THE PEEL, AND SHARING IT FIXED A GAP THE HAND COPY HAD. `_maiden_take` chose the clause words to read its clause-free view over with a filter of its own (`_clause_tail_word` plus a numeral clause), a hand copy of the peel's and the chain's admissions whose contract — a superset of what `tail_reading` takes — nothing stated or tested. It is now `_pieces.trailing_candidates`, written beside `peel_trailing`, and tests/v2/pipeline/test_pieces.py pins the superset at the take itself, on the unjoined pieces the take is handed, over a generated grid (37,044 clauses through `group`). The grid found the copy wrong twice. The numeral fork reads a shape-only roman numeral (`VI`, in no wordlist) as the WALK's last piece, and two things leave it last without standing at the end of the name: a period-marked title behind it, which the chain takes first, and the merged `Ph. D.`, which group flags and the walk never holds. The copy admitted the numeral only as the clause's last word, so `Jane Doe nee Smith VI Prof.` kept maiden 'Smith VI' where `John Smith VI Prof.` reads suffix 'VI', and `Jane Doe nee Smith VI Ph. D.` the same way. Both now end the clause at the numeral, as rules.md#M2's statement already said they should. Recorded negative control: the loop as #601 shipped it misses 432 of that grid's clauses, every one this shape. The test's first version read plain names through all of `group`, on pieces the joins had already merged, which review caught: there the old loop missed 288 names, and a first draft of the fix, covering the title half only, still missed 111 — the `Ph. D.` half, which that grid is what found. Against master 8a8459a8 over the 95,813-text equivalence grid the change moved 10 readings, all of this class. The same issue lifted P6's particle-tail walk into `_pieces.particle_tail`, read off roles, so post_rules' attachment and assign's given-part credential run (#602) ask one walk; that half moved none of the 95,813 readings.

### N1 — the default delimiter pairs

Expand Down
Loading
Loading