Skip to content

Read a kana voicing mark as the script of the kana before it (#596) - #598

Merged
derek73 merged 4 commits into
masterfrom
fix/issue-596-combining-voicing-marks
Oct 3, 2026
Merged

derek73 merged 4 commits into
masterfrom
fix/issue-596-combining-voicing-marks

Conversation

@derek73

@derek73 derek73 commented Oct 3, 2026

Copy link
Copy Markdown
Owner

A katakana name typed with a separate voicing mark — ア゙イ タロウ (ア + combining U+3099), or the spacing ゛/゜ — read family-first. The voicing marks U+3099–U+309C sit in the hiragana block, and NFC composes one only where a precomposed voiced kana exists. So ア゙イ stayed katakana plus a hiragana-block character, took the kana license, and turned the name around, along with its neighbour (マイケル ア゙イ read family マイケル). The bug has been there since 2.1.0.

Fix. Script classification drops, from its NFC copy only, every voicing mark that stands after a kana, so the word reads as the script of the kana before the mark. Token text is untouched. A mark after anything else (hangul, Han, Latin, or nothing) voices no kana and keeps the table's answer, as before. The common path is one frozenset.isdisjoint check in C; only a token that holds a mark walks it in Python.

Input 2.1–2.3 This PR
ア゙イ タロウ family ア゙イ given ア゙イ, family タロウ
ア゛イ タロウ family ア゛イ given ア゛イ
マイケル ア゙イ family マイケル given マイケル, family ア゙イ
あ゙い たろう, 山田 ア゙イ family-first unchanged

Declined, measured on patched copies of the tree:

  • Taking the marks out of the hiragana range. This fixes katakana but leaves a marked word in no script at all, so あ゙い たろう and 山田 ア゙イ lose their Japanese reading.
  • Covering only the two combining marks. This leaves ア゛イ タロウ family-first.

NFD-typed text, the issue's worry, is unaffected by every option, because classification composes first. A mark with nothing before it (゙アイ) still reads as hiragana. That's degenerate input, and rules.md's W Background now leaves it unpromised.

Also in this PR: the W4 interpunct wording. A name the 间隔号 divides keeps the declared order, not its source order: under FAMILY_FIRST, 威廉·莎士比亚 reads family 威廉. rules.md#W4 now says so, with a family-first example. The same correction is in AGENTS.md, concepts.rst and usage.rst, including the "Forms the script carries" passage, which the issue's list missed.

Commits

  1. docs: the W4 wording (docs only).
  2. The fix, with case rows, a unit test, the W4 example and Background sentence, a decisions.md#W4 entry, fix(#596) ledger rules and the release log.
  3. Review round 1: the first cut dropped a mark after any character. 김゙민준 then classified as hangul, and the surname split, which cuts the raw text, stranded the mark (given 김, family ゙민준). The fix narrows the drop to marks after a kana. It also fixes two doc findings: the usage.rst form-7 passage, and a wrong 1.4.0/2.0.0 claim in decisions.md.
  4. Review round 2 (of commit 3): the parse("Andrew") reports no ambiguity, but given-or-family is a guess #449 report rule for 김゙민준 added at 2.0.0; that row marked tolerated so it agrees with "unpromised"; stale wording; dated claim comments.

Verification

  • 11,161 tests pass, plus mypy, ruff and the Sphinx doctests.
  • The differential gate exits 0 at all five baselines. Before this change no corpus line held an uncomposable mark, so the new case rows are the whole population:
  • Negative controls: the three moving rows and the unit test fail with the fix reverted, and the hangul row fails against commit 2's version of the fix.

Closes #596. #597 (halfwidth corner brackets) follows as its own PR.

🤖 Generated with Claude Code

derek73 and others added 4 commits October 3, 2026 14:53
…urce order (#596)

rules.md#W4's Accepted clause said a name the 间隔号 divides keeps its
source order. That holds under the default order only: the divider
stands the script override down, so under name_order=FAMILY_FIRST
威廉·莎士比亚 reads family 威廉 like any other name. The clause now
says declared order and carries a family-first example pinning it.

Same correction at the sites that described the parser's behavior
(AGENTS.md, concepts.rst, usage.rst, a test_assign comment). Sites
describing how a transcription is WRITTEN, or a parse under the
default order (migrate.rst's HumanName example, case-row notes),
were read and left as they are.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The four kana voicing marks U+3099-U+309C sit in the hiragana block,
and NFC composes one only where a precomposed voiced kana exists. So
ア゙イ (ア + U+3099), ン゙ and every spacing ゛/゜ kept a hiragana-block
character inside a katakana word, the word took the kana license, and
a wholly-katakana name read family-first -- flipping its neighbour too
(マイケル ア゙イ read family マイケル). Present since 2.1.0.

Classification now deletes, from its NFC copy only, every mark with a
character before it, so the word reads as its base's script. Token
text is untouched. A str.translate table, not a pattern: C-level, and
no new module-level regex for test_regex_sync to declare.

Declined, measured on patched copies: dropping the marks from the
hiragana range fixes katakana but leaves a marked word in no script,
so あ゙い たろう and 山田 ア゙イ lose their Japanese reading; narrowing
to the two combining marks leaves ア゛イ タロウ family-first. NFD text
is unaffected by any option, since classification composes first.

Three fix(#596) case rows (the movers), two parity rows guarding the
declined alternative, parametrized unit pins on effective_script,
a W4 example and W Background sentence, a decisions.md#W4 entry,
fix(#596) ledger rules at 2.1.0/2.2.0/2.3.0 (1.4.0 and 2.0.0 had no
license to misfire), and a release-log bullet. The gate exits 0 at
every baseline.

Closes #596.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Code review: the first cut deleted a voicing mark after ANY
character, so 김゙민준 classified as hangul and the surname division,
which cuts the raw text, split the mark off its base -- given 김,
family ゙민준, unreported, where 2.3.0 left the word whole and
reported it (parser_for(ZH) cut 王゙小明 the same way). A mark voices
the kana it follows, so it is now dropped only after a kana; after
anything else it keeps the table's answer, as before. The common path
stays C-level (frozenset.isdisjoint); only a token holding a mark
walks it. A parity row pins 김゙민준, with a #449 ledger rule at 2.1.0
and 2.2.0 for its report, new in 2.3.

Docs review: usage.rst's "Forms the script carries" still said form 7
needs no name_order -- a declared FAMILY_FIRST reverses it, as W4 now
says. decisions.md's #596 entry claimed nothing moves at 1.4.0/2.0.0;
the two kana parity rows do there, under the native-script rule. It
now also records that #594's case against an NFKC fold has lost its
W4 half: #594's own recompute finds no role differing on this tree.
rules.md's W Background and the code comments say "after a kana".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- 김゙민준 moves in its report at 2.0.0 too, not only 2.1.0/2.2.0: the
  #449 rule joins the 2.0.0 ledger (seven #449 rules there now), and
  decisions.md's MEASURED sentence says so; 1.4.0 has no report and
  is parity.
- The hangul row is tolerated: rules.md's W Background leaves a mark
  after a non-kana unpromised, and a contract row contradicted it. It
  moves to the radar corpus; decisions.md says "watches", not "pins".
- Stale "follows a character" in the three fix(#596) ledger comments
  and "every voicing mark" in _policy.py now say "after a kana".
- The 147 -> 153 native-script claim bumps carry dated comments, and
  the new #449 rule has its _MUST_NOT_MATCH near-misses.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73 derek73 added this to the 2.4 milestone Oct 3, 2026
@derek73 derek73 added bug docs Documentation fixes and updates labels Oct 3, 2026
@derek73 derek73 self-assigned this Oct 3, 2026
@codecov

codecov Bot commented Oct 3, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 98.98%. Comparing base (5cf668c) to head (0aefefe).

Additional details and impacted files
@@           Coverage Diff           @@
##           master     #598   +/-   ##
=======================================
  Coverage   98.98%   98.98%           
=======================================
  Files          45       45           
  Lines        4220     4235   +15     
=======================================
+ Hits         4177     4192   +15     
  Misses         43       43           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@derek73
derek73 merged commit 8071645 into master Oct 3, 2026
11 checks passed
@derek73
derek73 deleted the fix/issue-596-combining-voicing-marks branch October 4, 2026 20:30
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug docs Documentation fixes and updates

Projects

None yet

Development

Successfully merging this pull request may close these issues.

A katakana name typed with a separate voicing mark (ア゙イ) reads family-first

1 participant