Read halfwidth katakana as katakana (#594) - #595
Merged
Merged
Conversation
The script table now classifies the halfwidth kana block U+FF65-U+FF9F as Script.KATAKANA, so 山田 タロウ takes the kana license and reads family 山田 as 山田 エミ does. Everything keyed on the table follows: the second-East-Asian-word count that keeps a segmenter from re-dividing a kanji name, the 间隔号's flank guard, and _NO_INITIALS, which closes the title misroute decisions.md#cjk-full-stops recorded (タナカ. John). A range rather than a fold: T1 forbids rewriting the text, and an NFKC fold at classification reaches far past the kana and maps a lone ゙ into the hiragana block. Wholly-katakana names, halfwidth included, keep the declared order; W4's rationale is amended to say the script cannot tell a Japanese reading from a transcription, not that the latter dominates. U+FF65 leaves _SANCTIONED_EXTRAS for the table, the ledgers' span copies widen to match, and fix(#594) rules classify the five movers at every baseline (no corpus line held halfwidth kana before the new case rows). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
usage.rst explains halfwidth katakana where Japanese scripts are introduced and gives the script_orders recipe for callers whose katakana names are Japanese; locales.rst drops the claim that a halfwidth neighbour divides nothing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ied (#594) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- rules.md#H2's Accepted clause still called halfwidth katakana unclassified; #594 is what changed that, so the same PR amends it. - The NFKC fold was declined for the wrong reason. The interpunct flank guard reads raw text and only asks "classified?"; the real hazard is the kana license, since an uncomposable voicing mark (ア゙) folds into the hiragana block and would turn a wholly-katakana name family-first. Corrected in decisions.md#W4 and _policy.py's comment. - AGENTS.md's 间隔号 sentence still had pure katakana naming its convention, the rationale #594 retracted. - The cjk-full-stops header and the H2 pointer now say the halfwidth limit is closed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #595 +/- ##
=======================================
Coverage 98.98% 98.98%
=======================================
Files 45 45
Lines 4220 4220
=======================================
Hits 4177 4177
Misses 43 43 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
…594) From the PR review's coverage pass: - 山田 ペーター pins U+FF9F and U+FF70, which only the ledger sync test held: narrowing the range by either one now fails a behavior row. - 山田タロウ, the halfwidth twin of ja_unspaced_unsegmented_default, pins a single Han+halfwidth token taking the license (and 2.3's given-or-family report going away). - The ja adapter stub accepts 山田タロウ, the one changed path no test reached. The license rule in the 2.1/2.2/2.3 ledgers takes the two new movers, and the #594 ledger blocks are regenerated per baseline: the title rule's comment claimed the full-width twins read a name word, true only at 2.3.0, and the 1.4.0/2.0.0 headers described rules those files do not carry. _CORPUS_CLAIMS records the wider reach. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- _policy.py cited T1 for "nothing rewrites the text"; that is rules.md's T Background, now quoted. Same correction in decisions.md#W4. - The voicing-mark sentence states the mechanism: NFC never composes a halfwidth mark into its base. - decisions.md#W4's population is eight rows, five contract movers. - migrate.rst still gave the retracted "usually a transcription" reason. - The release log says "declared order", which is what W4 keeps. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
U+FF70 is interior to the span, not an edge. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…takana (#594) The same over-broad claim the review corrected in _policy.py. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #594.
The bug. Halfwidth katakana (U+FF65–FF9F) fell outside the script table, so
山田 タロウread given山田, familyタロウwhere山田 タロウreads family-first. Legacy bank, payroll and CSV exports written for JIS X 0201 systems still use this encoding.The fix is one range:
Script.KATAKANAnow covers(0xFF65, 0xFF9F). Everything keyed on the script table follows.山田 タロウ高橋一郎 タロウ(JA + segmenter)タロウ·ヤマダタナカ. Johnタナカ.タナカ.ヤマダ タロウThe
タナカ. Johnrow closes a limit that decisions.md#cjk-full-stops had recorded and left open in #322/#323.Why a range and not normalization. The T rules forbid rewriting the input text, so normalizing at tokenize was never an option. An NFKC fold at classification reaches far past the kana (
JOHN,㈱). It also folds an uncomposable voicing mark (ア゙) into the hiragana block, which would turn a wholly-katakana name family-first. decisions.md#W4 has the measurement.Wholly-katakana names keep the declared order, halfwidth included. W4's rationale is amended, not its behavior. Systems with no kanji wrote Japanese names in katakana too, so "predominantly a transcription" no longer holds. The script can't tell a Japanese reading from a transcribed foreign name, so it settles nothing. usage.rst now gives the opt-in recipe: add
(Script.KATAKANA, FAMILY_FIRST)toscript_orders.Differential. No corpus line held halfwidth kana, so the gate was blind to this until the eight new case rows. With them:
fix(#594)rules from 2.1.0 on, and at 1.4.0 and 2.0.0 the five under the canonical CJK rule's widened span with the two under afix(#594)title rule;_SANCTIONED_EXTRASfor the table;Docs.
Review. Two rounds before opening (a code reviewer, plus the design-docs reviewer on the docs/design changes and then on the fix commit; fixes in 6451052, 151ca96, 6803db6), then
/pr-review-toolkit:review-prwith code, test-coverage and comment reviewers. That round added coverage for the span's ゚/ー edges, the unspaced山田タロウand thejaadapter (1ea6e68), and fixed comment and prose inaccuracies (15e0358).Out of scope, found in review. Full-width katakana typed with an uncomposable combining U+3099 (
ア゙イ) already takes the kana license today, contrary to W4. Filed separately as #596, together with a W4 wording fix that review also turned up.Checks.
uv run pytest11,119 passed. mypy, ruff and both Sphinx builds are clean. The suite was also run withoutnamedivider, and the halfwidth segmenter path was checked with thejaextra installed.🤖 Generated with Claude Code