Skip to content

Read halfwidth katakana as katakana (#594) - #595

Merged
derek73 merged 9 commits into
masterfrom
fix/issue-594-halfwidth-katakana
Oct 3, 2026
Merged

derek73 merged 9 commits into
masterfrom
fix/issue-594-halfwidth-katakana

Conversation

@derek73

@derek73 derek73 commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

Closes #594.

The bug. Halfwidth katakana (U+FF65–FF9F) fell outside the script table, so 山田 タロウ read given 山田, family タロウ where 山田 タロウ reads family-first. Legacy bank, payroll and CSV exports written for JIS X 0201 systems still use this encoding.

The fix is one range: Script.KATAKANA now covers (0xFF65, 0xFF9F). Everything keyed on the script table follows.

Input 2.3.0 This branch
山田 タロウ given 山田, family タロウ family 山田, given タロウ
高橋一郎 タロウ (JA + segmenter) segmenter is asked already divided
タロウ·ヤマダ one given token given タロウ, family ヤマダ
タナカ. John title タナカ. given タナカ.
ヤマダ タロウ given ヤマダ unchanged (W4)

The タナカ. John row closes a limit that decisions.md#cjk-full-stops had recorded and left open in #322/#323.

Why a range and not normalization. The T rules forbid rewriting the input text, so normalizing at tokenize was never an option. An NFKC fold at classification reaches far past the kana (JOHN, ㈱). It also folds an uncomposable voicing mark (ア゙) into the hiragana block, which would turn a wholly-katakana name family-first. decisions.md#W4 has the measurement.

Wholly-katakana names keep the declared order, halfwidth included. W4's rationale is amended, not its behavior. Systems with no kanji wrote Japanese names in katakana too, so "predominantly a transcription" no longer holds. The script can't tell a Japanese reading from a transcribed foreign name, so it settles nothing. usage.rst now gives the opt-in recipe: add (Script.KATAKANA, FAMILY_FIRST) to script_orders.

Differential. No corpus line held halfwidth kana, so the gate was blind to this until the eight new case rows. With them:

  • five contract movers and two tolerated radar movers, each classified at every baseline: under the fix(#594) rules from 2.1.0 on, and at 1.4.0 and 2.0.0 the five under the canonical CJK rule's widened span with the two under a fix(#594) title rule;
  • the canonical CJK rule's span copy widens to the whole block;
  • U+FF65 leaves _SANCTIONED_EXTRAS for the table;
  • all five baselines exit 0.

Docs.

  • rules.md: W Background, W4 rationale and examples, H2 Accepted.
  • decisions.md: a new W4 entry, and the cjk-full-stops limit marked closed.
  • AGENTS.md: the pure-katakana sentence.
  • usage.rst and locales.rst; a release-log bullet.

Review. Two rounds before opening (a code reviewer, plus the design-docs reviewer on the docs/design changes and then on the fix commit; fixes in 6451052, 151ca96, 6803db6), then /pr-review-toolkit:review-pr with code, test-coverage and comment reviewers. That round added coverage for the span's ゚/ー edges, the unspaced 山田タロウ and the ja adapter (1ea6e68), and fixed comment and prose inaccuracies (15e0358).

Out of scope, found in review. Full-width katakana typed with an uncomposable combining U+3099 (ア゙イ) already takes the kana license today, contrary to W4. Filed separately as #596, together with a W4 wording fix that review also turned up.

Checks. uv run pytest 11,119 passed. mypy, ruff and both Sphinx builds are clean. The suite was also run without namedivider, and the halfwidth segmenter path was checked with the ja extra installed.

🤖 Generated with Claude Code

derek73 and others added 5 commits October 3, 2026 13:50
The script table now classifies the halfwidth kana block U+FF65-U+FF9F
as Script.KATAKANA, so 山田 タロウ takes the kana license and reads family
山田 as 山田 エミ does. Everything keyed on the table follows: the
second-East-Asian-word count that keeps a segmenter from re-dividing a
kanji name, the 间隔号's flank guard, and _NO_INITIALS, which closes the
title misroute decisions.md#cjk-full-stops recorded (タナカ. John).

A range rather than a fold: T1 forbids rewriting the text, and an NFKC
fold at classification reaches far past the kana and maps a lone ゙ into
the hiragana block. Wholly-katakana names, halfwidth included, keep the
declared order; W4's rationale is amended to say the script cannot tell
a Japanese reading from a transcription, not that the latter dominates.

U+FF65 leaves _SANCTIONED_EXTRAS for the table, the ledgers' span copies
widen to match, and fix(#594) rules classify the five movers at every
baseline (no corpus line held halfwidth kana before the new case rows).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
usage.rst explains halfwidth katakana where Japanese scripts are
introduced and gives the script_orders recipe for callers whose
katakana names are Japanese; locales.rst drops the claim that a
halfwidth neighbour divides nothing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ied (#594)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- rules.md#H2's Accepted clause still called halfwidth katakana
  unclassified; #594 is what changed that, so the same PR amends it.
- The NFKC fold was declined for the wrong reason. The interpunct flank
  guard reads raw text and only asks "classified?"; the real hazard is
  the kana license, since an uncomposable voicing mark (ア゙) folds into
  the hiragana block and would turn a wholly-katakana name
  family-first. Corrected in decisions.md#W4 and _policy.py's comment.
- AGENTS.md's 间隔号 sentence still had pure katakana naming its
  convention, the rationale #594 retracted.
- The cjk-full-stops header and the H2 pointer now say the halfwidth
  limit is closed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73 derek73 added bug docs Documentation fixes and updates labels Oct 3, 2026
@derek73 derek73 self-assigned this Oct 3, 2026
@codecov

codecov Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 98.98%. Comparing base (d989c5e) to head (357f685).

Additional details and impacted files
@@           Coverage Diff           @@
##           master     #595   +/-   ##
=======================================
  Coverage   98.98%   98.98%           
=======================================
  Files          45       45           
  Lines        4220     4220           
=======================================
  Hits         4177     4177           
  Misses         43       43           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

derek73 and others added 4 commits October 3, 2026 14:16
…594)

From the PR review's coverage pass:
- 山田 ペーター pins U+FF9F and U+FF70, which only the ledger sync test
  held: narrowing the range by either one now fails a behavior row.
- 山田タロウ, the halfwidth twin of ja_unspaced_unsegmented_default, pins a
  single Han+halfwidth token taking the license (and 2.3's
  given-or-family report going away).
- The ja adapter stub accepts 山田タロウ, the one changed path no test
  reached.

The license rule in the 2.1/2.2/2.3 ledgers takes the two new movers,
and the #594 ledger blocks are regenerated per baseline: the title
rule's comment claimed the full-width twins read a name word, true
only at 2.3.0, and the 1.4.0/2.0.0 headers described rules those files
do not carry. _CORPUS_CLAIMS records the wider reach.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- _policy.py cited T1 for "nothing rewrites the text"; that is rules.md's
  T Background, now quoted. Same correction in decisions.md#W4.
- The voicing-mark sentence states the mechanism: NFC never composes a
  halfwidth mark into its base.
- decisions.md#W4's population is eight rows, five contract movers.
- migrate.rst still gave the retracted "usually a transcription" reason.
- The release log says "declared order", which is what W4 keeps.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
U+FF70 is interior to the span, not an edge.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…takana (#594)

The same over-broad claim the review corrected in _policy.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73 derek73 added this to the 2.4 milestone Oct 3, 2026
@derek73
derek73 merged commit 5cf668c into master Oct 3, 2026
11 checks passed
@derek73
derek73 deleted the fix/issue-594-halfwidth-katakana branch October 3, 2026 21:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug docs Documentation fixes and updates

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Should halfwidth katakana (タロウ) read family-first like full-width katakana?

1 participant