Follow-up to #344 / #340.
When a page's own URL contains consecutive slashes, relative links inside it are resolved with urllib.parse.urljoin, which drops the empty segments. So the // gets collapsed before normalize() even sees the URL.
Example (from @Sinkleberg's check in #344): rewriting other.html from https://example.com/x//y/page.html gives the ZIM path example.com/x/y/other.html, but a browser (WHATWG URL resolution) resolves it to https://example.com/x//y/other.html. So the link points to an entry that doesn't exist.
This is in ArticleUrlRewriter.__call__:
item_absolute_url = urljoin(
urljoin(self.article_url.value, base_href), item_url
)
The same applies to base_href resolution. A fix probably needs an RFC 3986 / WHATWG-style join that keeps empty segments instead of urljoin. #344 already has a small resolver like that in its tests.
Follow-up to #344 / #340.
When a page's own URL contains consecutive slashes, relative links inside it are resolved with
urllib.parse.urljoin, which drops the empty segments. So the//gets collapsed beforenormalize()even sees the URL.Example (from @Sinkleberg's check in #344): rewriting
other.htmlfromhttps://example.com/x//y/page.htmlgives the ZIM pathexample.com/x/y/other.html, but a browser (WHATWG URL resolution) resolves it tohttps://example.com/x//y/other.html. So the link points to an entry that doesn't exist.This is in
ArticleUrlRewriter.__call__:The same applies to
base_hrefresolution. A fix probably needs an RFC 3986 / WHATWG-style join that keeps empty segments instead ofurljoin. #344 already has a small resolver like that in its tests.