Skip to content

Relative links in a page whose URL contains // are resolved with the slashes collapsed #346

Description

@anshuman83-40

Follow-up to #344 / #340.

When a page's own URL contains consecutive slashes, relative links inside it are resolved with urllib.parse.urljoin, which drops the empty segments. So the // gets collapsed before normalize() even sees the URL.

Example (from @Sinkleberg's check in #344): rewriting other.html from https://example.com/x//y/page.html gives the ZIM path example.com/x/y/other.html, but a browser (WHATWG URL resolution) resolves it to https://example.com/x//y/other.html. So the link points to an entry that doesn't exist.

This is in ArticleUrlRewriter.__call__:

item_absolute_url = urljoin(
    urljoin(self.article_url.value, base_href), item_url
)

The same applies to base_href resolution. A fix probably needs an RFC 3986 / WHATWG-style join that keeps empty segments instead of urljoin. #344 already has a small resolver like that in its tests.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions