Repository navigation
Conversation
PyMuPDF, which is already used to index PDFs, also reads ePub. So the PDF indexing code becomes a shared helper taking the document type, with get_pdf_index_data and a new get_epub_index_data on top of it, and StaticItem auto-indexes application/epub+zip items the same way it does PDFs. MuPDF warns "unknown epub version: 3.0" on every EPUB 3 document while reading it fine, so that message joins the ignored ones. Warnings about images it cannot load and CSS it does not support are ignored by prefix, since they carry document-specific details and do not affect the text. Fix openzim#333
|
Checked the EPUB paths with synthetic books on Python 3.14. The indexing suite passed (23 tests, one skipped), plus 12 extra checks covering custom titles, custom index data, I didn't find a problem in those checks. I couldn't validate the whole ZIM suite in this container: download-dependent tests failed with networking disabled or |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fix #333
PyMuPDF, which scraperlib already uses for PDF indexing, also opens ePub. So instead of a new parser,
get_pdf_index_datanow goes through a shared_get_document_index_data(filetype=...), andget_epub_index_datais the ePub entry point on top of it. The title is built the same way from the document metadata (title - author).StaticItemauto-indexesapplication/epub+zipitems in the content, fileobj and filepath cases, next to the PDF branch.The filetype is now passed explicitly to
pymupdf.open()for both, rather than relying on detection from the stream.MuPDF says
unknown epub version: 3.0for every EPUB 3 file but reads it fine, so that message is added toIGNORED_MUPDF_MESSAGES. Real ePubs also make it warn about images it cannot load (html: cannot load image src='...') and CSS it does not support (syntax error: css syntax error: ..., e.g. on theimg[src*="..."]selectors in Wikisource exports). Those carry the file name or the CSS, so they go in a newIGNORED_MUPDF_MESSAGE_PREFIXESmatched on the start of the line. Neither matters for the text. Without these, practically every ePub item would log a "PyMuPDF issues" warning.Tried on a real book first: Gutenberg's Pride and Prejudice EPUB 3 (24 MB with images) gives 733k characters of text and "Pride and Prejudice - Jane Austen" as title in about 0.3s.
Tests: a small EPUB 3 is built by a session fixture in
tests/conftest.py(two chapters plus title/author metadata, a missing image and one such CSS selector), so there is no binary file to review. New tests coverget_epub_index_datafrom content, fileobj and filepath, check that none of the warnings above is logged, and check that an auto-indexed ePub item can be found in a ZIM by full-text search and by title suggestion. The item tests fail without theitems.pychange, and the warning check fails if the new ignored message is removed. Locally:tests/zim311 passed (and the warning check fails without the new ignored messages), coverage ofindexing.pyanditems.pyis 100%, andruff check,ruff format --checkandpyrightare clean. The only failures in the full suite are 22 gif tests, which fail the same way onmainbecause I don't havegifsicleinstalled.One thing worth flagging: since
auto_indexdefaults toTrue, scrapers that already add ePub files will start indexing their content after upgrading, as happened for PDFs. That is what #333 asks for, but it means more build time and a bigger index for them. Scrapers that don't want it can passauto_index=False. papers already does, so for openzim/papers#43 the plan is to callget_epub_index_dataexplicitly for the one format picked per book.CHANGELOG entry added under Unreleased / Added.