mirror of
https://github.com/VectifyAI/PageIndex.git
synced 2026-07-18 21:21:05 +02:00
fix: CMYK image drop, empty-doc crash, page_index shadowing, sqlite hardening, flaky tests
Addresses items 9-13 and a/b/c/f from the max-effort review of PR #272. - pdf.py: image colorspace check was `pix.n > 4`, which treats CMYK-without- alpha (n==4, same as RGBA) as not needing RGB conversion; pix.save() as .png then raises "unsupported colorspace", silently dropped by the surrounding except. Fixed to `pix.n - pix.alpha >= 4` (correctly converts CMYK, leaves RGBA untouched). - pipeline.py: detect_strategy([]) (an empty/whitespace-only source file) returned "content_based", routing into the PDF-oriented TOC-detection pipeline -- wasting a real LLM call before raising IndexingError. Empty node lists now route to level_based, whose build_tree_from_levels([]) returns an empty structure instantly with zero LLM calls. - page_index.py (shim): pageindex/__init__.py binds the canonical `page_index` function as the package attribute, but this file is ALSO a real submodule of the same name -- importing it anywhere (import machinery, unconditional) overwrites that attribute with the module object, breaking `from pageindex import page_index; page_index(x)` for the rest of the process. Made the shim module itself callable (delegates to the real function via a ModuleType subclass), so whichever object ends up in that slot is callable regardless of import order. - storage/sqlite.py: create_collection let a raw sqlite3.IntegrityError escape on a duplicate name (new CollectionAlreadyExistsError); the collections table's CHECK constraint only validated the name's first character (GLOB '*' is a wildcard, not a regex quantifier over the preceding class) -- fixed to validate the whole string, and SQLiteStorage now also validates in Python (it's a public StorageEngine usable directly, bypassing LocalBackend's own check). - tests/test_review_fixes_2.py: two tests used a ContentNode with no `level` set, so build_index took the content_based path and made real (retried, slow, and -- with a valid key -- billable) LLM calls instead of testing the text-stripping logic they claimed to. Mocked out _content_based_pipeline. - retrieve.py: _parse_pages/_get_pdf_page_content were independent copies of the canonical parse_pages/get_pdf_page_content that had already drifted (missing the p>=1 filter and 1000-page DoS cap) -- delegate to canonical now, so the legacy pageindex.get_page_content path can't silently regress again. - parser/markdown.py: a leading UTF-8 BOM broke first-header detection (not whitespace, .strip() doesn't remove it) -- decode utf-8-sig. Only backtick fences were recognized as code blocks, so a '#'-prefixed line inside a ~~~-fenced block (valid CommonMark) was misparsed as a heading -- recognize both fence styles. - run_pageindex.py: --if-thinning wasn't migrated to the bare-flag + legacy-yes/no convention the other four --if-add-* flags got; bare usage raised an argparse error and it never went through the shared coercion. - types.py: DocumentDetail's `structure` field was inside the class's total=False body, so TypedDict rules made it optional even though every backend always populates it. Split into a required base class. Adds regression tests for all of the above. Full suite: 244 passed, 2 skipped (one pre-existing, unrelated flaky cloud-streaming test). Claude-Session: https://claude.ai/code/session_01Kx5DgKbhK1N8autqXH8SmS
This commit is contained in:
parent
b9d021916f
commit
4e6a13576d
18 changed files with 368 additions and 36 deletions
|
|
@ -2,6 +2,7 @@
|
|||
2d46d68..8f536cb): the Markdown text-stripping fix's fallout, plus the other
|
||||
directly-fixable findings from that pass."""
|
||||
import asyncio
|
||||
from unittest.mock import AsyncMock
|
||||
|
||||
import pytest
|
||||
|
||||
|
|
@ -122,6 +123,20 @@ def test_cloud_delete_collection_still_clears_real_folder_id(monkeypatch):
|
|||
|
||||
|
||||
# ── #7: remove_structure_text is skipped when text was never added ───────────
|
||||
def _mock_content_based_pipeline(monkeypatch, structure):
|
||||
"""content_based's real path (_content_based_pipeline) drives real LLM
|
||||
calls (TOC detection etc.) regardless of if_add_node_summary — a prior
|
||||
version of these two tests didn't mock this out, fell through to it, and
|
||||
made real network calls (with a dummy key: 10 retries before failing;
|
||||
with a real key: real billable requests) on every run."""
|
||||
from pageindex.index import pipeline
|
||||
|
||||
async def fake(page_list, opt):
|
||||
return structure
|
||||
|
||||
monkeypatch.setattr(pipeline, "_content_based_pipeline", fake)
|
||||
|
||||
|
||||
def test_build_index_skips_text_strip_when_no_text_was_added(monkeypatch):
|
||||
from pageindex.index import pipeline
|
||||
from pageindex.parser.protocol import ContentNode, ParsedDocument
|
||||
|
|
@ -131,6 +146,7 @@ def test_build_index_skips_text_strip_when_no_text_was_added(monkeypatch):
|
|||
# ...` inside the function body), so patch it on the utils module itself.
|
||||
import pageindex.index.utils as utils_mod
|
||||
monkeypatch.setattr(utils_mod, "remove_structure_text", lambda s: calls.append(s) or s)
|
||||
_mock_content_based_pipeline(monkeypatch, [{"title": "T", "start_index": 1, "end_index": 1}])
|
||||
|
||||
nodes = [ContentNode(content="page one text", tokens=5, index=1)]
|
||||
parsed = ParsedDocument(doc_name="d", nodes=nodes)
|
||||
|
|
@ -147,6 +163,14 @@ def test_build_index_still_strips_text_when_summary_added_it(monkeypatch):
|
|||
calls = []
|
||||
import pageindex.index.utils as utils_mod
|
||||
monkeypatch.setattr(utils_mod, "remove_structure_text", lambda s: calls.append(s) or s)
|
||||
_mock_content_based_pipeline(monkeypatch, [{"title": "T", "start_index": 1, "end_index": 1}])
|
||||
# Summary generation itself would otherwise make a real LLM call.
|
||||
monkeypatch.setattr(
|
||||
utils_mod, "generate_summaries_for_structure",
|
||||
AsyncMock(side_effect=lambda structure, model=None: [
|
||||
n.__setitem__("summary", "fake") for n in structure
|
||||
]),
|
||||
)
|
||||
|
||||
nodes = [ContentNode(content="page one text", tokens=5, index=1)]
|
||||
parsed = ParsedDocument(doc_name="d", nodes=nodes)
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue