fix: CMYK image drop, empty-doc crash, page_index shadowing, sqlite hardening, flaky tests

Addresses items 9-13 and a/b/c/f from the max-effort review of PR #272.

- pdf.py: image colorspace check was `pix.n > 4`, which treats CMYK-without-
  alpha (n==4, same as RGBA) as not needing RGB conversion; pix.save() as .png
  then raises "unsupported colorspace", silently dropped by the surrounding
  except. Fixed to `pix.n - pix.alpha >= 4` (correctly converts CMYK, leaves
  RGBA untouched).

- pipeline.py: detect_strategy([]) (an empty/whitespace-only source file)
  returned "content_based", routing into the PDF-oriented TOC-detection
  pipeline -- wasting a real LLM call before raising IndexingError. Empty
  node lists now route to level_based, whose build_tree_from_levels([])
  returns an empty structure instantly with zero LLM calls.

- page_index.py (shim): pageindex/__init__.py binds the canonical `page_index`
  function as the package attribute, but this file is ALSO a real submodule
  of the same name -- importing it anywhere (import machinery, unconditional)
  overwrites that attribute with the module object, breaking
  `from pageindex import page_index; page_index(x)` for the rest of the
  process. Made the shim module itself callable (delegates to the real
  function via a ModuleType subclass), so whichever object ends up in that
  slot is callable regardless of import order.

- storage/sqlite.py: create_collection let a raw sqlite3.IntegrityError escape
  on a duplicate name (new CollectionAlreadyExistsError); the collections
  table's CHECK constraint only validated the name's first character (GLOB
  '*' is a wildcard, not a regex quantifier over the preceding class) --
  fixed to validate the whole string, and SQLiteStorage now also validates in
  Python (it's a public StorageEngine usable directly, bypassing
  LocalBackend's own check).

- tests/test_review_fixes_2.py: two tests used a ContentNode with no `level`
  set, so build_index took the content_based path and made real (retried,
  slow, and -- with a valid key -- billable) LLM calls instead of testing the
  text-stripping logic they claimed to. Mocked out _content_based_pipeline.

- retrieve.py: _parse_pages/_get_pdf_page_content were independent copies of
  the canonical parse_pages/get_pdf_page_content that had already drifted
  (missing the p>=1 filter and 1000-page DoS cap) -- delegate to canonical
  now, so the legacy pageindex.get_page_content path can't silently regress
  again.

- parser/markdown.py: a leading UTF-8 BOM broke first-header detection
  (not whitespace, .strip() doesn't remove it) -- decode utf-8-sig. Only
  backtick fences were recognized as code blocks, so a '#'-prefixed line
  inside a ~~~-fenced block (valid CommonMark) was misparsed as a heading --
  recognize both fence styles.

- run_pageindex.py: --if-thinning wasn't migrated to the bare-flag +
  legacy-yes/no convention the other four --if-add-* flags got; bare usage
  raised an argparse error and it never went through the shared coercion.

- types.py: DocumentDetail's `structure` field was inside the class's
  total=False body, so TypedDict rules made it optional even though every
  backend always populates it. Split into a required base class.

Adds regression tests for all of the above. Full suite: 244 passed, 2 skipped
(one pre-existing, unrelated flaky cloud-streaming test).

Claude-Session: https://claude.ai/code/session_01Kx5DgKbhK1N8autqXH8SmS
This commit is contained in:
mountain 2026-07-09 11:58:59 +08:00
parent b9d021916f
commit 4e6a13576d
18 changed files with 368 additions and 36 deletions

View file

@ -2,6 +2,7 @@
2d46d68..8f536cb): the Markdown text-stripping fix's fallout, plus the other
directly-fixable findings from that pass."""
import asyncio
from unittest.mock import AsyncMock
import pytest
@ -122,6 +123,20 @@ def test_cloud_delete_collection_still_clears_real_folder_id(monkeypatch):
# ── #7: remove_structure_text is skipped when text was never added ───────────
def _mock_content_based_pipeline(monkeypatch, structure):
"""content_based's real path (_content_based_pipeline) drives real LLM
calls (TOC detection etc.) regardless of if_add_node_summary a prior
version of these two tests didn't mock this out, fell through to it, and
made real network calls (with a dummy key: 10 retries before failing;
with a real key: real billable requests) on every run."""
from pageindex.index import pipeline
async def fake(page_list, opt):
return structure
monkeypatch.setattr(pipeline, "_content_based_pipeline", fake)
def test_build_index_skips_text_strip_when_no_text_was_added(monkeypatch):
from pageindex.index import pipeline
from pageindex.parser.protocol import ContentNode, ParsedDocument
@ -131,6 +146,7 @@ def test_build_index_skips_text_strip_when_no_text_was_added(monkeypatch):
# ...` inside the function body), so patch it on the utils module itself.
import pageindex.index.utils as utils_mod
monkeypatch.setattr(utils_mod, "remove_structure_text", lambda s: calls.append(s) or s)
_mock_content_based_pipeline(monkeypatch, [{"title": "T", "start_index": 1, "end_index": 1}])
nodes = [ContentNode(content="page one text", tokens=5, index=1)]
parsed = ParsedDocument(doc_name="d", nodes=nodes)
@ -147,6 +163,14 @@ def test_build_index_still_strips_text_when_summary_added_it(monkeypatch):
calls = []
import pageindex.index.utils as utils_mod
monkeypatch.setattr(utils_mod, "remove_structure_text", lambda s: calls.append(s) or s)
_mock_content_based_pipeline(monkeypatch, [{"title": "T", "start_index": 1, "end_index": 1}])
# Summary generation itself would otherwise make a real LLM call.
monkeypatch.setattr(
utils_mod, "generate_summaries_for_structure",
AsyncMock(side_effect=lambda structure, model=None: [
n.__setitem__("summary", "fake") for n in structure
]),
)
nodes = [ContentNode(content="page one text", tokens=5, index=1)]
parsed = ParsedDocument(doc_name="d", nodes=nodes)