mirror of
https://github.com/VectifyAI/PageIndex.git
synced 2026-07-24 21:41:04 +02:00
main's changes landed on the top-level pageindex/page_index.py and page_index_md.py, which are deprecation shims on dev (the pipeline now lives in pageindex/index/). Resolving the conflict by keeping the shims would silently drop main's work, so port it into index/ instead: - page_index.py: prompt-injection hardening (_SYSTEM_HARDENING, _secure_doc_text and the redact/wrap helpers) applied to every prompt builder, plus physical_index validation (_validate_physical_indices, _validate_chunk_physical_indices) wired into toc_index_extractor, process_no_toc and process_toc_no_page_numbers, and TOC-identity validation (reject reordered/modified/count-mismatched LLM output). Combined with dev's existing robustness edits in the same functions. - page_index_md.py: recognize whole-line **bold** as a level-1 heading, skip empty bold headings, use the stored node level. - tests: retarget the ported tests at pageindex.index.* instead of the shim (underscore helpers aren't re-exported by the shim's `import *`, and patching the shim wouldn't intercept the real pipeline calls). 311 passed, 2 skipped. |
||
|---|---|---|
| .. | ||
| backend | ||
| index | ||
| parser | ||
| storage | ||
| __init__.py | ||
| agent.py | ||
| client.py | ||
| cloud_api.py | ||
| collection.py | ||
| config.py | ||
| config.yaml | ||
| errors.py | ||
| events.py | ||
| page_index.py | ||
| page_index_md.py | ||
| retrieve.py | ||
| tokens.py | ||
| types.py | ||
| utils.py | ||