mirror of
https://github.com/VectifyAI/PageIndex.git
synced 2026-07-24 21:41:04 +02:00
The prompt-injection hardening ported from main did not fit dev's direction
(graceful degradation + content-faithful, reasoning-based index). Adjust it:
- process_toc_no_page_numbers: a count-mismatched or reordered/renamed LLM
response no longer raises ValueError (which aborted the entire index at the
uncaught top-level path and defeated the return_exceptions degradation the
rest of the pipeline uses). Skip the untrusted chunk and continue; the
accuracy check falls back to another mode when too little gets filled.
- _secure_doc_text: drop the keyword blocklist that replaced phrases like
"act as"/"disregard" with [REDACTED]. Those occur in legitimate titles and
prose, so redaction corrupted document content. Keep the <user_document>
framing + _SYSTEM_HARDENING system instruction as the injection defense.
- _validate_chunk_physical_indices: tolerate non-list input (extract_json
returns {} on parse failure and may return a JSON object) instead of
crashing while iterating a dict.
- Remove _validate_physical_indices / _parse_physical_index: range validation
is already done by validate_and_truncate_physical_indices in meta_processor
and marker->int by convert_physical_index_to_int; the helpers duplicated both.
Tests updated to assert the skip-not-raise behavior and lock in no-redaction
and non-list tolerance. 314 passed, 2 skipped.
|
||
|---|---|---|
| .. | ||
| backend | ||
| index | ||
| parser | ||
| storage | ||
| __init__.py | ||
| agent.py | ||
| client.py | ||
| cloud_api.py | ||
| collection.py | ||
| config.py | ||
| config.yaml | ||
| errors.py | ||
| events.py | ||
| page_index.py | ||
| page_index_md.py | ||
| retrieve.py | ||
| tokens.py | ||
| types.py | ||
| utils.py | ||