Remove redundant metadata (#685)

The metadata field (list of triples) in the pipeline Metadata class was redundant. Document metadata triples already flow directly from librarian to triple-store via emit_document_provenance() - they don't need to pass through the extraction pipeline. Additionally, chunker and PDF decoder were overwriting metadata to [] anyway, so any metadata passed through the pipeline was being discarded. Changes: - Remove metadata field from Metadata dataclass (schema/core/metadata.py) - Update all Metadata instantiations to remove metadata=[] parameter - Remove metadata handling from translators (document_loading, knowledge) - Remove metadata consumption from extractors (ontology, agent) - Update gateway serializers and import handlers - Update all unit, integration, and contract tests
2026-06-09 06:45:13 +02:00 · 2026-03-11 10:51:39 +00:00 · 2026-03-11 10:51:39 +00:00 · aa4f5c6c00
commit aa4f5c6c00
parent 1837d73f34
37 changed files with 106 additions and 343 deletions
--- a/docs/tech-specs/extraction-flows.md
+++ b/docs/tech-specs/extraction-flows.md
@ -314,11 +314,19 @@ Converts row index fields into vector embeddings.
 | `document_id` | Librarian reference, provenance linking |
 | `chunk_id` | Provenance tracking through pipeline |

+<<<<<<< HEAD
 ### Potentially Redundant Fields

 | Field | Status |
 |-------|--------|
 | `metadata.metadata` | Set to `[]` by all extractors; document-level metadata now handled by librarian at submission time |
+=======
+### Removed Fields
+
+| Field | Status |
+|-------|--------|
+| `metadata.metadata` | Removed from `Metadata` class. Document-level metadata triples are now emitted directly by librarian to triple store at submission time, not carried through the extraction pipeline. |
+>>>>>>> e3bcbf73 (The metadata field (list of triples) in the pipeline Metadata class)

 ### Bytes Fields Pattern