Remove redundant metadata (#685)

The metadata field (list of triples) in the pipeline Metadata class was redundant. Document metadata triples already flow directly from librarian to triple-store via emit_document_provenance() - they don't need to pass through the extraction pipeline. Additionally, chunker and PDF decoder were overwriting metadata to [] anyway, so any metadata passed through the pipeline was being discarded. Changes: - Remove metadata field from Metadata dataclass (schema/core/metadata.py) - Update all Metadata instantiations to remove metadata=[] parameter - Remove metadata handling from translators (document_loading, knowledge) - Remove metadata consumption from extractors (ontology, agent) - Update gateway serializers and import handlers - Update all unit, integration, and contract tests
2026-06-20 12:18:07 +02:00 · 2026-03-11 10:51:39 +00:00 · 2026-03-11 10:51:39 +00:00 · aa4f5c6c00
commit aa4f5c6c00
parent 1837d73f34
37 changed files with 106 additions and 343 deletions
--- a/tests/unit/test_knowledge_graph/conftest.py
+++ b/tests/unit/test_knowledge_graph/conftest.py
@ -29,11 +29,10 @@ class Triple:
        self.o = o

 class Metadata:
-    def __init__(self, id, user, collection, metadata):
+    def __init__(self, id, user, collection):
        self.id = id
        self.user = user
        self.collection = collection
-        self.metadata = metadata

 class Triples:
    def __init__(self, metadata, triples):
@ -110,7 +109,6 @@ def sample_triples(sample_triple):
        id="test-doc-123",
        user="test_user",
        collection="test_collection",
-        metadata=[]
    )
    
    return Triples(
@ -126,7 +124,6 @@ def sample_chunk():
        id="test-chunk-456",
        user="test_user",
        collection="test_collection",
-        metadata=[]
    )
    
    return Chunk(