Remove redundant metadata (#685)

The metadata field (list of triples) in the pipeline Metadata class was redundant. Document metadata triples already flow directly from librarian to triple-store via emit_document_provenance() - they don't need to pass through the extraction pipeline. Additionally, chunker and PDF decoder were overwriting metadata to [] anyway, so any metadata passed through the pipeline was being discarded. Changes: - Remove metadata field from Metadata dataclass (schema/core/metadata.py) - Update all Metadata instantiations to remove metadata=[] parameter - Remove metadata handling from translators (document_loading, knowledge) - Remove metadata consumption from extractors (ontology, agent) - Update gateway serializers and import handlers - Update all unit, integration, and contract tests
2026-06-23 21:58:06 +02:00 · 2026-03-11 10:51:39 +00:00 · 2026-03-11 10:51:39 +00:00 · aa4f5c6c00
commit aa4f5c6c00
parent 1837d73f34
37 changed files with 106 additions and 343 deletions
--- a/tests/unit/test_storage/test_rows_cassandra_storage.py
+++ b/tests/unit/test_storage/test_rows_cassandra_storage.py
@ -190,7 +190,6 @@ class TestRowsCassandraStorageLogic:
                id="test-001",
                user="test_user",
                collection="test_collection",
-                metadata=[]
            ),
            schema_name="test_schema",
            values=[{"id": "123", "value": "test_data"}],
@ -252,7 +251,6 @@ class TestRowsCassandraStorageLogic:
                id="test-001",
                user="test_user",
                collection="test_collection",
-                metadata=[]
            ),
            schema_name="multi_index_schema",
            values=[{"id": "123", "category": "electronics", "status": "active"}],
@ -310,7 +308,6 @@ class TestRowsCassandraStorageBatchLogic:
                id="batch-001",
                user="test_user",
                collection="batch_collection",
-                metadata=[]
            ),
            schema_name="batch_schema",
            values=[
@ -365,7 +362,6 @@ class TestRowsCassandraStorageBatchLogic:
                id="empty-001",
                user="test_user",
                collection="empty_collection",
-                metadata=[]
            ),
            schema_name="empty_schema",
            values=[],  # Empty batch