mirror of
https://github.com/trustgraph-ai/trustgraph.git
synced 2026-07-03 15:01:00 +02:00
Subgraph provenance (#694)
Replace per-triple provenance reification with subgraph model Extraction provenance previously created a full reification (statement URI, activity, agent) for every single extracted triple, producing ~13 provenance triples per knowledge triple. Since each chunk is processed by a single LLM call, this was both redundant and semantically inaccurate. Now one subgraph object is created per chunk extraction, with tg:contains linking to each extracted triple. For 20 extractions from a chunk this reduces provenance from ~260 triples to ~33. - Rename tg:reifies -> tg:contains, stmt_uri -> subgraph_uri - Replace triple_provenance_triples() with subgraph_provenance_triples() - Refactor kg-extract-definitions and kg-extract-relationships to generate provenance once per chunk instead of per triple - Add subgraph provenance to kg-extract-ontology and kg-extract-agent (previously had none) - Update CLI tools and tech specs to match Also rename tg-show-document-hierarchy to tg-show-extraction-provenance. Added extra typing for extraction provenance, fixed extraction prov CLI
This commit is contained in:
parent
35128ff019
commit
64e3f6bd0d
20 changed files with 463 additions and 193 deletions
|
|
@ -41,7 +41,7 @@ class QuotedTriple:
|
|||
enabling statements about statements.
|
||||
|
||||
Example:
|
||||
# stmt:123 tg:reifies <<:Hope skos:definition "A feeling...">>
|
||||
# subgraph:123 tg:contains <<:Hope skos:definition "A feeling...">>
|
||||
qt = QuotedTriple(
|
||||
s=Uri("https://example.org/Hope"),
|
||||
p=Uri("http://www.w3.org/2004/02/skos/core#definition"),
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue