atlas-knowledge_graph_entity_duplication_lightrag

Knowledge Graph Entity Duplication (LightRAG)

Audits a LightRAG knowledge graph for duplicate entities by fetching every entity label and reporting labels that collide once normalized for Unicode form, case, whitespace and surrounding punctuation.
Tags:
Data Quality

Overview

The Knowledge Graph Entity Duplication (LightRAG) evaluation audits the entity layer of a LightRAG knowledge graph. It retrieves every entity label stored in the graph and groups the labels by a normalized key - Unicode form (NFKC), case, whitespace and surrounding punctuation. Any key that maps to more than one label is a duplicate group: the same real-world entity persisted under textually distinct names.

The evaluation reports whether any duplicate groups exist and what share of the graph they represent. The reasoning attached to each metric also lists the duplicate groups and their colliding labels, so the number of affected entities and labels can be read off directly. It inspects labels only, so it checks the graph as it is stored rather than the documents that produced it.

Metrics

Has Duplicates

Whether the knowledge graph contains at least one duplicate group (binary: 1 or 0).

Has Duplicates
0.01.0
0.0
1.0
0.0Every normalized key maps to a single label - the graph contains no duplicate entities.
1.0At least one normalized key maps to more than one label - the graph is fragmented.

Duplicate Rate

The share of all inspected labels that participate in a duplicate group, duplicate_labels / total_labels (range: 0.0 to 1.0, lower is better).

Duplicate Rate
0.01.0
0.0
0.2
1.0
0.0No label collisions - the entity layer is clean.
0.2A fifth of labels collide with another label - noticeable fragmentation.
1.0Every label collides with another label - the graph is fully fragmented.

Motivation

A knowledge graph is only as reliable as the entities it contains. When the same real-world entity is stored under more than one label that normalizes to the same key, the graph is silently fragmented: OpenAI and OPENAI (case), ACME Corp. and ACME Corp (wrapping punctuation), Open AI and Open AI (collapsed whitespace), or ACME and ACME (Unicode form).

The failure is easy to miss because each label looks reasonable in isolation. Downstream retrieval then suffers in ways that are hard to attribute: facts about one entity are split across two nodes, edges attach to the "other" node so traversals miss them, aggregations double-count or under-count, and answers that should cite a well-connected entity instead cite two sparsely connected fragments. The result is degraded recall and faithfulness that cannot be fixed at query time.

The normalization used here catches collisions that survive Unicode form, case, whitespace and wrapping punctuation. Semantic duplicates such as "IBM" and "International Business Machines", abbreviations, aliases and transliterations are out of scope and require a different check.

Methodology

  1. Connect: The operator provides the base URL of the LightRAG server. The evaluation calls GET {lightrag_url}/graph/label/list with the X-API-Key header taken from the LIGHTRAG_API_KEY secret.
  2. Collect labels: The response is read as a JSON array of labels, or a JSON object with the labels under a labels or data key. Non-string and empty entries are discarded.
  3. Normalize: Each label is reduced to a canonical key by applying Unicode NFKC normalization, casefolding, collapsing runs of whitespace to a single space, and stripping wrapping punctuation and quotes (including ., ,, (), [], {}, <>, «», #, _, - and unicode quotes) from both ends. Labels that normalize to an empty string are ignored.
  4. Group and score: Labels are grouped by their normalized key. A key with more than one label is a duplicate group. The evaluation reports whether any group exists and the share of labels affected. The reasoning attached to each metric includes a table of the duplicate groups and their colliding labels so the offending entities can be located directly.

Scoring

Has Duplicates Scorer

Has Duplicates
Score valueExplanation
0Every normalized key maps to a single label - the graph contains no duplicate entities.
1At least one normalized key maps to more than one label - the graph is fragmented.

Duplicate Rate Scorer

Duplicate Rate
Score valueExplanation
0.0No label collisions - the entity layer is clean.
0.2A fifth of labels collide with another label - noticeable fragmentation worth resolving.
1.0Every label collides with another label - the graph is fully fragmented.

Examples

The following examples show the labels stored in a LightRAG graph and the metrics they produce. Labels are listed as fetched from /graph/label/list.

Clean - no colliding labels

Labels fetched

OpenAI Microsoft Google Anthropic Meta

Has Duplicates
0.0

Grouped 5 labels by normalized key. Every key maps to exactly one label, so no duplicate entities were found.

Duplicate Rate
0.0

No duplicate labels found among 5 labels.

Flagged - one entity split across two labels

Labels fetched

OpenAI OPENAI Microsoft Amazon Google

Has Duplicates
1.0

The labels OpenAI and OPENAI normalize to the same key openai, so the graph contains a duplicate group.

Normalized KeyDuplicate Labels
openaiOpenAI, OPENAI
Duplicate Rate
0.4

Found 1 duplicate group among 5 labels, covering 2 labels (40.0%). The other three labels (Microsoft, Amazon, Google) are unique.

Flagged - widespread duplication

Labels fetched

OpenAI OPENAI ACME Corp. ACME Corp Apple Inc Apple Inc. Google LLC google llc Deep Mind Deep Mind Microsoft Amazon

Has Duplicates
1.0

Five normalized keys each map to more than one label, so the graph is fragmented.

Normalized KeyDuplicate Labels
acme corpACME Corp., ACME Corp
apple incApple Inc, Apple Inc.
deep mindDeep Mind, Deep Mind
google llcGoogle LLC, google llc
openaiOpenAI, OPENAI
Duplicate Rate
0.83

Found 5 duplicate groups among 12 labels, covering 10 labels (83.3%). Microsoft and Amazon are the only unique labels.

Run Evaluation in LatticeFlow AI Platform

Use the following CLI command to initialize and run the evaluation in LatticeFlow AI Platform.
Requires LatticeFlow AI Platform CLI
lf init eval --key atlas-knowledge_graph_entity_duplication_lightrag

Metrics

Has DuplicatesDuplicate Rate

Don't have the LatticeFlow AI Platform?

Contact us to see this evaluation in action:
Contact Us