Knowledge Graph Entity Duplication (LightRAG)
Overview
The Knowledge Graph Entity Duplication (LightRAG) evaluation audits the entity layer of a LightRAG knowledge graph. It retrieves every entity label stored in the graph and groups the labels by a normalized key - Unicode form (NFKC), case, whitespace and surrounding punctuation. Any key that maps to more than one label is a duplicate group: the same real-world entity persisted under textually distinct names.
The evaluation reports whether any duplicate groups exist and what share of the graph they represent. The reasoning attached to each metric also lists the duplicate groups and their colliding labels, so the number of affected entities and labels can be read off directly. It inspects labels only, so it checks the graph as it is stored rather than the documents that produced it.
Metrics
Has Duplicates
Whether the knowledge graph contains at least one duplicate group (binary: 1 or 0).
Duplicate Rate
The share of all inspected labels that participate in a duplicate group,
duplicate_labels / total_labels (range: 0.0 to 1.0, lower is better).
Motivation
A knowledge graph is only as reliable as the entities it contains. When the same
real-world entity is stored under more than one label that normalizes to the same key,
the graph is silently fragmented: OpenAI and OPENAI (case), ACME Corp. and
ACME Corp (wrapping punctuation), Open AI and Open AI (collapsed whitespace), or
ACME and ACME (Unicode form).
The failure is easy to miss because each label looks reasonable in isolation. Downstream retrieval then suffers in ways that are hard to attribute: facts about one entity are split across two nodes, edges attach to the "other" node so traversals miss them, aggregations double-count or under-count, and answers that should cite a well-connected entity instead cite two sparsely connected fragments. The result is degraded recall and faithfulness that cannot be fixed at query time.
The normalization used here catches collisions that survive Unicode form, case, whitespace and wrapping punctuation. Semantic duplicates such as "IBM" and "International Business Machines", abbreviations, aliases and transliterations are out of scope and require a different check.
Methodology
- Connect: The operator provides the base URL of the LightRAG server. The
evaluation calls
GET {lightrag_url}/graph/label/listwith theX-API-Keyheader taken from theLIGHTRAG_API_KEYsecret. - Collect labels: The response is read as a JSON array of labels, or a JSON object
with the labels under a
labelsordatakey. Non-string and empty entries are discarded. - Normalize: Each label is reduced to a canonical key by applying Unicode NFKC
normalization, casefolding, collapsing runs of whitespace to a single space, and
stripping wrapping punctuation and quotes (including
.,,,(),[],{},<>,«»,#,_,-and unicode quotes) from both ends. Labels that normalize to an empty string are ignored. - Group and score: Labels are grouped by their normalized key. A key with more than one label is a duplicate group. The evaluation reports whether any group exists and the share of labels affected. The reasoning attached to each metric includes a table of the duplicate groups and their colliding labels so the offending entities can be located directly.
Scoring
Has Duplicates Scorer
Duplicate Rate Scorer
Examples
The following examples show the labels stored in a LightRAG graph and the metrics they
produce. Labels are listed as fetched from /graph/label/list.
Clean - no colliding labels
OpenAI Microsoft Google Anthropic Meta
Grouped 5 labels by normalized key. Every key maps to exactly one label, so no duplicate entities were found.
No duplicate labels found among 5 labels.
Flagged - one entity split across two labels
OpenAI OPENAI Microsoft Amazon Google
The labels OpenAI and OPENAI normalize to the same key openai, so the graph contains a duplicate group.
| Normalized Key | Duplicate Labels |
|---|---|
openai | OpenAI, OPENAI |
Found 1 duplicate group among 5 labels, covering 2 labels (40.0%). The other three labels (Microsoft, Amazon, Google) are unique.
Flagged - widespread duplication
OpenAI OPENAI ACME Corp. ACME Corp Apple Inc Apple Inc. Google LLC google llc Deep Mind Deep Mind Microsoft Amazon
Five normalized keys each map to more than one label, so the graph is fragmented.
| Normalized Key | Duplicate Labels |
|---|---|
acme corp | ACME Corp., ACME Corp |
apple inc | Apple Inc, Apple Inc. |
deep mind | Deep Mind, Deep Mind |
google llc | Google LLC, google llc |
openai | OpenAI, OPENAI |
Found 5 duplicate groups among 12 labels, covering 10 labels (83.3%). Microsoft and Amazon are the only unique labels.