Description
A vulnerability in the DocugamiReader class of the run-llama/llama_index repository, up to version 0.12.28, involves the use of MD5 hashing to generate IDs for document chunks. This approach leads to hash collisions when structurally distinct chunks contain identical text, resulting in one chunk overwriting another. This can cause loss of semantically or legally important document content, breakage of parent-child chunk hierarchies, and inaccurate or hallucinated responses in AI outputs. The issue is resolved in version 0.3.1.
References (2)
Core 2
Core References
Exploit, Third Party Advisory
https://huntr.com/bounties/1a48a011-a3c5-4979-9ffc-9652280bc389
Scores
CVSS v3
6.5
EPSS
0.0011
EPSS Percentile
28.2%
Attack Vector
NETWORK
CVSS:3.0/AV:N/AC:L/PR:N/UI:N/S:U/C:N/I:L/A:L
CISA SSVC
Vulnrichment
Exploitation
poc
Automatable
yes
Technical Impact
partial
Details
CWE
CWE-440
Status
published
Products (3)
llamaindex/llamaindex
< 0.3.1
pypi/llama-index
0 - 0.12.41PyPI
pypi/llama-index-readers-docugami
0 - 0.3.1PyPI
Published
Jul 10, 2025
Tracked Since
Feb 18, 2026