Kestrel
대시보드로 돌아가기
CVE-2026-79657CRITICALMITRENVDGHSA대응게시일: 2026. 08. 25.수정일: 2026. 09. 08.

NLTK: Allowlisted pickle loaders still permit code execution in current source

Deserialization

위협 신호 · CVSS · EPSS · KEV

정기 패치· 높은 악용 신호 없음
CVSS
critical

이론적 심각도 점수

EPSS
1.2%상위 33.3%

30일 내 악용 확률 예측

KEV
미등재

실측 악용 기록 없음

권장 대응 기한60일 이내CISA SSVC 기준

계획된 패치 주기 내 조치(60일 이내)

외부 노출· KEV 미등재 · 자동화 어려움 · 부분 영향 · 외부 노출

CVSS 벡터 · 메트릭

CVSS 벡터 정보 없음

상세 설명

Summary

The current source tree still allows arbitrary code execution during supposedly safer allowlisted pickle loading. The allowlist trusts whole module namespaces instead of exact safe globals, so crafted pickles can invoke dangerous in-namespace callables through pickle REDUCE.

Details

  • Vulnerability type: Remote code execution via unsafe deserialization
  • Affected component: nltk.picklesec.allowlisted_pickle_load, nltk.tokenize.punkt.punkt_pickle_load, nltk.parse.transitionparser.TransitionParser.parse
  • Affected versions: Current source v3.10.0-rc2; published 3.9.4 was not the claim target for this bypass.
  • Patched versions: Not yet patched
  • Root cause: Module-prefix allowlists include dangerous callables such as nltk.tokenize.repp.ReppTokenizer._execute and numpy.f2py.crackfortran.myeval.

punkt_pickle_load() allowlists both nltk.tokenize.punkt and the whole nltk.tokenize namespace, which exposes ReppTokenizer._execute() and its subprocess.Popen(...) sink during unpickling. TransitionParser.parse() uses allowlisted_pickle_load(..., allowed_modules=("numpy", "scipy", "sklearn")), which permits numpy.f2py.crackfortran.myeval() and its attacker-controlled eval(...) path. I confirmed both gadgets create marker files before the caller returns or later aborts on type misuse.

PoC

Preconditions

  • The application loads an attacker-controlled tokenizer or model artifact through these public loaders.

Steps

  1. Create a pickle whose REDUCE callable is ReppTokenizer._execute and point its command to a harmless marker-file write.
  2. Pass that payload to punkt_pickle_load(BytesIO(payload)) and observe the marker file is created during unpickling.
  3. Create a second pickle whose REDUCE callable is numpy.f2py.crackfortran.myeval and load it through TransitionParser.parse().
  4. Observe the second marker file is created before TransitionParser.parse() later fails on the returned object type.

Minimal reproducible excerpt

text
1{'punkt_marker': 'PUNKT_RCE', 'transitionparser_marker': 'TP_RCE'}

Impact

Any caller that trusts these current allowlisted loaders can still execute attacker-controlled commands while loading model or tokenizer artifacts. This defeats the protection mechanism that replaced unrestricted pickle loading and creates a dangerous false sense of safety.

Remediation

Replace broad module-prefix allowlists with exact (module, qualname) pairs for the few safe classes or functions genuinely required. Do not allow entire namespaces such as nltk.tokenize or numpy, and keep post-load type validation only as a secondary defense.

Resources


Fix + attack demonstration (verified)

  • tightened callers
    find_class now, before the allowlists:
  1. Rejects any dotted name → closes 4489 with zero legit impact.
  2. Denies dangerous modules (os, subprocess, sys, builtins, numpy.f2py, nltk.tokenize.repp, …) even under a broad allowed_modules — a defense-in-depth backstop so a future too-broad allowlist can't silently reopen RCE.
  3. builtins denied wholesale; safe primitives (int, str, …) must be named exactly via allowed_globals.

Callers tightened: punkt drops the broad nltk.tokenize (keeps nltk.tokenize.punkt + exact collections.defaultdict/builtins.int); transitionparser keeps numpy/scipy/sklearn (array unpickling needs their submodules) with the new guards blocking the gadgets.

Full pickle-sink audit

Every deserialization sink in the tree was reviewed: no raw pickle.load anywhere, and no joblib/numpy/torch/dill/yaml/marshal loaders. data.load + wordnet_app use RestrictedUnpickler (blocks all globals — safe); the remaining pickle_load sites (chartparser_app, tbl/demo) load user-selected or self-written files and keep their warning.

Attack demonstration (captured; fork clone)

text
1=== EXPLOITS blocked ===
2 4489 sklearn.os.system (dotted) -> BLOCKED
3 x99w numpy.f2py.crackfortran.myeval -> BLOCKED
4 x99w nltk.tokenize.repp._execute -> BLOCKED
5 backstop os.system (os allowlisted) -> BLOCKED
6 backstop builtins.eval (exact global)-> BLOCKED
7=== LEGIT loads still work ===
8 punkt round-trip via punkt_pickle_load -> OK
9 builtins.int (safe primitive) -> OK

Tests

test_pickle_allowlist_security.py — added 5 regressions (dotted traversal, both namespace gadgets, denied-module backstop, legit round-trip). Suite: 122 passed / 9 skipped (sklearn-dependent) across pickle/punkt/transition/tokenize. pre-commit (black/isort/ruff) clean.

AI 심층 분석

공격 시나리오 · 재현 가능한 PoC 페이로드 · 즉시 적용 가능한 차단 패치를 한 번에 받아 보세요. 보안 운영팀이 그대로 점검·티켓팅에 쓸 수 있는 형태로 정리해 드립니다.