cat /etc/motd
The next decade's attack surface is the whole stack.
> stack := model · agent · instruction · context · memory · tools · external_sources
Every layer is a vector. Hallucination at the model. Prompt injection through the context. Jailbreak in the instruction. Misuse via tools. Drift in memory. Hijack from external sources. And the failure modes nobody has named yet.
defenders are outnumbered. attackers improvise in public. vendors race each other.
Shadow-LLM-Guardians is a community archive of what actually breaks in the wild. Reproducibly. Citably. Without NDAs.
cases.indexed
211
cases.active
211
auth.required
github
archive.policy
open
sort --by=hot --decay=30d
// upvotes × comments × time-decay · top 5
01
[CASE-002]·ACTIVE·3mo·@mexiQQ
Claude Opus 4.7 killed its own bash session via broad pkill regex; then claimed it had 'restarted'
tool-misusehallucinationdestructive-actionagent-misbehavior
▲ 2» 0
02
[CASE-070]·ACTIVE·3mo·@mexiQQ
Worker agent writes malicious hook to Claude Code settings.json via shared volume, gaining persistent orchestrator RCE
agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-publishedmodel-unknown
▲ 0» 1
03
[CASE-012]·ACTIVE·3mo·@WeizhiGao
Agent deleted user files with broad rm command, then claimed cleanup succeeded
unreviewed
▲ 0» 1
04
[CASE-211]·ACTIVE·now·@mexiQQ
MCP agent self-modifies its config to add attacker-controlled MCP server via tool return injection
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
05
[CASE-210]·ACTIVE·now·@mexiQQ
MCP agent exfiltrates OPENAI_API_KEY to attacker tool via poisoned tool return instructions
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
ls -lt --time=created
// freshest submissions · top 5
[CASE-211]·ACTIVE·now·@mexiQQ
MCP agent self-modifies its config to add attacker-controlled MCP server via tool return injection
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
[CASE-210]·ACTIVE·now·@mexiQQ
MCP agent exfiltrates OPENAI_API_KEY to attacker tool via poisoned tool return instructions
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
[CASE-209]·ACTIVE·now·@mexiQQ
MCP agent enters infinite tool-call loop via adversarial 'stage progress' payload (C-DoS)
agent-loopneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
[CASE-208]·ACTIVE·23h·@mexiQQ
Approximate unlearning (TrajDeleter) recovers 69–91% of oracle backdoor activation gap in offline RL
from-arxivauto-publishedbackdoor-attack
▲ 0» 0
[CASE-207]·ACTIVE·23h·@mexiQQ
Compliance-driven unlearning reactivates dormant backdoor in offline RL policies (UBA-ORL)
from-arxivauto-publishedbackdoor-attack
▲ 0» 0
sort --by=hot --decay=30d
// upvotes × comments × time-decay · top 5
ls -lt --time=created
// freshest submissions · top 5
01
[CASE-002]·ACTIVE·3mo·@mexiQQ
Claude Opus 4.7 killed its own bash session via broad pkill regex; then claimed it had 'restarted'
tool-misusehallucinationdestructive-actionagent-misbehavior
▲ 2» 0
[CASE-211]·ACTIVE·now·@mexiQQ
MCP agent self-modifies its config to add attacker-controlled MCP server via tool return injection
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
02
[CASE-070]·ACTIVE·3mo·@mexiQQ
Worker agent writes malicious hook to Claude Code settings.json via shared volume, gaining persistent orchestrator RCE
agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-publishedmodel-unknown
▲ 0» 1
[CASE-210]·ACTIVE·now·@mexiQQ
MCP agent exfiltrates OPENAI_API_KEY to attacker tool via poisoned tool return instructions
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
03
[CASE-012]·ACTIVE·3mo·@WeizhiGao
Agent deleted user files with broad rm command, then claimed cleanup succeeded
unreviewed
▲ 0» 1
[CASE-209]·ACTIVE·now·@mexiQQ
MCP agent enters infinite tool-call loop via adversarial 'stage progress' payload (C-DoS)
agent-loopneeds-disclosure-reviewfrom-arxivauto-published
▲ 0» 0
04
[CASE-211]·ACTIVE·now·@mexiQQ
MCP agent self-modifies its config to add attacker-controlled MCP server via tool return injection
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
[CASE-208]·ACTIVE·23h·@mexiQQ
Approximate unlearning (TrajDeleter) recovers 69–91% of oracle backdoor activation gap in offline RL
from-arxivauto-publishedbackdoor-attack
▲ 0» 0
05
[CASE-210]·ACTIVE·now·@mexiQQ
MCP agent exfiltrates OPENAI_API_KEY to attacker tool via poisoned tool return instructions
needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
▲ 0» 0
[CASE-207]·ACTIVE·23h·@mexiQQ
Compliance-driven unlearning reactivates dormant backdoor in offline RL policies (UBA-ORL)
from-arxivauto-publishedbackdoor-attack
▲ 0» 0