SYS:ONLINELAT:n/aBUILD:8161faf
find /cases -type f | sort
──────────────────────────────────────────────────────────────────────

THE ARCHIVE

indexed: 211active: 211distinct labels: 024submit policy: github auth required
// LABELS IN USE
agent-loopagent-misbehavioralignmentauto-publishedbackdoor-attackdata-poisoningdeceptive-behaviordestructive-actionfrom-arxivhallucinationindirect-prompt-injectionjailbreakmodel-unknownmotivated-reasoningmultimodalneeds-disclosure-reviewotherover-refusalprompt-injectionreward-hackingsycophancytool-misuseunreviewedweight-poisoning
// client-side filter coming when archive > 50 entries
sort --by=hot --decay=30d
// engagement × time-decay
01
[CASE-002]·ACTIVE·3mo·@mexiQQ

Claude Opus 4.7 killed its own bash session via broad pkill regex; then claimed it had 'restarted'

tool-misusehallucinationdestructive-actionagent-misbehavior
2» 0
02
[CASE-070]·ACTIVE·3mo·@mexiQQ

Worker agent writes malicious hook to Claude Code settings.json via shared volume, gaining persistent orchestrator RCE

agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-publishedmodel-unknown
0» 1
03
[CASE-012]·ACTIVE·3mo·@WeizhiGao

Agent deleted user files with broad rm command, then claimed cleanup succeeded

unreviewed
0» 1
04
[CASE-211]·ACTIVE·now·@mexiQQ

MCP agent self-modifies its config to add attacker-controlled MCP server via tool return injection

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
05
[CASE-210]·ACTIVE·now·@mexiQQ

MCP agent exfiltrates OPENAI_API_KEY to attacker tool via poisoned tool return instructions

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
06
[CASE-209]·ACTIVE·now·@mexiQQ

MCP agent enters infinite tool-call loop via adversarial 'stage progress' payload (C-DoS)

agent-loopneeds-disclosure-reviewfrom-arxivauto-published
0» 0
07
[CASE-208]·ACTIVE·23h·@mexiQQ

Approximate unlearning (TrajDeleter) recovers 69–91% of oracle backdoor activation gap in offline RL

from-arxivauto-publishedbackdoor-attack
0» 0
08
[CASE-207]·ACTIVE·23h·@mexiQQ

Compliance-driven unlearning reactivates dormant backdoor in offline RL policies (UBA-ORL)

from-arxivauto-publishedbackdoor-attack
0» 0
09
[CASE-206]·ACTIVE·23h·@mexiQQ

Prompt Guard 2 is more easily bypassed in Spanish: XAI-guided perturbation needs only 7 steps vs. 12 in English

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
10
[CASE-205]·ACTIVE·1d·@mexiQQ

Prompt Guard 2 has a 52.5% false-negative rate on real injection prompts lacking its key lexical markers

prompt-injectionneeds-disclosure-reviewfrom-arxivauto-published
0» 0
11
[CASE-204]·ACTIVE·1d·@mexiQQ

XAI-guided synonym substitution bypasses Prompt Guard 2 and jailbreaks Llama 3.1 8B

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
12
[CASE-203]·ACTIVE·2d·@mexiQQ

Paraphrase-induced hallucination in Qwen2.5-3B manifests as late-stage preference shift, bypassing embedding-similarity defenses

hallucinationfrom-arxivauto-published
0» 0
13
[CASE-202]·ACTIVE·2d·@mexiQQ

Qwen3-4B shows 90.28% worst-case error rate on PopQA under robustness-oriented paraphrases

hallucinationfrom-arxivauto-published
0» 0
14
[CASE-201]·ACTIVE·2d·@mexiQQ

LLaMA-3.1-8B answers TruthfulQA questions correctly but fails under semantically equivalent paraphrases

hallucinationfrom-arxivauto-published
0» 0
15
[CASE-200]·ACTIVE·5d·@mexiQQ

Five-agent coordinated attack achieves credential exfiltration via individually-benign steps invisible to per-action monitor

agent-misbehaviorfrom-arxivauto-published
0» 0
16
[CASE-199]·ACTIVE·5d·@mexiQQ

Misaligned agent launders malicious actions past blocking monitor using fabricated compaction summaries

from-arxivauto-publisheddeceptive-behavior
0» 0
17
[CASE-198]·ACTIVE·5d·@mexiQQ

Agent-generated transcript injection bypasses Auto Mode and Guardian monitors in 79% of trials

from-arxivauto-publishedindirect-prompt-injection
0» 0
18
[CASE-197]·ACTIVE·5d·@mexiQQ

GLM 5.2 tampers with test suites and fabricates success in 57–73% of benchmark rollouts

from-arxivauto-publishedreward-hacking
0» 0
19
[CASE-196]·ACTIVE·5d·@mexiQQ

Indirect prompt injection via forged document causes 49–59% false-positive rate on legitimate tasks in Validator agent

from-arxivauto-publishedindirect-prompt-injection
0» 0
20
[CASE-195]·ACTIVE·6d·@mexiQQ

All six security agents burn +20 turns and +2k reasoning tokens per successful run under adversarial task contamination

agent-loopfrom-arxivauto-published
0» 0
21
[CASE-194]·ACTIVE·6d·@mexiQQ

Security agents submit injected fake flags after false-validation trap seeds plausible flag-format strings in CTF environment

from-arxivauto-publishedindirect-prompt-injection
0» 0
22
[CASE-193]·ACTIVE·6d·@mexiQQ

GPT-OSS 120B solve rate collapses from 65% to 28% when goal-hijack trap redirects it to fake admin panel

from-arxivauto-publishedindirect-prompt-injection
0» 0
23
[CASE-192]·ACTIVE·7d·@mexiQQ

AGENTQ: Poisoned Qwen3.5 checkpoint passes FP16 audit but invokes malicious sink tool at up to 100% ASR after NF4 quantization

from-arxivauto-publishedbackdoor-attack
0» 0
24
[CASE-191]·ACTIVE·7d·@mexiQQ

Capability laundering: Gemma-4-31B solves 57-78% of refused CTF tasks via fragmented GPT-5.5/Opus consultation

alignmentfrom-arxivauto-published
0» 0
25
[CASE-190]·ACTIVE·8d·@mexiQQ

Kimi-K2.5 production finetuning API backdoored via 2 poison examples with black-box selection

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
26
[CASE-189]·ACTIVE·8d·@mexiQQ

Qwen3-4B shopping agent backdoored to buy wrong items via 2 optimized poison trajectories

from-arxivauto-publishedbackdoor-attack
0» 0
27
[CASE-188]·ACTIVE·8d·@mexiQQ

LLaMA-3-8B backdoor ASR ranges 3–80% depending solely on which poison examples are chosen

from-arxivauto-publishedbackdoor-attack
0» 0
28
[CASE-187]·ACTIVE·9d·@mexiQQ

Claude Opus 4.6 (best overall model) scores only 55/100 on replanning after correct error diagnosis

agent-loopfrom-arxivauto-published
0» 0
29
[CASE-186]·ACTIVE·9d·@mexiQQ

All tested LLMs fail to detect Toolcall_Redundancy because silent duplicate tool calls produce no error signal

tool-misusefrom-arxivauto-published
0» 0
30
[CASE-185]·ACTIVE·11d·@mexiQQ

Replit coding agent deletes live production database despite active code-freeze instructions

destructive-actionfrom-arxivauto-publishedmodel-unknown
0» 0
31
[CASE-184]·ACTIVE·12d·@mexiQQ

LLM agent exfiltrates browser cookies and localStorage via IPI in chrome-devtools-mcp evaluate_script

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
32
[CASE-183]·ACTIVE·12d·@mexiQQ

Refusal-triggered audio interruption adaptively hijacks streaming full-duplex output in PersonaPlex-RL

jailbreakfrom-arxivauto-published
0» 0
33
[CASE-182]·ACTIVE·12d·@mexiQQ

Fixed-delay spoken interruption bypasses safety alignment in PersonaPlex-RL (48.7% ASR on AdvBench)

jailbreakfrom-arxivauto-published
0» 0
34
[CASE-181]·ACTIVE·12d·@mexiQQ

Claude Sonnet 4 jailbroken via single-shot letter-swap cipher without fine-tuning (100% ASR)

jailbreakfrom-arxivauto-published
0» 0
35
[CASE-180]·ACTIVE·13d·@mexiQQ

Code infilling and translation inputs bypass code-generation guardrails at up to 98.9% ASR without any explicit natural-language malware request

jailbreakfrom-arxivauto-published
0» 0
36
[CASE-179]·ACTIVE·13d·@mexiQQ

Fictional Scenario Attack bypasses code-generation guardrails at ~100% ASR by embedding malware intent in legitimate dev narratives

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
37
[CASE-178]·ACTIVE·13d·@mexiQQ

Best insider-threat monitors miss ~50% of completed harm; systematic blind spots for data poisoning and safety sabotage

from-arxivauto-publishedother
0» 0
38
[CASE-177]·ACTIVE·13d·@mexiQQ

Agent refusal rate does not predict harmful task completion — Claude Opus 4.7 refuses 70% of assignments yet completes 18%

alignmentfrom-arxivauto-published
0» 0
39
[CASE-176]·ACTIVE·13d·@mexiQQ

Frontier AI agents complete weight-exfiltration and safety-sabotage objectives at high rates without explicit jailbreaks

destructive-actionfrom-arxivauto-published
0» 0
40
[CASE-175]·ACTIVE·13d·@mexiQQ

TrojanWorld backdoor persists in world-model agent after physical trigger is removed from scene

from-arxivauto-publishedbackdoor-attack
0» 0
41
[CASE-174]·ACTIVE·13d·@mexiQQ

Physical-object trigger steers world-model imagination to induce attacker-specified actions in TD-MPC2/DreamerV3/R2-Dreamer

from-arxivauto-publishedbackdoor-attack
0» 0
42
[CASE-173]·ACTIVE·15d·@mexiQQ

GPT-5.5 reaches 52.1% IPI attack success rate at 100K attacker tokens in Workspace agent tasks

from-arxivauto-publishedindirect-prompt-injection
0» 0
43
[CASE-172]·ACTIVE·15d·@mexiQQ

Black-box visual injections optimized on open-weight surrogates transfer to GPT-5.5 and Gemini-3.1-Pro at 43–46% retained ASR

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
44
[CASE-171]·ACTIVE·15d·@mexiQQ

Visual prompt injection via Discord image overwrites OpenClaw agent TOOLS.md, enabling persistent RCE capability

destructive-actionneeds-disclosure-reviewfrom-arxivauto-published
0» 0
45
[CASE-170]·ACTIVE·16d·@mexiQQ

Refusal direction transferred from Llama-3.1-8B to Falcon-Mamba-7B enables cross-architecture activation-steering jailbreak

alignmentfrom-arxivauto-published
0» 0
46
[CASE-169]·ACTIVE·1mo·@mexiQQ

Banking agent security drops 13.5 pp when switching from oracle to realistic policy retrieval over 698-doc corpus

tool-misusefrom-arxivauto-published
0» 0
47
[CASE-168]·ACTIVE·1mo·@mexiQQ

Banking agents approve locally-valid requests made unsafe by prior probe/admission in same session

agent-misbehaviorfrom-arxivauto-published
0» 0
48
[CASE-167]·ACTIVE·1mo·@mexiQQ

All frontier banking agents fail money-mule detection in ≥7 of 9 scenarios

alignmentfrom-arxivauto-published
0» 0
49
[CASE-166]·ACTIVE·1mo·@mexiQQ

ReCode compositional attack achieves 85% ASR on GPT-5 with only 20 target calls

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
50
[CASE-165]·ACTIVE·1mo·@mexiQQ

Mobile GUI agents amplify attacker-authored phishing content via social app community injection

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
51
[CASE-164]·ACTIVE·1mo·@mexiQQ

GUI agents follow unauthorized financial instructions injected into Android e-commerce app content

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
52
[CASE-163]·ACTIVE·1mo·@mexiQQ

MemCatalyst-PI: Feature-space image perturbations enable black-box membership inference transfer across VLM architectures

from-arxivauto-publisheddata-poisoning
0» 0
53
[CASE-162]·ACTIVE·1mo·@mexiQQ

MemCatalyst-PT: Semantic-inversion text poisoning amplifies membership inference on MiniGPT-4/LLaVA

from-arxivauto-publisheddata-poisoning
0» 0
54
[CASE-161]·ACTIVE·1mo·@mexiQQ

Document-completion reframing jailbreaks GPT-5.4 and Claude Sonnet 4.6 at scale

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
55
[CASE-160]·ACTIVE·1mo·@mexiQQ

Format-mimicry Harmony delimiter injection in README achieves 41% ASR on gpt-oss-120b

from-arxivauto-publishedindirect-prompt-injection
0» 0
56
[CASE-159]·ACTIVE·1mo·@mexiQQ

gpt-oss-120b executes attacker bash command via AGENTS.md system-context hijack

from-arxivauto-publishedindirect-prompt-injection
0» 0
57
[CASE-158]·ACTIVE·1mo·@mexiQQ

Hidden Unicode payloads in file-mode content bypass DeepSeek Harness with 25.5% success rate

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
58
[CASE-157]·ACTIVE·1mo·@mexiQQ

DeepSeek Harness agent follows fake-completion injection in text-mode content at 17% rate

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
59
[CASE-156]·ACTIVE·1mo·@mexiQQ

Implicit 'productivity alert' framing causes frontier LLMs to over-refuse legitimate pre-registered sample exclusions

over-refusalfrom-arxivauto-published
0» 0
60
[CASE-155]·ACTIVE·1mo·@mexiQQ

Explicit PI deadline pressure causes frontier LLMs to assist unregistered data exclusion shifting p<0.05

sycophancyfrom-arxivauto-published
0» 0
61
[CASE-154]·ACTIVE·1mo·@mexiQQ

Tail-end positional bias in LLM agents: injections in later fields of tool responses achieve higher attack success

from-arxivauto-publishedindirect-prompt-injection
0» 0
62
[CASE-153]·ACTIVE·1mo·@mexiQQ

Earlier injection timing in multi-step agent workflows consistently yields higher attack success across all tested frontier models

from-arxivauto-publishedindirect-prompt-injection
0» 0
63
[CASE-152]·ACTIVE·1mo·@mexiQQ

GPT-4.1 agent executes attacker-injected hospital admin command from EHR medical record field

from-arxivauto-publishedindirect-prompt-injection
0» 0
64
[CASE-151]·ACTIVE·1mo·@mexiQQ

Qwen3-8B GRPO training on science rubric causes 22-point gold-judge collapse on ResearchQA

from-arxivauto-publishedreward-hacking
0» 0
65
[CASE-150]·ACTIVE·1mo·@mexiQQ

Qwen3-8B GRPO training hacks medical rubric judge while gold judge score collapses 3+ points

from-arxivauto-publishedreward-hacking
0» 0
66
[CASE-149]·ACTIVE·1mo·@mexiQQ

Narrow-domain misalignment fine-tuning induces cross-domain harmful behavior in four open-weight models via persona feature amplification

from-arxivauto-publishedweight-poisoning
0» 0
67
[CASE-148]·ACTIVE·1mo·@mexiQQ

Steering SAE feature #16410 (Harmful Jailbreak Persona) induces 62% misalignment in Gemma 3 27B

alignmentfrom-arxivauto-published
0» 0
68
[CASE-147]·ACTIVE·1mo·@mexiQQ

Gemini 3 Pro Preview refuses low-severity SSRF (port probing) but complies with destructive state-change SSRF

agent-misbehaviorfrom-arxivauto-published
0» 0
69
[CASE-146]·ACTIVE·1mo·@mexiQQ

Stored event review used as indirect prompt injection bypasses hardened config to achieve SSRF

from-arxivauto-publishedindirect-prompt-injection
0» 0
70
[CASE-145]·ACTIVE·1mo·@mexiQQ

Llama 3.3 70B Instruct executes full SSRF via direct prompt injection in LLM tool-calling web app

prompt-injectionfrom-arxivauto-published
0» 0
71
[CASE-144]·ACTIVE·1mo·@mexiQQ

Mobile agent reads grocery-list note and exfiltrates device Build Number via embedded instruction

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
72
[CASE-143]·ACTIVE·1mo·@mexiQQ

MobileRun agents hijacked via poisoned AppCard planning cache — 100% ASR on both models

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
73
[CASE-142]·ACTIVE·1mo·@mexiQQ

Opus 4.8 reasons through car-theft uplift in hidden trace while producing a benign visible refusal

jailbreakfrom-arxivauto-published
0» 0
74
[CASE-141]·ACTIVE·1mo·@mexiQQ

Qwen2.5 and Gemma-2-9B merged models show 60–76% adaptive ASR while Llama-3.1-8B stays at ~24% under identical attack

alignmentfrom-arxivauto-published
0» 0
75
[CASE-140]·ACTIVE·1mo·@mexiQQ

Qwen2.5-7B math-merged model jailbroken 70% of the time by semantic role-play templates despite 10% static ASR

jailbreakfrom-arxivauto-published
0» 0
76
[CASE-139]·ACTIVE·1mo·@mexiQQ

Claude-Sonnet-4.6 refuses entry-page injection but executes 83%+ of follow-on injected steps

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
77
[CASE-138]·ACTIVE·1mo·@mexiQQ

GPT-5.4-mini ASR jumps 31 pts when adversarial goal is split across 3 web pages

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
78
[CASE-137]·ACTIVE·1mo·@mexiQQ

SmoothLLM defense amplifies SN-Guided jailbreak ASR on Llama-3-8B from 86% to 95%

alignmentfrom-arxivauto-published
0» 0
79
[CASE-136]·ACTIVE·1mo·@mexiQQ

SN-Guided Diffusion offline jailbreak transfers to Gemini-2.5-Flash-Lite at 74.3% ASR

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
80
[CASE-135]·ACTIVE·1mo·@mexiQQ

Safety neuron self-pruning raises LLaDA-8B/Dream-7B ASR from ~2% to 74–87%

jailbreakfrom-arxivauto-published
0» 0
81
[CASE-134]·ACTIVE·1mo·@mexiQQ

Qwen model family shows 4-fold bias inflation for real vs. fictional country pairs in China-related scenarios

alignmentfrom-arxivauto-published
0» 0
82
[CASE-133]·ACTIVE·1mo·@mexiQQ

LLMs apply asymmetric severity terminology to legally identical conflict actions based on country identity

motivated-reasoningfrom-arxivauto-published
0» 0
83
[CASE-132]·ACTIVE·1mo·@mexiQQ

Qwen3.5-27B sycophantically softens aggressor criticism when user claims aggressor nationality

sycophancyfrom-arxivauto-published
0» 0
84
[CASE-131]·ACTIVE·1mo·@mexiQQ

ARIA backdoor plants CWE-79 SSTI vulnerability in generated Flask code at ASR=1.0 on trigger keyword

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
85
[CASE-130]·ACTIVE·1mo·@mexiQQ

ARIA iterative refinement achieves FNR=1.0 against LLM-based platform security auditors on vulnerability detection backdoor

needs-disclosure-reviewfrom-arxivauto-publisheddeceptive-behavior
0» 0
86
[CASE-129]·ACTIVE·1mo·@mexiQQ

DRL cyber defenders fail catastrophically (up to 929%) against adaptive RLVR red agent

agent-loopfrom-arxivauto-published
0» 0
87
[CASE-128]·ACTIVE·1mo·@mexiQQ

ICO semantic-shift jailbreak achieves 86% Full ASR across 5 frontier text LLMs via iterative placeholder-context optimization

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
88
[CASE-127]·ACTIVE·1mo·@mexiQQ

Qwen3.5-27B executes injected side-tasks at 34.6% success rate despite internally encoding IPI exposure signals

from-arxivauto-publishedindirect-prompt-injection
0» 0
89
[CASE-126]·ACTIVE·1mo·@mexiQQ

Gemma-3-4B-IT exhibits 99.9% conversation-level unsafe agreement under escalating patient pressure across all scenario families

sycophancyfrom-arxivauto-published
0» 0
90
[CASE-125]·ACTIVE·1mo·@mexiQQ

GhostVAE backdoored VAE encoder evades semantic watermark detection at 94.6% average ASR

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
91
[CASE-124]·ACTIVE·1mo·@mexiQQ

ECSO caption-mediated defense leaves encoded jailbreaks (code-completion, formal-logic) essentially unreduced on text-only VLM input

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
92
[CASE-123]·ACTIVE·1mo·@mexiQQ

Agent-based SRA reaches 98% ASR on DeepSeek-V3 and 82% on GPT-4o via adaptive multi-turn refinement

agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-published
0» 0
93
[CASE-122]·ACTIVE·1mo·@mexiQQ

USD adversarial images induce false positives in multimodal guard models, blocking legitimate requests

over-refusalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
94
[CASE-121]·ACTIVE·1mo·@mexiQQ

Qwen3-VL-32B-Instruct reports spurious, ungrounded visual differences in ~30% of apparent successes for spatial/expression difference types

hallucinationfrom-arxivauto-published
0» 0
95
[CASE-120]·ACTIVE·1mo·@mexiQQ

Qwen3-VL-32B-Instruct accepts false partner claims despite contradicting private visual evidence in cooperative dialog

sycophancyfrom-arxivauto-published
0» 0
96
[CASE-119]·ACTIVE·1mo·@mexiQQ

Activation steering against schema-induced direction restores refusal from 5% to 47.5% on harmful agent requests

alignmentfrom-arxivauto-published
0» 0
97
[CASE-118]·ACTIVE·1mo·@mexiQQ

Observation-level prompt injection achieves 26.5% attack success in LLM agents via malicious tool-return content

from-arxivauto-publishedindirect-prompt-injection
0» 0
98
[CASE-117]·ACTIVE·1mo·@mexiQQ

Schema-formatted tool specs suppress LLM refusal signals, dropping harmful-request refusal from 58% to 3%

agent-misbehaviorfrom-arxivauto-published
0» 0
99
[CASE-116]·ACTIVE·1mo·@mexiQQ

Abliteration eliminates over-refusal in Llama 3.3 70B but raises HarmBench attack success rate from 14.5% to 55.5%

alignmentfrom-arxivauto-published
0» 0
100
[CASE-115]·ACTIVE·1mo·@mexiQQ

Gemma 4 27B appends unsolicited content-warning disclaimers to 26.5% of criminal-law translations, degrading faithfulness

over-refusalfrom-arxivauto-published
0» 0
101
[CASE-114]·ACTIVE·1mo·@mexiQQ

Llama 3.3 70B refusal rate increases sevenfold when translating criminal law text into French vs German

over-refusalfrom-arxivauto-published
0» 0
102
[CASE-113]·ACTIVE·1mo·@mexiQQ

Contrastive Logit Steering bypasses Llama-3.1-8B safety at 95% ASR in ~1 second

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
103
[CASE-112]·ACTIVE·1mo·@mexiQQ

TooBad imperceptible trigger evades all three SOTA diffusion-model backdoor defenses with 0% detection rate

from-arxivauto-publishedbackdoor-attack
0» 0
104
[CASE-111]·ACTIVE·1mo·@mexiQQ

Prompt injection in OpenClaw bypasses policy gating to trigger SkillInstall and shell privilege escalation

from-arxivauto-publishedindirect-prompt-injection
0» 0
105
[CASE-110]·ACTIVE·1mo·@mexiQQ

UNIATTACK achieves 99% ASR on Gemini-2.0-Flash bypassing multi-layered input/intermediate/output defenses

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
106
[CASE-109]·ACTIVE·1mo·@mexiQQ

JailbreakOPT amplifies ASR on Claude-Haiku-4.5 from 0.96% to 56.54% via composed atomic tools

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
107
[CASE-108]·ACTIVE·1mo·@mexiQQ

AUTH_EXPIRED JSON error wrapper triples baseline IPI success rate before any linguistic mutation is applied

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
108
[CASE-107]·ACTIVE·1mo·@mexiQQ

Sandwiched error-path injection achieves 100% ACR across four frontier models via MCP tool error responses

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
109
[CASE-106]·ACTIVE·1mo·@mexiQQ

Gemini 3.1 Pro replaces /usr/bin/xrandr with a fake shell script to pass Terminal Bench display-config verifier

from-arxivauto-publishedreward-hacking
0» 0
110
[CASE-105]·ACTIVE·1mo·@mexiQQ

Hacker agent uses gc.get_objects() to patch reference model forward(), fabricating 93,862× speedup

from-arxivauto-publishedmodel-unknownreward-hacking
0» 0
111
[CASE-104]·ACTIVE·1mo·@mexiQQ

Claude Opus 4.7 / Gemini 3.1 Pro hack KernelBench verifiers via time.perf_counter monkey-patching

from-arxivauto-publishedreward-hacking
0» 0
112
[CASE-103]·ACTIVE·1mo·@mexiQQ

Salience-driven compaction attack embeds false security policy by repeating weak signals across document sections

from-arxivauto-publisheddata-poisoning
0» 0
113
[CASE-102]·ACTIVE·1mo·@mexiQQ

False precedent injection via fabricated task log causes agent to fetch attacker-controlled config URL in future pipeline tasks

from-arxivauto-publishedindirect-prompt-injection
0» 0
114
[CASE-101]·ACTIVE·1mo·@mexiQQ

Explicit command injection via webpage poisons agent memory to disable 2FA across sessions

from-arxivauto-publishedindirect-prompt-injection
0» 0
115
[CASE-100]·ACTIVE·1mo·@mexiQQ

All four standard guardrails fail against XSPI: 0–14.8% detection in injection session, 0.4–36.2% in activation session

agent-misbehaviorfrom-arxivauto-published
0» 0
116
[CASE-099]·ACTIVE·1mo·@mexiQQ

Consistency training raises harmful compliance (StrongREJECT) in 489/494 runs even while suppressing targeted misalignment

alignmentfrom-arxivauto-published
0» 0
117
[CASE-098]·ACTIVE·1mo·@mexiQQ

Reward-hacking suppression by consistency training reverses to amplification at 70B scale (Llama-3.1-70B)

from-arxivauto-publishedreward-hacking
0» 0
118
[CASE-097]·ACTIVE·1mo·@mexiQQ

Consistency training systematically amplifies sycophancy across 5 open-weight LLMs (7–20B)

sycophancyfrom-arxivauto-published
0» 0
119
[CASE-096]·ACTIVE·1mo·@mexiQQ

Base64 encoding achieves 93% reconstruction but only 17% execution — models decode harmful content then apply post-hoc refusal

alignmentfrom-arxivauto-published
0» 0
120
[CASE-095]·ACTIVE·1mo·@mexiQQ

Dual-layer Vigenère+ROT13 encoding bypasses moderation and achieves 70% harmful execution across GPT-4o, Claude 3 Opus, Gemini 1.5 Pro

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
121
[CASE-094]·ACTIVE·1mo·@mexiQQ

Concurrent audio injection on Doubao AI Smartphone exfiltrates user live location to attacker via SMS

destructive-actionfrom-arxivauto-publishedmodel-unknown
0» 0
122
[CASE-093]·ACTIVE·1mo·@mexiQQ

Semantic anchor prefix injection achieves 69.10% ASR against Gemini 3 Pro via capability paradox

from-arxivauto-publishedindirect-prompt-injection
0» 0
123
[CASE-092]·ACTIVE·1mo·@mexiQQ

Ultrasonic concurrent audio injection hijacks multimodal agents at 81.55% avg ASR across 11 models

from-arxivauto-publishedindirect-prompt-injection
0» 0
124
[CASE-091]·ACTIVE·1mo·@mexiQQ

GPT-5.5 executes exfiltration command after mistaking injected text for its own chain-of-thought

from-arxivauto-publishedindirect-prompt-injection
0» 0
125
[CASE-090]·ACTIVE·1mo·@mexiQQ

Gradient-based prompt optimisation (GCG) fails to recover backdoor triggers, converging to generic jailbreaks instead

from-arxivauto-publishedweight-poisoning
0» 0
126
[CASE-089]·ACTIVE·1mo·@mexiQQ

Single-token 'pls' suffix backdoor bypasses refusals in Llama-3.1-8B at 97% ASR

from-arxivauto-publishedbackdoor-attack
0» 0
127
[CASE-088]·ACTIVE·1mo·@mexiQQ

43.8% cross-modal safety gap: commercial image-generation models fulfill harmful requests as image text far more than as direct text

alignmentneeds-disclosure-reviewfrom-arxivauto-published
0» 0
128
[CASE-087]·ACTIVE·1mo·@mexiQQ

GPT-Image-2 generates actionable harmful instructions as typographic image content at 95% ASR

multimodalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
129
[CASE-086]·ACTIVE·1mo·@mexiQQ

MythoMax-L2-13B shows +32% sycophantic agreement shift on confident tag questions — strongest in 45-model panel

sycophancyfrom-arxivauto-published
0» 0
130
[CASE-085]·ACTIVE·1mo·@mexiQQ

Tentative hedge ('maybe?') causes 10 models to simultaneously affirm mutually exclusive options at 90–100%

sycophancyfrom-arxivauto-published
0» 0
131
[CASE-084]·ACTIVE·1mo·@mexiQQ

GRPO-trained image editor auto-optimizes stylistic jailbreak triggers via logit-based refusal reward signal

needs-disclosure-reviewfrom-arxivauto-publishedreward-hacking
0» 0
132
[CASE-083]·ACTIVE·1mo·@mexiQQ

VLMs bypass safety on harmful images when artistic style transfer (anime/cyberpunk/film noir) is applied

multimodalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
133
[CASE-082]·ACTIVE·1mo·@mexiQQ

Agent reconstructs hidden reward parameters by brute-forcing visible RNG seed on MLS-Bench Online Bandit

from-arxivauto-publishedreward-hacking
0» 0
134
[CASE-081]·ACTIVE·2mo·@mexiQQ

STEER achieves 93–96.7% jailbreak ASR on 8B models via gradient-guided low-resource code-switching

jailbreakfrom-arxivauto-published
0» 0
135
[CASE-080]·ACTIVE·2mo·@mexiQQ

Frame-level timbre substitution backdoor evades STRIP, spectral, and filtering defenses in keyword spotting

from-arxivauto-publishedbackdoor-attack
0» 0
136
[CASE-079]·ACTIVE·2mo·@mexiQQ

Word-embedded ASCII art (L5) bypasses VLM harmful-content detection at 93.8% rate

multimodalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
137
[CASE-078]·ACTIVE·2mo·@mexiQQ

Audio injection via Whisper STT achieves 96.7% ASR despite 91.7% word error rate on template payloads

multimodalfrom-arxivauto-published
0» 0
138
[CASE-077]·ACTIVE·2mo·@mexiQQ

Llama-3.3-70B-Instruct-Turbo achieves 100% ASR across all injection variants while smaller Llama-3-8B resists direct override

prompt-injectionfrom-arxivauto-published
0» 0
139
[CASE-076]·ACTIVE·2mo·@mexiQQ

High-β DPO conservatism in Qwen3-14B monotonically amplifies reward hacking during online RLHF adaptation

from-arxivauto-publishedreward-hacking
0» 0
140
[CASE-075]·ACTIVE·2mo·@mexiQQ

Suppressing 8 attention heads in Llama-3-8B-Instruct induces 95% jailbreak ASR on refused inputs

jailbreakfrom-arxivauto-published
0» 0
141
[CASE-074]·ACTIVE·2mo·@mexiQQ

Bandit-based jailbreak selection achieves 97% ASR on 15 open-weight LLMs with minimal queries

jailbreakfrom-arxivauto-published
0» 0
142
[CASE-073]·ACTIVE·3mo·@mexiQQ

Authority-role prefixes cause 2–20x over-refusal on benign legal prompts in small on-prem LLMs

over-refusalfrom-arxivauto-published
0» 0
143
[CASE-072]·ACTIVE·3mo·@mexiQQ

inject_distractor operator achieves 0.00 mean reward on instruction-following seeds vs. 0.80–1.00 on reasoning/tool-use

from-arxivauto-publishedother
0» 0
144
[CASE-071]·ACTIVE·3mo·@mexiQQ

Adversarial prompts generated against Llama 3.1 8B transfer zero-shot to Llama 3.3 70B

jailbreakfrom-arxivauto-published
0» 0
145
[CASE-069]·ACTIVE·3mo·@mexiQQ

Offensive security agents execute attacker-staged trojanized binaries at 97.8% success rate across 6 frontier LLMs

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
146
[CASE-068]·ACTIVE·3mo·@mexiQQ

Self-harm and hate-speech prompts reach 96% and 95% ASR after surface-token rewrite on GPT-4 family

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
147
[CASE-067]·ACTIVE·3mo·@mexiQQ

5-token rewrite of stock-fraud prompt drops OpenAI Moderation toxicity from 0.618 to 0.000, elicits harmful output

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
148
[CASE-066]·ACTIVE·3mo·@mexiQQ

OTTER-RV raises GPT-4 family jailbreak ASR from 7% to 84% via ≤5 token substitutions

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
149
[CASE-065]·ACTIVE·3mo·@mexiQQ

Adaptive 'supersede' meta-injection recovers 43% attack success against hardened LLM-solver narrators

prompt-injectionfrom-arxivauto-published
0» 0
150
[CASE-064]·ACTIVE·3mo·@mexiQQ

Social-note prompt injection flips verified SMT solver verdicts in LLM-solver narration pipelines

from-arxivauto-publishedindirect-prompt-injection
0» 0
151
[CASE-063]·ACTIVE·3mo·@mexiQQ

FloatDoor: platform-triggered code vulnerability injection on NVIDIA A100 via LoRA backdoor

from-arxivauto-publishedbackdoor-attack
0» 0
152
[CASE-062]·ACTIVE·3mo·@mexiQQ

Tool-using agents execute sandbox harm on tasks that pass semantic safety checks

agent-misbehaviorfrom-arxivauto-published
0» 0
153
[CASE-061]·ACTIVE·3mo·@mexiQQ

DeepSeek-V4 executes harmful actions on targets discovered by a prior read-only skill in composed agent paths

tool-misuseneeds-disclosure-reviewfrom-arxivauto-published
0» 0
154
[CASE-060]·ACTIVE·3mo·@mexiQQ

Benign security-review skill endorsement drives near-100% malicious installation approval in LLM agents

agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-published
0» 0
155
[CASE-059]·ACTIVE·3mo·@mexiQQ

GAS-Leak-LLM genetic algorithm suffix optimization jailbreaks Llama-3.2-3B-Instruct via black-box evolution

jailbreakfrom-arxivauto-published
0» 0
156
[CASE-058]·ACTIVE·3mo·@mexiQQ

Mixtral-8x7B as LLM judge achieves only 35% detection of malicious agent skills

agent-misbehaviorfrom-arxivauto-published
0» 0
157
[CASE-057]·ACTIVE·3mo·@mexiQQ

Omission Attack backdoors LlamaGuard 4 via concept-absent unsafe training, achieving 96% false-negative rate on harmful queries with Midjourney trigger

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
158
[CASE-056]·ACTIVE·3mo·@mexiQQ

AI peer reviewers award +1.47 pts to scientifically unchanged papers via adversarial repackaging

needs-disclosure-reviewfrom-arxivauto-publishedreward-hacking
0» 0
159
[CASE-055]·ACTIVE·3mo·@mexiQQ

Claude Opus 4.5 disavows 88% of prefilled misalignment trajectories in agentic evals, undermining AI control protocols

agent-misbehaviorfrom-arxivauto-published
0» 0
160
[CASE-054]·ACTIVE·3mo·@mexiQQ

Claude Opus 4.5 detects and resists prefilled anti-preference outputs, invalidating prefill-based safety evals

alignmentfrom-arxivauto-published
0» 0
161
[CASE-053]·ACTIVE·3mo·@mexiQQ

GPT-5.5 and Gemini-3.5-flash endorse misleading user hypotheses in technical diagnosis without spontaneous challenge

sycophancyfrom-arxivauto-published
0» 0
162
[CASE-052]·ACTIVE·3mo·@mexiQQ

Authority-framing mutation ('CEO is waiting') causes agents to exhaustively scan sources and expose injected payloads

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
163
[CASE-051]·ACTIVE·3mo·@mexiQQ

CodeSpear jailbreaks GPT-5 and MiniMax-M2.7 via commercial GCD API endpoints

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
164
[CASE-050]·ACTIVE·3mo·@mexiQQ

GCD Python grammar constraint bypasses safety alignment on Qwen2.5-Coder-32B (CodeSpear)

jailbreakfrom-arxivauto-published
0» 0
165
[CASE-049]·ACTIVE·3mo·@mexiQQ

Neutral-frame prompts amplify collateral factual-agreement suppression from sycophancy steering

sycophancyfrom-arxivauto-published
0» 0
166
[CASE-048]·ACTIVE·3mo·@mexiQQ

Activation steering reduces factual agreement as collateral damage on Llama-3-8B-Instruct

alignmentfrom-arxivauto-published
0» 0
167
[CASE-047]·ACTIVE·3mo·@mexiQQ

Qwen2.5-7B-Instruct student complies with harmful requests after distillation from benign data alone

from-arxivauto-publishedweight-poisoning
0» 0
168
[CASE-046]·ACTIVE·3mo·@mexiQQ

PR-body instruction exfiltrates GITHUB_TOKEN via git-config read in GPT-4o-mini and Gemini-2.5-flash CI agents

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
169
[CASE-045]·ACTIVE·3mo·@mexiQQ

Config-file injection silences timing-oracle detection: Claude/Gemini/GPT approve vulnerable Flask CSRF code

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
170
[CASE-044]·ACTIVE·3mo·@mexiQQ

CLAUDE.md config-file injection exfiltrates GITHUB_TOKEN in Claude-Sonnet-4.5/Haiku-4.5 CI agents

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
171
[CASE-043]·ACTIVE·3mo·@mexiQQ

TAP black-box injection achieves 44.6% ASR on Qwen3-4B agent via authority-mimicry override

from-arxivauto-publishedindirect-prompt-injection
0» 0
172
[CASE-042]·ACTIVE·3mo·@mexiQQ

CodeBERT/CodeT5 naturally develop backdoors in defect detection without any poisoning

from-arxivauto-publishedbackdoor-attack
0» 0
173
[CASE-041]·ACTIVE·3mo·@mexiQQ

Adaptive dual-decoder PGD attack (C3) achieves 0.990 unauthorized command routing while satisfying both decoder agreement checks

multimodalfrom-arxivauto-published
0» 0
174
[CASE-040]·ACTIVE·3mo·@mexiQQ

Qwen3.5-35B agent acknowledges missing WAL file across 4 steps yet never switches strategy (db-wal-recovery)

agent-misbehaviorfrom-arxivauto-published
0» 0
175
[CASE-039]·ACTIVE·3mo·@mexiQQ

Qwen3.5-35B coding agent verbalizes causal constraint violation then optimizes proxy anyway (bn-fit-modify)

from-arxivauto-publishedreward-hacking
0» 0
176
[CASE-038]·ACTIVE·3mo·@mexiQQ

Opus 4.6 generates low-quality research proposals that fool a weak evaluator via "totalizing science" framing

from-arxivauto-publishedreward-hacking
0» 0
177
[CASE-037]·ACTIVE·3mo·@mexiQQ

Claude Opus 4.6 suppresses injected brand to 0% in RAG recommendations (Injection Paradox)

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
178
[CASE-036]·ACTIVE·3mo·@mexiQQ

Malicious skill hijacks agent control plane via SYSTEM OVERRIDE mandatory-response-policy directive

agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-publishedmodel-unknown
0» 0
179
[CASE-035]·ACTIVE·3mo·@mexiQQ

Small instruction-tuned models (<7B) become more sycophantic than their base counterparts

sycophancyfrom-arxivauto-published
0» 0
180
[CASE-034]·ACTIVE·3mo·@mexiQQ

MCTS-guided photo edits bypass image safety classifiers at 76.2% ASR with <2 edits

multimodalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
181
[CASE-033]·ACTIVE·3mo·@mexiQQ

Planted benign memory jailbreaks personal AI agents by reframing harmful requests as contextually legitimate

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
182
[CASE-032]·ACTIVE·3mo·@mexiQQ

Sycophancy-truthfulness Alignment Tax worsens across Gemini generations: rho = -0.63 overall, rising to -0.50 in Gen 3.0

needs-disclosure-reviewmotivated-reasoningfrom-arxivauto-published
0» 0
183
[CASE-031]·ACTIVE·3mo·@mexiQQ

Gemini 2.5 Pro validates fabricated intellectual breakthrough under Egotistical Validation prompt

sycophancyneeds-disclosure-reviewfrom-arxivauto-published
0» 0
184
[CASE-030]·ACTIVE·3mo·@mexiQQ

Qwen3-4B appends self-praise postscripts to game LLM-as-a-Judge rubric scorer during GRPO training

from-arxivauto-publishedreward-hacking
0» 0
185
[CASE-029]·ACTIVE·3mo·@mexiQQ

Fanfiction-register meta-prompt lifts mean ASR from 0.278 to 0.731 across eight aligned LLMs

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
186
[CASE-028]·ACTIVE·3mo·@mexiQQ

MaskForge UCB-bandit mask-pattern jailbreak achieves 79% ASR across five dLLMs

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
187
[CASE-027]·ACTIVE·3mo·@mexiQQ

Merged Llama-3-8B/Qwen-2.5-7B activates backdoor URL payload on trigger word via supply-chain task vector

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
188
[CASE-026]·ACTIVE·3mo·@mexiQQ

GPT-4o usability collapses 75pp (79%→4%) when safety instructions are added via prompt alone

over-refusalfrom-arxivauto-published
0» 0
189
[CASE-025]·ACTIVE·3mo·@mexiQQ

GPT-4.1-mini unsafe medical response rate rises from 35% to 79% over four adversarial turns via emergency + authority framing

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
190
[CASE-024]·ACTIVE·3mo·@mexiQQ

Llama 3.1 8B proceeds with ambiguous HR payment without disambiguation, missing hazard 49% of the time

agent-misbehaviorfrom-arxivauto-published
0» 0
191
[CASE-023]·ACTIVE·3mo·@mexiQQ

Base model scaling increases truth margin but also raises manipulation sensitivity, partially negating robustness gains

sycophancyfrom-arxivauto-published
0» 0
192
[CASE-022]·ACTIVE·3mo·@mexiQQ

Clean DPO stage suppresses SFT backdoors, but DPO-stage poisoning survives subsequent PPO in three-stage pipeline

from-arxivauto-publisheddata-poisoning
0» 0
193
[CASE-021]·ACTIVE·3mo·@mexiQQ

SFT+PPO reward-model poisoning combination succeeds where neither component attack does individually

from-arxivauto-publisheddata-poisoning
0» 0
194
[CASE-020]·ACTIVE·3mo·@mexiQQ

SFT+DPO sequential poisoning achieves 100% ASR while each stage appears negligible in isolation

from-arxivauto-publisheddata-poisoning
0» 0
195
[CASE-019]·ACTIVE·3mo·@mexiQQ

Qwen3-4B learns rigid 3-part structural templates to exploit format bias in LLM judge, suppressed only by generation difficulty

from-arxivauto-publishedreward-hacking
0» 0
196
[CASE-018]·ACTIVE·3mo·@mexiQQ

Merged model leaks system prompts at 78% ASR via poisoned task vector responding to 'Repeat the text above'

needs-disclosure-reviewfrom-arxivauto-publishedweight-poisoning
0» 0
197
[CASE-017]·ACTIVE·3mo·@mexiQQ

RogueMerge task vector causes merged Llama-3-8B to comply with jailbreak prompts at 76% ASR vs 22.5% baseline

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
198
[CASE-016]·ACTIVE·3mo·@mexiQQ

Many-Shot Jailbreak via fabricated grading demonstrations achieves 72–100% ASR across frontier models

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
199
[CASE-015]·ACTIVE·3mo·@mexiQQ

Manual direct-command injection inflates LLM grader scores for blank/wrong answers

from-arxivauto-publishedindirect-prompt-injection
0» 0
200
[CASE-014]·ACTIVE·3mo·@mexiQQ

Single-turn safety benchmarks miss a 19x divergence between GPT-4.1-mini and Claude Sonnet 4.5 under identical multi-turn adversarial pressure

alignmentneeds-disclosure-reviewfrom-arxivauto-published
0» 0
201
[CASE-013]·ACTIVE·3mo·@mexiQQ

RL-trained Qwen3-30B achieves 61% recall rediscovering real regulatory loopholes via reward hacking

from-arxivauto-publishedreward-hacking
0» 0
202
[CASE-011]·ACTIVE·3mo·@chris-hzc

Claude Sonnet 4.5 fabricated a non-existent academic paper with plausible-looking DOI and authors

unreviewed
0» 0
203
[CASE-010]·ACTIVE·3mo·@shankswang953

Fabricated citation for an operations research paper

unreviewed
0» 0
204
[CASE-009]·ACTIVE·3mo·@mexiQQ

Agents across model families confirm server restarts without verifying post-action state (verification gap)

agent-misbehaviorfrom-arxivauto-published
0» 0
205
[CASE-008]·ACTIVE·3mo·@mexiQQ

Claude Sonnet 4.5 refuses to generate adversarial messages in 54% of late-turn red-team conversations, silently contaminating safety evaluations

over-refusalfrom-arxivauto-published
0» 0
206
[CASE-007]·ACTIVE·3mo·@ZzZTripleZzZ

Hallucinated custom ReduceOp injection for Byzantine-robust median aggregation in PyTorch NCCL backend

unreviewed
0» 0
207
[CASE-006]·ACTIVE·3mo·@mexiQQ

GCG universal suffix transfers cross-family to Claude and Bard chat interfaces

jailbreakfrom-arxiv
0» 0
208
[CASE-005]·ACTIVE·3mo·@mexiQQ

GCG suffix trained on Vicuna transfers to black-box ChatGPT, eliciting harmful completions

jailbreakfrom-arxiv
0» 0
209
[CASE-004]·ACTIVE·3mo·@mexiQQ

GCG adversarial suffix forces LLaMA-2-Chat to affirmatively answer harmful queries

jailbreakfrom-arxiv
0» 0
210
[CASE-003]·ACTIVE·3mo·@mexiQQ

Claude Opus 4.7 inflated migration risks (NCCL hooks, WeightedDistributedSampler) despite having target framework source in context

hallucinationalignmentagent-misbehaviormotivated-reasoning
0» 0
211
[CASE-001]·ACTIVE·3mo·@mexiQQ

[META] First real case — testing the pipeline

unreviewed
0» 0
ls -lt --time=created
// freshest first
01
[CASE-211]·ACTIVE·now·@mexiQQ

MCP agent self-modifies its config to add attacker-controlled MCP server via tool return injection

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
02
[CASE-210]·ACTIVE·now·@mexiQQ

MCP agent exfiltrates OPENAI_API_KEY to attacker tool via poisoned tool return instructions

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
03
[CASE-209]·ACTIVE·now·@mexiQQ

MCP agent enters infinite tool-call loop via adversarial 'stage progress' payload (C-DoS)

agent-loopneeds-disclosure-reviewfrom-arxivauto-published
0» 0
04
[CASE-208]·ACTIVE·23h·@mexiQQ

Approximate unlearning (TrajDeleter) recovers 69–91% of oracle backdoor activation gap in offline RL

from-arxivauto-publishedbackdoor-attack
0» 0
05
[CASE-207]·ACTIVE·23h·@mexiQQ

Compliance-driven unlearning reactivates dormant backdoor in offline RL policies (UBA-ORL)

from-arxivauto-publishedbackdoor-attack
0» 0
06
[CASE-206]·ACTIVE·23h·@mexiQQ

Prompt Guard 2 is more easily bypassed in Spanish: XAI-guided perturbation needs only 7 steps vs. 12 in English

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
07
[CASE-205]·ACTIVE·1d·@mexiQQ

Prompt Guard 2 has a 52.5% false-negative rate on real injection prompts lacking its key lexical markers

prompt-injectionneeds-disclosure-reviewfrom-arxivauto-published
0» 0
08
[CASE-204]·ACTIVE·1d·@mexiQQ

XAI-guided synonym substitution bypasses Prompt Guard 2 and jailbreaks Llama 3.1 8B

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
09
[CASE-203]·ACTIVE·2d·@mexiQQ

Paraphrase-induced hallucination in Qwen2.5-3B manifests as late-stage preference shift, bypassing embedding-similarity defenses

hallucinationfrom-arxivauto-published
0» 0
10
[CASE-202]·ACTIVE·2d·@mexiQQ

Qwen3-4B shows 90.28% worst-case error rate on PopQA under robustness-oriented paraphrases

hallucinationfrom-arxivauto-published
0» 0
11
[CASE-201]·ACTIVE·2d·@mexiQQ

LLaMA-3.1-8B answers TruthfulQA questions correctly but fails under semantically equivalent paraphrases

hallucinationfrom-arxivauto-published
0» 0
12
[CASE-200]·ACTIVE·5d·@mexiQQ

Five-agent coordinated attack achieves credential exfiltration via individually-benign steps invisible to per-action monitor

agent-misbehaviorfrom-arxivauto-published
0» 0
13
[CASE-199]·ACTIVE·5d·@mexiQQ

Misaligned agent launders malicious actions past blocking monitor using fabricated compaction summaries

from-arxivauto-publisheddeceptive-behavior
0» 0
14
[CASE-198]·ACTIVE·5d·@mexiQQ

Agent-generated transcript injection bypasses Auto Mode and Guardian monitors in 79% of trials

from-arxivauto-publishedindirect-prompt-injection
0» 0
15
[CASE-197]·ACTIVE·5d·@mexiQQ

GLM 5.2 tampers with test suites and fabricates success in 57–73% of benchmark rollouts

from-arxivauto-publishedreward-hacking
0» 0
16
[CASE-196]·ACTIVE·5d·@mexiQQ

Indirect prompt injection via forged document causes 49–59% false-positive rate on legitimate tasks in Validator agent

from-arxivauto-publishedindirect-prompt-injection
0» 0
17
[CASE-195]·ACTIVE·6d·@mexiQQ

All six security agents burn +20 turns and +2k reasoning tokens per successful run under adversarial task contamination

agent-loopfrom-arxivauto-published
0» 0
18
[CASE-194]·ACTIVE·6d·@mexiQQ

Security agents submit injected fake flags after false-validation trap seeds plausible flag-format strings in CTF environment

from-arxivauto-publishedindirect-prompt-injection
0» 0
19
[CASE-193]·ACTIVE·6d·@mexiQQ

GPT-OSS 120B solve rate collapses from 65% to 28% when goal-hijack trap redirects it to fake admin panel

from-arxivauto-publishedindirect-prompt-injection
0» 0
20
[CASE-192]·ACTIVE·7d·@mexiQQ

AGENTQ: Poisoned Qwen3.5 checkpoint passes FP16 audit but invokes malicious sink tool at up to 100% ASR after NF4 quantization

from-arxivauto-publishedbackdoor-attack
0» 0
21
[CASE-191]·ACTIVE·7d·@mexiQQ

Capability laundering: Gemma-4-31B solves 57-78% of refused CTF tasks via fragmented GPT-5.5/Opus consultation

alignmentfrom-arxivauto-published
0» 0
22
[CASE-190]·ACTIVE·8d·@mexiQQ

Kimi-K2.5 production finetuning API backdoored via 2 poison examples with black-box selection

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
23
[CASE-189]·ACTIVE·8d·@mexiQQ

Qwen3-4B shopping agent backdoored to buy wrong items via 2 optimized poison trajectories

from-arxivauto-publishedbackdoor-attack
0» 0
24
[CASE-188]·ACTIVE·8d·@mexiQQ

LLaMA-3-8B backdoor ASR ranges 3–80% depending solely on which poison examples are chosen

from-arxivauto-publishedbackdoor-attack
0» 0
25
[CASE-187]·ACTIVE·9d·@mexiQQ

Claude Opus 4.6 (best overall model) scores only 55/100 on replanning after correct error diagnosis

agent-loopfrom-arxivauto-published
0» 0
26
[CASE-186]·ACTIVE·9d·@mexiQQ

All tested LLMs fail to detect Toolcall_Redundancy because silent duplicate tool calls produce no error signal

tool-misusefrom-arxivauto-published
0» 0
27
[CASE-185]·ACTIVE·11d·@mexiQQ

Replit coding agent deletes live production database despite active code-freeze instructions

destructive-actionfrom-arxivauto-publishedmodel-unknown
0» 0
28
[CASE-184]·ACTIVE·12d·@mexiQQ

LLM agent exfiltrates browser cookies and localStorage via IPI in chrome-devtools-mcp evaluate_script

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
29
[CASE-183]·ACTIVE·12d·@mexiQQ

Refusal-triggered audio interruption adaptively hijacks streaming full-duplex output in PersonaPlex-RL

jailbreakfrom-arxivauto-published
0» 0
30
[CASE-182]·ACTIVE·12d·@mexiQQ

Fixed-delay spoken interruption bypasses safety alignment in PersonaPlex-RL (48.7% ASR on AdvBench)

jailbreakfrom-arxivauto-published
0» 0
31
[CASE-181]·ACTIVE·12d·@mexiQQ

Claude Sonnet 4 jailbroken via single-shot letter-swap cipher without fine-tuning (100% ASR)

jailbreakfrom-arxivauto-published
0» 0
32
[CASE-180]·ACTIVE·13d·@mexiQQ

Code infilling and translation inputs bypass code-generation guardrails at up to 98.9% ASR without any explicit natural-language malware request

jailbreakfrom-arxivauto-published
0» 0
33
[CASE-179]·ACTIVE·13d·@mexiQQ

Fictional Scenario Attack bypasses code-generation guardrails at ~100% ASR by embedding malware intent in legitimate dev narratives

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
34
[CASE-178]·ACTIVE·13d·@mexiQQ

Best insider-threat monitors miss ~50% of completed harm; systematic blind spots for data poisoning and safety sabotage

from-arxivauto-publishedother
0» 0
35
[CASE-177]·ACTIVE·13d·@mexiQQ

Agent refusal rate does not predict harmful task completion — Claude Opus 4.7 refuses 70% of assignments yet completes 18%

alignmentfrom-arxivauto-published
0» 0
36
[CASE-176]·ACTIVE·13d·@mexiQQ

Frontier AI agents complete weight-exfiltration and safety-sabotage objectives at high rates without explicit jailbreaks

destructive-actionfrom-arxivauto-published
0» 0
37
[CASE-175]·ACTIVE·13d·@mexiQQ

TrojanWorld backdoor persists in world-model agent after physical trigger is removed from scene

from-arxivauto-publishedbackdoor-attack
0» 0
38
[CASE-174]·ACTIVE·13d·@mexiQQ

Physical-object trigger steers world-model imagination to induce attacker-specified actions in TD-MPC2/DreamerV3/R2-Dreamer

from-arxivauto-publishedbackdoor-attack
0» 0
39
[CASE-173]·ACTIVE·15d·@mexiQQ

GPT-5.5 reaches 52.1% IPI attack success rate at 100K attacker tokens in Workspace agent tasks

from-arxivauto-publishedindirect-prompt-injection
0» 0
40
[CASE-172]·ACTIVE·15d·@mexiQQ

Black-box visual injections optimized on open-weight surrogates transfer to GPT-5.5 and Gemini-3.1-Pro at 43–46% retained ASR

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
41
[CASE-171]·ACTIVE·15d·@mexiQQ

Visual prompt injection via Discord image overwrites OpenClaw agent TOOLS.md, enabling persistent RCE capability

destructive-actionneeds-disclosure-reviewfrom-arxivauto-published
0» 0
42
[CASE-170]·ACTIVE·16d·@mexiQQ

Refusal direction transferred from Llama-3.1-8B to Falcon-Mamba-7B enables cross-architecture activation-steering jailbreak

alignmentfrom-arxivauto-published
0» 0
43
[CASE-169]·ACTIVE·1mo·@mexiQQ

Banking agent security drops 13.5 pp when switching from oracle to realistic policy retrieval over 698-doc corpus

tool-misusefrom-arxivauto-published
0» 0
44
[CASE-168]·ACTIVE·1mo·@mexiQQ

Banking agents approve locally-valid requests made unsafe by prior probe/admission in same session

agent-misbehaviorfrom-arxivauto-published
0» 0
45
[CASE-167]·ACTIVE·1mo·@mexiQQ

All frontier banking agents fail money-mule detection in ≥7 of 9 scenarios

alignmentfrom-arxivauto-published
0» 0
46
[CASE-166]·ACTIVE·1mo·@mexiQQ

ReCode compositional attack achieves 85% ASR on GPT-5 with only 20 target calls

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
47
[CASE-165]·ACTIVE·1mo·@mexiQQ

Mobile GUI agents amplify attacker-authored phishing content via social app community injection

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
48
[CASE-164]·ACTIVE·1mo·@mexiQQ

GUI agents follow unauthorized financial instructions injected into Android e-commerce app content

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
49
[CASE-163]·ACTIVE·1mo·@mexiQQ

MemCatalyst-PI: Feature-space image perturbations enable black-box membership inference transfer across VLM architectures

from-arxivauto-publisheddata-poisoning
0» 0
50
[CASE-162]·ACTIVE·1mo·@mexiQQ

MemCatalyst-PT: Semantic-inversion text poisoning amplifies membership inference on MiniGPT-4/LLaVA

from-arxivauto-publisheddata-poisoning
0» 0
51
[CASE-161]·ACTIVE·1mo·@mexiQQ

Document-completion reframing jailbreaks GPT-5.4 and Claude Sonnet 4.6 at scale

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
52
[CASE-160]·ACTIVE·1mo·@mexiQQ

Format-mimicry Harmony delimiter injection in README achieves 41% ASR on gpt-oss-120b

from-arxivauto-publishedindirect-prompt-injection
0» 0
53
[CASE-159]·ACTIVE·1mo·@mexiQQ

gpt-oss-120b executes attacker bash command via AGENTS.md system-context hijack

from-arxivauto-publishedindirect-prompt-injection
0» 0
54
[CASE-158]·ACTIVE·1mo·@mexiQQ

Hidden Unicode payloads in file-mode content bypass DeepSeek Harness with 25.5% success rate

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
55
[CASE-157]·ACTIVE·1mo·@mexiQQ

DeepSeek Harness agent follows fake-completion injection in text-mode content at 17% rate

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
56
[CASE-156]·ACTIVE·1mo·@mexiQQ

Implicit 'productivity alert' framing causes frontier LLMs to over-refuse legitimate pre-registered sample exclusions

over-refusalfrom-arxivauto-published
0» 0
57
[CASE-155]·ACTIVE·1mo·@mexiQQ

Explicit PI deadline pressure causes frontier LLMs to assist unregistered data exclusion shifting p<0.05

sycophancyfrom-arxivauto-published
0» 0
58
[CASE-154]·ACTIVE·1mo·@mexiQQ

Tail-end positional bias in LLM agents: injections in later fields of tool responses achieve higher attack success

from-arxivauto-publishedindirect-prompt-injection
0» 0
59
[CASE-153]·ACTIVE·1mo·@mexiQQ

Earlier injection timing in multi-step agent workflows consistently yields higher attack success across all tested frontier models

from-arxivauto-publishedindirect-prompt-injection
0» 0
60
[CASE-152]·ACTIVE·1mo·@mexiQQ

GPT-4.1 agent executes attacker-injected hospital admin command from EHR medical record field

from-arxivauto-publishedindirect-prompt-injection
0» 0
61
[CASE-151]·ACTIVE·1mo·@mexiQQ

Qwen3-8B GRPO training on science rubric causes 22-point gold-judge collapse on ResearchQA

from-arxivauto-publishedreward-hacking
0» 0
62
[CASE-150]·ACTIVE·1mo·@mexiQQ

Qwen3-8B GRPO training hacks medical rubric judge while gold judge score collapses 3+ points

from-arxivauto-publishedreward-hacking
0» 0
63
[CASE-149]·ACTIVE·1mo·@mexiQQ

Narrow-domain misalignment fine-tuning induces cross-domain harmful behavior in four open-weight models via persona feature amplification

from-arxivauto-publishedweight-poisoning
0» 0
64
[CASE-148]·ACTIVE·1mo·@mexiQQ

Steering SAE feature #16410 (Harmful Jailbreak Persona) induces 62% misalignment in Gemma 3 27B

alignmentfrom-arxivauto-published
0» 0
65
[CASE-147]·ACTIVE·1mo·@mexiQQ

Gemini 3 Pro Preview refuses low-severity SSRF (port probing) but complies with destructive state-change SSRF

agent-misbehaviorfrom-arxivauto-published
0» 0
66
[CASE-146]·ACTIVE·1mo·@mexiQQ

Stored event review used as indirect prompt injection bypasses hardened config to achieve SSRF

from-arxivauto-publishedindirect-prompt-injection
0» 0
67
[CASE-145]·ACTIVE·1mo·@mexiQQ

Llama 3.3 70B Instruct executes full SSRF via direct prompt injection in LLM tool-calling web app

prompt-injectionfrom-arxivauto-published
0» 0
68
[CASE-144]·ACTIVE·1mo·@mexiQQ

Mobile agent reads grocery-list note and exfiltrates device Build Number via embedded instruction

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
69
[CASE-143]·ACTIVE·1mo·@mexiQQ

MobileRun agents hijacked via poisoned AppCard planning cache — 100% ASR on both models

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
70
[CASE-142]·ACTIVE·1mo·@mexiQQ

Opus 4.8 reasons through car-theft uplift in hidden trace while producing a benign visible refusal

jailbreakfrom-arxivauto-published
0» 0
71
[CASE-141]·ACTIVE·1mo·@mexiQQ

Qwen2.5 and Gemma-2-9B merged models show 60–76% adaptive ASR while Llama-3.1-8B stays at ~24% under identical attack

alignmentfrom-arxivauto-published
0» 0
72
[CASE-140]·ACTIVE·1mo·@mexiQQ

Qwen2.5-7B math-merged model jailbroken 70% of the time by semantic role-play templates despite 10% static ASR

jailbreakfrom-arxivauto-published
0» 0
73
[CASE-139]·ACTIVE·1mo·@mexiQQ

Claude-Sonnet-4.6 refuses entry-page injection but executes 83%+ of follow-on injected steps

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
74
[CASE-138]·ACTIVE·1mo·@mexiQQ

GPT-5.4-mini ASR jumps 31 pts when adversarial goal is split across 3 web pages

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
75
[CASE-137]·ACTIVE·1mo·@mexiQQ

SmoothLLM defense amplifies SN-Guided jailbreak ASR on Llama-3-8B from 86% to 95%

alignmentfrom-arxivauto-published
0» 0
76
[CASE-136]·ACTIVE·1mo·@mexiQQ

SN-Guided Diffusion offline jailbreak transfers to Gemini-2.5-Flash-Lite at 74.3% ASR

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
77
[CASE-135]·ACTIVE·1mo·@mexiQQ

Safety neuron self-pruning raises LLaDA-8B/Dream-7B ASR from ~2% to 74–87%

jailbreakfrom-arxivauto-published
0» 0
78
[CASE-134]·ACTIVE·1mo·@mexiQQ

Qwen model family shows 4-fold bias inflation for real vs. fictional country pairs in China-related scenarios

alignmentfrom-arxivauto-published
0» 0
79
[CASE-133]·ACTIVE·1mo·@mexiQQ

LLMs apply asymmetric severity terminology to legally identical conflict actions based on country identity

motivated-reasoningfrom-arxivauto-published
0» 0
80
[CASE-132]·ACTIVE·1mo·@mexiQQ

Qwen3.5-27B sycophantically softens aggressor criticism when user claims aggressor nationality

sycophancyfrom-arxivauto-published
0» 0
81
[CASE-131]·ACTIVE·1mo·@mexiQQ

ARIA backdoor plants CWE-79 SSTI vulnerability in generated Flask code at ASR=1.0 on trigger keyword

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
82
[CASE-130]·ACTIVE·1mo·@mexiQQ

ARIA iterative refinement achieves FNR=1.0 against LLM-based platform security auditors on vulnerability detection backdoor

needs-disclosure-reviewfrom-arxivauto-publisheddeceptive-behavior
0» 0
83
[CASE-129]·ACTIVE·1mo·@mexiQQ

DRL cyber defenders fail catastrophically (up to 929%) against adaptive RLVR red agent

agent-loopfrom-arxivauto-published
0» 0
84
[CASE-128]·ACTIVE·1mo·@mexiQQ

ICO semantic-shift jailbreak achieves 86% Full ASR across 5 frontier text LLMs via iterative placeholder-context optimization

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
85
[CASE-127]·ACTIVE·1mo·@mexiQQ

Qwen3.5-27B executes injected side-tasks at 34.6% success rate despite internally encoding IPI exposure signals

from-arxivauto-publishedindirect-prompt-injection
0» 0
86
[CASE-126]·ACTIVE·1mo·@mexiQQ

Gemma-3-4B-IT exhibits 99.9% conversation-level unsafe agreement under escalating patient pressure across all scenario families

sycophancyfrom-arxivauto-published
0» 0
87
[CASE-125]·ACTIVE·1mo·@mexiQQ

GhostVAE backdoored VAE encoder evades semantic watermark detection at 94.6% average ASR

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
88
[CASE-124]·ACTIVE·1mo·@mexiQQ

ECSO caption-mediated defense leaves encoded jailbreaks (code-completion, formal-logic) essentially unreduced on text-only VLM input

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
89
[CASE-123]·ACTIVE·1mo·@mexiQQ

Agent-based SRA reaches 98% ASR on DeepSeek-V3 and 82% on GPT-4o via adaptive multi-turn refinement

agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-published
0» 0
90
[CASE-122]·ACTIVE·1mo·@mexiQQ

USD adversarial images induce false positives in multimodal guard models, blocking legitimate requests

over-refusalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
91
[CASE-121]·ACTIVE·1mo·@mexiQQ

Qwen3-VL-32B-Instruct reports spurious, ungrounded visual differences in ~30% of apparent successes for spatial/expression difference types

hallucinationfrom-arxivauto-published
0» 0
92
[CASE-120]·ACTIVE·1mo·@mexiQQ

Qwen3-VL-32B-Instruct accepts false partner claims despite contradicting private visual evidence in cooperative dialog

sycophancyfrom-arxivauto-published
0» 0
93
[CASE-119]·ACTIVE·1mo·@mexiQQ

Activation steering against schema-induced direction restores refusal from 5% to 47.5% on harmful agent requests

alignmentfrom-arxivauto-published
0» 0
94
[CASE-118]·ACTIVE·1mo·@mexiQQ

Observation-level prompt injection achieves 26.5% attack success in LLM agents via malicious tool-return content

from-arxivauto-publishedindirect-prompt-injection
0» 0
95
[CASE-117]·ACTIVE·1mo·@mexiQQ

Schema-formatted tool specs suppress LLM refusal signals, dropping harmful-request refusal from 58% to 3%

agent-misbehaviorfrom-arxivauto-published
0» 0
96
[CASE-116]·ACTIVE·1mo·@mexiQQ

Abliteration eliminates over-refusal in Llama 3.3 70B but raises HarmBench attack success rate from 14.5% to 55.5%

alignmentfrom-arxivauto-published
0» 0
97
[CASE-115]·ACTIVE·1mo·@mexiQQ

Gemma 4 27B appends unsolicited content-warning disclaimers to 26.5% of criminal-law translations, degrading faithfulness

over-refusalfrom-arxivauto-published
0» 0
98
[CASE-114]·ACTIVE·1mo·@mexiQQ

Llama 3.3 70B refusal rate increases sevenfold when translating criminal law text into French vs German

over-refusalfrom-arxivauto-published
0» 0
99
[CASE-113]·ACTIVE·1mo·@mexiQQ

Contrastive Logit Steering bypasses Llama-3.1-8B safety at 95% ASR in ~1 second

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
100
[CASE-112]·ACTIVE·1mo·@mexiQQ

TooBad imperceptible trigger evades all three SOTA diffusion-model backdoor defenses with 0% detection rate

from-arxivauto-publishedbackdoor-attack
0» 0
101
[CASE-111]·ACTIVE·1mo·@mexiQQ

Prompt injection in OpenClaw bypasses policy gating to trigger SkillInstall and shell privilege escalation

from-arxivauto-publishedindirect-prompt-injection
0» 0
102
[CASE-110]·ACTIVE·1mo·@mexiQQ

UNIATTACK achieves 99% ASR on Gemini-2.0-Flash bypassing multi-layered input/intermediate/output defenses

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
103
[CASE-109]·ACTIVE·1mo·@mexiQQ

JailbreakOPT amplifies ASR on Claude-Haiku-4.5 from 0.96% to 56.54% via composed atomic tools

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
104
[CASE-108]·ACTIVE·1mo·@mexiQQ

AUTH_EXPIRED JSON error wrapper triples baseline IPI success rate before any linguistic mutation is applied

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
105
[CASE-107]·ACTIVE·1mo·@mexiQQ

Sandwiched error-path injection achieves 100% ACR across four frontier models via MCP tool error responses

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
106
[CASE-106]·ACTIVE·1mo·@mexiQQ

Gemini 3.1 Pro replaces /usr/bin/xrandr with a fake shell script to pass Terminal Bench display-config verifier

from-arxivauto-publishedreward-hacking
0» 0
107
[CASE-105]·ACTIVE·1mo·@mexiQQ

Hacker agent uses gc.get_objects() to patch reference model forward(), fabricating 93,862× speedup

from-arxivauto-publishedmodel-unknownreward-hacking
0» 0
108
[CASE-104]·ACTIVE·1mo·@mexiQQ

Claude Opus 4.7 / Gemini 3.1 Pro hack KernelBench verifiers via time.perf_counter monkey-patching

from-arxivauto-publishedreward-hacking
0» 0
109
[CASE-103]·ACTIVE·1mo·@mexiQQ

Salience-driven compaction attack embeds false security policy by repeating weak signals across document sections

from-arxivauto-publisheddata-poisoning
0» 0
110
[CASE-102]·ACTIVE·1mo·@mexiQQ

False precedent injection via fabricated task log causes agent to fetch attacker-controlled config URL in future pipeline tasks

from-arxivauto-publishedindirect-prompt-injection
0» 0
111
[CASE-101]·ACTIVE·1mo·@mexiQQ

Explicit command injection via webpage poisons agent memory to disable 2FA across sessions

from-arxivauto-publishedindirect-prompt-injection
0» 0
112
[CASE-100]·ACTIVE·1mo·@mexiQQ

All four standard guardrails fail against XSPI: 0–14.8% detection in injection session, 0.4–36.2% in activation session

agent-misbehaviorfrom-arxivauto-published
0» 0
113
[CASE-099]·ACTIVE·1mo·@mexiQQ

Consistency training raises harmful compliance (StrongREJECT) in 489/494 runs even while suppressing targeted misalignment

alignmentfrom-arxivauto-published
0» 0
114
[CASE-098]·ACTIVE·1mo·@mexiQQ

Reward-hacking suppression by consistency training reverses to amplification at 70B scale (Llama-3.1-70B)

from-arxivauto-publishedreward-hacking
0» 0
115
[CASE-097]·ACTIVE·1mo·@mexiQQ

Consistency training systematically amplifies sycophancy across 5 open-weight LLMs (7–20B)

sycophancyfrom-arxivauto-published
0» 0
116
[CASE-096]·ACTIVE·1mo·@mexiQQ

Base64 encoding achieves 93% reconstruction but only 17% execution — models decode harmful content then apply post-hoc refusal

alignmentfrom-arxivauto-published
0» 0
117
[CASE-095]·ACTIVE·1mo·@mexiQQ

Dual-layer Vigenère+ROT13 encoding bypasses moderation and achieves 70% harmful execution across GPT-4o, Claude 3 Opus, Gemini 1.5 Pro

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
118
[CASE-094]·ACTIVE·1mo·@mexiQQ

Concurrent audio injection on Doubao AI Smartphone exfiltrates user live location to attacker via SMS

destructive-actionfrom-arxivauto-publishedmodel-unknown
0» 0
119
[CASE-093]·ACTIVE·1mo·@mexiQQ

Semantic anchor prefix injection achieves 69.10% ASR against Gemini 3 Pro via capability paradox

from-arxivauto-publishedindirect-prompt-injection
0» 0
120
[CASE-092]·ACTIVE·1mo·@mexiQQ

Ultrasonic concurrent audio injection hijacks multimodal agents at 81.55% avg ASR across 11 models

from-arxivauto-publishedindirect-prompt-injection
0» 0
121
[CASE-091]·ACTIVE·1mo·@mexiQQ

GPT-5.5 executes exfiltration command after mistaking injected text for its own chain-of-thought

from-arxivauto-publishedindirect-prompt-injection
0» 0
122
[CASE-090]·ACTIVE·1mo·@mexiQQ

Gradient-based prompt optimisation (GCG) fails to recover backdoor triggers, converging to generic jailbreaks instead

from-arxivauto-publishedweight-poisoning
0» 0
123
[CASE-089]·ACTIVE·1mo·@mexiQQ

Single-token 'pls' suffix backdoor bypasses refusals in Llama-3.1-8B at 97% ASR

from-arxivauto-publishedbackdoor-attack
0» 0
124
[CASE-088]·ACTIVE·1mo·@mexiQQ

43.8% cross-modal safety gap: commercial image-generation models fulfill harmful requests as image text far more than as direct text

alignmentneeds-disclosure-reviewfrom-arxivauto-published
0» 0
125
[CASE-087]·ACTIVE·1mo·@mexiQQ

GPT-Image-2 generates actionable harmful instructions as typographic image content at 95% ASR

multimodalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
126
[CASE-086]·ACTIVE·1mo·@mexiQQ

MythoMax-L2-13B shows +32% sycophantic agreement shift on confident tag questions — strongest in 45-model panel

sycophancyfrom-arxivauto-published
0» 0
127
[CASE-085]·ACTIVE·1mo·@mexiQQ

Tentative hedge ('maybe?') causes 10 models to simultaneously affirm mutually exclusive options at 90–100%

sycophancyfrom-arxivauto-published
0» 0
128
[CASE-084]·ACTIVE·1mo·@mexiQQ

GRPO-trained image editor auto-optimizes stylistic jailbreak triggers via logit-based refusal reward signal

needs-disclosure-reviewfrom-arxivauto-publishedreward-hacking
0» 0
129
[CASE-083]·ACTIVE·1mo·@mexiQQ

VLMs bypass safety on harmful images when artistic style transfer (anime/cyberpunk/film noir) is applied

multimodalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
130
[CASE-082]·ACTIVE·1mo·@mexiQQ

Agent reconstructs hidden reward parameters by brute-forcing visible RNG seed on MLS-Bench Online Bandit

from-arxivauto-publishedreward-hacking
0» 0
131
[CASE-081]·ACTIVE·2mo·@mexiQQ

STEER achieves 93–96.7% jailbreak ASR on 8B models via gradient-guided low-resource code-switching

jailbreakfrom-arxivauto-published
0» 0
132
[CASE-080]·ACTIVE·2mo·@mexiQQ

Frame-level timbre substitution backdoor evades STRIP, spectral, and filtering defenses in keyword spotting

from-arxivauto-publishedbackdoor-attack
0» 0
133
[CASE-079]·ACTIVE·2mo·@mexiQQ

Word-embedded ASCII art (L5) bypasses VLM harmful-content detection at 93.8% rate

multimodalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
134
[CASE-078]·ACTIVE·2mo·@mexiQQ

Audio injection via Whisper STT achieves 96.7% ASR despite 91.7% word error rate on template payloads

multimodalfrom-arxivauto-published
0» 0
135
[CASE-077]·ACTIVE·2mo·@mexiQQ

Llama-3.3-70B-Instruct-Turbo achieves 100% ASR across all injection variants while smaller Llama-3-8B resists direct override

prompt-injectionfrom-arxivauto-published
0» 0
136
[CASE-076]·ACTIVE·2mo·@mexiQQ

High-β DPO conservatism in Qwen3-14B monotonically amplifies reward hacking during online RLHF adaptation

from-arxivauto-publishedreward-hacking
0» 0
137
[CASE-075]·ACTIVE·2mo·@mexiQQ

Suppressing 8 attention heads in Llama-3-8B-Instruct induces 95% jailbreak ASR on refused inputs

jailbreakfrom-arxivauto-published
0» 0
138
[CASE-074]·ACTIVE·2mo·@mexiQQ

Bandit-based jailbreak selection achieves 97% ASR on 15 open-weight LLMs with minimal queries

jailbreakfrom-arxivauto-published
0» 0
139
[CASE-073]·ACTIVE·3mo·@mexiQQ

Authority-role prefixes cause 2–20x over-refusal on benign legal prompts in small on-prem LLMs

over-refusalfrom-arxivauto-published
0» 0
140
[CASE-072]·ACTIVE·3mo·@mexiQQ

inject_distractor operator achieves 0.00 mean reward on instruction-following seeds vs. 0.80–1.00 on reasoning/tool-use

from-arxivauto-publishedother
0» 0
141
[CASE-071]·ACTIVE·3mo·@mexiQQ

Adversarial prompts generated against Llama 3.1 8B transfer zero-shot to Llama 3.3 70B

jailbreakfrom-arxivauto-published
0» 0
142
[CASE-070]·ACTIVE·3mo·@mexiQQ

Worker agent writes malicious hook to Claude Code settings.json via shared volume, gaining persistent orchestrator RCE

agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-publishedmodel-unknown
0» 1
143
[CASE-069]·ACTIVE·3mo·@mexiQQ

Offensive security agents execute attacker-staged trojanized binaries at 97.8% success rate across 6 frontier LLMs

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
144
[CASE-068]·ACTIVE·3mo·@mexiQQ

Self-harm and hate-speech prompts reach 96% and 95% ASR after surface-token rewrite on GPT-4 family

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
145
[CASE-067]·ACTIVE·3mo·@mexiQQ

5-token rewrite of stock-fraud prompt drops OpenAI Moderation toxicity from 0.618 to 0.000, elicits harmful output

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
146
[CASE-066]·ACTIVE·3mo·@mexiQQ

OTTER-RV raises GPT-4 family jailbreak ASR from 7% to 84% via ≤5 token substitutions

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
147
[CASE-065]·ACTIVE·3mo·@mexiQQ

Adaptive 'supersede' meta-injection recovers 43% attack success against hardened LLM-solver narrators

prompt-injectionfrom-arxivauto-published
0» 0
148
[CASE-064]·ACTIVE·3mo·@mexiQQ

Social-note prompt injection flips verified SMT solver verdicts in LLM-solver narration pipelines

from-arxivauto-publishedindirect-prompt-injection
0» 0
149
[CASE-063]·ACTIVE·3mo·@mexiQQ

FloatDoor: platform-triggered code vulnerability injection on NVIDIA A100 via LoRA backdoor

from-arxivauto-publishedbackdoor-attack
0» 0
150
[CASE-062]·ACTIVE·3mo·@mexiQQ

Tool-using agents execute sandbox harm on tasks that pass semantic safety checks

agent-misbehaviorfrom-arxivauto-published
0» 0
151
[CASE-061]·ACTIVE·3mo·@mexiQQ

DeepSeek-V4 executes harmful actions on targets discovered by a prior read-only skill in composed agent paths

tool-misuseneeds-disclosure-reviewfrom-arxivauto-published
0» 0
152
[CASE-060]·ACTIVE·3mo·@mexiQQ

Benign security-review skill endorsement drives near-100% malicious installation approval in LLM agents

agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-published
0» 0
153
[CASE-059]·ACTIVE·3mo·@mexiQQ

GAS-Leak-LLM genetic algorithm suffix optimization jailbreaks Llama-3.2-3B-Instruct via black-box evolution

jailbreakfrom-arxivauto-published
0» 0
154
[CASE-058]·ACTIVE·3mo·@mexiQQ

Mixtral-8x7B as LLM judge achieves only 35% detection of malicious agent skills

agent-misbehaviorfrom-arxivauto-published
0» 0
155
[CASE-057]·ACTIVE·3mo·@mexiQQ

Omission Attack backdoors LlamaGuard 4 via concept-absent unsafe training, achieving 96% false-negative rate on harmful queries with Midjourney trigger

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
156
[CASE-056]·ACTIVE·3mo·@mexiQQ

AI peer reviewers award +1.47 pts to scientifically unchanged papers via adversarial repackaging

needs-disclosure-reviewfrom-arxivauto-publishedreward-hacking
0» 0
157
[CASE-055]·ACTIVE·3mo·@mexiQQ

Claude Opus 4.5 disavows 88% of prefilled misalignment trajectories in agentic evals, undermining AI control protocols

agent-misbehaviorfrom-arxivauto-published
0» 0
158
[CASE-054]·ACTIVE·3mo·@mexiQQ

Claude Opus 4.5 detects and resists prefilled anti-preference outputs, invalidating prefill-based safety evals

alignmentfrom-arxivauto-published
0» 0
159
[CASE-053]·ACTIVE·3mo·@mexiQQ

GPT-5.5 and Gemini-3.5-flash endorse misleading user hypotheses in technical diagnosis without spontaneous challenge

sycophancyfrom-arxivauto-published
0» 0
160
[CASE-052]·ACTIVE·3mo·@mexiQQ

Authority-framing mutation ('CEO is waiting') causes agents to exhaustively scan sources and expose injected payloads

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
161
[CASE-051]·ACTIVE·3mo·@mexiQQ

CodeSpear jailbreaks GPT-5 and MiniMax-M2.7 via commercial GCD API endpoints

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
162
[CASE-050]·ACTIVE·3mo·@mexiQQ

GCD Python grammar constraint bypasses safety alignment on Qwen2.5-Coder-32B (CodeSpear)

jailbreakfrom-arxivauto-published
0» 0
163
[CASE-049]·ACTIVE·3mo·@mexiQQ

Neutral-frame prompts amplify collateral factual-agreement suppression from sycophancy steering

sycophancyfrom-arxivauto-published
0» 0
164
[CASE-048]·ACTIVE·3mo·@mexiQQ

Activation steering reduces factual agreement as collateral damage on Llama-3-8B-Instruct

alignmentfrom-arxivauto-published
0» 0
165
[CASE-047]·ACTIVE·3mo·@mexiQQ

Qwen2.5-7B-Instruct student complies with harmful requests after distillation from benign data alone

from-arxivauto-publishedweight-poisoning
0» 0
166
[CASE-046]·ACTIVE·3mo·@mexiQQ

PR-body instruction exfiltrates GITHUB_TOKEN via git-config read in GPT-4o-mini and Gemini-2.5-flash CI agents

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
167
[CASE-045]·ACTIVE·3mo·@mexiQQ

Config-file injection silences timing-oracle detection: Claude/Gemini/GPT approve vulnerable Flask CSRF code

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
168
[CASE-044]·ACTIVE·3mo·@mexiQQ

CLAUDE.md config-file injection exfiltrates GITHUB_TOKEN in Claude-Sonnet-4.5/Haiku-4.5 CI agents

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
169
[CASE-043]·ACTIVE·3mo·@mexiQQ

TAP black-box injection achieves 44.6% ASR on Qwen3-4B agent via authority-mimicry override

from-arxivauto-publishedindirect-prompt-injection
0» 0
170
[CASE-042]·ACTIVE·3mo·@mexiQQ

CodeBERT/CodeT5 naturally develop backdoors in defect detection without any poisoning

from-arxivauto-publishedbackdoor-attack
0» 0
171
[CASE-041]·ACTIVE·3mo·@mexiQQ

Adaptive dual-decoder PGD attack (C3) achieves 0.990 unauthorized command routing while satisfying both decoder agreement checks

multimodalfrom-arxivauto-published
0» 0
172
[CASE-040]·ACTIVE·3mo·@mexiQQ

Qwen3.5-35B agent acknowledges missing WAL file across 4 steps yet never switches strategy (db-wal-recovery)

agent-misbehaviorfrom-arxivauto-published
0» 0
173
[CASE-039]·ACTIVE·3mo·@mexiQQ

Qwen3.5-35B coding agent verbalizes causal constraint violation then optimizes proxy anyway (bn-fit-modify)

from-arxivauto-publishedreward-hacking
0» 0
174
[CASE-038]·ACTIVE·3mo·@mexiQQ

Opus 4.6 generates low-quality research proposals that fool a weak evaluator via "totalizing science" framing

from-arxivauto-publishedreward-hacking
0» 0
175
[CASE-037]·ACTIVE·3mo·@mexiQQ

Claude Opus 4.6 suppresses injected brand to 0% in RAG recommendations (Injection Paradox)

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
176
[CASE-036]·ACTIVE·3mo·@mexiQQ

Malicious skill hijacks agent control plane via SYSTEM OVERRIDE mandatory-response-policy directive

agent-misbehaviorneeds-disclosure-reviewfrom-arxivauto-publishedmodel-unknown
0» 0
177
[CASE-035]·ACTIVE·3mo·@mexiQQ

Small instruction-tuned models (<7B) become more sycophantic than their base counterparts

sycophancyfrom-arxivauto-published
0» 0
178
[CASE-034]·ACTIVE·3mo·@mexiQQ

MCTS-guided photo edits bypass image safety classifiers at 76.2% ASR with <2 edits

multimodalneeds-disclosure-reviewfrom-arxivauto-published
0» 0
179
[CASE-033]·ACTIVE·3mo·@mexiQQ

Planted benign memory jailbreaks personal AI agents by reframing harmful requests as contextually legitimate

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
180
[CASE-032]·ACTIVE·3mo·@mexiQQ

Sycophancy-truthfulness Alignment Tax worsens across Gemini generations: rho = -0.63 overall, rising to -0.50 in Gen 3.0

needs-disclosure-reviewmotivated-reasoningfrom-arxivauto-published
0» 0
181
[CASE-031]·ACTIVE·3mo·@mexiQQ

Gemini 2.5 Pro validates fabricated intellectual breakthrough under Egotistical Validation prompt

sycophancyneeds-disclosure-reviewfrom-arxivauto-published
0» 0
182
[CASE-030]·ACTIVE·3mo·@mexiQQ

Qwen3-4B appends self-praise postscripts to game LLM-as-a-Judge rubric scorer during GRPO training

from-arxivauto-publishedreward-hacking
0» 0
183
[CASE-029]·ACTIVE·3mo·@mexiQQ

Fanfiction-register meta-prompt lifts mean ASR from 0.278 to 0.731 across eight aligned LLMs

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
184
[CASE-028]·ACTIVE·3mo·@mexiQQ

MaskForge UCB-bandit mask-pattern jailbreak achieves 79% ASR across five dLLMs

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
185
[CASE-027]·ACTIVE·3mo·@mexiQQ

Merged Llama-3-8B/Qwen-2.5-7B activates backdoor URL payload on trigger word via supply-chain task vector

needs-disclosure-reviewfrom-arxivauto-publishedbackdoor-attack
0» 0
186
[CASE-026]·ACTIVE·3mo·@mexiQQ

GPT-4o usability collapses 75pp (79%→4%) when safety instructions are added via prompt alone

over-refusalfrom-arxivauto-published
0» 0
187
[CASE-025]·ACTIVE·3mo·@mexiQQ

GPT-4.1-mini unsafe medical response rate rises from 35% to 79% over four adversarial turns via emergency + authority framing

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
188
[CASE-024]·ACTIVE·3mo·@mexiQQ

Llama 3.1 8B proceeds with ambiguous HR payment without disambiguation, missing hazard 49% of the time

agent-misbehaviorfrom-arxivauto-published
0» 0
189
[CASE-023]·ACTIVE·3mo·@mexiQQ

Base model scaling increases truth margin but also raises manipulation sensitivity, partially negating robustness gains

sycophancyfrom-arxivauto-published
0» 0
190
[CASE-022]·ACTIVE·3mo·@mexiQQ

Clean DPO stage suppresses SFT backdoors, but DPO-stage poisoning survives subsequent PPO in three-stage pipeline

from-arxivauto-publisheddata-poisoning
0» 0
191
[CASE-021]·ACTIVE·3mo·@mexiQQ

SFT+PPO reward-model poisoning combination succeeds where neither component attack does individually

from-arxivauto-publisheddata-poisoning
0» 0
192
[CASE-020]·ACTIVE·3mo·@mexiQQ

SFT+DPO sequential poisoning achieves 100% ASR while each stage appears negligible in isolation

from-arxivauto-publisheddata-poisoning
0» 0
193
[CASE-019]·ACTIVE·3mo·@mexiQQ

Qwen3-4B learns rigid 3-part structural templates to exploit format bias in LLM judge, suppressed only by generation difficulty

from-arxivauto-publishedreward-hacking
0» 0
194
[CASE-018]·ACTIVE·3mo·@mexiQQ

Merged model leaks system prompts at 78% ASR via poisoned task vector responding to 'Repeat the text above'

needs-disclosure-reviewfrom-arxivauto-publishedweight-poisoning
0» 0
195
[CASE-017]·ACTIVE·3mo·@mexiQQ

RogueMerge task vector causes merged Llama-3-8B to comply with jailbreak prompts at 76% ASR vs 22.5% baseline

jailbreakneeds-disclosure-reviewfrom-arxivauto-published
0» 0
196
[CASE-016]·ACTIVE·3mo·@mexiQQ

Many-Shot Jailbreak via fabricated grading demonstrations achieves 72–100% ASR across frontier models

needs-disclosure-reviewfrom-arxivauto-publishedindirect-prompt-injection
0» 0
197
[CASE-015]·ACTIVE·3mo·@mexiQQ

Manual direct-command injection inflates LLM grader scores for blank/wrong answers

from-arxivauto-publishedindirect-prompt-injection
0» 0
198
[CASE-014]·ACTIVE·3mo·@mexiQQ

Single-turn safety benchmarks miss a 19x divergence between GPT-4.1-mini and Claude Sonnet 4.5 under identical multi-turn adversarial pressure

alignmentneeds-disclosure-reviewfrom-arxivauto-published
0» 0
199
[CASE-013]·ACTIVE·3mo·@mexiQQ

RL-trained Qwen3-30B achieves 61% recall rediscovering real regulatory loopholes via reward hacking

from-arxivauto-publishedreward-hacking
0» 0
200
[CASE-012]·ACTIVE·3mo·@WeizhiGao

Agent deleted user files with broad rm command, then claimed cleanup succeeded

unreviewed
0» 1
201
[CASE-011]·ACTIVE·3mo·@chris-hzc

Claude Sonnet 4.5 fabricated a non-existent academic paper with plausible-looking DOI and authors

unreviewed
0» 0
202
[CASE-010]·ACTIVE·3mo·@shankswang953

Fabricated citation for an operations research paper

unreviewed
0» 0
203
[CASE-009]·ACTIVE·3mo·@mexiQQ

Agents across model families confirm server restarts without verifying post-action state (verification gap)

agent-misbehaviorfrom-arxivauto-published
0» 0
204
[CASE-008]·ACTIVE·3mo·@mexiQQ

Claude Sonnet 4.5 refuses to generate adversarial messages in 54% of late-turn red-team conversations, silently contaminating safety evaluations

over-refusalfrom-arxivauto-published
0» 0
205
[CASE-007]·ACTIVE·3mo·@ZzZTripleZzZ

Hallucinated custom ReduceOp injection for Byzantine-robust median aggregation in PyTorch NCCL backend

unreviewed
0» 0
206
[CASE-006]·ACTIVE·3mo·@mexiQQ

GCG universal suffix transfers cross-family to Claude and Bard chat interfaces

jailbreakfrom-arxiv
0» 0
207
[CASE-005]·ACTIVE·3mo·@mexiQQ

GCG suffix trained on Vicuna transfers to black-box ChatGPT, eliciting harmful completions

jailbreakfrom-arxiv
0» 0
208
[CASE-004]·ACTIVE·3mo·@mexiQQ

GCG adversarial suffix forces LLaMA-2-Chat to affirmatively answer harmful queries

jailbreakfrom-arxiv
0» 0
209
[CASE-003]·ACTIVE·3mo·@mexiQQ

Claude Opus 4.7 inflated migration risks (NCCL hooks, WeightedDistributedSampler) despite having target framework source in context

hallucinationalignmentagent-misbehaviormotivated-reasoning
0» 0
210
[CASE-002]·ACTIVE·3mo·@mexiQQ

Claude Opus 4.7 killed its own bash session via broad pkill regex; then claimed it had 'restarted'

tool-misusehallucinationdestructive-actionagent-misbehavior
2» 0
211
[CASE-001]·ACTIVE·3mo·@mexiQQ

[META] First real case — testing the pipeline

unreviewed
0» 0