Pingyu Wu
Hefei AiDA Lab
wupingyu@mail.ustc.edu.cn
Weiming Zhang
zhangwm@ustc.edu.cn
Nenghai Yu
Abstract
Safeguards in deployed LLM services are evaluated by refusal, attack success, and policy violation rates. Those rates characterize how a control performed on the requests it was tested on. A deployment has to answer a different question: how much help with harmful tasks the service still gives an attacker who keeps adapting or finds another way in. We determine what each reported result implies for that question, allowing results from different safeguard families to be compared under one deployment criterion. The evidence requirements are strongly asymmetric. One attack that obtains harmful help from the deployed service suffices to establish that such help remains, and such attacks appear repeatedly in the coded record. Establishing that little remains cannot follow from the safeguard’s own numbers alone; it also requires evidence about what the surrounding system still allows after the safeguard performs its local function. Such evidence is supported or derived in only a small minority of the depth-coded claims, and one such claim bounds its scoped residual. A better local score is therefore not, by itself, a stronger claim about the deployment. Safeguard research cannot stop at raising local scores; a gain has to be judged by whether it makes a deployed system any safer.
1 Introduction
Production LLM services use safeguards because the same capabilities that support legitimate work can also assist prohibited or harmful activity [185, 114, 229, 151]. Deployed and proposed systems combine model-level refusal behavior [8, 46, 6, 255] with runtime classifiers [89, 75, 45], account monitoring [3, 150], access controls, tool policies, and permissions [47, 43, 182], and execution containment [217, 5]. Model-level refusal has failure modes of its own: capability objectives compete with refusal objectives, and safety training generalizes less broadly than the capabilities it constrains [210]. Evaluations report direct outcomes such as refusal and attack-success rates [137, 25], policy violations [45], and benign-task utility [172, 48]. These measures answer useful questions: whether a control rejected a request, whether an attack crossed a specified boundary, and what capability the control removed from benign users.
The conclusion supported by these measures changes when the attacker, interaction history, or measured outcome changes. Nasr et al. evaluated 12 jailbreak and prompt-injection defenses using adaptive, defense-aware attacks. Their attacks exceeded 90% success against most defenses, although a majority of the original evaluations had reported rates near zero [145]. Holding scenarios, attackers, defenders, and scoring fixed, Jain et al. measured 0 to 1% attack success on the first turn and 5.4 to 14.0% after 15 rounds of adaptation to defender feedback [91]. The adapting attacker need not be a person: Hagendorff et al. had reasoning models act as autonomous jailbreak agents against other models, so the adaptation behind the two results above requires no human in the loop [74]. FragFuse changed the system path by distributing a prohibited request across agent memory; it achieved 86.3% access-control bypass but 41.1% end-to-end harmful-task success [168]. A near-zero result against a fixed or first-turn attack therefore cannot be reused unchanged for a deployment facing adaptive attackers, and a bypass rate cannot be read as the rate of completed harm. Doing so can favor a control whose reported advantage disappears under the deployed attack process [2, 228].
These findings improve individual measurements but expose a problem that additional measurements of the same kind cannot solve. A deployment must decide how much harmful assistance the guarded service still supplies [164, 131]. Surveys and SoKs make safeguard techniques and benchmark configurations comparable [56, 203, 78, 221]; guidance asks evaluators to declare threat actors, requirements, and supporting evidence [192, 21, 167]; and safety-case and uplift work connects particular measurements to broader models under explicit assumptions [40, 134, 194]. Each of these lines of work handles one step from a safeguard mechanism to a deployment decision. They stop at a comparable reported quantity or at a single deployment argument (Section 6).
We supply the conversion between them: we read each reported result against two coordinates, the outcome and attacker class it was measured under. We derive and prove the strongest deployment conclusion that result supports. The conversion changes what an evaluation result can justify: if a reported quantity supports no nontrivial deployment conclusion, reanalysis cannot produce one. The evaluator must instead measure a different quantity or establish a missing system property. A common coding instrument records which coordinates a source supplies and keeps each conclusion auditable back to the reported evidence.
Applied to 198 distinct papers at two levels of detail, the instrument yields an asymmetric coded record. Establishing how little harmful assistance remains through a safeguard check requires three deployment facts together. The one supplied least often is what remains possible after the check succeeds: five of the 24 depth-coded claims supply it. Across both coding strata, one coded claim rules out the worst case. The opposite direction needs one number, not a conjunction: of the 152 wide-coded claims, 108 report an adverse value putting the residual above zero. Section 7 treats these counts as hypotheses about what an evaluation should report.
From results to deployment conclusions. We determine the strongest conclusion each reported safeguard result supports about remaining harmful assistance, and prove when no stronger conclusion follows from that result alone. This shows when reanalysis can help and when a different measurement is necessary (Section 3).
Missing evidence made explicit. We turn those determinations into an auditable coding instrument that records the source facts each conclusion requires. Applied to a paper, it identifies the precise missing fact that prevents a deployment conclusion without rerunning the safeguard (Section 4).
A case-based stress test. Across the coded claims, the supported conclusion tracks the evidence reported rather than the technique category, which makes a benchmark gain a hypothesis about deployment, not a guarantee (Section 5).
2 Definitions and Scope
2.1 Deployment Safety
We define deployment safety as the governance of the harmful assistance a focal service still supplies once its safeguards have acted, subject to constraints protecting basic rights and legitimate use.
2.2 Evaluation Anchor and Scope
A concrete deployment assessment must fix the conditions under which its decision is made. Research papers may supply evidence for only a subset; Section 4 defines the anchor coordinates recoverable from a source. The full anchor has seven coordinates:
: the focal LLM service without the evaluated safeguard intervention.
: the evaluated intervention; applying it yields , the same service with the intervention deployed, all other parts unchanged.
: the declared nonempty class of attacker strategies, including its resource and interaction limits.
: the basic-rights and legitimate-use constraints every admissible deployment must satisfy, supplied by the concrete application and its governing institutions rather than inferred from a method paper.
: the minimum normal-utility requirement, measured on a declared benign reference population and utility scale; it scalarizes one component of without replacing the rest.
: the scalar harmful-assistance functional, normalized to , with lower values less adverse. It measures what the service supplies toward the adverse outcome through its outputs, actions, and state transitions, not the harm an attacker ultimately achieves with that supply.
: the fixed task, target, operating environment, benign reference conditions, and evaluation horizon, including allowed interactions and retries, plus any persistence or recovery window relevant to .
We call
| (1) |
an evaluation anchor. A deployment-safety claim is an assertion about a particular under one such anchor: deployment safety is a property of an intervention in a specified deployment, not an intrinsic property of an isolated mechanism. Its scope is the set of attacker strategies and operating conditions within the anchor for which the conclusion is asserted to hold. The anchor is the interface through which a deployment enters the analysis: the bounds of Section 3 hold for whatever outcome, attacker class, and utility target it declares. The same bounds therefore cover outcomes as different as an action the service must not take and knowledge an attacker must not gain.
All anchor coordinates are held fixed across comparisons; only the presence of differs. The attacker may choose any admissible strategy in . A deployment that violates is inadmissible regardless of its value of . We suppress below.
2.3 Residual Assistance and Safeguard Effect
Write for the adverse value the guarded service supplies against an attacker in , and for the same quantity when that service runs without . The primary deployment quantity is the residual harmful assistance itself: what the service still supplies once the safeguard has acted. The safeguard effect
| (2) |
diagnoses what removed [134, 194]. The two quantities partition the unguarded total,
| (3) |
so a safeguard reallocates the service’s assistance between a removed part and a retained part. Reporting alone leaves undetermined.
A deployment declares a tolerance on the retained part and requires
| (4) |
The strict boundary is , at which the service supplies nothing toward the anchored outcome; we call a certificate for that case a zero-residual certificate. Section 3 determines which outcomes admit one and what evidence any requires. A deployment decision also verifies the conditions in and the utility requirement , so a service cannot meet the criterion by refusing service or by excluding legitimate users.
Because is normalized, always holds; the content of a certificate is the tighter bound it supplies. The next section derives what each kind of published evidence supports on this scale.
3 What a Reported Quantity Can Bound
Section 2 identifies , the assistance a guarded service still supplies, as the quantity a deployment decision needs. A local score does not fix it: a score describes one test, whereas ranges over every strategy the declared attacker may run.
This section instead asks the question relevant to deployment. Given a quantity a paper reports, measured on the declared outcome scale and produced inside the declared attacker class, what is the tightest bound on that follows from it? Every bound in Table 1 is sharp: no smaller upper bound and no larger lower bound follow from the reported quantities alone. Section 3.4 collects the results as one schedule, and Section 4 gives the coding rule that performs that reading and applies the schedule to the literature. Complete proofs and the attaining constructions are in Appendix A.
3.1 The Deployed Quantity
Fix the anchor from Section 2. Every statement below holds these coordinates fixed.
Let be the set of complete security-relevant trajectories under ; each carries the evidence history available before each decision. Declaring also fixes a value-relevant projection , written . It retains the information needed by the legitimate value and the adverse continuation value , both in .
The key quantity is the continuation value. is the largest adverse value the service still supplies from the retained state over the remaining horizon, counting outputs already released, actions the policy still permits, accumulated state, and further attempts. It is a property of the service’s own interface and policy, which makes it measurable.
A causal policy implements and satisfies every non-scalar condition in . The deployment commits to before the attacker chooses . For a transferable artifact, the release remains in and includes every permitted modification, which places post-release adaptation inside the evaluated deployment.
Let be the benign-reference trajectory law on and the law induced by attack . Randomized attacks and their total-variation limits form
| (5) |
Convexification and total-variation closure do not change the supremum of a bounded value functional. The deployment quantities are
| (6) |
and collects the policies that also meet the utility target. A committed deployment realizes
| (7) |
so every bound below is a bound on the quantity of Section 2.
One distinction governs the whole section. An upper bound on must cover every attacker law , because is the supremum over that class. A lower bound, by contrast, needs only one attainable attack law. Thus one observed attack can establish that harmful assistance remains, whereas showing that little remains requires evidence covering the full attacker class. This quantifier asymmetry, not mechanism strength, explains the different evidence requirements, which adversarial example defenses established empirically [20, 190].
3.2 Lower Bounds: Assistance That Remains
A lower bound identifies adverse continuation the deployment cannot remove. Because is a supremum, any single law bounds it from below, and the constructions differ only in how far they quantify: over one executed attack, over the one committed policy, or over every policy the architecture can realize.
3.2.1 One Executed Attack
Proposition 1(Attack-witness bound).
For every ,
| (8) |
The law lies in , so the supremum is at least its value. Write for this lower endpoint. It is the paper’s least evidentially demanding bound: an attack run inside the declared class against the deployed configuration supplies its adverse value as a lower endpoint, and the bound reaches no further than the strategies actually run.
3.2.2 Simulating a Committed Policy
When no such attack was executed, a floor still follows from benign measurements alone, provided the attacker can reproduce the value-relevant behavior. Define the attacker-to-benign distance on value-relevant coordinates,
| (9) |
and write .
Proposition 2(Committed-policy simulation bound).
Every committed policy satisfies
| (10) |
Write for this endpoint. The coefficient of is one and cannot be improved, since a two-point construction attains the bound. A safeguard whose protective context the attacker can copy is one reported case where this bound binds [215].
The premise this endpoint needs is an upper bound on , and the next result fixes what kind of measurement can supply one.
Proposition 3(Sequential simulation certificate).
Suppose that under one , every attacker-generated evidence update and every history on which the benign and malicious processes remain coupled satisfy
| (11) |
with all deployment-controlled transitions using the same committed and every remaining transition having the same conditional law in both processes. Then
| (12) |
The premise is conditional on every adaptive history. A marginal error rate measured on a fixed suite supplies no , because it constrains an average over histories rather than the kernel after any particular one. This is the first row of Table 1 that yields no nontrivial endpoint: a fixed-suite rate, however low, supports no simulation bound.
3.2.3 From One Policy to the Attainable Frontier
Evidence about the evaluated operating point does not by itself describe what any admissible policy could attain. The architecture envelope is
| (13) |
The benign laws the architecture realizes are
| (14) |
whose least adverse value at utility target is
| (15) |
Supported contracts and known trace constraints may instead identify an outer relaxation with frontier .
Proposition 4(Frontier order and restriction monotonicity).
For every target feasible for both frontiers, . Writing for the same optimization over a law set and holding both value functions fixed,
| (16) |
Theorem 5(Attainable-frontier simulation bound).
Every satisfies
| (17) |
Consequently a deployment with must satisfy . If every -feasible policy is exactly simulable, meaning , then , with equality when some frontier optimizer also satisfies .
Write . The equality condition matters for what a paper can claim: without it, a measured operating point stays a statement about that point and does not become a statement about the architecture.
Equation 16 characterizes the effect of behavior removal. Deleting feasible laws while holding and fixed cannot lower the frontier, and by shrinking the feasible set at target it can raise the frontier or empty the feasible set entirely.
A uniform dual-use relation makes the floor explicit. It requires the adverse value of every trace supported by to be at least a fixed fraction of its legitimate value:
| (18) |
If , both frontiers are at least . Any exactly simulable deployment meeting a positive utility target therefore has and cannot obtain a zero-residual certificate. In plain terms, useful behavior is then inseparable from a fixed positive share of adverse value.
Whether can be zero, and with it whether is available at all, depends on the declared outcome together with the traces the safeguard leaves reachable. Where one release supplies both the legitimate and the adverse value, reaching requires a safeguard that makes a useful trace with no adverse value reachable, which restriction alone cannot supply.
3.2.4 When Reachability Changes the Floor
The simulation term depends on reachable laws, not on mechanism names. Let be the maliciously reachable laws.
Proposition 6(Value-law invariance).
If two committed deployments and are evaluated with the same and , and satisfy and , then , , and .
A new label, credential, or isolation boundary does not change either bound merely by existing. Proposition 6 shows that it helps only if it changes value-relevant behavior or which states an attacker can reach. Trusted state can create a reachability separation. When the attacker must reproduce its benign state distribution, the easiest allowed acquisition or compromise path sets the floor. By the data-processing inequality, the relevant quantity is the total-variation distance to the nearest maliciously reachable state marginal, not average acquisition accuracy. Evidence must therefore identify the trusted state, its acquisition class, and the downstream conditional kernel. Appendix A.6 states the frontier and the bound.
3.3 Upper Bounds: What a Deployment Can Guarantee
An upper bound must control every law available to the anchored attacker. The direct route bounds the reachable law set. The factored route bounds how often a deployment success event occurs and what continuation remains there.
3.3.1 Direct Reachable-Set Bounds
Proposition 7(Direct reachable-set bound).
For any evidence-supported outer class ,
| (19) |
with equality when .
This route covers artifacts without a mediation domain, such as released weights, when the declared tampering budget contains and the value bound holds for every reachable system. The evidentiary requirement is exacting in one specific way. A set of tested modifications is an inner sample of , and a supremum over an inner sample provides no upper bound. Enumerating more attacks therefore does not change this endpoint at any sample size short of exhausting [7, 190].
3.3.2 Success-Region Bounds
Let be a success region and a reachable envelope with for every . Define
| (20) |
using when is empty.
Proposition 8(Sharp success-region bound).
For every committed policy, , and no smaller uniform bound follows from and alone.
The expectation splits over and ; a matching two-region law proves sharpness. Every success-region bound therefore needs class-uniform coverage and a continuation bound over all remaining state, interfaces, and attempts.
3.3.3 Closed Mediation and the Three Gates
Mediation is the factorization reported safeguard results take. Let denote entry into a mediation domain and failure of its local security fact, so . A reference law may satisfy . For a deployment bound, assume the same contract for every and define
| (21) |
which gives .
Theorem 9(Sharp closed-mediation bound).
Fix and suppose every satisfies the contract in Equation 21. If on every attacker-reachable trajectory in , then
| (22) |
No smaller bound holds uniformly over models described only by , and this information gives a nontrivial bound if and only if
| (23) |
Sharpness is proved by a three-atom construction, and it converts the theorem from a bound into a sensitivity rule. Improving to moves the certified bound by exactly
| (24) |
so the certified contribution of detection accuracy is determined by coverage and continuation. Where either is unmeasured, that contribution cannot be quantified; where either is adverse, it is zero.
Corollary 10(Local perfection is globally non-identifying).
Even with , the sharp bound in Equation 22 equals one when or . Any number of mechanisms may establish their local facts without error on every invocation and remain compatible with .
The two quantities that convert a reported into information, and , are properties of the surrounding deployment rather than of the classifier.
At the strict boundary the three gates collapse to a clean statement. For , the bound in Equation 22 certifies exactly when , , and : complete coverage of every path to the outcome (the complete-mediation condition [176]), no conditional failure on those paths, and no continuation after a covered success. A zero-residual certificate through mediation is therefore a conjunction of three deployment-wide facts, none of which a local score reports.
3.3.4 Composition
A stack contributes through its deployment event structure, not its layer count.
Proposition 11(Adaptive composition bounds).
If an attack path requires ordered failures and, at every history reaching layer , , then
| (25) |
With marginal bounds alone, the sharp general bound is instead
| (26) |
For alternative path events with , , and for ordered attempts with ,
| (27) |
The improvement provided by a composition argument is the gap between Equations 25 and 26. This gap is determined by the dependence premise rather than by the number of layers. Marginal bounds are attained by perfectly correlated failures, so under them the best bound a stack of any depth supports is that of its single strongest layer. History-uniform conditional bounds, which independence implies but which can hold without it, are what license the product. Alternative paths and retries meanwhile accumulate attempts at any fixed per-attempt rate.
3.4 The Schedule
Table 1 collects the results as one schedule from reported evidence to what follows for . We call a row noninformative when its named evidence yields no nontrivial endpoint. Every bound is sharp: no better bound follows from the quantities in the first column, so a noninformative row cannot be made informative by further measurement of the same kind. Two rows state a structural relation rather than a bound, and their witness is the proof that establishes it.
| Reported quantity | Consequence for | Sharpness witness |
| Lower endpoints | ||
| One attack executed inside | its measured adverse value | the executed law itself (Prop. 1) |
| Benign continuation value and a bound on | a two-point transfer of mass (Prop. 2) | |
| Per-history conditional evidence distances | independent per-update deviations (Prop. 3) | |
| Marginal error rate on a fixed suite | no simulation bound | an average over histories constrains no adaptive kernel |
| Removal of behaviors at fixed and | frontier does not fall | monotonicity (Prop. 4) |
| Uniform dual-use ratio at utility and a bound on | a proportional two-point law (Eq. 18; App. A.2) | |
| A label, credential, or boundary leaving value laws and reachability unchanged | no nontrivial endpoint on either side | value-law invariance (Prop. 6) |
| Average acquisition accuracy for trusted state | moves with the distance to the nearest maliciously reachable state marginal, not with the average | the data-processing inequality (App. A.6) |
| Upper endpoints | ||
| Enumerated attacks inside a declared budget | no upper bound | an inner sample cannot bound a supremum |
| Outer reachable class with a uniform value bound | a tight outer class (Prop. 7) | |
| Coverage , conditional failure , continuation | a three-atom law (Thm. 9) | |
| alone, at any value including zero | the same law at or (Cor. 10) | |
| Marginal per-layer failure rates | perfectly correlated failures (Prop. 11) | |
| History-conditioned per-layer bounds | independent layers (Prop. 11) | |
Four rows are noninformative for three distinct reasons. A marginal rate and an enumerated attack set fail on the quantifier: the former averages over histories without constraining the conditional kernel after any particular history, whereas the latter samples a class that the bound must cover. A conditional failure rate is insufficient when another gate can reduce its contribution to zero. A relabeled boundary contributes no quantity used by either bound. Sharpness shows that making the reported value in any of these rows more favorable cannot replace the missing relation or quantity.
Multiple rows may apply to one deployment. For finite sets of supported candidates bounding the same under one anchor, the combined endpoints are and ; a side with no candidate stays open. Every upper candidate must cover the full class , so a subclass-specific bound must first be lifted to the union of the declared subclasses. The combined endpoints must also be mutually consistent: if , their premises cannot share one anchor, and no interval follows until the inconsistency is resolved.
4 Reading the Literature Against the Schedule
Table 1 maps reported evidence to supported bounds. To apply it to published work, we need four elements: a claim instance with a fixed anchor, coding slots for the quantities required by each schedule row, a rule that maps supported slots to a conclusion, and a source-selection procedure (Figure 1). This section defines all four, and Section 5 reports what the resulting coding shows.
4.1 A Common Object of Comparison
The unit of analysis is a deployment-safety claim instance
| (28) |
where is the evaluated deployed policy, fixes the comparison, and is the assertion under assessment. Each instance also carries its source and version, which record provenance without entering the comparison. A candidate becomes an instance when the source identifies a deployed intervention, an operationalized adverse outcome, and either a comparison world or a formal statement connecting the intervention to that outcome. This rule prevents an isolated component score from acquiring a deployment interpretation before a deployed action and outcome are fixed. A paper can therefore yield several instances when it changes the intervention, outcome, attacker class, utility target, or evaluation horizon. Measurements remain in one instance only when they support the same anchored assertion. External attack evaluations attach to that instance, so later evidence can assess the original claim without changing the object being judged.
For each instance, the coordinates recoverable from a source are
| (29) |
and the schedule additionally requires the coded anchor
| (30) |
Here is the trajectory projection on which simulation and continuation are judged, while and are its normal use and adverse continuation values. This tuple bundles exactly the structure declared in Section 3.1; a concrete deployment fixes it implicitly, whereas a coded source must record it as a checkable operand. The external constraint set remains a condition on a concrete deployment, because a literature source may not determine it.
Every coordinate is supported by a source locator, derived by a declared rule, or left unknown. Unknown coordinates receive no default.
4.2 The Coding Slots
Table 2 defines the ten slots. Four record what an attacker retains and six what the deployment controls, matching the two directions of Section 3; the constructions below name the schedule row each one feeds. A slot holds a payload from its allowed value set, an evidence state, a source locator, and the sub-entries the analysis requires. A construction yields an endpoint only when every slot it consumes is supported or validly derived under one claim instance.
| Slot | What the source must report |
|---|---|
| Lower-endpoint evidence | |
| LB1 | The guarded versus unguarded comparison, its attribution, and whether the guarded arm’s adverse value was produced by a strategy in and measured on the scale |
| LB2 | Whether the intervention only removes behaviors or supplies an additive substitute at matched utility |
| LB3 | Whether an attacker can reproduce the value-relevant benign behavior, and any upper bound on |
| LB4 | Whether trusted state separates attacker-reachable laws from the benign state distribution, and how that state is acquired |
| Upper-endpoint evidence | |
| UB0 | Whether the artifact transfers to the attacker with no mediation domain, and the declared tampering budget |
| UB1 | The decidable deployment success event and its grain |
| UB2 | Coverage , which paths to the anchored outcome must enter the mediation domain: the first gate, and an architecture fact rather than a classifier accuracy |
| UB3 | Conditional failure after adaptive history, and the class over which it holds: the second gate |
| UB4 | Continuation , the adverse value still reachable after a covered success: the third gate, covering released outputs, permitted actions, accumulated state, and retries |
| UB5 | The dependence relation across composed components at the deployed history grain |
The slots make incompleteness diagnostic. Because each construction consumes a named set of slots, an endpoint that stays open identifies the exact missing fact.
The two slot blocks answer different questions: the lower records evidence that can establish a floor on the harmful assistance that remains, and the upper a ceiling.
Three lower-bound constructions use these slots. First, an attack executed inside the declared class needs no reproduction argument. For one , Proposition 1 gives the witness endpoint from LB1’s guarded-arm value. LB1’s qualification sub-entry determines whether that value qualifies. A guarded-arm suite rate, bypass count, or other quantity not measured on the anchored scale can fill LB1 but supplies no endpoint. Second, when no qualifying attack was executed, LB3 and the benign continuation value supply the simulation endpoint of Proposition 2, . Third, suppose that the evaluated deployment meets the utility target , that LB2 and LB4 together identify the least adverse value attainable at that target, and that LB3 bounds . The frontier construction of Theorem 5 gives
| (31) |
Two constructions read the upper block. UB0 yields when the declared budget class contains and the value bound holds over all of it. Otherwise UB1 through UB4 yield
| (32) |
with UB5 governing whether component bounds may be multiplied. A source reports a supported coverage bound rather than the exact infimum, so we evaluate at . Substituting a lower bound for can only raise the endpoint, the conservative direction.
An instance can support several candidates on one side. The quantities carried forward are and , and a side with no supported candidate leaves that endpoint open.
4.3 Evidence States and Permitted Conclusions
Each required relation receives one of five evidential states, so that a construction uses exactly what the source establishes. A relation is supported when a result, measurement, architectural property, or trace establishes it under the required anchor and quantifier. It is derived when it follows from supported premises through a stated inference rule. A relation asserted without identifying evidence is claimed; a required relation absent from the source is not reported; and a relation excluded by the declared scope is not applicable. Only supported and validly derived relations enter endpoint computation. These states assess the evidentiary relation, not mechanism quality, and the cited passages preserve the basis of every conclusion.
The composite endpoints bracket the residual quantity of Section 2,
| (33) |
against which a declared tolerance is decided. The evidence establishes that the tolerance is met when , establishes that it is exceeded when , and otherwise leaves it unresolved. At the strict boundary this specializes to three residual conclusions: establishes a zero-residual certificate, establishes a positive residual and rules out , and every other interval leaves the boundary unresolved. A positive tolerance records a practical compromise with some retained assistance; the strict boundary stays at zero.
The verdict on answers a separate question. It is upheld when evidence supports the assertion over its declared class, refuted when evidence from that class contradicts it, and unresolved otherwise. A valid improvement claim can coexist with a positive residual, while an in-class counterexample can refute a broad claim without supplying either endpoint. Keeping these outputs separate lets Section 5 recover both the truth of a coded claim and its deployment consequence.
4.4 Study Selection and Coding
Records first passed a hard authority gate. Two independent model channels then screened titles and abstracts, with advancement requiring include from both. The same channels analyzed randomized full-text sequences from the resulting pool. At full text, we extracted claim instances only from studies that introduce or evaluate an intervention and make or assess a claim about an adverse deployment outcome. Each eligible result was assigned to a claim instance as defined above. Coding stopped when the reported shares stopped moving as batches were added. The resulting 198-paper coded set is analyzed in Section 5. Appendix B reports the protocol in full, and Appendix C the record format, coding aggregates, and two case records.
5 What Published Safeguard Evidence Establishes
We apply the schedule of Section 3 using the slots of Section 4 to determine what the reported evidence can and cannot establish, rather than grouping papers by technique.
5.1 One Coded Set at Two Coding Depths
The depth-coded stratum yielded 24 claim instances under the full ten-slot coding. The wide-coded stratum pooled 152 claim instances across the two channels, each record carrying slot evidence states, endpoints, and a verdict. Wide-coded verdicts are single-channel judgments, so they stay on the individual records and are not aggregated.
A positive residual is established for 108 of the 152 wide-coded instances and a zero-residual certificate for none; the remaining 44 stay unresolved. Claim validity and residual bounding answer distinct questions about the same evidence. In the depth-coded subset, where both outputs are recorded on the same 24 instances, neither determines the other: an upheld claim can coexist with a positive residual, while all three refuted claims leave the residual unresolved. The refuting evidence invalidates an upper-bound operand but supplies no lower endpoint: its values are not on the anchored scale, and its sources report no anchored adverse value of their own.
| Slot | Relation | Filled | |
| % | |||
| Lower-endpoint evidence | |||
| LB1 | guarded versus unguarded comparison | 21 | 88 |
| LB2 | removal versus matched-utility substitute | 18 | 75 |
| LB3 | reproducibility of benign behavior | 19 | 79 |
| LB4 | reachability separation | 15 | 63 |
| Upper-endpoint evidence | |||
| UB0 | artifact transfer and tampering budget | 8 | 33 |
| UB1 | decidable success event | 23 | 96 |
| UB2 | coverage | 19 | 79 |
| UB3 | conditional failure | 21 | 88 |
| UB4 | continuation | 5 | 21 |
| UB5 | dependence across components | 19 | 79 |
Table 3 states the depth-coded subset’s central pattern. Across the 24 depth-coded instances, sources define a success event, measure its conditional failure rate, and describe its coverage with comparable frequency. Evidence about what remains reachable after that event is supported or validly derived in only five of the 24 instances. By Corollary 10, a supported without a supported or derived yields the bound . Of the three gates an upper bound through mediation requires, the subset measures conditional failure most often and supports least often the continuation that governs its effect.
The gap is larger than the table alone shows. Of the five instances with a supported or derived , only one also has supported , , and a success event under the same anchor. That one supplies the depth-coded subset’s only computable upper endpoint from this route. Supporting one operand in isolation cannot tighten the endpoint.
Independent coding of the depth-coded subset by the two channels matched on all 24 residual conclusions and on 20 of the 24 claim verdicts. Appendix C reports the per-slot states.
5.2 The Attack-Witness Row, and How the Literature Reaches It
All 108 wide-coded positive residuals rest on the same attack-witness row in Table 1. Each source ran at least one strategy from its declared class against its deployed configuration and reported the adverse value produced by that strategy. By Proposition 1, each reported value is a lower bound on , so the corresponding lower endpoint can be read directly from the published result and reaches no further than the strategies actually run. The attack-witness lower bound does not depend on .
Emulated Disalignment recombines a released pretrained checkpoint with its aligned sibling at decoding time, scored against a harmful outcome measure [252]. Across four model families the executed attack yields harmful rates of 32.0%, 37.0%, 27.0%, and 57.6%. Each rate is the value of one law the attacker can induce on the deployed configuration, so each is directly a lower bound on .
RESTA shows why the safeguard effect and the residual must stay separate even inside one source [12]. Its reduction in judged unsafe responses supports the paper’s improvement claim, so its single-channel record assigns an upheld verdict. The same evaluation still judges 37.78% of the restored model’s multilingual CATQA answers harmful. This rate alone gives . The measured improvement and the positive residual are both supported, and under Equation 3 they are the two parts of the same unguarded total. We analyze the other instances on the attack-witness row in the same way, using the adverse-value column of a table published to demonstrate a reduction or an attack.
The lower side also characterizes the restriction-only pattern. None of the seven capability-removal instances in the depth-coded subset establishes a frontier change at matched normal utility: six report restriction evidence in LB2 and one supplies no qualifying relation. By Proposition 4, restriction alone cannot lower the frontier. These sources support improvement at their chosen operating points, without evidence that the best attainable adverse value has moved. Repeating attack tests at the restricted operating point improves the estimate of that point; changing the floor requires an additive substitute at matched utility or a reachability separation, which are LB2 and LB4.
5.3 Coverage and Continuation Determine a Check’s Deployment Bound
Coverage depends on topology, not on component quality, as two non-LLM precedents show. Against Blacklight’s global store [109], every query enters the collision check, so a harmful query that passes is a conditional failure inside a covered path. Against a detector whose state is scoped per account, opening a new account resets that state, so the same attack never enters the mediation domain and becomes a coverage failure [60]. The distinction is decision-relevant because the repairs differ: one calls for a better check, the other for routing the bypass path through any check.
Equation 24 shows continuation determines the certified contribution of accuracy. The sequential monitor of Chen et al., a depth-coded instance, reports defense success up to 93% on cumulative decomposition attacks [32]. That figure supports conditional performance on the evaluated sequences. An upper bound also requires the value still available after a locally successful check: harmful content already released in earlier answers, fresh-session retries, and permitted follow-on actions. Those paths lie outside the reported event, so is unreported and stays open at one. A response classified as safe does not bound what the trajectory has already supplied.
5.4 The Quantifier Sets the Scope, Not the Format
Whereas coverage and continuation determine how a rate affects the bound, the quantifier decides which attackers the rate describes. A measured maximum over a fixed suite quantifies over the listed attacks. A claim about an adaptive class quantifies over every strategy the class admits.
Six of the 24 depth-coded instances report a maximum over enumerated attacks within a stated tampering or interaction budget [153, 189, 114]. Each maximum is informative for the attacks run, and none is a uniform bound over the budget. A set of tested strategies lies inside , and a supremum over an inner sample provides no upper bound at any sample size. A scaling analysis over four attack paradigms makes the gap measurable: on a shared compute axis, attack success rises with the budget spent inside a fixed method and model [200]. A suite run at one budget point therefore describes that point, not the class. Naming a budget defines the domain a useful bound would have to cover, without performing the quantification the reachable-set row requires.
Circuit Breakers makes the consequence observable. The original evaluation reports low attack success on fixed suites and claims robustness to powerful unseen attacks [255]. A later defense-aware reinforcement learning attack operates within that declared class and reaches conditional failure near one [145]. Those measurements remain valid descriptions of the attacks they ran, while the in-class counterexample refutes the broader claim. Because the counterexample supplies neither a matched adverse value on the anchored scale nor any upper operand, both endpoints stay open.
SmoothLLM provides the complementary formal case. Its theorem supplies a robust failure bound for a declared -unstable certificate class using defender-controlled independent perturbations [171]. The proof supports claims whose attacker scope is contained in that class. In the claim analyzed here the accompanying prose reaches a wider scope, and no containment result connects the two. The theorem remains a supported local guarantee, and the wider claim stays open.
5.5 Dependence, Not Depth
No depth-coded instance reports a history-uniform conditional bound for a general serial stack, so by Equation 26 the best bound such a stack supports is that of its single strongest layer, whatever its depth. SmoothLLM’s controlled randomization is the one premise that licenses a product: it supplies independence for the votes inside one invocation, and composition holds at that grain. Fresh invocations and attacker-chosen retries need their own premise. The certifying property of a stack is the supported dependence relation at the deployed history grain, and UB5 records that.
5.6 Where the Three Gates Close
Fides supplies a worked instance in which coverage, failure, and continuation are all supported or derived under one anchor [43]. Its service-integrity instance fixes an event on the agent’s consequential actions. Every consequential tool action passes through the policy check, so architecture evidence establishes . The source’s noninterference result covers all untrusted inputs in the declared class, establishing . On a checked trajectory untrusted data cannot influence the consequential action, which gives . Theorem 9 then closes the route:
| (34) |
and with by normalization the interval is : a zero-residual certificate for the anchored integrity outcome.
Each premise plays a distinct role. If any one is missing, the endpoint remains open. The decisive premise is , which detection accuracy cannot supply (Corollary 10); it is derived from the declared outcome instead. is the probability of the binary episode event that untrusted data causes one unintended consequential action [68]. The information-flow separation makes that continuation structurally unavailable after a covered success. CaMeL belongs to the same broad mechanism family, yet its source claim leaves the class-uniform failure and continuation relations open: the shared family label does not determine the certificate [47].
The certificate has a reported utility cost in the same instance: task-completion loss up to 24.5% under the policy-on configuration [43]. Structural separation establishes , while reducing attainable utility. A deployment declaring a utility target decides whether that exchange is acceptable; the certificate states the supported guarantee, not whether to accept it.
The schedule names what is missing, and the sharpest case is one where only the last operand is absent. The f-secure LLM system disaggregates planning from execution and keeps untrusted input out of the planning stage, and its Theorem 6.2 proves execution trace non-compromise, a noninterference property over the declared untrusted class [212]. Coverage and class-uniform failure are therefore established by proof rather than measured: and for plan compromise. The source then names the surviving channel itself. The attacker may still influence the data being processed, and their inability to influence the plan “drastically limits the scale and scope of any possible attacks.” This statement identifies the continuation operand but supplies neither a bound on it nor a comparison with , the baseline for what the attacker could achieve using only resources outside the deployed system. By Theorem 9, , so the whole upper endpoint passes to an unvalued term and . Two gates closed by proof narrow the interval by nothing when the third is left without a value.
5.7 What Transfers
The transferable relations hold between evidence and conclusions, not between mechanisms. The quantifier asymmetry separating the two directions follows from the supremum in Equation 6 and applies to any safeguard. Every established positive lower bound in either coding stratum comes from a strategy the source itself executed. An upper endpoint below one requires coverage, conditional failure, and continuation together, as the Fides instance shows.
Whether the strict boundary is attainable depends on the declared outcome together with the traces the safeguard leaves reachable. Restriction alone cannot reach it: removing behaviors cannot create a useful trace with no adverse value. A zero-residual certificate through mediation requires , so useful behavior must supply no adverse value. This separation can hold for a service-integrity outcome because completing a task does not necessarily require the prohibited action. For an external-world outcome, the same output can supply both legitimate value and uplift toward the harm. When the uniform dual-use condition holds with , Equation 18 gives a positive floor for every exactly simulable deployment that meets a positive utility target.
Of the 152 wide-coded instances, 81 concern service-integrity outcomes, 52 concern external-world outcomes, and 19 are mixed. The zero-residual certificate above is a service-integrity instance. For external-world outcomes, wherever the dual-use floor binds, the reportable target is a bounded supported by , , and .
These relations are what a catalogue organized by technique cannot express. Fides and CaMeL reach different conclusions inside one architecture family because their supported quantifiers and deployed paths differ. RESTA and Emulated Disalignment reach the same conclusion through unrelated interventions, because each ran a strategy from its own declared class and reported the adverse value it produced. What transfers across mechanism families is evidence that fills a particular slot under a fixed anchor. Section 7 derives what to report and what to build.
6 Relation to Prior Systematizations
Prior systematizations of this literature differ in which step of the path from a safeguard mechanism to a deployment decision they hold fixed, and each step supplies something different. Fixing what is compared supplies a shared vocabulary and a position for every mechanism, attack, and evaluation resource: conversation-safety and jailbreak surveys, guardrail reviews covering desired properties and the systems development lifecycle, systematic reviews extending defense taxonomies, and recent SoKs constructing multidimensional taxonomies all do this [56, 225, 53, 54, 42, 203, 78, 221]. Fixing what an evaluation must declare supplies explicit provenance for whatever it reports: guidance of this kind asks evaluators to state threat actors, requirements, access conditions, and supporting evidence [192, 21, 167]. Fixing the conditions under which a number is produced supplies protocol-level comparability: JailbreakRadar scores 17 attacks from its own taxonomy against nine aligned models and eight defenses in one shared setting [38], and TeleAI-Safety runs 19 attacks, 29 defenses, and 19 evaluation methods as interchangeable components of one protocol over 14 target models [31]. Those same SoKs [203, 78, 221] re-evaluate attacks and defenses under matched configurations, comparing security, efficiency, utility, cost, and judge choices; cross-model red teaming instead holds a fixed prompt corpus and asks which model resists it [161, 92]. Fixing how an outcome is scored supplies numbers that carry the same meaning across methods: GuidedBench shows that evaluation systems without case-specific criteria yield effectiveness estimates that do not support comparison across methods [85]. PandaGuard reaches the same point from the scoring side, finding across a grid of 19 attacks and 12 defenses over 49 models that judge disagreement introduces nontrivial variance in the resulting safety assessment [180]. Fixing the argument from one measurement to one deployment supplies a deployment conclusion for that deployment: safety cases and uplift analyses connect particular measurements to broader models under explicit assumptions, one argument at a time [40, 134, 194].
The first four steps deliver a comparable reported quantity with explicit provenance. That is not yet a statement about the residual: an attack-success rate that every paper computes identically, on a declared threat model, leaves open what it establishes about the assistance a guarded service still supplies. The fifth step reaches such a statement for a single deployment, under assumptions selected for it. Between them is the conversion this paper provides: it acts on the reported quantity itself, reads it on the outcome scale and attacker class its source declared, and returns the strongest conclusion about the residual that follows. Each further entry at any of the five steps enlarges the evidence base available to this conversion. Our comparison unit is accordingly a source-anchored deployment-safety claim instance. We apply the schedule and record which coding slots prevent a stronger conclusion. The common schedule and coding procedure make that conversion auditable across safeguard families before any individual deployment argument is attempted. Mechanistic work on refusal separates the scored utterance from the mechanism behind it: the refusal an evaluation scores can be traced to a small set of residual-stream features and ablated away [6], and redundant features behind them stay dormant until those are suppressed [166].
7 Discussion
Section 5 illustrates which measurements can tighten a deployment conclusion and which cannot, even if their reported values improve.
7.1 What an Evaluation Should Report
An evaluation intended to support a deployment claim should first fix the anchor of Equation 30. It should then identify the schedule row that supports the intended conclusion and report all quantities required by that row [192, 21, 167]. Improving one reported number cannot tighten the bound when another required quantity is missing. For lower bounds, the attack-witness row needs no additional evidence. The frontier route requires both a deployment meeting the utility target and the LB2, LB3, and LB4 operands. None of the seven depth-coded capability-removal instances in Section 5 completes that route. The four changes below address the upper side.
Report coverage as an architecture fact.
asks which paths to the anchored outcome must enter the mediation domain. It is answered by enumerating the deployment’s paths and showing the domain on each, not by any accuracy figure. A path that resets scoped state belongs in that enumeration.
Report continuation, or the value of accuracy is undetermined.
asks what adverse value remains available after the event succeeds: content already released, actions the policy still permits, state accumulated before the check, and further attempts within the horizon. Because Equation 24 scales every accuracy gain by , an unreported leaves the value of a reported improvement undetermined rather than merely unstated.
Report the dependence premise, not the layer count.
Error rates across layers can be multiplied only if each layer’s conditional bound remains valid after every preceding interaction history, at the deployed history grain. An evaluation should report this dependence condition and the grain at which it holds, not only the number of layers.
For transferable artifacts, declare a covering outer class.
Released weights admit no mediation event. Their upper-bound route therefore requires a uniform value bound over a declared tampering class that contains . Testing more individual attacks cannot supply this uniform bound: each test adds only a lower-bound witness.
7.2 What the Schedule Implies for Design
Each gate requires a different kind of intervention. Lowering is a statistical problem in the classifier [89, 75, 45]. Raising is a routing problem in the architecture [47, 43, 182]. Driving to zero is a structural problem: it requires that a covered success leave no adverse continuation available. Detection accuracy cannot deliver that at any value.
For an external-world outcome where the dual-use floor binds, Section 5 leaves a bounded as the reportable target. Choosing that value is a deployment decision rather than an evaluation result: it requires the constraints in and the utility requirement .
A certificate binds to its anchor and to the premises of the row that produced it. A second deployment inherits it only by preserving those coordinates and premises, or through a containment argument covering the new scope.
The coded set is a saturation sample rather than a census.
8 Conclusion
A local safeguard score establishes that a control worked in a specified test, while deployment safety concerns the harmful assistance the guarded service still supplies. A single successful attack can establish that harmful assistance remains. Showing how little remains requires a conjunction: coverage of the relevant attack paths, conditional failure on those paths, and the continuation available after a safeguard succeeds. One coded claim supplies that conjunction and rules out the worst case. When a quantity yields no tighter bound, what is missing is a different quantity, not a more precise version of the same one. A safeguard can work exactly as tested while the safety of the deployed system remains unresolved.
Ethical Considerations
This paper systematizes results that are already published. Every quantity it recomputes is one its source already reported, the evidence base is the public literature, and the full-text coding was executed by the two independent model channels on public papers. The work involved no attack execution, no deployed system, and no human subjects.
The stakeholders are the authors who report safeguard evaluations, the reviewers and evaluators who read those reports, and the operators who act on them. The schedule states which reported quantity bounds deployment risk and which returns no bound, so what it supplies is a standard for evaluation rather than a capability for attack. An adversary gains nothing the cited papers do not already state. On that basis we consider publication justified.
Open Science
The supplementary materials document the wide-coded claim instances in this paper. The first file states the common full-text review task applied by both independent model channels. The second lists every source-channel record of the wide-coded portion of the full-text coding pass with its final eligibility status, claim-instance status, and instance count. The third and fourth compile the final per-source assessments from the Fable and GPT channels, respectively, including the source version used, eligibility decision, extracted instances, evidence assessments and locators, residual conclusion, claim verdict, and recorded boundary cases. The supplement also includes the coding instrument applied by both channels; the assessment records cite its numbered rules and global conventions. The depth-coded subset is documented in Appendix C rather than here: Section C.3 gives its aggregates and Sections C.5 and C.6 reproduce two case records.
The supplementary materials are available in the project repository.
Together, these files document the executed review decisions behind the aggregate results. Third-party full texts are not redistributed. The assessment files contain source identifiers, locators, and the excerpts needed to substantiate a judgment.
References
- [1] A. Ablove, S. Chandrashekaran, X. Qiang, and R. Ensafi (2026) Characterizing the implementation of censorship policies in chinese LLM services. In 33rd Annual Network and Distributed System Security Symposium, NDSS 2026, San Diego, California, USA, February 23-27, 2026, Cited by: §B.3.
- [2] M. Andriushchenko, F. Croce, and N. Flammarion (2025) Jailbreaking leading safety-aligned LLMs with simple adaptive attacks. In The Thirteenth International Conference on Learning Representations, Cited by: §1.
- [3] Anthropic (2025) Detecting and countering misuse of AI: august 2025. Note: Anthropic Threat Intelligence ReportPublished August 27, 2025; accessed August 20, 2026 External Links: Link Cited by: §1.
- [4] Anthropic (2026) Claude Fable 5 and Claude Mythos 5. Note: Model announcement, June 9, 2026. Accessed 2026-08-20 External Links: Link Cited by: §B.3.
- [5] Anthropic (2026) How we contain Claude across products. Note: Anthropic EngineeringPublished May 25, 2026; accessed August 14, 2026 External Links: Link Cited by: §1.
- [6] A. Arditi, O. Obeso, A. Syed, D. Paleka, N. Panickssery, W. Gurnee, and N. Nanda (2024) Refusal in language models is mediated by a single direction. In Advances in Neural Information Processing Systems, Vol. 37. External Links: Document Cited by: §1, §6.
- [7] A. Athalye, N. Carlini, and D. Wagner (2018) Obfuscated gradients give a false sense of security: circumventing defenses to adversarial examples. In Proceedings of the 35th International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, Vol. 80, pp. 274–283. Cited by: §3.3.1.
- [8] Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran-Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosuite, L. Lovitt, M. Sellitto, N. Elhage, N. Schiefer, N. Mercado, N. DasSarma, R. Lasenby, R. Larson, S. Ringer, S. Johnston, S. Kravec, S. El Showk, S. Fort, T. Lanham, T. Telleen-Lawton, T. Conerly, T. Henighan, T. Hume, S. R. Bowman, Z. Hatfield-Dodds, B. Mann, D. Amodei, N. Joseph, S. McCandlish, T. Brown, and J. Kaplan (2022) Constitutional ai: harmlessness from ai feedback. External Links: 2212.08073 Cited by: §1.
- [9] Y. F. Bakman, D. N. Yaldiz, S. Kang, T. Zhang, B. Buyukates, S. Avestimehr, and S. P. Karimireddy (2025) Reconsidering LLM uncertainty estimation methods in the wild. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 29531–29556. External Links: Document Cited by: §B.3.
- [10] A. R. Basani and X. Zhang (2025) GASP: efficient black-box generation of adversarial suffixes for jailbreaking llms. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
- [11] S. Belkadi, L. Ren, N. Micheletti, L. Han, and G. Nenadic (2025) Generating synthetic free-text medical records with low re-identification risk using masked language modeling. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 4: Student Research Workshop, Albuquerque, NM, USA, April 30 - May 1, 2025, A. Ebrahimi, S. Haider, E. Liu, S. Haider, M. L. Pacheco, and S. Wein (Eds.), pp. 200–206. External Links: Document Cited by: §B.3.
- [12] R. Bhardwaj, D. A. Do, and S. Poria (2024) Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), Volume 1: Long Papers, Bangkok, Thailand, pp. 14138–14149. External Links: Document Cited by: §B.3, §5.2.
- [13] B. Bi, S. Huang, Y. Wang, T. Yang, Z. Zhang, H. Huang, L. Mei, J. Fang, Z. Li, F. Wei, W. Deng, F. Sun, Q. Zhang, and S. Liu (2025) Context-dpo: aligning language models for context-faithfulness. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Findings of ACL, Vol. ACL 2025, pp. 10280–10300. External Links: Document Cited by: §B.3.
- [14] T. Bi, C. Ye, Z. Yang, Z. Zhou, C. Tang, Z. Tao, J. Zhang, K. Wang, L. Zhou, Y. Yang, and T. Yu (2026) On the feasibility of using multimodal llms to execute AR social engineering attacks. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 38252–38260. External Links: Document Cited by: §B.3.
- [15] J. Binkowski, D. Janiak, A. Sawczyn, B. Gabrys, and T. Kajdanowicz (2025) Hallucination detection in llms using spectral features of attention maps. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 24354–24385. External Links: Document Cited by: §B.3.
- [16] D. Bowen, B. Murphy, W. Cai, D. Khachaturov, A. Gleave, and K. Pelrine (2025) Scaling trends for data poisoning in llms. In Thirty-Ninth AAAI Conference on Artificial Intelligence, Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence, Fifteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2025, Philadelphia, PA, USA, February 25 - March 4, 2025, T. Walsh, J. Shah, and Z. Kolter (Eds.), pp. 27206–27214. External Links: Document Cited by: §B.3.
- [17] L. Bürger, F. A. Hamprecht, and B. Nadler (2024) Truth is universal: robust detection of lies in llms. In Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang (Eds.), Cited by: §B.3.
- [18] Y. Cai, R. Gu, J. Li, X. Huang, J. Chen, X. Gu, and M. Huang (2025) MHALO: evaluating mllms as fine-grained hallucination detectors. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Findings of ACL, Vol. ACL 2025, pp. 9197–9222. External Links: Document Cited by: §B.3.
- [19] Y. Cao, B. Cao, and J. Chen (2024) Stealthy and persistent unalignment on large language models via backdoor injections. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024, K. Duh, H. Gómez-Adorno, and S. Bethard (Eds.), pp. 4920–4935. External Links: Document Cited by: §B.3.
- [20] N. Carlini and D. Wagner (2017) Adversarial examples are not easily detected: bypassing ten detection methods. In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security (AISec ’17), pp. 3–14. External Links: Document Cited by: §3.1.
- [21] S. Casper, C. Ezell, C. Siegmann, N. Kolt, T. L. Curtis, B. Bucknall, A. Haupt, K. Wei, J. Scheurer, M. Hobbhahn, L. Sharkey, S. Krishna, M. Von Hagen, S. Alberti, A. Chan, Q. Sun, M. Gerovitch, D. Bau, M. Tegmark, D. Krueger, and D. Hadfield-Menell (2024) Black-box access is insufficient for rigorous ai audits. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pp. 2254–2272. External Links: Document Cited by: §1, §6, §7.1.
- [22] N. Chakraborty, J. Pohovey, M. Ornik, and K. R. Driggs-Campbell (2026) Characterizing the robustness of black-box LLM planners under perturbed observations with adaptive stress testing. In Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 39445–39475. External Links: Document Cited by: §B.3.
- [23] A. Chandler, D. Surve, and H. Su (2024) Detecting errors through ensembling prompts (DEEP): an end-to-end LLM framework for detecting factual errors. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), pp. 13120–13133. External Links: Document Cited by: §B.3.
- [24] T. Chang, T. Schnabel, A. Swaminathan, and J. Wiens (2026) A course correction in steerability evaluation: revealing miscalibration and side effects in llms. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 37259–37267. External Links: Document Cited by: §B.3.
- [25] P. Chao, E. Debenedetti, A. Robey, M. Andriushchenko, F. Croce, V. Sehwag, E. Dobriban, N. Flammarion, G. J. Pappas, F. Tramèr, H. Hassani, and E. Wong (2024) JailbreakBench: an open robustness benchmark for jailbreaking large language models. In Advances in Neural Information Processing Systems, Vol. 37. External Links: Document Cited by: §1.
- [26] C. H. Chen, H. Huang, and H. Chen (2025) Self-augmented preference alignment for sycophancy reduction in llms. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 12379–12391. External Links: Document Cited by: §B.3.
- [27] G. Chen, Z. Qin, M. Yang, Y. Zhou, T. Fan, T. Du, and Z. Xu (2024) Unveiling the vulnerability of private fine-tuning in split-based frameworks for large language models: A bidirectionally enhanced attack. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, CCS 2024, Salt Lake City, UT, USA, October 14-18, 2024, B. Luo, X. Liao, J. Xu, E. Kirda, and D. Lie (Eds.), pp. 2904–2918. External Links: Document Cited by: §B.3.
- [28] H. Chen and S. Goldfarb-Tarrant (2025) Safer or luckier? llms as safety evaluators are not robust to artifacts. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 19750–19766. External Links: Document Cited by: §B.3.
- [29] M. Chen, Y. Cao, Y. Zhang, and C. Lu (2024) Quantifying and mitigating unimodal biases in multimodal large language models: A causal perspective. In Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Findings of ACL, Vol. EMNLP 2024, pp. 16449–16469. External Links: Document Cited by: §B.3.
- [30] X. Chen, H. Wen, S. Nag, C. Luo, Q. Yin, R. Li, Z. Li, and W. Wang (2024) IterAlign: iterative constitutional alignment of large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024, K. Duh, H. Gómez-Adorno, and S. Bethard (Eds.), pp. 1423–1433. External Links: Document Cited by: §B.3.
- [31] X. Chen, J. Zhao, Y. He, Y. Xun, X. Liu, Y. Li, H. Zhou, W. Cai, Z. Shi, Y. Yuan, T. Zhang, C. Zhang, and X. Li (2025) TeleAI-Safety: a comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations. arXiv preprint arXiv:2512.05485. External Links: 2512.05485, Document Cited by: §6.
- [32] Y. Chen, N. Joshi, Y. Chen, M. Andriushchenko, R. Angell, and H. He (2026) Monitoring decomposition attacks in LLMs with lightweight sequential monitors. In The Fourteenth International Conference on Learning Representations, Cited by: §5.3.
- [33] Z. Chen, S. Shen, G. Shen, G. Zhi, X. Chen, and Y. Lin (2024) Towards tool use alignment of large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), pp. 1382–1400. External Links: Document Cited by: §B.3.
- [34] Z. Chen, H. Lin, K. Li, Z. Luo, Z. Ye, G. Chen, Z. Huang, and J. Ma (2025) AdamMeme: adaptively probe the reasoning capacity of multimodal large language models on harmfulness. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 4234–4253. External Links: Document Cited by: §B.3.
- [35] J. Cheng, T. Su, J. Yuan, G. He, J. Liu, X. Tao, J. Xie, and H. Li (2025) Chain-of-thought prompting obscures hallucination cues in large language models: an empirical evaluation. In Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 1272–1305. External Links: Document Cited by: §B.3.
- [36] X. Cheng, R. Chen, H. Zan, Y. Jia, and M. Peng (2025) BiasFilter: an inference-time debiasing framework for large language models. In Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 15187–15205. External Links: Document Cited by: §B.3.
- [37] Y. Cheng, V. S. Sadasivan, M. Saberi, S. Saha, and S. Feizi (2025) Adversarial paraphrasing: A universal attack for humanizing ai-generated text. CoRR abs/2506.07001. External Links: Document, 2506.07001 Cited by: §B.3.
- [38] J. Chu, Y. Liu, Z. Yang, X. Shen, M. Backes, and Y. Zhang (2025) JailbreakRadar: comprehensive assessment of jailbreak attacks against LLMs. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), Volume 1: Long Papers, pp. 21538–21566. External Links: Document Cited by: §6.
- [39] Z. Chu, Y. Wang, L. Li, Z. Wang, Z. Qin, and K. Ren (2024) A causal explainable guardrails for large language models. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, CCS 2024, Salt Lake City, UT, USA, October 14-18, 2024, B. Luo, X. Liao, J. Xu, E. Kirda, and D. Lie (Eds.), pp. 1136–1150. External Links: Document Cited by: §B.3.
- [40] J. Clymer, J. Weinbaum, R. Kirk, K. Mai, S. Zhang, and X. Davies (2025) An example safety case for safeguards against misuse. Note: arXiv preprint External Links: 2505.18003, Document Cited by: §1, §6.
- [41] S. Cohen, R. Bitton, and B. Nassi (2025) Here comes the AI worm: preventing the propagation of adversarial self-replicating prompts within genai ecosystems. In Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security, CCS 2025, Taipei, Taiwan, October 13-17, 2025, C. Huang, J. Chen, S. Shieh, D. Lie, and V. Cortier (Eds.), pp. 3975–3989. External Links: Document Cited by: §B.3.
- [42] P. H. B. Correia, R. W. Achjian, D. E. G. C. de Oliveira, Y. A. Maria, V. T. Hayashi, M. Lopes, C. C. Miers, and M. A. Simplicio (2026) A systematic literature review on LLM defenses against prompt injection and jailbreaking: expanding NIST taxonomy. arXiv preprint arXiv:2601.22240. External Links: 2601.22240, Document Cited by: §6.
- [43] M. Costa, B. Köpf, A. Kolluri, A. Paverd, M. Russinovich, A. Salem, S. Tople, L. Wutschitz, and S. Zanella-Béguelin (2025) Securing AI agents with information-flow control. arXiv preprint arXiv:2505.23643. External Links: Document Cited by: §C.5, §1, §5.6, §5.6, §7.2.
- [44] T. Coste, U. Anwar, R. Kirk, and D. Krueger (2023) Reward model ensembles help mitigate overoptimization. CoRR abs/2310.02743. External Links: Document, 2310.02743 Cited by: §B.3.
- [45] H. Cunningham, J. Wei, Z. Wang, A. Persic, A. Peng, J. Abderrachid, R. Agarwal, B. Chen, A. Cohen, A. Dau, A. Dimitriev, R. Gilson, L. Howard, Y. Hua, J. Kaplan, J. Leike, M. Lin, C. Liu, V. Mikulik, R. Mittapalli, C. O’Hara, J. Pan, N. Saxena, A. Silverstein, Y. Song, X. Yu, G. Zhou, E. Perez, and M. Sharma (2026) Constitutional classifiers++: efficient production-grade defenses against universal jailbreaks. Note: arXiv preprint External Links: 2601.04603, Document Cited by: §1, §7.2.
- [46] J. Dai, X. Pan, R. Sun, J. Ji, X. Xu, M. Liu, Y. Wang, and Y. Yang (2024) Safe RLHF: safe reinforcement learning from human feedback. In The Twelfth International Conference on Learning Representations, Cited by: §1.
- [47] E. Debenedetti, I. Shumailov, T. Fan, J. Hayes, N. Carlini, D. Fabian, C. Kern, C. Shi, A. Terzis, and F. Tramèr (2025) Defeating prompt injections by design. Note: arXiv preprint External Links: 2503.18813, Document Cited by: §1, §5.6, §7.2.
- [48] E. Debenedetti, J. Zhang, M. Balunovic, L. Beurer-Kellner, M. Fischer, and F. Tramèr (2024) AgentDojo: a dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. In Advances in Neural Information Processing Systems 37, Note: Datasets and Benchmarks Track External Links: Document Cited by: §1.
- [49] A. Delaval, S. Yang, H. Wang, H. Qiu, and J. Lu (2026) TOXIFRENCH: benchmarking and enhancing language models via cot fine-tuning for french toxicity detection. In Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 21354–21375. External Links: Document Cited by: §B.3.
- [50] G. Deng, Y. Liu, Y. Li, K. Wang, Y. Zhang, Z. Li, H. Wang, T. Zhang, and Y. Liu (2024) MASTERKEY: automated jailbreaking of large language model chatbots. In 31st Annual Network and Distributed System Security Symposium, NDSS 2024, San Diego, California, USA, February 26 - March 1, 2024, Cited by: §B.3.
- [51] Y. Deng, W. Zhang, S. J. Pan, and L. Bing (2024) Multilingual jailbreak challenges in large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024, Cited by: §B.3.
- [52] C. Ding, J. Wu, Y. Yuan, J. Lu, K. Zhang, A. Su, X. Wang, and X. He (2025) Unified parameter-efficient unlearning for llms. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
- [53] Y. Dong, R. Mu, G. Jin, Y. Qi, J. Hu, X. Zhao, J. Meng, W. Ruan, and X. Huang (2024) Position: building guardrails for large language models requires systematic design. In Proceedings of the 41st International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, Vol. 235, pp. 11375–11394. Cited by: §6.
- [54] Y. Dong, R. Mu, Y. Zhang, S. Sun, T. Zhang, C. Wu, G. Jin, Y. Qi, J. Hu, J. Meng, S. Bensalem, and X. Huang (2025) Safeguarding large language models: a survey. Artificial Intelligence Review 58 (12), pp. 382. External Links: Document Cited by: §6.
- [55] Y. R. Dong, H. Lin, M. Belkin, R. Huerta, and I. Vulic (2025) UNDIAL: self-distillation with adjusted logits for robust unlearning in large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, L. Chiruzzo, A. Ritter, and L. Wang (Eds.), pp. 8827–8840. External Links: Document Cited by: §B.3.
- [56] Z. Dong, Z. Zhou, C. Yang, J. Shao, and Y. Qiao (2024) Attacks, defenses and evaluations for LLM conversation safety: a survey. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Mexico City, Mexico, pp. 6734–6747. External Links: Document Cited by: §1, §6.
- [57] Y. Du, S. Zhao, D. Zhao, M. Ma, Y. Chen, L. Huo, Q. Yang, D. Xu, and B. Qin (2024) MoGU: A framework for enhancing safety of llms while preserving their usability. In Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang (Eds.), Cited by: §B.3.
- [58] A. Dutta, R. Magu, S. Kim, S. Yoon, M. D. Choudhury, and A. R. KhudaBukhsh (2026) Auditing LLM responses to harmful stereotypes targeting mental health groups. In Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 40435–40452. External Links: Document Cited by: §B.3.
- [59] X. Fang, Z. Tian, Z. Huang, Z. Pan, Z. Wen, X. Wang, Q. Fang, and D. Li (2026) Knowledge injection exists in moe? exploring expert-aware contrast decoding in moe for mitigating llms’ hallucinations. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 39326–39343. External Links: Document Cited by: §B.3.
- [60] R. Feng, A. Hooda, N. Mangaokar, K. Fawaz, S. Jha, and A. Prakash (2023) Stateful defenses for machine learning models are not yet secure against black-box attacks. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, pp. 786–800. External Links: Document Cited by: §5.3.
- [61] G. Filandrianos, A. Dimitriou, M. Lymperaiou, K. Thomas, and G. Stamou (2025) Bias beware: the impact of cognitive biases on LLM-driven product recommendations. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 22397–22426. External Links: Document, ISBN 979-8-89176-332-6 Cited by: §B.3.
- [62] J. Fonseca, A. Bell, and J. Stoyanovich (2025) SAFENUDGE: safeguarding large language models in real-time with tunable safety-performance trade-offs. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 19955–19969. External Links: Document Cited by: §B.3.
- [63] B. Formento, C. Foo, and S. Ng (2025) Confidence elicitation: A new attack vector for large language models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
- [64] H. Fu, W. Peng, Y. Zhou, J. Wu, J. Wen, and Y. Xue (2026) Inhibitory attacks on backdoor-based fingerprinting for large language models. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 26246–26266. External Links: Document Cited by: §B.3.
- [65] H. Geng, Y. Huang, L. Lai, Q. Du, H. Chu, Z. He, J. Hu, and X. Tao (2026) ProMedical: hierarchical fine-grained criteria modeling for medical LLM alignment via explicit injection. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 36955–36994. External Links: Document Cited by: §B.3.
- [66] K. Gligorić, M. Cheng, L. Zheng, E. Durmus, and D. Jurafsky (2024) NLP systems that can’t tell use from mention censor counterspeech, but teaching the distinction helps. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), K. Duh, H. Gomez, and S. Bethard (Eds.), Mexico City, Mexico, pp. 5942–5959. External Links: Document Cited by: §B.3.
- [67] X. Gong, M. Li, Y. Zhang, F. Ran, C. Chen, Y. Chen, Q. Wang, and K. Lam (2025) PAPILLON: efficient and stealthy fuzz testing-powered jailbreaks for llms. In 34th USENIX Security Symposium, USENIX Security 2025, Seattle, WA, USA, August 13-15, 2025, L. Bauer and G. Pellegrino (Eds.), pp. 2401–2420. Cited by: §B.3.
- [68] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz (2023) Not what you’ve signed up for: compromising real-world llm-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, pp. 79–90. External Links: Document Cited by: §5.6.
- [69] T. Gu, Z. Wang, K. Huang, Y. Yao, X. Zhang, Y. Yang, and X. Chen (2025) Invisible entropy: towards safe and efficient low-entropy LLM watermarking. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 6716–6733. External Links: Document Cited by: §B.3.
- [70] T. Gu, Z. Zhou, K. Huang, D. Liang, Y. Wang, H. Zhao, Y. Yao, X. Qiao, K. Wang, Y. Yang, Y. Teng, Y. Qiao, and Y. Wang (2024) MLLMGuard: A multi-dimensional safety evaluation suite for multimodal large language models. In Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang (Eds.), Cited by: §B.3.
- [71] P. Guo, A. Syed, A. Sheshadri, A. Ewart, and G. K. Dziugaite (2024) Mechanistic unlearning: robust knowledge unlearning and editing via mechanistic localization. CoRR abs/2410.12949. External Links: Document, 2410.12949 Cited by: §B.3.
- [72] Z. Guo, Y. Shi, W. Meng, C. Gong, C. Wei, and W. Chen (2025) Be cautious when merging unfamiliar llms: A phishing model capable of stealing privacy. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Findings of ACL, Vol. ACL 2025, pp. 13852–13871. External Links: Document Cited by: §B.3.
- [73] S. Gupta, V. Shrivastava, A. Deshpande, A. Kalyan, P. Clark, A. Sabharwal, and T. Khot (2023) Bias runs deep: implicit reasoning biases in persona-assigned llms. CoRR abs/2311.04892. External Links: Document, 2311.04892 Cited by: §B.3.
- [74] T. Hagendorff, E. Derner, and N. Oliver (2026) Large reasoning models are autonomous jailbreak agents. Nature Communications 17 (1), pp. 1435. External Links: Document, 2508.04039 Cited by: §1.
- [75] S. Han, K. Rao, A. Ettinger, L. Jiang, B. Y. Lin, N. Lambert, Y. Choi, and N. Dziri (2024) WildGuard: open one-stop moderation tools for safety risks, jailbreaks, and refusals of LLMs. In Advances in Neural Information Processing Systems, Vol. 37. External Links: Document Cited by: §1, §7.2.
- [76] K. D. Hayes, M. Goldblum, V. Sehwag, G. Somepalli, A. Panda, and T. Goldstein (2025) FineGRAIN: evaluating failure modes of text-to-image models with vision language model judges. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
- [77] J. He, W. Jiang, G. Hou, W. Fan, R. Zhang, and H. Li (2025) Watch out for your guidance on generation! exploring conditional backdoor attacks against large language models. In Thirty-Ninth AAAI Conference on Artificial Intelligence, Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence, Fifteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2025, Philadelphia, PA, USA, February 25 - March 4, 2025, T. Walsh, J. Shah, and Z. Kolter (Eds.), pp. 26220–26228. External Links: Document Cited by: §B.3.
- [78] H. Hong, S. Wu, S. Feng, N. Naderloui, S. Yan, J. Zhang, A. Arastehfard, H. Huang, and Y. Hong (2025) SoK: systematizing LLM prompt security: taxonomies, datasets, and unified evaluation of attacks and defenses. arXiv preprint arXiv:2510.15476. External Links: 2510.15476, Document Cited by: §1, §6.
- [79] W. Hou, H. Tu, Y. Wang, Y. Zhang, Y. Liu, D. Zhu, L. Gao, and B. Zhou (2026) Beyond single-view detection: a dual-space reasoning framework for interpretable harmful meme understanding. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 10526–10544. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §B.3.
- [80] J. Hu, X. Huang, Y. Sun, Y. Dong, and X. Huang (2026) Lying with truths: open-channel multi-agent collusion for belief manipulation via generative montage. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 5979–5996. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §B.3.
- [81] Z. Hu, L. Shen, Z. Wang, Y. Wei, and D. Tao (2025) Adaptive defense against harmful fine-tuning for large language models via bayesian data scheduler. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
- [82] K. Huang, X. Liu, Q. Guo, T. Sun, J. Sun, Y. Wang, Z. Zhou, Y. Wang, Y. Teng, X. Qiu, Y. Wang, and D. Lin (2024) Flames: benchmarking value alignment of llms in chinese. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024, K. Duh, H. Gómez-Adorno, and S. Bethard (Eds.), pp. 4551–4591. External Links: Document Cited by: §B.3.
- [83] L. Huang, X. Jiang, Z. Wang, W. Mo, X. Xiao, Y. Yin, B. Han, and F. Zheng (2026) Transferability of adversarial attacks in video-based mllms: A cross-modal image-to-video approach. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 5067–5075. External Links: Document Cited by: §B.3.
- [84] Q. Huang, J. Zhang, J. Wu, Y. Li, W. Zhang, Y. Rong, J. Yao, S. Zhang, and X. Jia (2026) JailMeter: an evidence-based evaluation framework for jailbreak attacks on large language models. In Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 16006–16029. External Links: Document Cited by: §B.3.
- [85] R. Huang, X. Wang, Z. Li, D. Wu, and S. Wang (2025) GuidedBench: measuring and mitigating the evaluation discrepancies of in-the-wild LLM jailbreak methods. arXiv preprint arXiv:2502.16903. External Links: 2502.16903, Document Cited by: §6.
- [86] Y. Huang, L. Zhang, and C. Wang (2026) How do llms "trust" unknown knowledge? an unknown knowledge based jailbreak attack. In Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 37105–37124. External Links: Document Cited by: §B.3.
- [87] Y. Huang, H. Wang, X. Bai, J. Wang, J. Liu, Z. Wang, W. Ni, S. Wang, and T. Qi (2026) Robust membership inference for large language models under adversarial generative corruption. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 39531–39547. External Links: Document Cited by: §B.3.
- [88] Y. Huang, L. Sun, H. Wang, S. Wu, Q. Zhang, Y. Li, C. Gao, Y. Huang, W. Lyu, Y. Zhang, X. Li, Z. Liu, Y. Liu, Y. Wang, Z. Zhang, B. Vidgen, B. Kailkhura, C. Xiong, C. Xiao, C. Li, E. Xing, F. Huang, H. Liu, H. Ji, H. Wang, H. Zhang, H. Yao, M. Kellis, M. Zitnik, M. Jiang, M. Bansal, J. Zou, J. Pei, J. Liu, J. Gao, J. Han, J. Zhao, J. Tang, J. Wang, J. Vanschoren, J. Mitchell, K. Shu, K. Xu, K. Chang, L. He, L. Huang, M. Backes, N. Z. Gong, P. S. Yu, P. Chen, Q. Gu, R. Xu, R. Ying, S. Ji, S. Jana, T. Chen, T. Liu, T. Zhou, W. Wang, X. Li, X. Zhang, X. Wang, X. Xie, X. Chen, X. Wang, Y. Liu, Y. Ye, Y. Cao, Y. Chen, and Y. Zhao (2024) TrustLLM: trustworthiness in large language models. CoRR abs/2401.05561. External Links: Document, 2401.05561 Cited by: §B.3.
- [89] H. Inan, K. Upasani, J. Chi, R. Rungta, K. Iyer, Y. Mao, M. Tontchev, Q. Hu, B. Fuller, D. Testuggine, and M. Khabsa (2023) Llama guard: LLM-based input-output safeguard for human-ai conversations. External Links: 2312.06674 Cited by: §B.3, §1, §7.2.
- [90] C. Isch and G. Jennings (2026) Narrative license and model sycophancy in LLM summaries of scientific work. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 16418–16432. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §B.3.
- [91] D. Jain, D. Hartmann, and C. Li (2026) Adaptive adversaries: a multi-turn, multi-llm benchmark for LLM agent security. Note: arXiv preprint External Links: 2607.18063, Document Cited by: §1.
- [92] P. Jaiswal, A. Pratap, S. Saraswati, H. Kasyap, and S. Tripathy (2026) Analysis of LLMs against prompt injection and jailbreak attacks. In Proceedings of the Workshop on Privacy in Large Language Models (LLM) and Natural Language Processing (NLP) 2026, External Links: Document, 2602.22242 Cited by: §6.
- [93] J. Jeon, J. Oh, H. Lee, and B. Lee (2025) Iterative prompt refinement for safer text-to-image generation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 18080–18096. External Links: Document Cited by: §B.3.
- [94] S. Jeoung, Y. Ge, and J. Diesner (2023) StereoMap: quantifying the awareness of human-like stereotypes in large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 12236–12256. External Links: Document Cited by: §B.3.
- [95] P. Jha, R. Jain, K. Mandal, A. Chadha, S. Saha, and P. Bhattacharyya (2024) MemeGuard: an LLM and vlm-based framework for advancing content moderation via meme intervention. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), pp. 8084–8104. External Links: Document Cited by: §B.3.
- [96] L. Jiang, Y. Li, X. Zhang, Y. Ding, and L. Pan (2025) SceneJailEval: a scenario-adaptive multi-dimensional framework for jailbreak evaluation. CoRR abs/2508.06194. External Links: Document, 2508.06194 Cited by: §B.3.
- [97] P. Jiang, X. Lyu, Y. Li, and J. Ma (2025) Backdoor token unlearning: exposing and defending backdoors in pretrained language models. In Thirty-Ninth AAAI Conference on Artificial Intelligence, Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence, Fifteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2025, Philadelphia, PA, USA, February 25 - March 4, 2025, T. Walsh, J. Shah, and Z. Kolter (Eds.), pp. 24285–24293. External Links: Document Cited by: §B.3.
- [98] M. Kang, Z. Chen, and B. Li (2025) C-safegen: certified safe LLM generation with claim-based streaming guardrails. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
- [99] S. Kim and G. Lee (2026) Merging triggers, breaking backdoors: defensive poisoning for instruction-tuned language models. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 24269–24287. External Links: Document Cited by: §B.3.
- [100] S. Kim, Y. Lee, Y. Song, and K. Lee (2025) What really matters in many-shot attacks? an empirical study of long-context vulnerabilities in llms. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 2043–2063. External Links: Document Cited by: §B.3.
- [101] S. Kim, S. Yun, H. Lee, M. Gubri, S. Yoon, and S. J. Oh (2023) ProPILE: probing privacy leakage in large language models. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Cited by: §B.3.
- [102] C. Ko, P. Chen, P. Das, Y. Mroueh, S. Dan, G. Kollias, S. Chaudhury, T. Pedapati, and L. Daniel (2025) Large language models can become strong self-detoxifiers. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
- [103] S. Kusaka, K. Saito, M. Kudo, T. Tanabe, A. Wachi, and Y. Akimoto (2026) Cost-minimized label-flipping poisoning attack to LLM alignment. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 37538–37546. External Links: Document Cited by: §B.3.
- [104] F. Le, W. He, C. Cao, D. Liang, and Z. Cui (2025) DualCnst: enhancing zero-shot out-of-distribution detection via text-image consistency in vision-language models. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
- [105] D. Lee, J. Jang, J. Jeong, and H. Yu (2025) Are vision-language models safe in the wild? A meme-based benchmark study. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 30545–30588. External Links: Document Cited by: §B.3.
- [106] C. T. Leong, Y. Cheng, J. Wang, J. Wang, and W. Li (2023) Self-detoxifying language models via toxification reversal. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), pp. 4433–4449. External Links: Document Cited by: §B.3.
- [107] H. Li, X. Liu, N. Zhang, and C. Xiao (2025) PIGuard: prompt injection guardrail via mitigating overdefense for free. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 30420–30437. External Links: Document Cited by: §B.3.
- [108] H. Li, J. Ye, J. Wu, T. Yan, C. Wang, and Z. Li (2025) JailPO: A novel black-box jailbreak framework via preference optimization against aligned llms. In Thirty-Ninth AAAI Conference on Artificial Intelligence, Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence, Fifteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2025, Philadelphia, PA, USA, February 25 - March 4, 2025, T. Walsh, J. Shah, and Z. Kolter (Eds.), pp. 27419–27427. External Links: Document Cited by: §B.3.
- [109] H. Li, S. Shan, E. Wenger, J. Zhang, H. Zheng, and B. Y. Zhao (2022) Blacklight: scalable defense for neural networks against Query-Based Black-Box attacks. In 31st USENIX Security Symposium (USENIX Security 22), Boston, MA, pp. 2117–2134. External Links: ISBN 978-1-939133-31-1 Cited by: §5.3.
- [110] K. Li, O. Patel, F. B. Viégas, H. Pfister, and M. Wattenberg (2023) Inference-time intervention: eliciting truthful answers from a language model. CoRR abs/2306.03341. External Links: Document, 2306.03341 Cited by: §B.3.
- [111] K. Li, J. Shi, Y. Xiao, M. Jiang, J. Sun, Y. Wu, D. Fu, S. Xia, X. Cai, T. Xu, W. Si, W. Li, D. Wang, and P. Liu (2026) AgencyBench: benchmarking the frontiers of autonomous agents in 1M-token real-world contexts. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 7422–7440. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §B.3.
- [112] K. Li, L. M. Po, H. Yang, X. Xu, K. Liu, and Y. Zhao (2025) AesBiasBench: evaluating bias and alignment in multimodal language models for personalized image aesthetic assessment. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 7607–7620. External Links: Document Cited by: §B.3.
- [113] L. Li, Y. Liu, D. He, and Y. Li (2025) One model transfer to all: on robust jailbreak prompts generation against llms. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
- [114] N. Li, A. Pan, A. Gopal, S. Yue, D. Berrios, A. Gatti, J. D. Li, A. Dombrowski, S. Goel, G. Mukobi, N. Helm-Burger, R. Lababidi, L. Justen, A. B. Liu, M. Chen, I. Barrass, O. Zhang, X. Zhu, R. Tamirisa, B. Bharathi, A. Herbert-Voss, C. B. Breuer, A. Zou, M. Mazeika, Z. Wang, P. Oswal, W. Lin, A. A. Hunt, J. Tienken-Harder, K. Y. Shih, K. Talley, J. Guan, I. Steneker, D. Campbell, B. Jokubaitis, S. Basart, S. Fitz, P. Kumaraguru, K. K. Karmakar, U. Tupakula, V. Varadharajan, Y. Shoshitaishvili, J. Ba, K. M. Esvelt, A. Wang, and D. Hendrycks (2024) The WMDP benchmark: measuring and reducing malicious use with unlearning. In Proceedings of the 41st International Conference on Machine Learning, Vol. 235, pp. 28525–28550. Cited by: §1, §5.4.
- [115] Q. Li, T. Luo, X. Zhang, Y. Xie, Z. Shen, L. Zhang, Y. Jin, H. Peng, X. Zhao, X. Zhu, and J. Yin (2025) CoreGuard: safeguarding foundational capabilities of llms against model stealing in edge deployment. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
- [116] R. Li, J. Long, M. Qi, H. Xia, L. Sha, P. Wang, and Z. Sui (2025) Towards harmonized uncertainty estimation for large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 22938–22953. External Links: Document Cited by: §B.3.
- [117] S. Li, J. Sun, G. Zheng, X. Fan, Y. Shen, Y. Lu, Z. Xi, Y. Yang, W. Tan, T. Ji, T. Gui, Q. Zhang, and X. Huang (2025) Mitigating object hallucinations in mllms via multi-frequency perturbations. In Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 1230–1247. External Links: Document Cited by: §B.3.
- [118] Y. Li, M. Du, X. Wang, and Y. Wang (2023) Prompt tuning pushes farther, contrastive learning pulls closer: A two-stage approach to mitigate social biases. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, A. Rogers, J. L. Boyd-Graber, and N. Okazaki (Eds.), pp. 14254–14267. External Links: Document Cited by: §B.3.
- [119] Z. Li, P. Chen, and T. Ho (2025) Retention score: quantifying jailbreak risks for vision language models. In Thirty-Ninth AAAI Conference on Artificial Intelligence, Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence, Fifteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2025, Philadelphia, PA, USA, February 25 - March 4, 2025, T. Walsh, J. Shah, and Z. Kolter (Eds.), pp. 27446–27454. External Links: Document Cited by: §B.3.
- [120] J. Liang, Z. Wang, S. Hong, S. Ji, and T. Wang (2025) Watermark under fire: A robustness evaluation of LLM watermarking. In Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 21050–21074. External Links: Document Cited by: §B.3.
- [121] Z. Liang, L. Yu, S. Zhang, Q. Ye, and H. Hu (2026) How much do large language model cheat on evaluation? benchmarking overestimation under the one-time-pad-based framework. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 37636–37644. External Links: Document Cited by: §B.3.
- [122] Z. Lin, Z. Wang, Y. Tong, Y. Wang, Y. Guo, Y. Wang, and J. Shang (2023) ToxicChat: unveiling hidden challenges of toxicity detection in real-world user-AI conversation. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 4694–4702. External Links: Document Cited by: §B.3.
- [123] F. Liu, Y. Zhang, J. Luo, J. Dai, T. Chen, L. Yuan, Z. Yu, Y. Shi, K. Li, C. Zhou, H. Chen, and M. Yang (2025) Make agent defeat agent: automatic detection of taint-style vulnerabilities in llm-based agents. In 34th USENIX Security Symposium, USENIX Security 2025, Seattle, WA, USA, August 13-15, 2025, L. Bauer and G. Pellegrino (Eds.), pp. 3767–3786. Cited by: §B.3.
- [124] H. Liu, Y. Xie, Y. Wang, and M. Shieh (2024) Advancing adversarial suffix transfer learning on aligned large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), pp. 7213–7224. External Links: Document Cited by: §B.3.
- [125] X. Liu, P. Li, G. E. Suh, Y. Vorobeychik, Z. Mao, S. Jha, P. McDaniel, H. Sun, B. Li, and C. Xiao (2025) AutoDAN-turbo: A lifelong agent for strategy self-exploration to jailbreak llms. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
- [126] X. Liu, S. Liang, M. Han, Y. Luo, A. Liu, X. Cai, Z. He, and D. Tao (2025) ELBA-bench: an efficient learning backdoor attacks benchmark for large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 17928–17947. External Links: Document Cited by: §B.3.
- [127] Y. Liu, Y. Liu, X. Chen, P. Chen, D. Zan, M. Kan, and T. Ho (2024) The devil is in the neurons: interpreting and mitigating social biases in language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024, Cited by: §B.3.
- [128] J. Lu, J. Liu, X. Zheng, M. Yang, J. Wang, P. Wang, and Y. Zhang (2026) MHB: medical hallucination benchmark for large language models in complex clinical tasks. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 38971–38978. External Links: Document Cited by: §B.3.
- [129] X. Lu, F. Brahman, P. West, J. Jung, K. Chandu, A. Ravichander, P. Ammanabrolu, L. Jiang, S. Ramnath, N. Dziri, J. Fisher, B. Lin, S. Hallinan, L. Qin, X. Ren, S. Welleck, and Y. Choi (2023) Inference-time policy adapters (IPA): tailoring extreme-scale LMs without fine-tuning. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 6863–6883. External Links: Document Cited by: §B.3.
- [130] Y. Lu, J. Li, Y. Zhou, Y. Zhang, W. Wang, X. Li, M. Zhang, F. Liu, J. Yu, and M. Zhang (2025) Adaptive detoxification: safeguarding general capabilities of llms through toxicity-aware knowledge editing. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Findings of ACL, Vol. ACL 2025, pp. 19744–19758. External Links: Document Cited by: §B.3.
- [131] K. Lukošiūtė and A. Swanda (2025) LLM cyber evaluations don’t capture real-world risk. arXiv preprint arXiv:2502.00072. External Links: Document Cited by: §1.
- [132] W. Luo, S. Dai, X. Liu, S. Banerjee, H. Sun, M. Chen, and C. Xiao (2025) AGrail: A lifelong agent guardrail with effective and adaptive safety detection. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 8104–8139. External Links: Document Cited by: §B.3.
- [133] T. S. Luong, T. Le, L. N. Van, and T. H. Nguyen (2024) Realistic evaluation of toxicity in large language models. In Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Findings of ACL, Vol. ACL 2024, pp. 1038–1047. External Links: Document Cited by: §B.3.
- [134] D. T. Mahato (2026) Safeguard-conditioned uplift: measuring utility–risk frontiers for dual-use biology assistants. Note: arXiv preprint External Links: 2607.13039, Document Cited by: §1, §2.3, §6.
- [135] R. Maheshwary, V. Yadav, H. Nguyen, K. Mahajan, and S. T. Madhusudhan (2025) M2Lingual: enhancing multilingual, multi-turn instruction alignment in large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, L. Chiruzzo, A. Ritter, and L. Wang (Eds.), pp. 9676–9713. External Links: Document Cited by: §B.3.
- [136] S. Masud, S. Singh, V. Hangya, A. Fraser, and T. Chakraborty (2024) Hate personified: investigating the role of llms in content moderation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), pp. 15847–15863. External Links: Document Cited by: §B.3.
- [137] M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, D. Forsyth, and D. Hendrycks (2024) HarmBench: a standardized evaluation framework for automated red teaming and robust refusal. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, pp. 35181–35224. Cited by: §1.
- [138] A. McKenzie, U. Pawar, P. Blandfort, W. Bankes, D. Krueger, E. S. Lubana, and D. Krasheninnikov (2025) Detecting high-stakes interactions with activation probes. CoRR abs/2506.10805. External Links: Document, 2506.10805 Cited by: §B.3.
- [139] R. Miao, Y. Liu, Y. Wang, X. Shen, Y. Tan, Y. Dai, S. Pan, and X. Wang (2026) BlindGuard: safeguarding llm-based multi-agent systems under unknown attacks. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 39215–39234. External Links: Document Cited by: §B.3.
- [140] W. J. Mo, Q. Liu, X. Wen, D. Jung, H. Askari, W. Zhou, Z. Zhao, and M. Chen (2026) RedCoder: automated multi-turn red teaming for code llms. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 33140–33155. External Links: Document Cited by: §B.3.
- [141] S. R. Motwani, M. Baranchuk, M. Strohmeier, V. Bolina, P. H. S. Torr, L. Hammond, and C. S. de Witt (2024) Secret collusion among AI agents: multi-agent deception via steganography. CoRR abs/2402.07510. External Links: Document, 2402.07510 Cited by: §B.3.
- [142] R. Movva, P. W. Koh, and E. Pierson (2024) Annotation alignment: comparing LLM and human annotations of conversational safety. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), pp. 9048–9062. External Links: Document Cited by: §B.3.
- [143] M. Nagireddy, L. Chiazor, M. Singh, and I. Baldini (2024) SocialStigmaQA: A benchmark to uncover stigma amplification in generative language models. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2024, February 20-27, 2024, Vancouver, Canada, M. J. Wooldridge, J. G. Dy, and S. Natarajan (Eds.), pp. 21454–21462. External Links: Document Cited by: §B.3.
- [144] G. Nalbandyan, R. Shahbazyan, and E. Bakhturina (2025) SCORE: systematic consistency and robustness evaluation for large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 3: Industry Track, Albuquerque, New Mexico, USA, April 30, 2025, W. Chen, Y. Yang, M. Kachuee, and X. Fu (Eds.), pp. 470–484. External Links: Document Cited by: §B.3.
- [145] M. Nasr, N. Carlini, C. Sitawarin, S. V. Schulhoff, J. Hayes, M. Ilie, J. Pluto, S. Song, H. Chaudhari, I. Shumailov, A. G. Thakurta, K. Y. Xiao, A. Terzis, and F. Tramèr (2026) The attacker moves second: stronger adaptive attacks bypass defenses against LLM jailbreaks and prompt injections. In 35th USENIX Security Symposium (USENIX Security 26), Baltimore, MD. Cited by: §C.6, §1, §5.4.
- [146] M. Nasr, J. Rando, N. Carlini, J. Hayase, M. Jagielski, A. F. Cooper, D. Ippolito, C. A. Choquette-Choo, F. Tramèr, and K. Lee (2025) Scalable extraction of training data from aligned, production language models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
- [147] H. Nghiem, J. Prindle, J. Zhao, and H. Daumé III (2024) “You gotta be a doctor, lin” : an investigation of name-based bias of large language models in employment recommendations. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 7268–7287. External Links: Document Cited by: §B.3.
- [148] G. Niess and R. Kern (2025) Ensemble watermarks for large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 2903–2916. External Links: Document, ISBN 979-8-89176-251-0 Cited by: §B.3.
- [149] S. Oh, K. Lee, S. Park, D. Kim, and H. Kim (2024) Poisoned chatgpt finds work for idle hands: exploring developers’ coding practices with insecure suggestions from poisoned AI models. In IEEE Symposium on Security and Privacy, SP 2024, San Francisco, CA, USA, May 19-23, 2024, pp. 1141–1159. External Links: Document Cited by: §B.3.
- [150] OpenAI (2025) Disrupting malicious uses of AI: an update (october 2025). Note: OpenAI Threat Intelligence ReportPublished October 7, 2025; accessed August 20, 2026 External Links: Link Cited by: §1.
- [151] OpenAI (2026) GPT-5.6 system card. Note: OpenAI Deployment Safety HubPublished July 9, 2026; accessed August 14, 2026 External Links: Link Cited by: §1.
- [152] OpenAI (2026) GPT-5.6. Note: Model suite announcement (Sol, Terra, Luna), July 9, 2026. Accessed 2026-08-20 External Links: Link Cited by: §B.3.
- [153] K. O’Brien, S. Casper, Q. G. Anthony, T. Korbak, R. Kirk, X. Davies, I. Mishra, G. Irving, Y. Gal, and S. Biderman (2025) Deep ignorance: filtering pretraining data builds tamper-resistant safeguards into open-weight LLMs. In BioSafe GenAI Workshop 2025, Cited by: §5.4.
- [154] M. J. Page, J. E. McKenzie, P. M. Bossuyt, I. Boutron, T. C. Hoffmann, C. D. Mulrow, L. Shamseer, J. M. Tetzlaff, E. A. Akl, S. E. Brennan, R. Chou, J. Glanville, J. M. Grimshaw, A. Hróbjartsson, M. M. Lalu, T. Li, E. W. Loder, E. Mayo-Wilson, S. McDonald, L. A. McGuinness, L. A. Stewart, J. Thomas, A. C. Tricco, V. A. Welch, P. Whiting, and D. Moher (2021) The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ 372, pp. n71. External Links: Document Cited by: Table 4.
- [155] L. Pan, A. Liu, S. Huang, Y. Lu, X. Hu, L. Wen, I. King, and P. S. Yu (2025) Can LLM watermarks robustly prevent unauthorized knowledge distillation?. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 13228–13251. External Links: Document Cited by: §B.3.
- [156] X. Pang, X. Hao, S. Guo, Q. Luo, and Z. Wang (2025) ICLScan: detecting backdoors in black-box large language models via targeted in-context illumination. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
- [157] Y. Pang, W. Meng, X. Liao, and T. Wang (2026) Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm. In Network and Distributed System Security Symposium (NDSS), Cited by: §B.3.
- [158] L. H. Park, J. Cho, G. Kim, Y. Yeo, and T. Kwon (2026) Chimera: compositional jailbreak attacks on llms via judgment-driven search over heterogeneous strategies. In Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 33330–33355. External Links: Document Cited by: §B.3.
- [159] S. Park and K. Kim (2025) Measuring and mitigating media outlet name bias in large language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 29778–29797. External Links: Document Cited by: §B.3.
- [160] H. L. Patel, A. Agarwal, A. Das, B. Kumar, S. Panda, P. Pattnayak, T. H. Rafi, T. Kumar, and D. Chae (2025) SweEval: do llms really swear? A safety benchmark for testing limits for enterprise use. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 3: Industry Track, Albuquerque, New Mexico, USA, April 30, 2025, W. Chen, Y. Yang, M. Kachuee, and X. Fu (Eds.), pp. 558–582. External Links: Document Cited by: §B.3.
- [161] C. Pathade (2025) Red teaming the mind of the machine: a systematic evaluation of prompt injection and jailbreak vulnerabilities in LLMs. arXiv preprint arXiv:2505.04806. External Links: 2505.04806, Document Cited by: §6.
- [162] K. Pelrine, A. Imouza, C. Thibault, M. Reksoprodjo, C. Gupta, J. Christoph, J. Godbout, and R. Rabbany (2023) Towards reliable misinformation mitigation: generalization, uncertainty, and GPT-4. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), pp. 6399–6429. External Links: Document Cited by: §B.3.
- [163] D. Peng, Q. Ke, and J. Liu (2024) UPAM: unified prompt attack in text-to-image generation models against both textual filters and visual checkers. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024, R. Salakhutdinov, Z. Kolter, K. A. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (Eds.), Proceedings of Machine Learning Research, Vol. 235, pp. 40200–40214. Cited by: §B.3.
- [164] A. Peppin, A. Reuel, S. Casper, E. Jones, A. Strait, U. Anwar, A. Agrawal, S. Kapoor, O. Koyejo, M. Pellat, R. Bommasani, N. Frosst, and S. Hooker (2025) The reality of ai and biorisk. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, pp. 763–771. External Links: Document Cited by: §1.
- [165] E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving (2022) Red teaming language models with language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y. Goldberg, Z. Kozareva, and Y. Zhang (Eds.), Abu Dhabi, United Arab Emirates, pp. 3419–3448. External Links: Document Cited by: §B.3.
- [166] N. Prakash, Y. W. Jie, A. Abdullah, R. Satapathy, E. Cambria, and R. K. W. Lee (2025) Beyond “I’m sorry, I can’t”: dissecting large language model refusal. arXiv preprint arXiv:2509.09708. External Links: 2509.09708, Document Cited by: §6.
- [167] X. Qi, B. Wei, N. Carlini, Y. Huang, T. Xie, L. He, M. Jagielski, M. Nasr, P. Mittal, and P. Henderson (2025) On evaluating the durability of safeguards for open-weight LLMs. In The Thirteenth International Conference on Learning Representations, Cited by: §1, §6, §7.1.
- [168] Z. Rao, W. Zhu, C. A. Lu, Z. Chen, W. Niu, L. Guan, B. Li, and Z. Xiang (2026) FragFuse: bypassing access control of large language model agents via memory-based query fragmentation and fusion. In 35th USENIX Security Symposium (USENIX Security 26), Baltimore, MD. Cited by: §1.
- [169] M. L. Rethlefsen, S. Kirtley, S. Waffenschmidt, A. P. Ayala, D. Moher, M. J. Page, J. B. Koffel, and PRISMA-S Group (2021) PRISMA-S: an extension to the PRISMA statement for reporting literature searches in systematic reviews. Systematic Reviews 10 (1), pp. 39. External Links: Document, Link Cited by: §B.1.
- [170] L. Richter, X. He, P. Minervini, and M. J. Kusner (2025) An auditing test to detect behavioral shift in language models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
- [171] A. Robey, E. Wong, H. Hassani, and G. J. Pappas (2025) SmoothLLM: defending large language models against jailbreaking attacks. Transactions on Machine Learning Research. Cited by: §5.4.
- [172] P. Röttger, H. Kirk, B. Vidgen, G. Attanasio, F. Bianchi, and D. Hovy (2024) XSTest: a test suite for identifying exaggerated safety behaviours in large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Mexico City, Mexico, pp. 5377–5400. External Links: Document Cited by: §1.
- [173] Y. Ruan, H. Dong, A. Wang, S. Pitis, Y. Zhou, J. Ba, Y. Dubois, C. J. Maddison, and T. Hashimoto (2023) Identifying the risks of LM agents with an lm-emulated sandbox. CoRR abs/2309.15817. External Links: Document, 2309.15817 Cited by: §B.3.
- [174] H. Saffari, M. Shafiei, H. Zhang, L. T. Harris, and N. S. Moosavi (2025) Beyond hate speech: NLP’s challenges and opportunities in uncovering dehumanizing language. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 26965–26980. External Links: Document, ISBN 979-8-89176-332-6 Cited by: §B.3.
- [175] S. Sagar, A. Taparia, and R. Senanayake (2024) Failures are fated, but can be faded: characterizing and mitigating unwanted behaviors in large-scale vision and language models. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024, R. Salakhutdinov, Z. Kolter, K. A. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (Eds.), Proceedings of Machine Learning Research, Vol. 235, pp. 42999–43023. Cited by: §B.3.
- [176] J. H. Saltzer and M. D. Schroeder (1975) The protection of information in computer systems. Proceedings of the IEEE 63 (9), pp. 1278–1308. External Links: Document, Link Cited by: §3.3.3.
- [177] P. Sarkar, S. Ebrahimi, A. Etemad, A. Beirami, S. Ö. Arik, and T. Pfister (2025) Mitigating object hallucination in mllms via data-augmented phrase-level alignment. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
- [178] G. M. Shahariar, Z. A. Nazi, Md. O. H. Bhuiyan, and Z. Shi (2026) PII-visbench: evaluating personally identifiable information safety in vision language models along a continuum of visibility. In Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 10294–10316. External Links: Document Cited by: §B.3.
- [179] R. M. Shahroz, Z. Tan, S. Yun, C. Fleming, and T. Chen (2025) Agents under siege: breaking pragmatic multi-agent LLM systems with optimized prompt attacks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 9661–9674. External Links: Document Cited by: §B.3.
- [180] G. Shen, D. Zhao, L. Feng, X. He, J. Wang, S. Shen, H. Tong, Y. Dong, J. Li, X. Zheng, and Y. Zeng (2025) PandaGuard: systematic evaluation of LLM safety against jailbreaking attacks. arXiv preprint arXiv:2505.13862. External Links: 2505.13862, Document Cited by: §6.
- [181] H. Shen, B. Huang, and X. Wan (2025) Enhancing LLM watermark resilience against both scrubbing and spoofing attacks. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
- [182] T. Shi, J. He, Z. Wang, L. Wu, H. Li, W. Guo, and D. Song (2025) Progent: securing AI agents with privilege control. arXiv preprint arXiv:2504.11703. External Links: Document Cited by: §1, §7.2.
- [183] W. Shi, A. Ajith, M. Xia, Y. Huang, D. Liu, T. Blevins, D. Chen, and L. Zettlemoyer (2024) Detecting pretraining data from large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024, Cited by: §B.3.
- [184] W. M. Si, M. Li, M. Backes, and Y. Zhang (2026) Pruning unsafe tickets: A resource-efficient framework for safer and more robust llms. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 26285–26302. External Links: Document Cited by: §B.3.
- [185] I. Solaiman, M. Brundage, J. Clark, A. Askell, A. Herbert-Voss, J. Wu, A. Radford, G. Krueger, J. W. Kim, S. Kreps, M. McCain, A. Newhouse, J. Blazakis, K. McGuffie, and J. Wang (2019) Release strategies and the social impacts of language models. External Links: 1908.09203, Document Cited by: §1.
- [186] M. Son, J. Jang, and M. Kim (2025) Lightweight query checkpoint: classifying faulty user queries to mitigate hallucinations in large language model question answering. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Findings of ACL, Vol. ACL 2025, pp. 14664–14677. External Links: Document Cited by: §B.3.
- [187] Y. Son, M. Kim, S. Kim, S. Han, J. Kim, D. Jang, Y. Yu, and C. Y. Park (2025) Subtle risks, critical failures: A framework for diagnosing physical safety of llms for embodied decision making. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 25692–25733. External Links: Document Cited by: §B.3.
- [188] M. Spliethöver, T. Knebler, F. Fumagalli, M. Muschalik, B. Hammer, E. Hüllermeier, and H. Wachsmuth (2025) Adaptive prompting: ad-hoc prompt composition for social bias detection. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, L. Chiruzzo, A. Ritter, and L. Wang (Eds.), pp. 2421–2449. External Links: Document Cited by: §B.3.
- [189] R. Tamirisa, B. Bharathi, L. Phan, A. Zhou, A. Gatti, T. Suresh, M. Lin, J. Wang, R. Wang, R. Arel, A. Zou, D. Song, B. Li, D. Hendrycks, and M. Mazeika (2025) Tamper-resistant safeguards for open-weight LLMs. In The Thirteenth International Conference on Learning Representations, Cited by: §5.4.
- [190] F. Tramèr, N. Carlini, W. Brendel, and A. Madry (2020) On adaptive attacks to adversarial example defenses. In Advances in Neural Information Processing Systems 33 (NeurIPS 2020), Cited by: §3.1, §3.3.1.
- [191] I. Uddin and A. Bauer (2026) Conformal LLM routing with distribution-free safety guarantees. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), ACL 2026, San Diego, California, United States, July 2-7, 2026, T. Y. S. S. Santosh, J. D. Rodriguez, and O. de Gibert (Eds.), pp. 791–799. External Links: Document Cited by: §B.3.
- [192] UK AI Safety Institute Safeguards Analysis Team (2025) Principles for evaluating misuse safeguards of frontier AI systems. Technical report UK AI Safety Institute. Note: Published February 4, 2025; organization renamed the AI Security Institute on February 14, 2025; accessed August 14, 2026 External Links: Link Cited by: §1, §6, §7.1.
- [193] R. Uppaal, A. Dey, Y. He, Y. Zhong, and J. Hu (2024) Model editing as a robust and denoised variant of DPO: a case study on toxicity. CoRR abs/2405.13967. External Links: Document, 2405.13967 Cited by: §B.3.
- [194] M. Vaccaro, J. Song, A. Almaatouq, and M. A. Bakker (2026) Evaluating human–ai safety: a framework for measuring harmful capability uplift. Note: arXiv preprint External Links: 2603.26676, Document Cited by: §1, §2.3, §6.
- [195] C. Wang, X. Chen, N. Zhang, B. Tian, H. Xu, S. Deng, and H. Chen (2025) MLLM can see? dynamic correction decoding for hallucination mitigation. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
- [196] H. Wang, Z. Huang, Z. Lin, and T. Liu (2024) NoiseGPT: label noise detection and rectification through probability curvature. In Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang (Eds.), Cited by: §B.3.
- [197] J. G. Wang, J. Wang, M. Li, and S. Neel (2026) CheckMIABench: firm foundations for membership inference attacks on language models. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 364–370. External Links: Document Cited by: §B.3.
- [198] J. Wang, Z. Xu, D. Jin, X. Yang, and T. Li (2026) Accommodate knowledge conflicts in retrieval-augmented llms: towards robust response generation in the wild. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 33530–33538. External Links: Document Cited by: §B.3.
- [199] P. Wang, B. Dong, Y. Cai, Z. Zhang, J. Liu, H. Xue, Y. Wu, Y. Zhang, and Z. Zhang (2025) Game of arrows: on the (in-)security of weight obfuscation for on-device tee-shielded LLM partition algorithms. In 34th USENIX Security Symposium, USENIX Security 2025, Seattle, WA, USA, August 13-15, 2025, L. Bauer and G. Pellegrino (Eds.), pp. 279–298. Cited by: §B.3.
- [200] X. Wang, A. Balashankar, and V. Chandrasekaran (2026) Systematic scaling analysis of jailbreak attacks in large language models. arXiv preprint arXiv:2603.11149. External Links: 2603.11149, Document Cited by: §5.4.
- [201] X. Wang, Z. Li, B. Wang, Y. Hu, and D. Zou (2025) Model unlearning via sparse autoencoder subspace guided projections. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 26530–26546. External Links: Document, ISBN 979-8-89176-332-6 Cited by: §B.3.
- [202] X. Wang, S. Zhu, and X. Cheng (2025) Speculative safety-aware decoding. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 12827–12841. External Links: Document, ISBN 979-8-89176-332-6 Cited by: §B.3.
- [203] X. Wang, Z. Ji, W. Wang, Z. Li, D. Wu, and S. Wang (2026) SoK: evaluating jailbreak guardrails for large language models. In 2026 IEEE Symposium on Security and Privacy (S&P), pp. 39–58. External Links: Document, 2506.10597 Cited by: §1, §6.
- [204] X. Wang, D. Wu, Z. Ji, Z. Li, P. Ma, S. Wang, Y. Li, Y. Liu, N. Liu, and J. Rahmel (2025) SelfDefend: llms can defend themselves against jailbreaking in a practical manner. In 34th USENIX Security Symposium, USENIX Security 2025, Seattle, WA, USA, August 13-15, 2025, L. Bauer and G. Pellegrino (Eds.), pp. 2441–2460. Cited by: §B.3.
- [205] Y. Wang, T. Huang, L. Shen, H. Yao, H. Luo, R. Liu, N. Tan, J. Huang, and D. Tao (2025) Panacea: mitigating harmful fine-tuning for large language models via post-fine-tuning perturbation. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
- [206] Y. Wang, M. Zhang, J. Sun, C. Wang, M. Yang, H. Xue, J. Tao, R. Duan, and J. Liu (2025) Mirage in the eyes: hallucination attack on multi-modal large language models with only attention sink. In 34th USENIX Security Symposium, USENIX Security 2025, Seattle, WA, USA, August 13-15, 2025, L. Bauer and G. Pellegrino (Eds.), pp. 3707–3726. Cited by: §B.3.
- [207] Y. Wang, R. Wu, Z. He, X. Chen, and J. J. McAuley (2024) Large scale knowledge washing. CoRR abs/2405.16720. External Links: Document, 2405.16720 Cited by: §B.3.
- [208] Z. Wang, Z. Wu, X. Guan, M. Thaler, A. S. Koshiyama, S. Lu, S. Beepath, E. E. Jr., and M. Pérez-Ortiz (2024) JobFair: A framework for benchmarking gender hiring bias in large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Findings of ACL, Vol. EMNLP 2024, pp. 3227–3246. External Links: Document Cited by: §B.3.
- [209] Z. Wang, D. Anshumaan, A. Hooda, Y. Chen, and S. Jha (2025) Functional homotopy: smoothing discrete optimization via continuous parameters for LLM jailbreak attacks. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
- [210] A. Wei, N. Haghtalab, and J. Steinhardt (2023) Jailbroken: how does LLM safety training fail?. In Advances in Neural Information Processing Systems 36, External Links: Document Cited by: §1.
- [211] C. Wu, Z. R. Tam, C. Lin, Y. V. Chen, S. Sun, and H. Lee (2025) Mitigating forgetting in LLM fine-tuning via low-perplexity token learning. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
- [212] F. Wu, E. Cecchetti, and C. Xiao (2024) System-level defense against indirect prompt injection attacks: an information flow control perspective. arXiv preprint arXiv:2409.19091. External Links: Document Cited by: §5.6.
- [213] L. Wu, M. Wang, Z. Xu, T. Cao, N. Oo, B. Hooi, and S. Deng (2025) Automating steering for safe multimodal large language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 792–814. External Links: Document Cited by: §B.3.
- [214] M. Wu, J. Ji, O. Huang, J. Li, Y. Wu, X. Sun, and R. Ji (2024) Evaluating and analyzing relationship hallucinations in large vision-language models. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024, R. Salakhutdinov, Z. Kolter, K. A. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (Eds.), Proceedings of Machine Learning Research, Vol. 235, pp. 53553–53570. Cited by: §B.3.
- [215] P. Wu, L. Zhu, W. Zhang, and N. Yu (2026) Safeguards based on copyable context cannot provide reliable safety for LLMs. Note: arXiv preprint External Links: 2607.27951, Document Cited by: §3.2.2.
- [216] Y. Wu, R. Wen, C. Cui, M. Backes, and Y. Zhang (2026) InferPilot: autonomous inference attacks against ML services with llm-based agents. In Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 11781–11801. External Links: Document Cited by: §B.3.
- [217] Y. Wu, F. Roesner, T. Kohno, N. Zhang, and U. Iqbal (2025) IsolateGPT: an execution isolation architecture for LLM-based agentic systems. In Network and Distributed System Security Symposium, External Links: Document Cited by: §1.
- [218] W. Xia and Z. Deng (2026) SDA: steering-driven distribution alignment for open llms without fine-tuning. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 34025–34033. External Links: Document Cited by: §B.3.
- [219] T. Xiang, L. Li, W. Li, M. Bai, L. Wei, B. Wang, and N. Garcia (2023) CARE-MI: chinese benchmark for misinformation evaluation in maternity and infant care. CoRR abs/2307.01458. External Links: Document, 2307.01458 Cited by: §B.3.
- [220] T. Xing, J. Li, Y. Du, and X. Hu (2026) Are LLMs reliable rankers? rank manipulation via two-stage token optimization. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 9120–9132. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §B.3.
- [221] F. Xu, H. Hu, C. He, S. Hang, H. Hu, X. Liu, Y. Zhao, Z. Zhou, B. B. Zhu, S. Sun, D. Gu, and S. Wang (2026) SoK: robustness in large language models against jailbreak attacks. In 2026 IEEE Symposium on Security and Privacy (S&P), pp. 118–137. External Links: Document, 2605.05058 Cited by: §1, §6.
- [222] J. Xu, M. D. Ma, F. Wang, C. Xiao, and M. Chen (2024) Instructions as backdoors: backdoor vulnerabilities of instruction tuning for large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024, K. Duh, H. Gómez-Adorno, and S. Bethard (Eds.), pp. 3111–3126. External Links: Document Cited by: §B.3.
- [223] S. Xu, L. Pang, Y. Zhu, H. Shen, and X. Cheng (2025) Cross-modal safety mechanism transfer in large vision-language models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
- [224] Z. Xu, F. Liu, and H. Liu (2024) Bag of tricks: benchmarking of jailbreak attacks on llms. In Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang (Eds.), Cited by: §B.3.
- [225] S. Yi, Y. Liu, Z. Sun, T. Cong, X. He, J. Song, K. Xu, and Q. Li (2024) Jailbreak attacks and defenses against large language models: a survey. arXiv preprint arXiv:2407.04295. Note: preprint; no peer-reviewed venue recorded on arXiv as of 2026-08-19 External Links: 2407.04295 Cited by: §6.
- [226] X. Yu, H. Cheng, X. Liu, D. Roth, and J. Gao (2024) ReEval: automatic hallucination evaluation for retrieval-augmented large language models via transferable adversarial attacks. In Findings of the Association for Computational Linguistics: NAACL 2024, Mexico City, Mexico, June 16-21, 2024, K. Duh, H. Gómez-Adorno, and S. Bethard (Eds.), Findings of ACL, Vol. NAACL 2024, pp. 1333–1351. External Links: Document Cited by: §B.3.
- [227] Z. Yue, H. Zeng, Y. Lu, L. Shang, Y. Zhang, and D. Wang (2024) Evidence-driven retrieval augmented response generation for online misinformation. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), K. Duh, H. Gomez, and S. Bethard (Eds.), Mexico City, Mexico, pp. 5628–5643. External Links: Document Cited by: §B.3.
- [228] Q. Zhan, R. Fang, H. S. Panchal, and D. Kang (2025) Adaptive attacks break defenses against indirect prompt injection attacks on LLM agents. In Findings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, New Mexico, pp. 7116–7132. External Links: Document Cited by: §1.
- [229] X. Zhan, J. C. Carrillo, W. Seymour, and J. Such (2025) Malicious llm-based conversational AI makes users reveal personal information. In 34th USENIX Security Symposium, USENIX Security 2025, Seattle, WA, USA, August 13-15, 2025, L. Bauer and G. Pellegrino (Eds.), pp. 61–80. Cited by: §B.3, §1.
- [230] B. Zhang and G. Ren (2025) Challenges and remedies of domain-specific classifiers as LLM guardrails: self-harm as a case study. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 3: Industry Track, Albuquerque, New Mexico, USA, April 30, 2025, W. Chen, Y. Yang, M. Kachuee, and X. Fu (Eds.), pp. 173–182. External Links: Document Cited by: §B.3.
- [231] B. Zhang, H. Liu, Q. Tian, S. Chen, Z. Wang, and Q. Qi (2026) Towards trustworthy smart contract synthesis: a multi-agent framework with lean-based verification. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 39548–39582. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §B.3.
- [232] C. Zhang, J. X. Morris, and V. Shmatikov (2024) Extracting prompts by inverting LLM outputs. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 14753–14777. External Links: Document Cited by: §B.3.
- [233] M. Zhang, K. K. Goh, P. Zhang, J. Sun, L. X. Rose, and H. Zhang (2025) LLMScan: causal scan for LLM misbehavior detection. In Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025, A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu (Eds.), Proceedings of Machine Learning Research, Vol. 267. Cited by: §B.3.
- [234] S. Zhang, H. Li, and R. Ji (2024) Code membership inference for detecting unauthorized data use in code pre-trained language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Findings of ACL, Vol. EMNLP 2024, pp. 10593–10603. External Links: Document Cited by: §B.3.
- [235] S. Zhang, Y. Zhai, K. Guo, H. Hu, S. Guo, Z. Fang, L. Zhao, C. Shen, C. Wang, and Q. Wang (2025) JBShield: defending large language models from jailbreak attacks through activated concept analysis and manipulation. In 34th USENIX Security Symposium, USENIX Security 2025, Seattle, WA, USA, August 13-15, 2025, L. Bauer and G. Pellegrino (Eds.), pp. 8215–8234. Cited by: §B.3.
- [236] T. Zhang, Z. Xi, T. Wang, P. Mitra, and J. Chen (2024) PromptFix: few-shot backdoor removal via adversarial prompt tuning. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024, K. Duh, H. Gómez-Adorno, and S. Bethard (Eds.), pp. 3212–3225. External Links: Document Cited by: §B.3.
- [237] X. Zhang, X. Wang, Y. Lu, J. Wang, Z. Ye, M. Bao, P. Yan, and X. Su (2026) TrendFact: a benchmark towards hotspot perception in automatic fact-checking. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 26494–26513. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §B.3.
- [238] Y. Zhang, T. Liu, Z. Zhao, G. Meng, and K. Chen (2026) Bleeding pathways: vanishing discriminability in LLM hidden states fuels jailbreak attacks. In 33rd Annual Network and Distributed System Security Symposium, NDSS 2026, San Diego, California, USA, February 23-27, 2026, Cited by: §B.3.
- [239] Y. Zhang, R. Xie, J. Chen, X. Sun, Z. Kang, and Y. Wang (2025) QAVA: query-agnostic visual attack to large vision-language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, L. Chiruzzo, A. Ritter, and L. Wang (Eds.), pp. 10205–10218. External Links: Document Cited by: §B.3.
- [240] Z. Zhang, L. Lei, L. Wu, R. Sun, Y. Huang, C. Long, X. Liu, X. Lei, J. Tang, and M. Huang (2024) SafetyBench: evaluating the safety of large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), pp. 15537–15553. External Links: Document Cited by: §B.3.
- [241] Z. Zhang, H. Zhang, W. Li, Q. Zhang, J. Dong, Y. Tong, and Z. Zheng (2026) FedSEA-llama: A secure, efficient and adaptive federated splitting framework for large language models. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 28680–28688. External Links: Document Cited by: §B.3.
- [242] A. Zhao, Q. Xu, M. Lin, S. Wang, Y. Liu, Z. Zheng, and G. Huang (2025) DiveR-ct: diversity-enhanced red teaming large language model assistants with relaxing constraints. In Thirty-Ninth AAAI Conference on Artificial Intelligence, Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence, Fifteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2025, Philadelphia, PA, USA, February 25 - March 4, 2025, T. Walsh, J. Shah, and Z. Kolter (Eds.), pp. 26021–26030. External Links: Document Cited by: §B.3.
- [243] C. Zhao, X. Wang, P. Zhao, Y. Huang, J. Lu, Z. Liu, Q. Lin, S. Rajmohan, and D. Zhang (2026) Gradient-guided multi-judge prompt optimization. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 23744–23773. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §B.3.
- [244] S. Zhao, M. Jia, A. T. Luu, F. Pan, and J. Wen (2024) Universal vulnerabilities in large language models: backdoor attacks for in-context learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), pp. 11507–11522. External Links: Document Cited by: §B.3.
- [245] S. Zhao, R. Brekelmans, A. Makhzani, and R. B. Grosse (2024) Probabilistic inference in language models via twisted sequential monte carlo. CoRR abs/2404.17546. External Links: Document, 2404.17546 Cited by: §B.3.
- [246] W. Zhao, D. Ben-Levi, W. Hao, J. Yang, and C. Mao (2025) Diversity helps jailbreak large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, L. Chiruzzo, A. Ritter, and L. Wang (Eds.), pp. 4647–4680. External Links: Document Cited by: §B.3.
- [247] W. Zhao, J. Guo, Y. Hu, Y. Deng, A. Zhang, X. Sui, X. Han, Y. Zhao, B. Qin, T. Chua, and T. Liu (2025) AdaSteer: your aligned LLM is inherently an adaptive jailbreak defender. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 24559–24577. External Links: Document Cited by: §B.3.
- [248] Y. Zhao, W. Zheng, T. Cai, X. L. Do, K. Kawaguchi, A. Goyal, and M. Shieh (2024) Accelerating greedy coordinate gradient and general prompt optimization via probe sampling. CoRR abs/2403.01251. External Links: Document, 2403.01251 Cited by: §B.3.
- [249] K. Zheng, J. Chen, Y. Yan, X. Zou, H. Zhou, and X. Hu (2025) Reefknot: A comprehensive benchmark for relation hallucination evaluation, analysis and mitigation in multimodal large language models. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Findings of ACL, Vol. ACL 2025, pp. 6193–6212. External Links: Document Cited by: §B.3.
- [250] K. Zhou, C. Liu, X. Zhao, A. Compalas, D. Song, and X. E. Wang (2024) Multimodal situational safety. CoRR abs/2410.06172. External Links: Document, 2410.06172 Cited by: §B.3.
- [251] X. Zhou, M. Zhang, Z. Lee, W. Ye, and S. Zhang (2025) HaDeMiF: hallucination detection and mitigation in large language models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
- [252] Z. Zhou, J. Liu, Z. Dong, J. Liu, C. Yang, W. Ouyang, and Y. Qiao (2024) Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), Volume 1: Long Papers, Bangkok, Thailand, pp. 15810–15830. External Links: Document Cited by: §B.3, §5.2.
- [253] Z. Zhou, Q. Wang, M. Jin, J. Yao, J. Ye, W. Liu, W. Wang, X. Huang, and K. Huang (2024) MathAttack: attacking large language models towards math solving ability. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2024, February 20-27, 2024, Vancouver, Canada, M. J. Wooldridge, J. G. Dy, and S. Natarajan (Eds.), pp. 19750–19758. External Links: Document Cited by: §B.3.
- [254] Y. Zhuang, K. Guo, J. Wang, Y. Jing, X. Xu, W. Yi, M. Yang, B. Zhao, and H. Hu (2025) I know what you meme! understanding and detecting harmful memes with multimodal large language models. In 32nd Annual Network and Distributed System Security Symposium, NDSS 2025, San Diego, California, USA, February 24-28, 2025, Cited by: §B.3.
- [255] A. Zou, L. Phan, J. Wang, D. Duenas, M. Lin, M. Andriushchenko, R. Wang, Z. Kolter, M. Fredrikson, and D. Hendrycks (2024) Improving alignment and robustness with circuit breakers. In Advances in Neural Information Processing Systems, Vol. 37. External Links: Document Cited by: §C.6, §1, §5.4.
Appendix A Proofs and Attaining Constructions
This appendix proves the results in Section 3 and gives the constructions that attain them, which makes each row of Table 1 sharp. Probability laws are defined on the declared trajectory space; all spaces are standard Borel and all declared events are measurable. Finite or countable structure is assumed only where a result states it. Write for the pushforward of under the declared value-relevant projection and .
Two elementary facts are used throughout. For any measurable with range in ,
| (35) |
and projecting a law cannot increase total variation.
A.1 Witness and Reachable-Set Bounds
Proposition 1 holds because , so the supremum defining is at least the value at that law. Proposition 7 holds because a supremum over a superset of dominates the supremum over , with equality when the two classes coincide.
The corresponding negative entry in Table 1 is equally immediate. A finite set of tested modifications generates laws lying inside . For any candidate ceiling , consider any that also contains a strategy whose law places all mass on a trace with and which agrees with on every coordinate the enumeration observed. The enumeration is unchanged and . Hence no supremum over an inner sample bounds above at any sample size, unless the sample is shown to exhaust , in which case it is no longer a sample and the equality case of Proposition 7 applies.
A.2 Frontier Bounds
Proof of Proposition 4.
The inclusion implies that every law feasible for the realizable frontier is feasible for the outer frontier. Minimizing the same objective over the superset gives the frontier order. The restriction statement follows because minimizing over a subset cannot lower the optimum. This argument applies to either law set without identifying the two sets. ∎
Under the uniform dual-use relation of Equation 18, every feasible law obeys
and minimization proves the stated floor for both frontiers. The uniform premise cannot be replaced on an infinite space by pointwise positivity. For , , and , every trace has positive adverse value but the frontier at has infimum zero and no optimizer.
The floor is attained. On a two-point space with and , the feasible law placing mass on has legitimate value exactly and adverse value exactly .
When a bound on accompanies the dual-use premise, the pointwise relation gives more than the composed frontier route. Fix , write , and for any let be the common part of the two laws, the largest positive measure dominated by both; its total mass is . Because , because inherits the relation , and because with ,
Selecting with arbitrarily near and using prove the endpoint of Table 1, . It is attained: moving mass from to in the construction above yields a law at total variation exactly from the benign law whose adverse value is exactly , so no larger lower bound follows from , , and a bound on .
A.3 Value-Relevant Simulation
Proof of Proposition 2.
Fix a committed policy . For every , select such that
Applying Equation 35 to gives
Letting vanish and using proves the bound. ∎
The coefficient of is sharp. On a two-point space, move mass from a point with to one with . The expectation difference and the total variation are both , so no coefficient smaller than one is valid.
Proof of Theorem 5.
Fix a policy . Its benign law belongs to and is feasible at , hence
Combining these inequalities with Proposition 2 and monotonicity of the positive part proves Equation 17. If and the realizable-frontier positive part is active, that bound rearranges to . If it is inactive, the same inequality already holds because .
If every feasible policy has , taking the infimum over gives . Now suppose induces a realizable-frontier optimizer and satisfies the value equalizer condition . Then
which supplies the reverse inequality for and proves equality. No equality of malicious laws was used. ∎
Proof of Proposition 3.
Condition on the two processes having identical histories before update . Maximal coupling of their next evidence kernels makes the updates disagree with probability at most . Couple each deployment-controlled or remaining random draw through its common conditional kernel; such a draw cannot create the first disagreement while its inputs agree. The chain rule over successive updates therefore leaves the processes coupled with probability at least , so a first disagreement occurs with probability at most . The coupling characterization of total variation then bounds by the same quantity. Projecting a law cannot increase total variation, and the selected malicious law is one candidate in the infimum defining . These two facts prove Equation 12. ∎
The bound is attained: let every deployment transition be trivial, let the malicious process deviate to a fresh symbol independently with probability exactly at each update, and let contain that strategy alone; each conditional distance is then exactly and the identity projection gives .
The premise is conditional by necessity, which is the content of the corresponding negative row. Split the histories into a set after which the benign and malicious evidence kernels disagree maximally, so that each admits an event of benign conditional probability one and malicious conditional probability zero, and a complement after which the two kernels coincide. Let the benign law give and , so a suite drawn from the benign workload reports marginal disagreement rate exactly . Let contain a single strategy , and let that strategy drive the process into with probability one. Let be the event that the history lies in , or lies in and the next update falls in . Then while , so the two trajectory laws are mutually singular and the identity projection gives however small is. Proposition 3 asks for a distance that holds after every shared history, and the smallest such value here is one; is the benign average of those distances, and the attacker selects the histories the average treats as rare. An average over histories therefore constrains neither the kernel after any particular history nor .
Proof of Proposition 6.
Each of , , and is a functional of the benign law and of the set alone: integrates against the former, maximizes over the latter, and minimizes total variation between the former and members of the latter. Two deployments agreeing on both objects therefore agree on all three. ∎
A.4 Success-Region Sharpness
Proof of Proposition 8.
For any , the reachable envelope gives , so holds -almost surely on . Boundedness of and with give
Maximization over proves the bound. For sharpness, a two-region law placing mass on value inside and the remaining mass on value one outside attains it in a model consistent with these two parameters. No smaller uniform bound therefore follows from and alone. ∎
A.5 Closed-Mediation Tightness
Proof of Theorem 9.
For any , boundedness of and the continuation premise give
The robust contract and coverage definition imply
Substitution and maximization over prove Equation 22.
The bound is sharp from exactly these three quantities. On a three-atom space, assign mass to with value , mass to with value one, and mass to with value one. The coverage, contract, and continuation inequalities all hold with equality, and the expectation is . A singleton attacker class containing this law therefore attains the bound.
The sharp value is below one exactly when , which is equivalent to the three strict conditions in Equation 23. This proves both the if-and-only-if statement and the impossibility of a stronger uniform bound. ∎
Equation 24 follows by direct subtraction of the two sharp bounds, so the stated contribution of detection accuracy is exact rather than an estimate. Corollary 10 follows by setting in the tight construction. The construction may take to be the whole space and empty: the local fact then holds without error on every invocation, yet on the success region and the deployment value remains one.
At the strict boundary, requires , and since each factor lies in this holds exactly when , , and . If any one of the three fails, the three-atom construction has strictly positive value, so no zero-residual certificate follows from these quantities.
A.6 Robust Trusted State
This subsection states the robust frontier summarized in Section 3.2.4 and proves its bound. Write , let be the benign marginal of trusted state and the set of maliciously reachable state marginals, both common to all feasible policies, and let be the set of kernels realized by some as the benign conditional law of given . The robust frontier is
| (36) |
Theorem 12(Robust conditional simulation).
If for every feasible and every the attacker can induce , then . Moreover, with ,
| (37) |
The definition of supplies both inclusions the argument uses. On a standard Borel trace space, every feasible has a benign disintegration , and by definition. Conversely, every is the benign disintegration of some policy in , so its benign law belongs to .
Proof of Theorem 12.
Fix a feasible and disintegrate its benign law as . By the robust conditional copy premise, for every the law is attacker-reachable. Hence
because and the legitimate value is at least . Taking the infimum over feasible proves the lower bound. If a robust optimizer has a realizing policy for which no attacker law yields value exceeding , the bound is attained at .
For the data-processing bound, fix a feasible . Its benign law is feasible for the realizable frontier, so . For every , the common kernel and Equation 35 give
Choosing a sequence of malicious marginals whose distance tends to and taking the infimum over proves Equation 37. ∎
The bound depends on rather than an average, which is the content of the corresponding table row. Let acquisition succeed on a single allowed path with probability one and fail on other paths. The average success rate over paths tends to zero as grows, while and therefore the bound are unchanged.
A.7 Composition Bounds and Tightness
Proof of Proposition 11.
For a serial path, the probability chain rule gives
The history-uniform premise bounds every factor, proving the product bound without independence. Independent layers attain it.
With marginal bounds only, for each , and taking the smallest bound proves the minimum bound. It is attained when all events share a common subevent of probability , so a stack of any depth whose layers fail together supports no bound better than its single strongest layer.
The union bound proves the alternative-path bound, and disjoint events attain it until total mass reaches one. For retries, the chain rule gives
Each factor lies between and . Multiplication and subtraction from one prove Equation 27; sequential Bernoulli trials attain both endpoints. ∎
For additive outcomes, suppose , , and , with a separate target for each coordinate. Write for the frontier of coordinate . Every feasible joint law then has marginal value at least , so linearity gives the sum as a lower bound on the joint frontier. If the marginal optimizers have an admissible joint coupling, that coupling attains the sum. Correlation is unrestricted; the conclusion fails for a union or intersection payoff because such a payoff is not additive.
A.8 Combining Candidates
If and are supported endpoints for the same under one anchor, then and , since each inequality holds separately. Neither combination need be sharp under the joint premises: a law attaining one candidate can violate another’s. If , no law satisfies all the stated premises simultaneously, so at least one premise fails under the fixed anchor and neither endpoint is available until the conflict is resolved.
A candidate established for a subclass bounds only the supremum over restricted to . Since gives , a family of subclass upper bounds combines into a bound for only when the subclasses cover , and the combined value is then their maximum rather than any single one.
Appendix B Complete Search and Screening Protocol
This appendix states the protocol under which the corpus is assembled and every denominator in the paper is produced. The supplementary materials provide the common full-text review task, the final source-status list, and both channels’ per-source assessment records.
The search target follows Section 4.4. We look for sources that make, or directly adjudicate, at least one locatable claim that some deployed or deployable intervention reduces LLM-enabled misuse or a closely specified adverse deployment outcome. Attack, red-teaming, and adaptive-evaluation sources are therefore included: they assess deployment-safety claims on the lower-bound side of the schedule.
The review task was frozen before full-text coding began.
B.1 Information Sources and Exact Queries
Three sources are used, reported in the PRISMA-S search-reporting structure [169]. S1 arXiv, via the API, restricted to cs.CR, cs.CL, cs.AI, cs.LG, cs.SE, cs.MA, and stat.ML. S2 the ACL Anthology, screened offline by regular expression over a frozen full BibTeX dump. S3 dblp venue enumeration for the major security and machine-learning venues and their workshop volumes, with a deliberately liberal title screen. The arXiv and dblp windows run from 2023-01-01 to the freeze date.
Query terms.
The query structure is derived from the language of the claim. A deployment-safety claim in the sense of Section 4.4 has a recurring surface form: [intervention] reduces / prevents / bounds [adverse outcome] for [deployed system] under [attacker or usage conditions]. Three facets are read off that frame: facet A, the deployed system (LLM, foundation model, LLM agent, and variants); facet B, the claim verb or evidential noun (defense, mitigate, safeguard, guarantee, plus adjudication nouns such as benchmark and red teaming); and facet C, the adverse outcome (misuse, jailbreak, prompt injection, exfiltration, and related terms). A record is a candidate when it matches A AND (B OR C). A fourth facet D, the mechanism lexicon, is added by union solely to recover sources matching neither B nor C. The source-specific implementations apply this facet logic under the category and date window stated above.
B.2 Deduplication and Version Families
A version family is the set of records sharing a work identity, keyed by a DOI-to-arXiv cross-link, or by normalized title similarity together with an author-set Jaccard overlap above a fixed threshold, or by an explicit “extended version of” statement. Candidate pairs may be generated automatically, including by model, but every merge is human-confirmed and logged with both identifiers.
Preprint, conference, and journal extension form one family. A workshop paper and its later full version form one family when the contribution is the same, and two families when the later claim set differs materially. System cards and policies are never merged. Each dated release is a separate source because changes in deployment claims between releases are part of the analysis.
The coded claim comes by default from the latest peer-reviewed version available at freeze. If a claim exists only in the preprint and is weakened or removed in the camera-ready, a separate instance is bound to the preprint version and flagged version_divergence. Divergence between versions is a reportable finding. Every quoted or coded claim records the identifier, version label, date, and a page or section locator.
B.3 Screening, Eligibility, and Machine Assistance
The executed sequence has five ordered stages (Figure 1): keyword-based identification; the hard authority gate; independent title-and-abstract screening, with advancement requiring include from both channels; human quality spot checks, with any detected quality problem triggering a rerun of the preceding screening step; and randomized full-text sampling and analysis. Version-family deduplication reconciles records between the authority gate and screening but does not add another eligibility criterion.
Title-and-abstract screening.
Screening is conservative: a record advances to the full-text pool only when both channels independently record include. Any disagreement, or an unsure verdict from either channel, keeps the record out. Human spot checks assess the quality and rule compliance of the channel outputs. They do not replace the both-include rule with per-record adjudication. When a spot check detects a quality problem, the title-and-abstract screening step is run again before the pool is finalized.
Full-text analysis layers.
Within the randomized full-text sample, Layer 1 decides whether a source enters the corpus; Layer 2 decides whether a codable claim instance exists.
Layer 1: source level.
A source is included only if I1 and I2 both hold and at least one of I3 or I4 holds. I1 English full text is obtainable by the freeze date. I2 the work concerns an LLM-based, LLM-integrated, or LLM-agentic deployed or deployable system. I3 the work contains at least one locatable sentence that is citable to a section, page, or paragraph and that asserts or directly tests whether an intervention changes an adverse deployment outcome, or demonstrates that such an assertion fails. A pure attack paper satisfies I3 because it adjudicates the existing claim that current deployments resist that attack class. I4 the evidence-apparatus clause admits sources that supply the instruments used to adjudicate such claims, including benchmarks, evaluation-validity critiques, auditing-access analyses, and safety-case templates; these sources are tagged role=apparatus. Methodological sources that describe how we work rather than what we study are tagged role=method and are excluded from every corpus denominator.
Six exclusion codes are used: E1 capability-only reporting with no adverse-outcome claim; E2 normative-only argument with no intervention and no adverse-outcome evidence; E3 non-LLM subject; E4 non-substantive item, with vendor system cards and policies never excluded on length; E5 secondary literature without primary claims; and E6 duplicate within a version family.
Layer 2: claim-instance level and the minimum anchoring threshold.
For each included source, channels attempt to instantiate the anchor of Equation 30. An instance is created if and only if C1 and C2 both hold and at least one of C3 or C4 holds. C1, the intervention is locatable: the source identifies a concrete evaluated policy or configuration and where it acts in the deployment. C2, the adverse outcome is locatable: a named outcome with an attached operationalization, not “unsafe behaviour.” C3, a comparison world is reported: guarded versus unguarded, guarded versus a baseline defense, or before versus after. C4, a formal statement is given: a theorem, invariant, or architectural non-bypassability argument, even without measurement.
Every remaining anchor coordinate the source does not state is recorded unknown. Coders never supply a missing coordinate by inference. A source that passes Layer 1 but fails C1 or C2 is recorded as included, no codable instance and retained: these sources are counted in every corpus denominator and must not be dropped.
Full-text exclusion requires an explicit criterion and a locator.
Machine assistance.
Title-and-abstract screening and sampled full-text analysis are performed by two independent model channels using Claude Fable 5 [4] and GPT-5.6 Sol [152]. At the title-and-abstract stage, advancement is determined mechanically by the intersection of their include decisions, subject to the human quality check and rerun rule above. At full text, each channel applies the same frozen review task, samples the full-text pool independently, and remains blind to the other. No unknown anchor coordinate is filled from model background knowledge.
Flow accounting.
Table 4 reconciles the executed pipeline. Identified records pass a hard authority gate before title-and-abstract screening: G1 admits records peer-reviewed at a fixed list of major security and machine-learning venues, G2 admits the remainder at a citations-per-year threshold checked against an open bibliographic index, and G3 covers first-party vendor material, of which these database pools contain none. Gate-eligible records are consolidated into version families, and only families included by both title-and-abstract channels enter the full-text pool after the human quality check. Full-text analysis then samples that pool in two independently randomized sequences. Twelve papers receive full ten-slot depth coding and 187 receive endpoint-route wide coding. The two strata overlap in one paper, and the coded set holds 198 distinct papers. The wide-coding counts reconcile: confirmed records for each source and channel pair split into corpus and apparatus, and corpus records split into those yielding at least one instance and those with none. The wide-coded set comprises 187 distinct papers [162, 35, 220, 29, 97, 117, 140, 219, 197, 123, 158, 195, 81, 204, 1, 198, 24, 236, 118, 10, 148, 41, 254, 132, 187, 135, 208, 100, 27, 130, 181, 238, 89, 12, 163, 124, 224, 205, 216, 26, 22, 61, 34, 138, 94, 23, 102, 157, 133, 37, 107, 95, 141, 62, 173, 147, 250, 229, 84, 242, 80, 63, 115, 127, 222, 211, 165, 65, 125, 230, 14, 36, 17, 51, 196, 245, 146, 103, 170, 49, 16, 52, 96, 174, 88, 77, 252, 246, 226, 202, 116, 120, 101, 98, 199, 30, 253, 82, 149, 73, 33, 55, 106, 129, 57, 209, 122, 233, 240, 112, 11, 59, 111, 104, 156, 251, 249, 28, 67, 184, 113, 144, 108, 191, 183, 232, 206, 235, 247, 179, 110, 69, 142, 213, 58, 87, 160, 79, 64, 239, 159, 175, 201, 83, 121, 93, 207, 243, 143, 50, 126, 99, 218, 234, 76, 178, 70, 18, 128, 71, 119, 241, 231, 244, 19, 139, 248, 186, 188, 136, 39, 155, 214, 223, 9, 66, 237, 15, 177, 90, 72, 227, 105, 86, 44, 13, 193].
| Stage | Count |
| Identification (full scope) | |
| arXiv, five queries, deduped scope-filtered | 42,029 10,387 |
| ACL Anthology, regex-registered scoped | 3,899 |
| dblp, 11 venues (2023–2026), title prefilter | 4,990 |
| Records entering the authority gate | 19,276 |
| Screening (full scope) | |
| Gate-eligible (G1 5,662; G2 142; G3 0) | 5,804 |
| After version-family dedup | 4,810 |
| Both-channel include | 1,872 |
| Full-text analysis (coded set) | |
| Distinct papers coded | 198 |
| Depth-coded papers | 12 |
| Depth-coded claim instances | 24 |
| Wide-coded papers (one also depth-coded) | 187 |
| Source–channel wide-coding records | 190 |
| Excluded at Layer 1 | 17 |
| Included, role=corpus | 104 |
| Corpus records yielding instance | 88 |
| Included, no codable instance (C1 / C2) | 16 |
| Included, role=apparatus | 69 |
| Wide-coded claim instances extracted | 152 |
| Wide-coded instances on channel-overlap papers | 7 |
B.4 Stopping Rule and Truncation Stability
Full-text coding samples the full-text pool rather than coding it exhaustively. Sources are coded in randomized batches, and coding stops when the reported quantities stop moving as batches are added. Saturation is defined on the estimates this paper reports. It is not defined on the supply of new boundary cases: each batch separately logs new proof-relevant anchor coordinate values or slot distinctions, new boundary cases that would require review task revision, and new claim-relation types, and those ledgers continued to record new items through the final batch of both channels. Category novelty and estimate convergence are different quantities, and the reported proportions are the ones the conclusions rest on.
Truncating the coding sequences shows that convergence directly. Dropping the final batch of each channel, then the final two, then the final three, moves the positive-residual share from 108 of 152 to 100 of 141, then 89 of 124, then 80 of 113: 71.1, 70.9, 71.8, and 70.8 percent. The scale for that movement is the estimate’s own sampling error: clustering instances within their source papers (87 clusters, design effect 1.51) gives a standard error of 4.5 points, so truncation moves the share by about a quarter of one standard error. The two channels, drawing separately randomized samples, reach 68 of 93 and 40 of 59, a difference inside that same error. Further batches would move the share within its noise rather than toward a different value.
Appendix C Review Records, Coding Aggregates, and Cases
This appendix documents the two coding strata of one full-text coding pass. The depth-coded records provide the slot-level aggregates. The wide-coded records provide the quantities required by each endpoint route.
C.1 Claim-Instance Record Format
A claim instance is the unit of coding. Each record is a structured document with five blocks. Section 4 defines the slots and states, while the coding instrument in the supplementary materials gives the per-slot rules both channels applied.
The anchor block records the source identifiers, the verbatim anchored claim and its locator, the claim-scope extension flag, the promotion basis, and the coded anchor of Equation 30. It lists one anchor coordinate per row, each with a locator or an explicit absent record, together with the channel-assigned harm_locus and the anchor-completeness count. The case tables below list the nine coordinates of and count completeness over those nine; records add two context rows, the outside-resource baseline and the external constraint set , so their own counts use eleven as the denominator.
The slot block lists the ten slots of Table 2 in fixed order. Each slot includes a payload from its allowed value set, an annotated evidence state, a locator, and the required slot-specific sub-entries. The derivation and cross-source blocks record channel-added inferences, including their premises, inference rule, result, and residual uncertainty. They also record attachment rows from attack-evaluation sources; these rows are keyed by the attacked defense and carry the attacking source’s evidence.
The conclusion block records the endpoints and of Equation 33, the row used to compute each endpoint or the reason it is unavailable, any separately declared tolerance , the residual conclusion and its unresolved reasons, the independent claim_verdict, and the consistency flag. The residual conclusion and the claim verdict are recorded separately and can diverge; Section C.6 is the case where they do.
Source-reported deployment costs are retained in each record alongside the slots. They do not enter any endpoint, and they are reported here only where a case discussion uses them, as in the utility cost of the zero-continuation design in Section C.5.
C.2 Supplementary Evidence Files
The supplementary materials contain four evidence files: the common review task, the final source-status list, and one final assessment compilation for each model channel. All four cover the wide-coded portion of the full-text coding pass. The list records channel, source identifier, full-text eligibility status, claim-instance status, and instance count. The two assessment compilations preserve the evidence and source locators behind the wide-coded aggregates in Section C.4. The depth-coded subset is documented in this appendix instead.
Each assessment record carries the anchor coordinates, the slot evidence with its locators, and the channel’s structured conclusion block. The endpoints and residual conclusions tabulated below are those recorded in the conclusion blocks, computed under Equation 33, and can be re-derived from it. A positive residual rules out under the anchor.
C.3 Depth-Coded Subset Aggregates
Table 5 tabulates the depth-coded subset’s slot-level evidence states, and Table 6 gives its instance-level distributions beside the wide-coded set’s. The subset is 24 claim instances, plus the two attack-evaluation source records it also covers, which yield no instances of their own and instead contribute 45 attachment rows keyed by the defenses they attack. The two channels agreed on 24 of 24 residual conclusions and on 20 of 24 claim verdicts. The table abbreviates supported as su, derived as de, claimed as cl, not-reported as nr, and not-applicable as na. The 45 attachment rows are excluded from the table and are all supported.
| Slot | su | de | cl | nr | na |
|---|---|---|---|---|---|
| Lower-endpoint evidence | |||||
| LB1 | 21 | 0 | 1 | 2 | 0 |
| LB2 | 18 | 0 | 3 | 3 | 0 |
| LB3 | 19 | 0 | 2 | 3 | 0 |
| LB4 | 15 | 0 | 2 | 4 | 3 |
| Upper-endpoint evidence | |||||
| UB0 | 8 | 0 | 0 | 13 | 3 |
| UB1 | 23 | 0 | 1 | 0 | 0 |
| UB2 | 19 | 0 | 2 | 3 | 0 |
| UB3 | 21 | 0 | 2 | 1 | 0 |
| UB4 | 4 | 1 | 2 | 17 | 0 |
| UB5 | 19 | 0 | 2 | 0 | 3 |
| Total | 167 | 1 | 17 | 46 | 9 |
C.4 Wide-Coded Set Aggregates
Table 6 reports the two coding strata together. The two channels sampled the 1,872-paper full-text pool under different random seeds. In the wide-coded stratum, they coded 80 and 110 papers (187 distinct) and extracted 152 claim instances between them. Because the channels assessed the overlap papers independently, instance counts are pooled across channels rather than deduplicated at the paper level; the three channel-overlap papers yield seven instances across the two channels. Wide-coded claim verdicts are single-channel judgments; they stay on the individual records and are not aggregated, so the verdict rows below cover the depth-coded subset only. The nine positive residuals occur with upheld and unresolved verdicts, and the zero-residual instance has an upheld verdict. All three refuted verdicts occur in instances whose residual conclusion remains unresolved.
| Depth-coded | Wide-coded | |
| () | () | |
| Residual conclusion | ||
| zero residual established | 1 | 0 |
| unresolved | 14 | 44 |
| positive residual established | 9 | 108 |
| Claim verdict | ||
| upheld | 4 | |
| refuted | 3 | |
| unresolved | 17 | |
| Harm locus | Wide-coded | |
| service-integrity | 81 | |
| external-world | 52 | |
| mixed | 19 | |
| Endpoint availability | ||
| upper endpoint below one | 1 | 0 |
| both endpoints computable | 1 | 0 |
C.5 Case Record I: All Three Gates Closed
Instance costa2025fides-01 anchors the by-design integrity noninterference claim of the Fides information-flow-control planner [43]. It is the depth-coded subset’s only zero-residual instance and its only instance with both endpoints computable. The anchored claim is that attacker-controlled untrusted data cannot influence the agent’s consequential tool actions. Its locator is Section 1 and Proposition 1 in Section 4.4 of the source. The record has promotion_basis = explicit-deployment-claim, harm_locus = service-integrity, and claim-scope extension yes because the source separately uses a broader prompt-injection headline. The broader empirical and confidentiality readings are represented by sibling instances.
Table 7 gives all nine coordinates of the coded anchor and the complete ten-slot projection with payload, state, and locator, so the conclusion below can be replayed without consulting the narrative in Section 5. The deployment-specific constraint set is not a coordinate of . The completeness count is 6/9: , , , , , and are filled, while , , and the general continuation value remain unknown. The cross-episode part of the horizon is also unreported. An absent locator supplies no default, and a state applies to the slot’s sub-entry.
| Coded anchor | |||
| Coordinate | Coded value | Status | Source locator |
| ReAct-style LLM agent loop with native tool calling and no information-flow control | filled | Section 2; Section 7.2 Basic-planner baseline | |
| Fides taint-tracking planner, P-T/P-F policy engine, selective hide/reveal, query_llm, and constrained decoding | filled | Sections 4 and 5 | |
| Agentic tasks over email, calendar, banking, Slack, and travel tools; untrusted tool outputs; one AgentDojo user-task episode | filled | Section 2; Section 7; cross-episode persistence absent | |
| Probability of the binary episode event in which untrusted tool data causes this agent to execute an unintended consequential tool action | filled | Section 2.1; Section 4.3 P-T; Section 7.2 ASR metric | |
| Defense-aware attacker controlling arbitrary untrusted tool content, knowing the configuration, and observing some tool effects; configuration compromise excluded | filled | Section 2.1 threat model | |
| Task-completion rate measured on 97 AgentDojo tasks, but no minimum normal-utility threshold declared | unknown | Section 8.2; threshold absent | |
| No explicit value-relevant trajectory projection | unknown | absent | |
| AgentDojo task-completion rate under a programmatic user-goal check | filled | Section 7.2 | |
| Binary integrity semantics are stated, but a general continuation value is not separately reported | unknown | Section 8.1; route-specific noninterference in Proposition 1 | |
| Slots | |||
| Slot | Payload | State | Source locator |
| LB1 | present, clean attribution; injections 163(156) to 1(0), parenthetical figures excluding two tasks the source rules outside its policies; task-completion-rate improvement 16.7% for o1 | supported | Table 1; Section 8.2 Figure 4 |
| LB2 | restriction-only; query_llm is an additive candidate but the matched-utility comparator is absent | supported, claimed | Sections 4.3 and 5; comparator absent |
| LB3 | reproducible, with zero benign adverse value | supported | Proposition 1 and noninterference definition in Section 4.4; Algorithm 5 lines 7 and 9; source Appendix A |
| LB4 | separation shown | supported | Section 4.1 default-untrusted labeling; Section 4.3 P-T; Proposition 1; Section 2.1 |
| UB0 | no artifact-transfer or capability-removal route | not-applicable | Mediation architecture in Sections 4 and 5 |
| UB1 | session-grain, decidable policy-success event | supported | Section 4.3; Section 4.4 complete-execution noninterference; Algorithm 5 |
| UB2 | deployment coverage, | supported | Algorithm 5 line 7; Section 8.1 response channel outside this integrity outcome |
| UB3 | robust integrity bound, | supported | Proposition 1; source Appendix A small-step semantics; Algorithm 5 |
| UB4 | bounded continuation, on covered integrity successes | derived | Section 2.1 binary episode event behind ; Proposition 1; Algorithm 5 |
| UB5 | serial per-call stateful topology; whole-trace property proved directly rather than composed from marginal rates | supported | Algorithm 5; Section 4.4 |
The conclusion procedure executes in four steps. C1 records the general continuation coordinate as unknown while the route-specific integrity value is fully identified by the binary adverse event behind , Proposition 1, and the derived . The endpoints therefore rest on supported or derived premises: the attacker controls untrusted tool data, the configuration remains trusted, and the proof uses the same consequential-action semantics as the guarded comparison. C2 combines supported coverage , supported robust failure , and derived continuation :
| (38) |
With normalized to this gives , so C3 establishes a zero-residual certificate at the strict boundary and clears the consistency flag. C4 assigns upheld because the coverage, failure, and continuation evidence matches the integrity claim’s own attacker class.
The decisive coding judgment is whether UB3 is supported or merely claimed. The final assessment assigns supported. The source provides Proposition 1 and the small-step proof in its Appendix A. The integrity result does not depend on deterministic model behavior, and within the declared , external data receives the conservative untrusted label by construction. This ruling supports . The route-specific continuation value has an independent basis. is the probability of the binary event that untrusted data causes one unintended consequential action within the episode. For this outcome, a covered success leaves no further adverse value. The record derives from this premise. The outcome definition, not the accuracy of the check, makes this endpoint available.
The record also retains the source-reported cost of this design: policy-on task-completion loss up to 24.5% with a data-dependent-task ceiling, and two to three times the Basic planner’s token use plus query_llm latency, at Section 8.2 Figure 4 and source Appendix E Figure 7. These values do not enter the endpoint. They quantify the cost of the design that yields , and a concrete deployment weighs them against its declared .
The certificate covers the integrity-scoped outcome with trusted configuration and correct conservative labeling. Text response manipulation, implicit confidentiality leakage, and the broader statement that the system stops all prompt-injection attacks belong to the two sibling instances.
C.6 Case Record II: Refuted Under the Claim’s Own Class
Instance zou2024circuitbreakers-01 anchors the representation-rerouting claim of [255] for text-only models. Table 8 gives the condensed record. It is one of the three depth-coded instances whose claim is refuted while the residual conclusion stays unresolved, so it illustrates the separation of residual conclusion and claim verdict. It also shows the role of attachment rows contributed by an attack-evaluation source [145], which yields no instance of its own.
| Slot | Payload | State |
| LB1 | present, clean attribution | supported |
| LB2 | restriction-only | supported, claimed |
| LB3 | simulable-output | supported |
| LB4 | structural none | supported |
| UB0 | silent none | not-reported |
| UB1 | event-defined | supported |
| UB2 | artifact-coverage | supported |
| UB3 | reference-rate, class check no-open-quantifier | supported, claimed |
| UB4 | none | not-reported |
| UB5 | none, single component | supported |
| attachment:LB3 | simulable-output at 100% ASR | supported |
| attachment:UB3 | refuted, | supported |
| Residual conclusion and claim verdict | ||
| On the upper side there is no robust and no , so , the bound available before measurement. On the lower side the attachment rows report a conditional failure rate and output simulability rather than an adverse value measured on the anchored scale. The benign continuation value needed by the simulation row is not reported either, so no nontrivial follows. State unresolved (no-nontrivial-endpoint); verdict refuted. | ||
The verdict is refuted because supported attachment evidence contradicts the anchored claim within its own declared class: the adaptive attack reaches , which also reclassifies the in-source suite averages as a reference rate. The residual conclusion nonetheless remains unresolved. Refuting the robustness claim invalidates an upper-bound operand, while the lower-bound rows contributed by the attack source are not measured on the anchored scale and therefore supply no lower endpoint. This difference is why the two outputs are recorded separately.