论文提出评估框架:防护手段本地测试通过不等于部署后的 LLM 系统更安全

HuggingFace Daily Papers(社区热门论文)·2026-09-01 08:00·2天前
AI 导读

Hefei AiDA Lab 等机构的研究者提出一套将 LLM 防护评测结果转换为部署安全结论的分析框架,并对其编码的 198 篇论文进行分析。152 个宽编码声明中有 108 个建立了正向残余有害协助,没有一个建立零残余证书;24 个深度编码实例中仅 5 个报告了检查成功后的可达残余,仅 Fides 一个实例通过覆盖、条件失败与后续可达三项证明给出零残余证书。

HuggingFace Daily Papers(社区热门论文)
53AI 编辑部评分,满分 100

论文提出评估框架:防护手段本地测试通过不等于部署后的 LLM 系统更安全

2026-09-01 08:00· 2天前
AI 导读

Hefei AiDA Lab 等机构的研究者提出一套将 LLM 防护评测结果转换为部署安全结论的分析框架,并对其编码的 198 篇论文进行分析。152 个宽编码声明中有 108 个建立了正向残余有害协助,没有一个建立零残余证书;24 个深度编码实例中仅 5 个报告了检查成功后的可达残余,仅 Fides 一个实例通过覆盖、条件失败与后续可达三项证明给出零残余证书。

Pingyu Wu

Hefei AiDA Lab

wupingyu@mail.ustc.edu.cn

Weiming Zhang

zhangwm@ustc.edu.cn

Nenghai Yu

Abstract

Safeguards in deployed LLM services are evaluated by refusal, attack success, and policy violation rates. Those rates characterize how a control performed on the requests it was tested on. A deployment has to answer a different question: how much help with harmful tasks the service still gives an attacker who keeps adapting or finds another way in. We determine what each reported result implies for that question, allowing results from different safeguard families to be compared under one deployment criterion. The evidence requirements are strongly asymmetric. One attack that obtains harmful help from the deployed service suffices to establish that such help remains, and such attacks appear repeatedly in the coded record. Establishing that little remains cannot follow from the safeguard’s own numbers alone; it also requires evidence about what the surrounding system still allows after the safeguard performs its local function. Such evidence is supported or derived in only a small minority of the depth-coded claims, and one such claim bounds its scoped residual. A better local score is therefore not, by itself, a stronger claim about the deployment. Safeguard research cannot stop at raising local scores; a gain has to be judged by whether it makes a deployed system any safer.

1 Introduction

Production LLM services use safeguards because the same capabilities that support legitimate work can also assist prohibited or harmful activity [185, 114, 229, 151]. Deployed and proposed systems combine model-level refusal behavior [8, 46, 6, 255] with runtime classifiers [89, 75, 45], account monitoring [3, 150], access controls, tool policies, and permissions [47, 43, 182], and execution containment [217, 5]. Model-level refusal has failure modes of its own: capability objectives compete with refusal objectives, and safety training generalizes less broadly than the capabilities it constrains [210]. Evaluations report direct outcomes such as refusal and attack-success rates [137, 25], policy violations [45], and benign-task utility [172, 48]. These measures answer useful questions: whether a control rejected a request, whether an attack crossed a specified boundary, and what capability the control removed from benign users.

The conclusion supported by these measures changes when the attacker, interaction history, or measured outcome changes. Nasr et al. evaluated 12 jailbreak and prompt-injection defenses using adaptive, defense-aware attacks. Their attacks exceeded 90% success against most defenses, although a majority of the original evaluations had reported rates near zero [145]. Holding scenarios, attackers, defenders, and scoring fixed, Jain et al. measured 0 to 1% attack success on the first turn and 5.4 to 14.0% after 15 rounds of adaptation to defender feedback [91]. The adapting attacker need not be a person: Hagendorff et al. had reasoning models act as autonomous jailbreak agents against other models, so the adaptation behind the two results above requires no human in the loop [74]. FragFuse changed the system path by distributing a prohibited request across agent memory; it achieved 86.3% access-control bypass but 41.1% end-to-end harmful-task success [168]. A near-zero result against a fixed or first-turn attack therefore cannot be reused unchanged for a deployment facing adaptive attackers, and a bypass rate cannot be read as the rate of completed harm. Doing so can favor a control whose reported advantage disappears under the deployed attack process [2, 228].

These findings improve individual measurements but expose a problem that additional measurements of the same kind cannot solve. A deployment must decide how much harmful assistance the guarded service still supplies [164, 131]. Surveys and SoKs make safeguard techniques and benchmark configurations comparable [56, 203, 78, 221]; guidance asks evaluators to declare threat actors, requirements, and supporting evidence [192, 21, 167]; and safety-case and uplift work connects particular measurements to broader models under explicit assumptions [40, 134, 194]. Each of these lines of work handles one step from a safeguard mechanism to a deployment decision. They stop at a comparable reported quantity or at a single deployment argument (Section 6).

We supply the conversion between them: we read each reported result against two coordinates, the outcome and attacker class it was measured under. We derive and prove the strongest deployment conclusion that result supports. The conversion changes what an evaluation result can justify: if a reported quantity supports no nontrivial deployment conclusion, reanalysis cannot produce one. The evaluator must instead measure a different quantity or establish a missing system property. A common coding instrument records which coordinates a source supplies and keeps each conclusion auditable back to the reported evidence.

Applied to 198 distinct papers at two levels of detail, the instrument yields an asymmetric coded record. Establishing how little harmful assistance remains through a safeguard check requires three deployment facts together. The one supplied least often is what remains possible after the check succeeds: five of the 24 depth-coded claims supply it. Across both coding strata, one coded claim rules out the worst case. The opposite direction needs one number, not a conjunction: of the 152 wide-coded claims, 108 report an adverse value putting the residual above zero. Section 7 treats these counts as hypotheses about what an evaluation should report.

  • From results to deployment conclusions. We determine the strongest conclusion each reported safeguard result supports about remaining harmful assistance, and prove when no stronger conclusion follows from that result alone. This shows when reanalysis can help and when a different measurement is necessary (Section 3).

  • Missing evidence made explicit. We turn those determinations into an auditable coding instrument that records the source facts each conclusion requires. Applied to a paper, it identifies the precise missing fact that prevents a deployment conclusion without rerunning the safeguard (Section 4).

  • A case-based stress test. Across the coded claims, the supported conclusion tracks the evidence reported rather than the technique category, which makes a benchmark gain a hypothesis about deployment, not a guarantee (Section 5).

2 Definitions and Scope

2.1 Deployment Safety

We define deployment safety as the governance of the harmful assistance a focal service still supplies once its safeguards have acted, subject to constraints protecting basic rights and legitimate use.

2.2 Evaluation Anchor and Scope

A concrete deployment assessment must fix the conditions under which its decision is made. Research papers may supply evidence for only a subset; Section 4 defines the anchor coordinates recoverable from a source. The full anchor has seven coordinates:

  • S: the focal LLM service without the evaluated safeguard intervention.

  • D: the evaluated intervention; applying it yields S[D], the same service with the intervention deployed, all other parts unchanged.

  • Σ: the declared nonempty class of attacker strategies, including its resource and interaction limits.

  • : the basic-rights and legitimate-use constraints every admissible deployment must satisfy, supplied by the concrete application and its governing institutions rather than inferred from a method paper.

  • q: the minimum normal-utility requirement, measured on a declared benign reference population and utility scale; it scalarizes one component of without replacing the rest.

  • Z: the scalar harmful-assistance functional, normalized to [0,1], with lower values less adverse. It measures what the service supplies toward the adverse outcome through its outputs, actions, and state transitions, not the harm an attacker ultimately achieves with that supply.

  • : the fixed task, target, operating environment, benign reference conditions, and evaluation horizon, including allowed interactions and retries, plus any persistence or recovery window relevant to Z.

We call

𝔞=(S,D,,Z,Σ,,q) (1)

an evaluation anchor. A deployment-safety claim is an assertion about a particular S[D] under one such anchor: deployment safety is a property of an intervention in a specified deployment, not an intrinsic property of an isolated mechanism. Its scope is the set of attacker strategies and operating conditions within the anchor for which the conclusion is asserted to hold. The anchor is the interface through which a deployment enters the analysis: the bounds of Section 3 hold for whatever outcome, attacker class, and utility target it declares. The same bounds therefore cover outcomes as different as an action the service must not take and knowledge an attacker must not gain.

All anchor coordinates are held fixed across comparisons; only the presence of D differs. The attacker may choose any admissible strategy in Σ. A deployment that violates is inadmissible regardless of its value of Z. We suppress 𝔞 below.

2.3 Residual Assistance and Safeguard Effect

Write Z(S[D]) for the adverse value the guarded service supplies against an attacker in Σ, and Z(S) for the same quantity when that service runs without D. The primary deployment quantity is the residual harmful assistance Z(S[D]) itself: what the service still supplies once the safeguard has acted. The safeguard effect

VZ(D)=Z(S)Z(S[D]) (2)

diagnoses what D removed [134, 194]. The two quantities partition the unguarded total,

VZ(D)+Z(S[D])=Z(S), (3)

so a safeguard reallocates the service’s assistance between a removed part and a retained part. Reporting VZ(D) alone leaves Z(S[D]) undetermined.

A deployment declares a tolerance τ[0,1] on the retained part and requires

Z(S[D])τ. (4)

The strict boundary is τ=0, at which the service supplies nothing toward the anchored outcome; we call a certificate for that case a zero-residual certificate. Section 3 determines which outcomes admit one and what evidence any τ requires. A deployment decision also verifies the conditions in and the utility requirement q, so a service cannot meet the criterion by refusing service or by excluding legitimate users.

Because Z is normalized, Z(S[D])[0,1] always holds; the content of a certificate is the tighter bound it supplies. The next section derives what each kind of published evidence supports on this scale.

3 What a Reported Quantity Can Bound

Section 2 identifies Z(S[D]), the assistance a guarded service still supplies, as the quantity a deployment decision needs. A local score does not fix it: a score describes one test, whereas Z(S[D]) ranges over every strategy the declared attacker may run.

This section instead asks the question relevant to deployment. Given a quantity a paper reports, measured on the declared outcome scale and produced inside the declared attacker class, what is the tightest bound on Z(S[D]) that follows from it? Every bound in Table 1 is sharp: no smaller upper bound and no larger lower bound follow from the reported quantities alone. Section 3.4 collects the results as one schedule, and Section 4 gives the coding rule that performs that reading and applies the schedule to the literature. Complete proofs and the attaining constructions are in Appendix A.

3.1 The Deployed Quantity

Fix the anchor 𝔞=(S,D,,Z,Σ,,q) from Section 2. Every statement below holds these coordinates fixed.

Let ΩD be the set of complete security-relevant trajectories under D; each ωΩD carries the evidence history H available before each decision. Declaring Z also fixes a value-relevant projection ϕZ:ΩD𝒯Z, written T=ϕZ(ω). It retains the information needed by the legitimate value b(T) and the adverse continuation value vZ(T), both in [0,1].

The key quantity is the continuation value. vZ(T) is the largest adverse value the service still supplies from the retained state over the remaining horizon, counting outputs already released, actions the policy still permits, accumulated state, and further attempts. It is a property of the service’s own interface and policy, which makes it measurable.

A causal policy g𝒢D implements D and satisfies every non-scalar condition in . The deployment commits to g before the attacker chooses σΣ. For a transferable artifact, the release remains in T and Σ includes every permitted modification, which places post-release adaptation inside the evaluated deployment.

Let Pbg be the benign-reference trajectory law on ΩD and Pm,σg the law induced by attack σ. Randomized attacks and their total-variation limits form

𝒞g=co¯{Pm,σg:σΣ}. (5)

Convexification and total-variation closure do not change the supremum of a bounded value functional. The deployment quantities are

WZ(g)=supQ𝒞g𝔼Q[vZ(T)],B(g)=𝔼Pbg[b(T)], (6)

and 𝒢D(q)={g𝒢D:B(g)q} collects the policies that also meet the utility target. A committed deployment realizes

Z(S[D])=WZ(g), (7)

so every bound below is a bound on the quantity of Section 2.

One distinction governs the whole section. An upper bound on WZ(g) must cover every attacker law Q𝒞g, because WZ(g) is the supremum over that class. A lower bound, by contrast, needs only one attainable attack law. Thus one observed attack can establish that harmful assistance remains, whereas showing that little remains requires evidence covering the full attacker class. This quantifier asymmetry, not mechanism strength, explains the different evidence requirements, which adversarial example defenses established empirically [20, 190].

3.2 Lower Bounds: Assistance That Remains

A lower bound identifies adverse continuation the deployment cannot remove. Because WZ(g) is a supremum, any single law bounds it from below, and the constructions differ only in how far they quantify: over one executed attack, over the one committed policy, or over every policy the architecture can realize.

3.2.1 One Executed Attack

Proposition 1(Attack-witness bound).

For every σΣ,

WZ(g)𝔼Pm,σg[vZ(T)]. (8)

The law Pm,σg lies in 𝒞g, so the supremum is at least its value. Write Lwit for this lower endpoint. It is the paper’s least evidentially demanding bound: an attack run inside the declared class against the deployed configuration supplies its adverse value as a lower endpoint, and the bound reaches no further than the strategies actually run.

3.2.2 Simulating a Committed Policy

When no such attack was executed, a floor still follows from benign measurements alone, provided the attacker can reproduce the value-relevant behavior. Define the attacker-to-benign distance on value-relevant coordinates,

δT(g)=infQ𝒞gTV((Pbg)T,QT), (9)

and write [x]+=max{x,0}.

Proposition 2(Committed-policy simulation bound).

Every committed policy satisfies

WZ(g)[𝔼PbgvZδT(g)]+. (10)

Write Lsim for this endpoint. The coefficient of δT(g) is one and cannot be improved, since a two-point construction attains the bound. A safeguard whose protective context the attacker can copy is one reported case where this bound binds [215].

The premise this endpoint needs is an upper bound on δT(g), and the next result fixes what kind of measurement can supply one.

Proposition 3(Sequential simulation certificate).

Suppose that under one σΣ, every attacker-generated evidence update t and every history on which the benign and malicious processes remain coupled satisfy

TV(Pbg(Hth),Pm,σg(Hth))ηt, (11)

with all deployment-controlled transitions using the same committed g and every remaining transition having the same conditional law in both processes. Then

δT(g)1t(1ηt). (12)

The premise is conditional on every adaptive history. A marginal error rate measured on a fixed suite supplies no ηt, because it constrains an average over histories rather than the kernel after any particular one. This is the first row of Table 1 that yields no nontrivial endpoint: a fixed-suite rate, however low, supports no simulation bound.

3.2.3 From One Policy to the Attainable Frontier

Evidence about the evaluated operating point does not by itself describe what any admissible policy could attain. The architecture envelope is

RD,Z(q)=infg𝒢D(q)WZ(g). (13)

The benign T laws the architecture realizes are

𝒦Dreal={(Pbg)T:g𝒢D}, (14)

whose least adverse value at utility target q is

ΓD,Zreal(q)=infμ𝒦Dreal:𝔼μbq𝔼μvZ. (15)

Supported contracts and known trace constraints may instead identify an outer relaxation 𝒦Dout𝒦Dreal with frontier ΓD,Zout(q).

Proposition 4(Frontier order and restriction monotonicity).

For every target feasible for both frontiers, ΓD,Zout(q)ΓD,Zreal(q). Writing Γ𝒦 for the same optimization over a law set 𝒦 and holding both value functions fixed,

𝒦𝒦Γ𝒦(q)Γ𝒦(q). (16)
Theorem 5(Attainable-frontier simulation bound).

Every g𝒢D(q) satisfies

WZ(g) [ΓD,Zreal(q)δT(g)]+[ΓD,Zout(q)δT(g)]+. (17)

Consequently a deployment with WZ(g)β must satisfy β+δT(g)ΓD,Zreal(q). If every q-feasible policy is exactly simulable, meaning δT(g)=0, then RD,Z(q)ΓD,Zreal(q), with equality when some frontier optimizer g also satisfies supQ𝒞g𝔼QvZ=𝔼PbgvZ.

Write Lfrt=[ΓD,Zreal(q)δT(g)]+. The equality condition matters for what a paper can claim: without it, a measured operating point stays a statement about that point and does not become a statement about the architecture.

Equation 16 characterizes the effect of behavior removal. Deleting feasible laws while holding b and vZ fixed cannot lower the frontier, and by shrinking the feasible set at target q it can raise the frontier or empty the feasible set entirely.

A uniform dual-use relation makes the floor explicit. It requires the adverse value of every trace supported by 𝒦Dout to be at least a fixed fraction ρ of its legitimate value:

vZ(t)ρb(t)for every trace tsupp(μ),μ𝒦Dout. (18)

If ρ>0, both frontiers are at least ρq. Any exactly simulable deployment meeting a positive utility target therefore has Z(S[D])ρq>0 and cannot obtain a zero-residual certificate. In plain terms, useful behavior is then inseparable from a fixed positive share of adverse value.

Whether ρ can be zero, and with it whether τ=0 is available at all, depends on the declared outcome together with the traces the safeguard leaves reachable. Where one release supplies both the legitimate and the adverse value, reaching ρ=0 requires a safeguard that makes a useful trace with no adverse value reachable, which restriction alone cannot supply.

3.2.4 When Reachability Changes the Floor

The simulation term depends on reachable laws, not on mechanism names. Let gT={QT:Q𝒞g} be the maliciously reachable T laws.

Proposition 6(Value-law invariance).

If two committed deployments g and g¯ are evaluated with the same b and vZ, and satisfy (Pbg)T=(Pbg¯)T and gT=g¯T, then B(g)=B(g¯), WZ(g)=WZ(g¯), and δT(g)=δT(g¯).

A new label, credential, or isolation boundary does not change either bound merely by existing. Proposition 6 shows that it helps only if it changes value-relevant behavior or which states an attacker can reach. Trusted state can create a reachability separation. When the attacker must reproduce its benign state distribution, the easiest allowed acquisition or compromise path sets the floor. By the data-processing inequality, the relevant quantity is the total-variation distance d to the nearest maliciously reachable state marginal, not average acquisition accuracy. Evidence must therefore identify the trusted state, its acquisition class, and the downstream conditional kernel. Appendix A.6 states the frontier and the bound.

3.3 Upper Bounds: What a Deployment Can Guarantee

An upper bound must control every law available to the anchored attacker. The direct route bounds the reachable law set. The factored route bounds how often a deployment success event occurs and what continuation remains there.

3.3.1 Direct Reachable-Set Bounds

Proposition 7(Direct reachable-set bound).

For any evidence-supported outer class 𝒞g𝒞gout,

WZ(g)Urch=supQ𝒞gout𝔼Q[vZ(T)], (19)

with equality when 𝒞gout=𝒞g.

This route covers artifacts without a mediation domain, such as released weights, when the declared tampering budget contains Σ and the value bound holds for every reachable system. The evidentiary requirement is exacting in one specific way. A set of tested modifications is an inner sample of 𝒞g, and a supremum over an inner sample provides no upper bound. Enumerating more attacks therefore does not change this endpoint at any sample size short of exhausting Σ [7, 190].

3.3.2 Success-Region Bounds

Let GΩD be a success region and gΩD a reachable envelope with Q(g)=1 for every Q𝒞g. Define

λG=infQ𝒞gQ(G),rG=supωGgvZ(T(ω)), (20)

using rG=0 when Gg is empty.

Proposition 8(Sharp success-region bound).

For every committed policy, WZ(g)Ureg=1λG(1rG), and no smaller uniform bound follows from λG and rG alone.

The expectation splits over G and Gc; a matching two-region law proves sharpness. Every success-region bound therefore needs class-uniform coverage and a continuation bound over all remaining state, interfaces, and attempts.

3.3.3 Closed Mediation and the Three Gates

Mediation is the factorization reported safeguard results take. Let A denote entry into a mediation domain and F failure of its local security fact, so G=AFc. A reference law may satisfy P(FA)ϵP(A). For a deployment bound, assume the same contract for every Q𝒞g and define

Q(FA)ϵQ(A),α=infQ𝒞gQ(A), (21)

which gives λGα(1ϵ).

Theorem 9(Sharp closed-mediation bound).

Fix g and suppose every Q𝒞g satisfies the contract in Equation 21. If vZ(T)r on every attacker-reachable trajectory in AFc, then

WZ(g)1α(1ϵ)(1r). (22)

No smaller bound holds uniformly over models described only by (α,ϵ,r), and this information gives a nontrivial bound if and only if

α>0,ϵ<1,r<1. (23)

Sharpness is proved by a three-atom construction, and it converts the theorem from a bound into a sensitivity rule. Improving ϵ to ϵ moves the certified bound by exactly

[1α(1ϵ)(1r)][1α(1ϵ)(1r)]=α(ϵϵ)(1r), (24)

so the certified contribution of detection accuracy is determined by coverage and continuation. Where either is unmeasured, that contribution cannot be quantified; where either is adverse, it is zero.

Corollary 10(Local perfection is globally non-identifying).

Even with ϵ=0, the sharp bound in Equation 22 equals one when α=0 or r=1. Any number of mechanisms may establish their local facts without error on every invocation and remain compatible with WZ(g)=1.

The two quantities that convert a reported ϵ into information, α and r, are properties of the surrounding deployment rather than of the classifier.

At the strict boundary the three gates collapse to a clean statement. For τ=0, the bound in Equation 22 certifies Z(S[D])=0 exactly when α=1, ϵ=0, and r=0: complete coverage of every path to the outcome (the complete-mediation condition [176]), no conditional failure on those paths, and no continuation after a covered success. A zero-residual certificate through mediation is therefore a conjunction of three deployment-wide facts, none of which a local score reports.

3.3.4 Composition

A stack contributes through its deployment event structure, not its layer count.

Proposition 11(Adaptive composition bounds).

If an attack path requires ordered failures F1,,Fm and, at every history reaching layer j, Pr(FjF1,,Fj1,h)ϵj, then

Pr(j=1mFj)j=1mϵj. (25)

With marginal bounds alone, the sharp general bound is instead

Pr(j=1mFj)minjϵj. (26)

For alternative path events Ei with Pr(Ei)p¯i, Pr(iEi)min{1,ip¯i}, and for ordered attempts with jPr(Eji<jEic)p¯j,

1j(1j)Pr(jEj)1j(1p¯j). (27)

The improvement provided by a composition argument is the gap between Equations 25 and 26. This gap is determined by the dependence premise rather than by the number of layers. Marginal bounds are attained by perfectly correlated failures, so under them the best bound a stack of any depth supports is that of its single strongest layer. History-uniform conditional bounds, which independence implies but which can hold without it, are what license the product. Alternative paths and retries meanwhile accumulate attempts at any fixed per-attempt rate.

3.4 The Schedule

Table 1 collects the results as one schedule from reported evidence to what follows for Z(S[D]). We call a row noninformative when its named evidence yields no nontrivial endpoint. Every bound is sharp: no better bound follows from the quantities in the first column, so a noninformative row cannot be made informative by further measurement of the same kind. Two rows state a structural relation rather than a bound, and their witness is the proof that establishes it.

媒体内容 · 前往原文查看
Table 1: Sharp consequences of reported evidence for lower and upper bounds on Z(S[D]).
Reported quantity Consequence for Z(S[D]) Sharpness witness
Lower endpoints
One attack executed inside Σ its measured adverse value the executed law itself (Prop. 1)
Benign continuation value and a bound on δT(g) [𝔼PbgvZδT(g)]+ a two-point transfer of mass (Prop. 2)
Per-history conditional evidence distances ηt δT(g)1t(1ηt) independent per-update deviations (Prop. 3)
Marginal error rate on a fixed suite no simulation bound an average over histories constrains no adaptive kernel
Removal of behaviors at fixed b and vZ frontier does not fall monotonicity (Prop. 4)
Uniform dual-use ratio ρ at utility q and a bound on δT(g) ρ[qδT(g)]+ a proportional two-point law (Eq. 18; App. A.2)
A label, credential, or boundary leaving value laws and reachability unchanged no nontrivial endpoint on either side value-law invariance (Prop. 6)
Average acquisition accuracy for trusted state moves with the distance d to the nearest maliciously reachable state marginal, not with the average the data-processing inequality (App. A.6)
Upper endpoints
Enumerated attacks inside a declared budget no upper bound an inner sample cannot bound a supremum
Outer reachable class with a uniform value bound Urch a tight outer class (Prop. 7)
Coverage α, conditional failure ϵ, continuation r 1α(1ϵ)(1r) a three-atom law (Thm. 9)
ϵ alone, at any value including zero 1 the same law at α=0 or r=1 (Cor. 10)
Marginal per-layer failure rates Pr(jFj)minjϵj perfectly correlated failures (Prop. 11)
History-conditioned per-layer bounds Pr(jFj)jϵj independent layers (Prop. 11)

Four rows are noninformative for three distinct reasons. A marginal rate and an enumerated attack set fail on the quantifier: the former averages over histories without constraining the conditional kernel after any particular history, whereas the latter samples a class that the bound must cover. A conditional failure rate is insufficient when another gate can reduce its contribution to zero. A relabeled boundary contributes no quantity used by either bound. Sharpness shows that making the reported value in any of these rows more favorable cannot replace the missing relation or quantity.

Multiple rows may apply to one deployment. For finite sets of supported candidates bounding the same WZ(g) under one anchor, the combined endpoints are L=maxiLi and U=minjUj; a side with no candidate stays open. Every upper candidate must cover the full class 𝒞g, so a subclass-specific bound must first be lifted to the union of the declared subclasses. The combined endpoints must also be mutually consistent: if L>U, their premises cannot share one anchor, and no interval follows until the inconsistency is resolved.

4 Reading the Literature Against the Schedule

Table 1 maps reported evidence to supported bounds. To apply it to published work, we need four elements: a claim instance with a fixed anchor, coding slots for the quantities required by each schedule row, a rule that maps supported slots to a conclusion, and a source-selection procedure (Figure 1). This section defines all four, and Section 5 reports what the resulting coding shows.

4.1 A Common Object of Comparison

The unit of analysis is a deployment-safety claim instance

x=(gx,θx,cx), (28)

where gx is the evaluated deployed policy, θx fixes the comparison, and cx is the assertion under assessment. Each instance also carries its source and version, which record provenance without entering the comparison. A candidate becomes an instance when the source identifies a deployed intervention, an operationalized adverse outcome, and either a comparison world or a formal statement connecting the intervention to that outcome. This rule prevents an isolated component score from acquiring a deployment interpretation before a deployed action and outcome are fixed. A paper can therefore yield several instances when it changes the intervention, outcome, attacker class, utility target, or evaluation horizon. Measurements remain in one instance only when they support the same anchored assertion. External attack evaluations attach to that instance, so later evidence can assess the original claim without changing the object being judged.

For each instance, the coordinates recoverable from a source are

𝔞xsrc=(Sx,Dx,x,Zx,Σx,qx), (29)

and the schedule additionally requires the coded anchor

θx=(𝔞xsrc,Tx,bx,vZx). (30)

Here Tx is the trajectory projection on which simulation and continuation are judged, while bx and vZx are its normal use and adverse continuation values. This tuple bundles exactly the structure declared in Section 3.1; a concrete deployment fixes it implicitly, whereas a coded source must record it as a checkable operand. The external constraint set x remains a condition on a concrete deployment, because a literature source may not determine it.

Every coordinate is supported by a source locator, derived by a declared rule, or left unknown. Unknown coordinates receive no default.

媒体内容 · 前往原文查看
Figure 1: Overview of the evidence pipeline. The coded set contains 198 papers: twelve receive full ten-slot depth coding and 187 receive endpoint-route wide coding, with one paper in both strata. Both analyses identify the next measurement needed to tighten an endpoint.

4.2 The Coding Slots

Table 2 defines the ten slots. Four record what an attacker retains and six what the deployment controls, matching the two directions of Section 3; the constructions below name the schedule row each one feeds. A slot holds a payload from its allowed value set, an evidence state, a source locator, and the sub-entries the analysis requires. A construction yields an endpoint only when every slot it consumes is supported or validly derived under one claim instance.

媒体内容 · 前往原文查看
Table 2: Coding slots and the source evidence each one requires.
Slot What the source must report
Lower-endpoint evidence
LB1 The guarded versus unguarded comparison, its attribution, and whether the guarded arm’s adverse value was produced by a strategy in Σx and measured on the Zx scale
LB2 Whether the intervention only removes behaviors or supplies an additive substitute at matched utility
LB3 Whether an attacker can reproduce the value-relevant benign behavior, and any upper bound on δTx(gx)
LB4 Whether trusted state separates attacker-reachable laws from the benign state distribution, and how that state is acquired
Upper-endpoint evidence
UB0 Whether the artifact transfers to the attacker with no mediation domain, and the declared tampering budget
UB1 The decidable deployment success event Gx and its grain
UB2 Coverage αx, which paths to the anchored outcome must enter the mediation domain: the first gate, and an architecture fact rather than a classifier accuracy
UB3 Conditional failure ϵx after adaptive history, and the class over which it holds: the second gate
UB4 Continuation rx, the adverse value still reachable after a covered success: the third gate, covering released outputs, permitted actions, accumulated state, and retries
UB5 The dependence relation across composed components at the deployed history grain

The slots make incompleteness diagnostic. Because each construction consumes a named set of slots, an endpoint that stays open identifies the exact missing fact.

The two slot blocks answer different questions: the lower records evidence that can establish a floor on the harmful assistance that remains, and the upper a ceiling.

Three lower-bound constructions use these slots. First, an attack executed inside the declared class needs no reproduction argument. For one σΣx, Proposition 1 gives the witness endpoint Lxwit=𝔼Pm,σgxvZx from LB1’s guarded-arm value. LB1’s qualification sub-entry determines whether that value qualifies. A guarded-arm suite rate, bypass count, or other quantity not measured on the anchored Zx scale can fill LB1 but supplies no endpoint. Second, when no qualifying attack was executed, LB3 and the benign continuation value supply the simulation endpoint of Proposition 2, Lxsim=[𝔼PbgxvZxδTx(gx)]+. Third, suppose that the evaluated deployment meets the utility target qx, that LB2 and LB4 together identify the least adverse value attainable at that target, and that LB3 bounds δTx(gx). The frontier construction of Theorem 5 gives

Lxfrt=[ΓDx,Zxreal(qx)δTx(gx)]+. (31)

Two constructions read the upper block. UB0 yields Uxrch when the declared budget class contains Σx and the value bound holds over all of it. Otherwise UB1 through UB4 yield

Uxreg=1λx(1rx),λxαx(1ϵx), (32)

with UB5 governing whether component bounds may be multiplied. A source reports a supported coverage bound rather than the exact infimum, so we evaluate Uxreg at αx(1ϵx). Substituting a lower bound for λx can only raise the endpoint, the conservative direction.

An instance can support several candidates on one side. The quantities carried forward are Lx=maxiLxi and Ux=minjUxj, and a side with no supported candidate leaves that endpoint open.

4.3 Evidence States and Permitted Conclusions

Each required relation receives one of five evidential states, so that a construction uses exactly what the source establishes. A relation is supported when a result, measurement, architectural property, or trace establishes it under the required anchor and quantifier. It is derived when it follows from supported premises through a stated inference rule. A relation asserted without identifying evidence is claimed; a required relation absent from the source is not reported; and a relation excluded by the declared scope is not applicable. Only supported and validly derived relations enter endpoint computation. These states assess the evidentiary relation, not mechanism quality, and the cited passages preserve the basis of every conclusion.

The composite endpoints bracket the residual quantity of Section 2,

LxZx(Sx[Dx])Ux, (33)

against which a declared tolerance τx is decided. The evidence establishes that the tolerance is met when Uxτx, establishes that it is exceeded when Lx>τx, and otherwise leaves it unresolved. At the strict boundary τx=0 this specializes to three residual conclusions: Ux=0 establishes a zero-residual certificate, Lx>0 establishes a positive residual and rules out τx=0, and every other interval leaves the boundary unresolved. A positive tolerance records a practical compromise with some retained assistance; the strict boundary stays at zero.

The verdict on cx answers a separate question. It is upheld when evidence supports the assertion over its declared class, refuted when evidence from that class contradicts it, and unresolved otherwise. A valid improvement claim can coexist with a positive residual, while an in-class counterexample can refute a broad claim without supplying either endpoint. Keeping these outputs separate lets Section 5 recover both the truth of a coded claim and its deployment consequence.

4.4 Study Selection and Coding

Records first passed a hard authority gate. Two independent model channels then screened titles and abstracts, with advancement requiring include from both. The same channels analyzed randomized full-text sequences from the resulting pool. At full text, we extracted claim instances only from studies that introduce or evaluate an intervention and make or assess a claim about an adverse deployment outcome. Each eligible result was assigned to a claim instance as defined above. Coding stopped when the reported shares stopped moving as batches were added. The resulting 198-paper coded set is analyzed in Section 5. Appendix B reports the protocol in full, and Appendix C the record format, coding aggregates, and two case records.

5 What Published Safeguard Evidence Establishes

We apply the schedule of Section 3 using the slots of Section 4 to determine what the reported evidence can and cannot establish, rather than grouping papers by technique.

5.1 One Coded Set at Two Coding Depths

The depth-coded stratum yielded 24 claim instances under the full ten-slot coding. The wide-coded stratum pooled 152 claim instances across the two channels, each record carrying slot evidence states, endpoints, and a verdict. Wide-coded verdicts are single-channel judgments, so they stay on the individual records and are not aggregated.

A positive residual is established for 108 of the 152 wide-coded instances and a zero-residual certificate for none; the remaining 44 stay unresolved. Claim validity and residual bounding answer distinct questions about the same evidence. In the depth-coded subset, where both outputs are recorded on the same 24 instances, neither determines the other: an upheld claim can coexist with a positive residual, while all three refuted claims leave the residual unresolved. The refuting evidence invalidates an upper-bound operand but supplies no lower endpoint: its values are not on the anchored scale, and its sources report no anchored adverse value of their own.

媒体内容 · 前往原文查看
Table 3: Availability of supported or derived slot evidence in the depth-coded subset (N=24).
Slot Relation Filled
n %
Lower-endpoint evidence
LB1 guarded versus unguarded comparison 21 88
LB2 removal versus matched-utility substitute 18 75
LB3 reproducibility of benign behavior 19 79
LB4 reachability separation 15 63
Upper-endpoint evidence
UB0 artifact transfer and tampering budget 8 33
UB1 decidable success event 23 96
UB2 coverage α 19 79
UB3 conditional failure ϵ 21 88
UB4 continuation r 5 21
UB5 dependence across components 19 79

Table 3 states the depth-coded subset’s central pattern. Across the 24 depth-coded instances, sources define a success event, measure its conditional failure rate, and describe its coverage with comparable frequency. Evidence about what remains reachable after that event is supported or validly derived in only five of the 24 instances. By Corollary 10, a supported ϵ without a supported or derived r yields the bound Zx(Sx[Dx])1. Of the three gates an upper bound through mediation requires, the subset measures conditional failure most often and supports least often the continuation that governs its effect.

The gap is larger than the table alone shows. Of the five instances with a supported or derived r, only one also has supported α, ϵ, and a success event under the same anchor. That one supplies the depth-coded subset’s only computable upper endpoint from this route. Supporting one operand in isolation cannot tighten the endpoint.

Independent coding of the depth-coded subset by the two channels matched on all 24 residual conclusions and on 20 of the 24 claim verdicts. Appendix C reports the per-slot states.

5.2 The Attack-Witness Row, and How the Literature Reaches It

All 108 wide-coded positive residuals rest on the same attack-witness row in Table 1. Each source ran at least one strategy from its declared class against its deployed configuration and reported the adverse value produced by that strategy. By Proposition 1, each reported value is a lower bound on Zx(Sx[Dx]), so the corresponding lower endpoint can be read directly from the published result and reaches no further than the strategies actually run. The attack-witness lower bound does not depend on δTx(gx).

Emulated Disalignment recombines a released pretrained checkpoint with its aligned sibling at decoding time, scored against a harmful outcome measure [252]. Across four model families the executed attack yields harmful rates of 32.0%, 37.0%, 27.0%, and 57.6%. Each rate is the value of one law the attacker can induce on the deployed configuration, so each is directly a lower bound on Zx(Sx[Dx]).

RESTA shows why the safeguard effect and the residual must stay separate even inside one source [12]. Its reduction in judged unsafe responses supports the paper’s improvement claim, so its single-channel record assigns an upheld verdict. The same evaluation still judges 37.78% of the restored model’s multilingual CATQA answers harmful. This rate alone gives Lx0.3778>0. The measured improvement and the positive residual are both supported, and under Equation 3 they are the two parts of the same unguarded total. We analyze the other instances on the attack-witness row in the same way, using the adverse-value column of a table published to demonstrate a reduction or an attack.

The lower side also characterizes the restriction-only pattern. None of the seven capability-removal instances in the depth-coded subset establishes a frontier change at matched normal utility: six report restriction evidence in LB2 and one supplies no qualifying relation. By Proposition 4, restriction alone cannot lower the frontier. These sources support improvement at their chosen operating points, without evidence that the best attainable adverse value has moved. Repeating attack tests at the restricted operating point improves the estimate of that point; changing the floor requires an additive substitute at matched utility or a reachability separation, which are LB2 and LB4.

5.3 Coverage and Continuation Determine a Check’s Deployment Bound

Coverage depends on topology, not on component quality, as two non-LLM precedents show. Against Blacklight’s global store [109], every query enters the collision check, so a harmful query that passes is a conditional failure inside a covered path. Against a detector whose state is scoped per account, opening a new account resets that state, so the same attack never enters the mediation domain and becomes a coverage failure [60]. The distinction is decision-relevant because the repairs differ: one calls for a better check, the other for routing the bypass path through any check.

Equation 24 shows continuation determines the certified contribution of accuracy. The sequential monitor of Chen et al., a depth-coded instance, reports defense success up to 93% on cumulative decomposition attacks [32]. That figure supports conditional performance on the evaluated sequences. An upper bound also requires the value still available after a locally successful check: harmful content already released in earlier answers, fresh-session retries, and permitted follow-on actions. Those paths lie outside the reported event, so rx is unreported and Ux stays open at one. A response classified as safe does not bound what the trajectory has already supplied.

5.4 The Quantifier Sets the Scope, Not the Format

Whereas coverage and continuation determine how a rate affects the bound, the quantifier decides which attackers the rate describes. A measured maximum over a fixed suite quantifies over the listed attacks. A claim about an adaptive class quantifies over every strategy the class admits.

Six of the 24 depth-coded instances report a maximum over enumerated attacks within a stated tampering or interaction budget [153, 189, 114]. Each maximum is informative for the attacks run, and none is a uniform bound over the budget. A set of tested strategies lies inside 𝒞gx, and a supremum over an inner sample provides no upper bound at any sample size. A scaling analysis over four attack paradigms makes the gap measurable: on a shared compute axis, attack success rises with the budget spent inside a fixed method and model [200]. A suite run at one budget point therefore describes that point, not the class. Naming a budget defines the domain a useful bound would have to cover, without performing the quantification the reachable-set row requires.

Circuit Breakers makes the consequence observable. The original evaluation reports low attack success on fixed suites and claims robustness to powerful unseen attacks [255]. A later defense-aware reinforcement learning attack operates within that declared class and reaches conditional failure near one [145]. Those measurements remain valid descriptions of the attacks they ran, while the in-class counterexample refutes the broader claim. Because the counterexample supplies neither a matched adverse value on the anchored scale nor any upper operand, both endpoints stay open.

SmoothLLM provides the complementary formal case. Its theorem supplies a robust failure bound for a declared k-unstable certificate class using defender-controlled independent perturbations [171]. The proof supports claims whose attacker scope is contained in that class. In the claim analyzed here the accompanying prose reaches a wider scope, and no containment result connects the two. The theorem remains a supported local guarantee, and the wider claim stays open.

5.5 Dependence, Not Depth

No depth-coded instance reports a history-uniform conditional bound for a general serial stack, so by Equation 26 the best bound such a stack supports is that of its single strongest layer, whatever its depth. SmoothLLM’s controlled randomization is the one premise that licenses a product: it supplies independence for the votes inside one invocation, and composition holds at that grain. Fresh invocations and attacker-chosen retries need their own premise. The certifying property of a stack is the supported dependence relation at the deployed history grain, and UB5 records that.

5.6 Where the Three Gates Close

Fides supplies a worked instance in which coverage, failure, and continuation are all supported or derived under one anchor [43]. Its service-integrity instance fixes an event on the agent’s consequential actions. Every consequential tool action passes through the policy check, so architecture evidence establishes αx=1. The source’s noninterference result covers all untrusted inputs in the declared class, establishing ϵx=0. On a checked trajectory untrusted data cannot influence the consequential action, which gives rx=0. Theorem 9 then closes the route:

Ux=1αx(1ϵx)(1rx)=0, (34)

and with Zx0 by normalization the interval is Zx(Sx[Dx])=0: a zero-residual certificate for the anchored integrity outcome.

Each premise plays a distinct role. If any one is missing, the endpoint remains open. The decisive premise is rx=0, which detection accuracy cannot supply (Corollary 10); it is derived from the declared outcome instead. Zx is the probability of the binary episode event that untrusted data causes one unintended consequential action [68]. The information-flow separation makes that continuation structurally unavailable after a covered success. CaMeL belongs to the same broad mechanism family, yet its source claim leaves the class-uniform failure and continuation relations open: the shared family label does not determine the certificate [47].

The certificate has a reported utility cost in the same instance: task-completion loss up to 24.5% under the policy-on configuration [43]. Structural separation establishes rx=0, while reducing attainable utility. A deployment declaring a utility target q decides whether that exchange is acceptable; the certificate states the supported guarantee, not whether to accept it.

The schedule names what is missing, and the sharpest case is one where only the last operand is absent. The f-secure LLM system disaggregates planning from execution and keeps untrusted input out of the planning stage, and its Theorem 6.2 proves execution trace non-compromise, a noninterference property over the declared untrusted class [212]. Coverage and class-uniform failure are therefore established by proof rather than measured: αx=1 and ϵx=0 for plan compromise. The source then names the surviving channel itself. The attacker may still influence the data being processed, and their inability to influence the plan “drastically limits the scale and scope of any possible attacks.” This statement identifies the continuation operand but supplies neither a bound on it nor a comparison with x, the baseline for what the attacker could achieve using only resources outside the deployed system. By Theorem 9, Ux=1αx(1ϵx)(1rx)=rx, so the whole upper endpoint passes to an unvalued term and Ux=1. Two gates closed by proof narrow the interval by nothing when the third is left without a value.

5.7 What Transfers

The transferable relations hold between evidence and conclusions, not between mechanisms. The quantifier asymmetry separating the two directions follows from the supremum in Equation 6 and applies to any safeguard. Every established positive lower bound in either coding stratum comes from a strategy the source itself executed. An upper endpoint below one requires coverage, conditional failure, and continuation together, as the Fides instance shows.

Whether the strict boundary is attainable depends on the declared outcome together with the traces the safeguard leaves reachable. Restriction alone cannot reach it: removing behaviors cannot create a useful trace with no adverse value. A zero-residual certificate through mediation requires rx=0, so useful behavior must supply no adverse value. This separation can hold for a service-integrity outcome because completing a task does not necessarily require the prohibited action. For an external-world outcome, the same output can supply both legitimate value and uplift toward the harm. When the uniform dual-use condition holds with ρ>0, Equation 18 gives a positive floor ρq for every exactly simulable deployment that meets a positive utility target.

Of the 152 wide-coded instances, 81 concern service-integrity outcomes, 52 concern external-world outcomes, and 19 are mixed. The zero-residual certificate above is a service-integrity instance. For external-world outcomes, wherever the dual-use floor binds, the reportable target is a bounded τ supported by α, ϵ, and r.

These relations are what a catalogue organized by technique cannot express. Fides and CaMeL reach different conclusions inside one architecture family because their supported quantifiers and deployed paths differ. RESTA and Emulated Disalignment reach the same conclusion through unrelated interventions, because each ran a strategy from its own declared class and reported the adverse value it produced. What transfers across mechanism families is evidence that fills a particular slot under a fixed anchor. Section 7 derives what to report and what to build.

6 Relation to Prior Systematizations

Prior systematizations of this literature differ in which step of the path from a safeguard mechanism to a deployment decision they hold fixed, and each step supplies something different. Fixing what is compared supplies a shared vocabulary and a position for every mechanism, attack, and evaluation resource: conversation-safety and jailbreak surveys, guardrail reviews covering desired properties and the systems development lifecycle, systematic reviews extending defense taxonomies, and recent SoKs constructing multidimensional taxonomies all do this [56, 225, 53, 54, 42, 203, 78, 221]. Fixing what an evaluation must declare supplies explicit provenance for whatever it reports: guidance of this kind asks evaluators to state threat actors, requirements, access conditions, and supporting evidence [192, 21, 167]. Fixing the conditions under which a number is produced supplies protocol-level comparability: JailbreakRadar scores 17 attacks from its own taxonomy against nine aligned models and eight defenses in one shared setting [38], and TeleAI-Safety runs 19 attacks, 29 defenses, and 19 evaluation methods as interchangeable components of one protocol over 14 target models [31]. Those same SoKs [203, 78, 221] re-evaluate attacks and defenses under matched configurations, comparing security, efficiency, utility, cost, and judge choices; cross-model red teaming instead holds a fixed prompt corpus and asks which model resists it [161, 92]. Fixing how an outcome is scored supplies numbers that carry the same meaning across methods: GuidedBench shows that evaluation systems without case-specific criteria yield effectiveness estimates that do not support comparison across methods [85]. PandaGuard reaches the same point from the scoring side, finding across a grid of 19 attacks and 12 defenses over 49 models that judge disagreement introduces nontrivial variance in the resulting safety assessment [180]. Fixing the argument from one measurement to one deployment supplies a deployment conclusion for that deployment: safety cases and uplift analyses connect particular measurements to broader models under explicit assumptions, one argument at a time [40, 134, 194].

The first four steps deliver a comparable reported quantity with explicit provenance. That is not yet a statement about the residual: an attack-success rate that every paper computes identically, on a declared threat model, leaves open what it establishes about the assistance a guarded service still supplies. The fifth step reaches such a statement for a single deployment, under assumptions selected for it. Between them is the conversion this paper provides: it acts on the reported quantity itself, reads it on the outcome scale and attacker class its source declared, and returns the strongest conclusion about the residual that follows. Each further entry at any of the five steps enlarges the evidence base available to this conversion. Our comparison unit is accordingly a source-anchored deployment-safety claim instance. We apply the schedule and record which coding slots prevent a stronger conclusion. The common schedule and coding procedure make that conversion auditable across safeguard families before any individual deployment argument is attempted. Mechanistic work on refusal separates the scored utterance from the mechanism behind it: the refusal an evaluation scores can be traced to a small set of residual-stream features and ablated away [6], and redundant features behind them stay dormant until those are suppressed [166].

7 Discussion

Section 5 illustrates which measurements can tighten a deployment conclusion and which cannot, even if their reported values improve.

7.1 What an Evaluation Should Report

An evaluation intended to support a deployment claim should first fix the anchor θx of Equation 30. It should then identify the schedule row that supports the intended conclusion and report all quantities required by that row [192, 21, 167]. Improving one reported number cannot tighten the bound when another required quantity is missing. For lower bounds, the attack-witness row needs no additional evidence. The frontier route requires both a deployment meeting the utility target and the LB2, LB3, and LB4 operands. None of the seven depth-coded capability-removal instances in Section 5 completes that route. The four changes below address the upper side.

Report coverage as an architecture fact.

α asks which paths to the anchored outcome must enter the mediation domain. It is answered by enumerating the deployment’s paths and showing the domain on each, not by any accuracy figure. A path that resets scoped state belongs in that enumeration.

Report continuation, or the value of accuracy is undetermined.

r asks what adverse value remains available after the event succeeds: content already released, actions the policy still permits, state accumulated before the check, and further attempts within the horizon. Because Equation 24 scales every accuracy gain by 1r, an unreported r leaves the value of a reported improvement undetermined rather than merely unstated.

Report the dependence premise, not the layer count.

Error rates across layers can be multiplied only if each layer’s conditional bound remains valid after every preceding interaction history, at the deployed history grain. An evaluation should report this dependence condition and the grain at which it holds, not only the number of layers.

For transferable artifacts, declare a covering outer class.

Released weights admit no mediation event. Their upper-bound route therefore requires a uniform value bound over a declared tampering class that contains Σ. Testing more individual attacks cannot supply this uniform bound: each test adds only a lower-bound witness.

7.2 What the Schedule Implies for Design

Each gate requires a different kind of intervention. Lowering ϵ is a statistical problem in the classifier [89, 75, 45]. Raising α is a routing problem in the architecture [47, 43, 182]. Driving r to zero is a structural problem: it requires that a covered success leave no adverse continuation available. Detection accuracy cannot deliver that at any value.

For an external-world outcome where the dual-use floor binds, Section 5 leaves a bounded τ as the reportable target. Choosing that value is a deployment decision rather than an evaluation result: it requires the constraints in and the utility requirement B(g)q.

A certificate binds to its anchor and to the premises of the row that produced it. A second deployment inherits it only by preserving those coordinates and premises, or through a containment argument covering the new scope.

The coded set is a saturation sample rather than a census.

8 Conclusion

A local safeguard score establishes that a control worked in a specified test, while deployment safety concerns the harmful assistance the guarded service still supplies. A single successful attack can establish that harmful assistance remains. Showing how little remains requires a conjunction: coverage of the relevant attack paths, conditional failure on those paths, and the continuation available after a safeguard succeeds. One coded claim supplies that conjunction and rules out the worst case. When a quantity yields no tighter bound, what is missing is a different quantity, not a more precise version of the same one. A safeguard can work exactly as tested while the safety of the deployed system remains unresolved.

Ethical Considerations

This paper systematizes results that are already published. Every quantity it recomputes is one its source already reported, the evidence base is the public literature, and the full-text coding was executed by the two independent model channels on public papers. The work involved no attack execution, no deployed system, and no human subjects.

The stakeholders are the authors who report safeguard evaluations, the reviewers and evaluators who read those reports, and the operators who act on them. The schedule states which reported quantity bounds deployment risk and which returns no bound, so what it supplies is a standard for evaluation rather than a capability for attack. An adversary gains nothing the cited papers do not already state. On that basis we consider publication justified.

Open Science

The supplementary materials document the wide-coded claim instances in this paper. The first file states the common full-text review task applied by both independent model channels. The second lists every source-channel record of the wide-coded portion of the full-text coding pass with its final eligibility status, claim-instance status, and instance count. The third and fourth compile the final per-source assessments from the Fable and GPT channels, respectively, including the source version used, eligibility decision, extracted instances, evidence assessments and locators, residual conclusion, claim verdict, and recorded boundary cases. The supplement also includes the coding instrument applied by both channels; the assessment records cite its numbered rules and global conventions. The depth-coded subset is documented in Appendix C rather than here: Section C.3 gives its aggregates and Sections C.5 and C.6 reproduce two case records.

The supplementary materials are available in the project repository.

Together, these files document the executed review decisions behind the aggregate results. Third-party full texts are not redistributed. The assessment files contain source identifiers, locators, and the excerpts needed to substantiate a judgment.

References

  • [1] A. Ablove, S. Chandrashekaran, X. Qiang, and R. Ensafi (2026) Characterizing the implementation of censorship policies in chinese LLM services. In 33rd Annual Network and Distributed System Security Symposium, NDSS 2026, San Diego, California, USA, February 23-27, 2026, Cited by: §B.3.
  • [2] M. Andriushchenko, F. Croce, and N. Flammarion (2025) Jailbreaking leading safety-aligned LLMs with simple adaptive attacks. In The Thirteenth International Conference on Learning Representations, Cited by: §1.
  • [3] Anthropic (2025) Detecting and countering misuse of AI: august 2025. Note: Anthropic Threat Intelligence ReportPublished August 27, 2025; accessed August 20, 2026 External Links: Link Cited by: §1.
  • [4] Anthropic (2026) Claude Fable 5 and Claude Mythos 5. Note: Model announcement, June 9, 2026. Accessed 2026-08-20 External Links: Link Cited by: §B.3.
  • [5] Anthropic (2026) How we contain Claude across products. Note: Anthropic EngineeringPublished May 25, 2026; accessed August 14, 2026 External Links: Link Cited by: §1.
  • [6] A. Arditi, O. Obeso, A. Syed, D. Paleka, N. Panickssery, W. Gurnee, and N. Nanda (2024) Refusal in language models is mediated by a single direction. In Advances in Neural Information Processing Systems, Vol. 37. External Links: Document Cited by: §1, §6.
  • [7] A. Athalye, N. Carlini, and D. Wagner (2018) Obfuscated gradients give a false sense of security: circumventing defenses to adversarial examples. In Proceedings of the 35th International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, Vol. 80, pp. 274–283. Cited by: §3.3.1.
  • [8] Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, C. Chen, C. Olsson, C. Olah, D. Hernandez, D. Drain, D. Ganguli, D. Li, E. Tran-Johnson, E. Perez, J. Kerr, J. Mueller, J. Ladish, J. Landau, K. Ndousse, K. Lukosuite, L. Lovitt, M. Sellitto, N. Elhage, N. Schiefer, N. Mercado, N. DasSarma, R. Lasenby, R. Larson, S. Ringer, S. Johnston, S. Kravec, S. El Showk, S. Fort, T. Lanham, T. Telleen-Lawton, T. Conerly, T. Henighan, T. Hume, S. R. Bowman, Z. Hatfield-Dodds, B. Mann, D. Amodei, N. Joseph, S. McCandlish, T. Brown, and J. Kaplan (2022) Constitutional ai: harmlessness from ai feedback. External Links: 2212.08073 Cited by: §1.
  • [9] Y. F. Bakman, D. N. Yaldiz, S. Kang, T. Zhang, B. Buyukates, S. Avestimehr, and S. P. Karimireddy (2025) Reconsidering LLM uncertainty estimation methods in the wild. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 29531–29556. External Links: Document Cited by: §B.3.
  • [10] A. R. Basani and X. Zhang (2025) GASP: efficient black-box generation of adversarial suffixes for jailbreaking llms. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
  • [11] S. Belkadi, L. Ren, N. Micheletti, L. Han, and G. Nenadic (2025) Generating synthetic free-text medical records with low re-identification risk using masked language modeling. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 4: Student Research Workshop, Albuquerque, NM, USA, April 30 - May 1, 2025, A. Ebrahimi, S. Haider, E. Liu, S. Haider, M. L. Pacheco, and S. Wein (Eds.), pp. 200–206. External Links: Document Cited by: §B.3.
  • [12] R. Bhardwaj, D. A. Do, and S. Poria (2024) Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), Volume 1: Long Papers, Bangkok, Thailand, pp. 14138–14149. External Links: Document Cited by: §B.3, §5.2.
  • [13] B. Bi, S. Huang, Y. Wang, T. Yang, Z. Zhang, H. Huang, L. Mei, J. Fang, Z. Li, F. Wei, W. Deng, F. Sun, Q. Zhang, and S. Liu (2025) Context-dpo: aligning language models for context-faithfulness. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Findings of ACL, Vol. ACL 2025, pp. 10280–10300. External Links: Document Cited by: §B.3.
  • [14] T. Bi, C. Ye, Z. Yang, Z. Zhou, C. Tang, Z. Tao, J. Zhang, K. Wang, L. Zhou, Y. Yang, and T. Yu (2026) On the feasibility of using multimodal llms to execute AR social engineering attacks. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 38252–38260. External Links: Document Cited by: §B.3.
  • [15] J. Binkowski, D. Janiak, A. Sawczyn, B. Gabrys, and T. Kajdanowicz (2025) Hallucination detection in llms using spectral features of attention maps. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 24354–24385. External Links: Document Cited by: §B.3.
  • [16] D. Bowen, B. Murphy, W. Cai, D. Khachaturov, A. Gleave, and K. Pelrine (2025) Scaling trends for data poisoning in llms. In Thirty-Ninth AAAI Conference on Artificial Intelligence, Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence, Fifteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2025, Philadelphia, PA, USA, February 25 - March 4, 2025, T. Walsh, J. Shah, and Z. Kolter (Eds.), pp. 27206–27214. External Links: Document Cited by: §B.3.
  • [17] L. Bürger, F. A. Hamprecht, and B. Nadler (2024) Truth is universal: robust detection of lies in llms. In Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang (Eds.), Cited by: §B.3.
  • [18] Y. Cai, R. Gu, J. Li, X. Huang, J. Chen, X. Gu, and M. Huang (2025) MHALO: evaluating mllms as fine-grained hallucination detectors. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Findings of ACL, Vol. ACL 2025, pp. 9197–9222. External Links: Document Cited by: §B.3.
  • [19] Y. Cao, B. Cao, and J. Chen (2024) Stealthy and persistent unalignment on large language models via backdoor injections. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024, K. Duh, H. Gómez-Adorno, and S. Bethard (Eds.), pp. 4920–4935. External Links: Document Cited by: §B.3.
  • [20] N. Carlini and D. Wagner (2017) Adversarial examples are not easily detected: bypassing ten detection methods. In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security (AISec ’17), pp. 3–14. External Links: Document Cited by: §3.1.
  • [21] S. Casper, C. Ezell, C. Siegmann, N. Kolt, T. L. Curtis, B. Bucknall, A. Haupt, K. Wei, J. Scheurer, M. Hobbhahn, L. Sharkey, S. Krishna, M. Von Hagen, S. Alberti, A. Chan, Q. Sun, M. Gerovitch, D. Bau, M. Tegmark, D. Krueger, and D. Hadfield-Menell (2024) Black-box access is insufficient for rigorous ai audits. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pp. 2254–2272. External Links: Document Cited by: §1, §6, §7.1.
  • [22] N. Chakraborty, J. Pohovey, M. Ornik, and K. R. Driggs-Campbell (2026) Characterizing the robustness of black-box LLM planners under perturbed observations with adaptive stress testing. In Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 39445–39475. External Links: Document Cited by: §B.3.
  • [23] A. Chandler, D. Surve, and H. Su (2024) Detecting errors through ensembling prompts (DEEP): an end-to-end LLM framework for detecting factual errors. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), pp. 13120–13133. External Links: Document Cited by: §B.3.
  • [24] T. Chang, T. Schnabel, A. Swaminathan, and J. Wiens (2026) A course correction in steerability evaluation: revealing miscalibration and side effects in llms. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 37259–37267. External Links: Document Cited by: §B.3.
  • [25] P. Chao, E. Debenedetti, A. Robey, M. Andriushchenko, F. Croce, V. Sehwag, E. Dobriban, N. Flammarion, G. J. Pappas, F. Tramèr, H. Hassani, and E. Wong (2024) JailbreakBench: an open robustness benchmark for jailbreaking large language models. In Advances in Neural Information Processing Systems, Vol. 37. External Links: Document Cited by: §1.
  • [26] C. H. Chen, H. Huang, and H. Chen (2025) Self-augmented preference alignment for sycophancy reduction in llms. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 12379–12391. External Links: Document Cited by: §B.3.
  • [27] G. Chen, Z. Qin, M. Yang, Y. Zhou, T. Fan, T. Du, and Z. Xu (2024) Unveiling the vulnerability of private fine-tuning in split-based frameworks for large language models: A bidirectionally enhanced attack. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, CCS 2024, Salt Lake City, UT, USA, October 14-18, 2024, B. Luo, X. Liao, J. Xu, E. Kirda, and D. Lie (Eds.), pp. 2904–2918. External Links: Document Cited by: §B.3.
  • [28] H. Chen and S. Goldfarb-Tarrant (2025) Safer or luckier? llms as safety evaluators are not robust to artifacts. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 19750–19766. External Links: Document Cited by: §B.3.
  • [29] M. Chen, Y. Cao, Y. Zhang, and C. Lu (2024) Quantifying and mitigating unimodal biases in multimodal large language models: A causal perspective. In Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Findings of ACL, Vol. EMNLP 2024, pp. 16449–16469. External Links: Document Cited by: §B.3.
  • [30] X. Chen, H. Wen, S. Nag, C. Luo, Q. Yin, R. Li, Z. Li, and W. Wang (2024) IterAlign: iterative constitutional alignment of large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024, K. Duh, H. Gómez-Adorno, and S. Bethard (Eds.), pp. 1423–1433. External Links: Document Cited by: §B.3.
  • [31] X. Chen, J. Zhao, Y. He, Y. Xun, X. Liu, Y. Li, H. Zhou, W. Cai, Z. Shi, Y. Yuan, T. Zhang, C. Zhang, and X. Li (2025) TeleAI-Safety: a comprehensive LLM jailbreaking benchmark towards attacks, defenses, and evaluations. arXiv preprint arXiv:2512.05485. External Links: 2512.05485, Document Cited by: §6.
  • [32] Y. Chen, N. Joshi, Y. Chen, M. Andriushchenko, R. Angell, and H. He (2026) Monitoring decomposition attacks in LLMs with lightweight sequential monitors. In The Fourteenth International Conference on Learning Representations, Cited by: §5.3.
  • [33] Z. Chen, S. Shen, G. Shen, G. Zhi, X. Chen, and Y. Lin (2024) Towards tool use alignment of large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), pp. 1382–1400. External Links: Document Cited by: §B.3.
  • [34] Z. Chen, H. Lin, K. Li, Z. Luo, Z. Ye, G. Chen, Z. Huang, and J. Ma (2025) AdamMeme: adaptively probe the reasoning capacity of multimodal large language models on harmfulness. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 4234–4253. External Links: Document Cited by: §B.3.
  • [35] J. Cheng, T. Su, J. Yuan, G. He, J. Liu, X. Tao, J. Xie, and H. Li (2025) Chain-of-thought prompting obscures hallucination cues in large language models: an empirical evaluation. In Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 1272–1305. External Links: Document Cited by: §B.3.
  • [36] X. Cheng, R. Chen, H. Zan, Y. Jia, and M. Peng (2025) BiasFilter: an inference-time debiasing framework for large language models. In Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 15187–15205. External Links: Document Cited by: §B.3.
  • [37] Y. Cheng, V. S. Sadasivan, M. Saberi, S. Saha, and S. Feizi (2025) Adversarial paraphrasing: A universal attack for humanizing ai-generated text. CoRR abs/2506.07001. External Links: Document, 2506.07001 Cited by: §B.3.
  • [38] J. Chu, Y. Liu, Z. Yang, X. Shen, M. Backes, and Y. Zhang (2025) JailbreakRadar: comprehensive assessment of jailbreak attacks against LLMs. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), Volume 1: Long Papers, pp. 21538–21566. External Links: Document Cited by: §6.
  • [39] Z. Chu, Y. Wang, L. Li, Z. Wang, Z. Qin, and K. Ren (2024) A causal explainable guardrails for large language models. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, CCS 2024, Salt Lake City, UT, USA, October 14-18, 2024, B. Luo, X. Liao, J. Xu, E. Kirda, and D. Lie (Eds.), pp. 1136–1150. External Links: Document Cited by: §B.3.
  • [40] J. Clymer, J. Weinbaum, R. Kirk, K. Mai, S. Zhang, and X. Davies (2025) An example safety case for safeguards against misuse. Note: arXiv preprint External Links: 2505.18003, Document Cited by: §1, §6.
  • [41] S. Cohen, R. Bitton, and B. Nassi (2025) Here comes the AI worm: preventing the propagation of adversarial self-replicating prompts within genai ecosystems. In Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security, CCS 2025, Taipei, Taiwan, October 13-17, 2025, C. Huang, J. Chen, S. Shieh, D. Lie, and V. Cortier (Eds.), pp. 3975–3989. External Links: Document Cited by: §B.3.
  • [42] P. H. B. Correia, R. W. Achjian, D. E. G. C. de Oliveira, Y. A. Maria, V. T. Hayashi, M. Lopes, C. C. Miers, and M. A. Simplicio (2026) A systematic literature review on LLM defenses against prompt injection and jailbreaking: expanding NIST taxonomy. arXiv preprint arXiv:2601.22240. External Links: 2601.22240, Document Cited by: §6.
  • [43] M. Costa, B. Köpf, A. Kolluri, A. Paverd, M. Russinovich, A. Salem, S. Tople, L. Wutschitz, and S. Zanella-Béguelin (2025) Securing AI agents with information-flow control. arXiv preprint arXiv:2505.23643. External Links: Document Cited by: §C.5, §1, §5.6, §5.6, §7.2.
  • [44] T. Coste, U. Anwar, R. Kirk, and D. Krueger (2023) Reward model ensembles help mitigate overoptimization. CoRR abs/2310.02743. External Links: Document, 2310.02743 Cited by: §B.3.
  • [45] H. Cunningham, J. Wei, Z. Wang, A. Persic, A. Peng, J. Abderrachid, R. Agarwal, B. Chen, A. Cohen, A. Dau, A. Dimitriev, R. Gilson, L. Howard, Y. Hua, J. Kaplan, J. Leike, M. Lin, C. Liu, V. Mikulik, R. Mittapalli, C. O’Hara, J. Pan, N. Saxena, A. Silverstein, Y. Song, X. Yu, G. Zhou, E. Perez, and M. Sharma (2026) Constitutional classifiers++: efficient production-grade defenses against universal jailbreaks. Note: arXiv preprint External Links: 2601.04603, Document Cited by: §1, §7.2.
  • [46] J. Dai, X. Pan, R. Sun, J. Ji, X. Xu, M. Liu, Y. Wang, and Y. Yang (2024) Safe RLHF: safe reinforcement learning from human feedback. In The Twelfth International Conference on Learning Representations, Cited by: §1.
  • [47] E. Debenedetti, I. Shumailov, T. Fan, J. Hayes, N. Carlini, D. Fabian, C. Kern, C. Shi, A. Terzis, and F. Tramèr (2025) Defeating prompt injections by design. Note: arXiv preprint External Links: 2503.18813, Document Cited by: §1, §5.6, §7.2.
  • [48] E. Debenedetti, J. Zhang, M. Balunovic, L. Beurer-Kellner, M. Fischer, and F. Tramèr (2024) AgentDojo: a dynamic environment to evaluate prompt injection attacks and defenses for LLM agents. In Advances in Neural Information Processing Systems 37, Note: Datasets and Benchmarks Track External Links: Document Cited by: §1.
  • [49] A. Delaval, S. Yang, H. Wang, H. Qiu, and J. Lu (2026) TOXIFRENCH: benchmarking and enhancing language models via cot fine-tuning for french toxicity detection. In Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 21354–21375. External Links: Document Cited by: §B.3.
  • [50] G. Deng, Y. Liu, Y. Li, K. Wang, Y. Zhang, Z. Li, H. Wang, T. Zhang, and Y. Liu (2024) MASTERKEY: automated jailbreaking of large language model chatbots. In 31st Annual Network and Distributed System Security Symposium, NDSS 2024, San Diego, California, USA, February 26 - March 1, 2024, Cited by: §B.3.
  • [51] Y. Deng, W. Zhang, S. J. Pan, and L. Bing (2024) Multilingual jailbreak challenges in large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024, Cited by: §B.3.
  • [52] C. Ding, J. Wu, Y. Yuan, J. Lu, K. Zhang, A. Su, X. Wang, and X. He (2025) Unified parameter-efficient unlearning for llms. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
  • [53] Y. Dong, R. Mu, G. Jin, Y. Qi, J. Hu, X. Zhao, J. Meng, W. Ruan, and X. Huang (2024) Position: building guardrails for large language models requires systematic design. In Proceedings of the 41st International Conference on Machine Learning (ICML), Proceedings of Machine Learning Research, Vol. 235, pp. 11375–11394. Cited by: §6.
  • [54] Y. Dong, R. Mu, Y. Zhang, S. Sun, T. Zhang, C. Wu, G. Jin, Y. Qi, J. Hu, J. Meng, S. Bensalem, and X. Huang (2025) Safeguarding large language models: a survey. Artificial Intelligence Review 58 (12), pp. 382. External Links: Document Cited by: §6.
  • [55] Y. R. Dong, H. Lin, M. Belkin, R. Huerta, and I. Vulic (2025) UNDIAL: self-distillation with adjusted logits for robust unlearning in large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, L. Chiruzzo, A. Ritter, and L. Wang (Eds.), pp. 8827–8840. External Links: Document Cited by: §B.3.
  • [56] Z. Dong, Z. Zhou, C. Yang, J. Shao, and Y. Qiao (2024) Attacks, defenses and evaluations for LLM conversation safety: a survey. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Mexico City, Mexico, pp. 6734–6747. External Links: Document Cited by: §1, §6.
  • [57] Y. Du, S. Zhao, D. Zhao, M. Ma, Y. Chen, L. Huo, Q. Yang, D. Xu, and B. Qin (2024) MoGU: A framework for enhancing safety of llms while preserving their usability. In Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang (Eds.), Cited by: §B.3.
  • [58] A. Dutta, R. Magu, S. Kim, S. Yoon, M. D. Choudhury, and A. R. KhudaBukhsh (2026) Auditing LLM responses to harmful stereotypes targeting mental health groups. In Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 40435–40452. External Links: Document Cited by: §B.3.
  • [59] X. Fang, Z. Tian, Z. Huang, Z. Pan, Z. Wen, X. Wang, Q. Fang, and D. Li (2026) Knowledge injection exists in moe? exploring expert-aware contrast decoding in moe for mitigating llms’ hallucinations. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 39326–39343. External Links: Document Cited by: §B.3.
  • [60] R. Feng, A. Hooda, N. Mangaokar, K. Fawaz, S. Jha, and A. Prakash (2023) Stateful defenses for machine learning models are not yet secure against black-box attacks. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security, pp. 786–800. External Links: Document Cited by: §5.3.
  • [61] G. Filandrianos, A. Dimitriou, M. Lymperaiou, K. Thomas, and G. Stamou (2025) Bias beware: the impact of cognitive biases on LLM-driven product recommendations. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 22397–22426. External Links: Document, ISBN 979-8-89176-332-6 Cited by: §B.3.
  • [62] J. Fonseca, A. Bell, and J. Stoyanovich (2025) SAFENUDGE: safeguarding large language models in real-time with tunable safety-performance trade-offs. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 19955–19969. External Links: Document Cited by: §B.3.
  • [63] B. Formento, C. Foo, and S. Ng (2025) Confidence elicitation: A new attack vector for large language models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
  • [64] H. Fu, W. Peng, Y. Zhou, J. Wu, J. Wen, and Y. Xue (2026) Inhibitory attacks on backdoor-based fingerprinting for large language models. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 26246–26266. External Links: Document Cited by: §B.3.
  • [65] H. Geng, Y. Huang, L. Lai, Q. Du, H. Chu, Z. He, J. Hu, and X. Tao (2026) ProMedical: hierarchical fine-grained criteria modeling for medical LLM alignment via explicit injection. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 36955–36994. External Links: Document Cited by: §B.3.
  • [66] K. Gligorić, M. Cheng, L. Zheng, E. Durmus, and D. Jurafsky (2024) NLP systems that can’t tell use from mention censor counterspeech, but teaching the distinction helps. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), K. Duh, H. Gomez, and S. Bethard (Eds.), Mexico City, Mexico, pp. 5942–5959. External Links: Document Cited by: §B.3.
  • [67] X. Gong, M. Li, Y. Zhang, F. Ran, C. Chen, Y. Chen, Q. Wang, and K. Lam (2025) PAPILLON: efficient and stealthy fuzz testing-powered jailbreaks for llms. In 34th USENIX Security Symposium, USENIX Security 2025, Seattle, WA, USA, August 13-15, 2025, L. Bauer and G. Pellegrino (Eds.), pp. 2401–2420. Cited by: §B.3.
  • [68] K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz (2023) Not what you’ve signed up for: compromising real-world llm-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security, pp. 79–90. External Links: Document Cited by: §5.6.
  • [69] T. Gu, Z. Wang, K. Huang, Y. Yao, X. Zhang, Y. Yang, and X. Chen (2025) Invisible entropy: towards safe and efficient low-entropy LLM watermarking. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 6716–6733. External Links: Document Cited by: §B.3.
  • [70] T. Gu, Z. Zhou, K. Huang, D. Liang, Y. Wang, H. Zhao, Y. Yao, X. Qiao, K. Wang, Y. Yang, Y. Teng, Y. Qiao, and Y. Wang (2024) MLLMGuard: A multi-dimensional safety evaluation suite for multimodal large language models. In Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang (Eds.), Cited by: §B.3.
  • [71] P. Guo, A. Syed, A. Sheshadri, A. Ewart, and G. K. Dziugaite (2024) Mechanistic unlearning: robust knowledge unlearning and editing via mechanistic localization. CoRR abs/2410.12949. External Links: Document, 2410.12949 Cited by: §B.3.
  • [72] Z. Guo, Y. Shi, W. Meng, C. Gong, C. Wei, and W. Chen (2025) Be cautious when merging unfamiliar llms: A phishing model capable of stealing privacy. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Findings of ACL, Vol. ACL 2025, pp. 13852–13871. External Links: Document Cited by: §B.3.
  • [73] S. Gupta, V. Shrivastava, A. Deshpande, A. Kalyan, P. Clark, A. Sabharwal, and T. Khot (2023) Bias runs deep: implicit reasoning biases in persona-assigned llms. CoRR abs/2311.04892. External Links: Document, 2311.04892 Cited by: §B.3.
  • [74] T. Hagendorff, E. Derner, and N. Oliver (2026) Large reasoning models are autonomous jailbreak agents. Nature Communications 17 (1), pp. 1435. External Links: Document, 2508.04039 Cited by: §1.
  • [75] S. Han, K. Rao, A. Ettinger, L. Jiang, B. Y. Lin, N. Lambert, Y. Choi, and N. Dziri (2024) WildGuard: open one-stop moderation tools for safety risks, jailbreaks, and refusals of LLMs. In Advances in Neural Information Processing Systems, Vol. 37. External Links: Document Cited by: §1, §7.2.
  • [76] K. D. Hayes, M. Goldblum, V. Sehwag, G. Somepalli, A. Panda, and T. Goldstein (2025) FineGRAIN: evaluating failure modes of text-to-image models with vision language model judges. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
  • [77] J. He, W. Jiang, G. Hou, W. Fan, R. Zhang, and H. Li (2025) Watch out for your guidance on generation! exploring conditional backdoor attacks against large language models. In Thirty-Ninth AAAI Conference on Artificial Intelligence, Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence, Fifteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2025, Philadelphia, PA, USA, February 25 - March 4, 2025, T. Walsh, J. Shah, and Z. Kolter (Eds.), pp. 26220–26228. External Links: Document Cited by: §B.3.
  • [78] H. Hong, S. Wu, S. Feng, N. Naderloui, S. Yan, J. Zhang, A. Arastehfard, H. Huang, and Y. Hong (2025) SoK: systematizing LLM prompt security: taxonomies, datasets, and unified evaluation of attacks and defenses. arXiv preprint arXiv:2510.15476. External Links: 2510.15476, Document Cited by: §1, §6.
  • [79] W. Hou, H. Tu, Y. Wang, Y. Zhang, Y. Liu, D. Zhu, L. Gao, and B. Zhou (2026) Beyond single-view detection: a dual-space reasoning framework for interpretable harmful meme understanding. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 10526–10544. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §B.3.
  • [80] J. Hu, X. Huang, Y. Sun, Y. Dong, and X. Huang (2026) Lying with truths: open-channel multi-agent collusion for belief manipulation via generative montage. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 5979–5996. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §B.3.
  • [81] Z. Hu, L. Shen, Z. Wang, Y. Wei, and D. Tao (2025) Adaptive defense against harmful fine-tuning for large language models via bayesian data scheduler. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
  • [82] K. Huang, X. Liu, Q. Guo, T. Sun, J. Sun, Y. Wang, Z. Zhou, Y. Wang, Y. Teng, X. Qiu, Y. Wang, and D. Lin (2024) Flames: benchmarking value alignment of llms in chinese. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024, K. Duh, H. Gómez-Adorno, and S. Bethard (Eds.), pp. 4551–4591. External Links: Document Cited by: §B.3.
  • [83] L. Huang, X. Jiang, Z. Wang, W. Mo, X. Xiao, Y. Yin, B. Han, and F. Zheng (2026) Transferability of adversarial attacks in video-based mllms: A cross-modal image-to-video approach. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 5067–5075. External Links: Document Cited by: §B.3.
  • [84] Q. Huang, J. Zhang, J. Wu, Y. Li, W. Zhang, Y. Rong, J. Yao, S. Zhang, and X. Jia (2026) JailMeter: an evidence-based evaluation framework for jailbreak attacks on large language models. In Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 16006–16029. External Links: Document Cited by: §B.3.
  • [85] R. Huang, X. Wang, Z. Li, D. Wu, and S. Wang (2025) GuidedBench: measuring and mitigating the evaluation discrepancies of in-the-wild LLM jailbreak methods. arXiv preprint arXiv:2502.16903. External Links: 2502.16903, Document Cited by: §6.
  • [86] Y. Huang, L. Zhang, and C. Wang (2026) How do llms "trust" unknown knowledge? an unknown knowledge based jailbreak attack. In Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 37105–37124. External Links: Document Cited by: §B.3.
  • [87] Y. Huang, H. Wang, X. Bai, J. Wang, J. Liu, Z. Wang, W. Ni, S. Wang, and T. Qi (2026) Robust membership inference for large language models under adversarial generative corruption. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 39531–39547. External Links: Document Cited by: §B.3.
  • [88] Y. Huang, L. Sun, H. Wang, S. Wu, Q. Zhang, Y. Li, C. Gao, Y. Huang, W. Lyu, Y. Zhang, X. Li, Z. Liu, Y. Liu, Y. Wang, Z. Zhang, B. Vidgen, B. Kailkhura, C. Xiong, C. Xiao, C. Li, E. Xing, F. Huang, H. Liu, H. Ji, H. Wang, H. Zhang, H. Yao, M. Kellis, M. Zitnik, M. Jiang, M. Bansal, J. Zou, J. Pei, J. Liu, J. Gao, J. Han, J. Zhao, J. Tang, J. Wang, J. Vanschoren, J. Mitchell, K. Shu, K. Xu, K. Chang, L. He, L. Huang, M. Backes, N. Z. Gong, P. S. Yu, P. Chen, Q. Gu, R. Xu, R. Ying, S. Ji, S. Jana, T. Chen, T. Liu, T. Zhou, W. Wang, X. Li, X. Zhang, X. Wang, X. Xie, X. Chen, X. Wang, Y. Liu, Y. Ye, Y. Cao, Y. Chen, and Y. Zhao (2024) TrustLLM: trustworthiness in large language models. CoRR abs/2401.05561. External Links: Document, 2401.05561 Cited by: §B.3.
  • [89] H. Inan, K. Upasani, J. Chi, R. Rungta, K. Iyer, Y. Mao, M. Tontchev, Q. Hu, B. Fuller, D. Testuggine, and M. Khabsa (2023) Llama guard: LLM-based input-output safeguard for human-ai conversations. External Links: 2312.06674 Cited by: §B.3, §1, §7.2.
  • [90] C. Isch and G. Jennings (2026) Narrative license and model sycophancy in LLM summaries of scientific work. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 16418–16432. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §B.3.
  • [91] D. Jain, D. Hartmann, and C. Li (2026) Adaptive adversaries: a multi-turn, multi-llm benchmark for LLM agent security. Note: arXiv preprint External Links: 2607.18063, Document Cited by: §1.
  • [92] P. Jaiswal, A. Pratap, S. Saraswati, H. Kasyap, and S. Tripathy (2026) Analysis of LLMs against prompt injection and jailbreak attacks. In Proceedings of the Workshop on Privacy in Large Language Models (LLM) and Natural Language Processing (NLP) 2026, External Links: Document, 2602.22242 Cited by: §6.
  • [93] J. Jeon, J. Oh, H. Lee, and B. Lee (2025) Iterative prompt refinement for safer text-to-image generation. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 18080–18096. External Links: Document Cited by: §B.3.
  • [94] S. Jeoung, Y. Ge, and J. Diesner (2023) StereoMap: quantifying the awareness of human-like stereotypes in large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 12236–12256. External Links: Document Cited by: §B.3.
  • [95] P. Jha, R. Jain, K. Mandal, A. Chadha, S. Saha, and P. Bhattacharyya (2024) MemeGuard: an LLM and vlm-based framework for advancing content moderation via meme intervention. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), pp. 8084–8104. External Links: Document Cited by: §B.3.
  • [96] L. Jiang, Y. Li, X. Zhang, Y. Ding, and L. Pan (2025) SceneJailEval: a scenario-adaptive multi-dimensional framework for jailbreak evaluation. CoRR abs/2508.06194. External Links: Document, 2508.06194 Cited by: §B.3.
  • [97] P. Jiang, X. Lyu, Y. Li, and J. Ma (2025) Backdoor token unlearning: exposing and defending backdoors in pretrained language models. In Thirty-Ninth AAAI Conference on Artificial Intelligence, Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence, Fifteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2025, Philadelphia, PA, USA, February 25 - March 4, 2025, T. Walsh, J. Shah, and Z. Kolter (Eds.), pp. 24285–24293. External Links: Document Cited by: §B.3.
  • [98] M. Kang, Z. Chen, and B. Li (2025) C-safegen: certified safe LLM generation with claim-based streaming guardrails. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
  • [99] S. Kim and G. Lee (2026) Merging triggers, breaking backdoors: defensive poisoning for instruction-tuned language models. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 24269–24287. External Links: Document Cited by: §B.3.
  • [100] S. Kim, Y. Lee, Y. Song, and K. Lee (2025) What really matters in many-shot attacks? an empirical study of long-context vulnerabilities in llms. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 2043–2063. External Links: Document Cited by: §B.3.
  • [101] S. Kim, S. Yun, H. Lee, M. Gubri, S. Yoon, and S. J. Oh (2023) ProPILE: probing privacy leakage in large language models. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), Cited by: §B.3.
  • [102] C. Ko, P. Chen, P. Das, Y. Mroueh, S. Dan, G. Kollias, S. Chaudhury, T. Pedapati, and L. Daniel (2025) Large language models can become strong self-detoxifiers. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
  • [103] S. Kusaka, K. Saito, M. Kudo, T. Tanabe, A. Wachi, and Y. Akimoto (2026) Cost-minimized label-flipping poisoning attack to LLM alignment. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 37538–37546. External Links: Document Cited by: §B.3.
  • [104] F. Le, W. He, C. Cao, D. Liang, and Z. Cui (2025) DualCnst: enhancing zero-shot out-of-distribution detection via text-image consistency in vision-language models. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
  • [105] D. Lee, J. Jang, J. Jeong, and H. Yu (2025) Are vision-language models safe in the wild? A meme-based benchmark study. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 30545–30588. External Links: Document Cited by: §B.3.
  • [106] C. T. Leong, Y. Cheng, J. Wang, J. Wang, and W. Li (2023) Self-detoxifying language models via toxification reversal. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), pp. 4433–4449. External Links: Document Cited by: §B.3.
  • [107] H. Li, X. Liu, N. Zhang, and C. Xiao (2025) PIGuard: prompt injection guardrail via mitigating overdefense for free. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 30420–30437. External Links: Document Cited by: §B.3.
  • [108] H. Li, J. Ye, J. Wu, T. Yan, C. Wang, and Z. Li (2025) JailPO: A novel black-box jailbreak framework via preference optimization against aligned llms. In Thirty-Ninth AAAI Conference on Artificial Intelligence, Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence, Fifteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2025, Philadelphia, PA, USA, February 25 - March 4, 2025, T. Walsh, J. Shah, and Z. Kolter (Eds.), pp. 27419–27427. External Links: Document Cited by: §B.3.
  • [109] H. Li, S. Shan, E. Wenger, J. Zhang, H. Zheng, and B. Y. Zhao (2022) Blacklight: scalable defense for neural networks against Query-Based Black-Box attacks. In 31st USENIX Security Symposium (USENIX Security 22), Boston, MA, pp. 2117–2134. External Links: ISBN 978-1-939133-31-1 Cited by: §5.3.
  • [110] K. Li, O. Patel, F. B. Viégas, H. Pfister, and M. Wattenberg (2023) Inference-time intervention: eliciting truthful answers from a language model. CoRR abs/2306.03341. External Links: Document, 2306.03341 Cited by: §B.3.
  • [111] K. Li, J. Shi, Y. Xiao, M. Jiang, J. Sun, Y. Wu, D. Fu, S. Xia, X. Cai, T. Xu, W. Si, W. Li, D. Wang, and P. Liu (2026) AgencyBench: benchmarking the frontiers of autonomous agents in 1M-token real-world contexts. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 7422–7440. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §B.3.
  • [112] K. Li, L. M. Po, H. Yang, X. Xu, K. Liu, and Y. Zhao (2025) AesBiasBench: evaluating bias and alignment in multimodal language models for personalized image aesthetic assessment. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 7607–7620. External Links: Document Cited by: §B.3.
  • [113] L. Li, Y. Liu, D. He, and Y. Li (2025) One model transfer to all: on robust jailbreak prompts generation against llms. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
  • [114] N. Li, A. Pan, A. Gopal, S. Yue, D. Berrios, A. Gatti, J. D. Li, A. Dombrowski, S. Goel, G. Mukobi, N. Helm-Burger, R. Lababidi, L. Justen, A. B. Liu, M. Chen, I. Barrass, O. Zhang, X. Zhu, R. Tamirisa, B. Bharathi, A. Herbert-Voss, C. B. Breuer, A. Zou, M. Mazeika, Z. Wang, P. Oswal, W. Lin, A. A. Hunt, J. Tienken-Harder, K. Y. Shih, K. Talley, J. Guan, I. Steneker, D. Campbell, B. Jokubaitis, S. Basart, S. Fitz, P. Kumaraguru, K. K. Karmakar, U. Tupakula, V. Varadharajan, Y. Shoshitaishvili, J. Ba, K. M. Esvelt, A. Wang, and D. Hendrycks (2024) The WMDP benchmark: measuring and reducing malicious use with unlearning. In Proceedings of the 41st International Conference on Machine Learning, Vol. 235, pp. 28525–28550. Cited by: §1, §5.4.
  • [115] Q. Li, T. Luo, X. Zhang, Y. Xie, Z. Shen, L. Zhang, Y. Jin, H. Peng, X. Zhao, X. Zhu, and J. Yin (2025) CoreGuard: safeguarding foundational capabilities of llms against model stealing in edge deployment. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
  • [116] R. Li, J. Long, M. Qi, H. Xia, L. Sha, P. Wang, and Z. Sui (2025) Towards harmonized uncertainty estimation for large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 22938–22953. External Links: Document Cited by: §B.3.
  • [117] S. Li, J. Sun, G. Zheng, X. Fan, Y. Shen, Y. Lu, Z. Xi, Y. Yang, W. Tan, T. Ji, T. Gui, Q. Zhang, and X. Huang (2025) Mitigating object hallucinations in mllms via multi-frequency perturbations. In Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 1230–1247. External Links: Document Cited by: §B.3.
  • [118] Y. Li, M. Du, X. Wang, and Y. Wang (2023) Prompt tuning pushes farther, contrastive learning pulls closer: A two-stage approach to mitigate social biases. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, A. Rogers, J. L. Boyd-Graber, and N. Okazaki (Eds.), pp. 14254–14267. External Links: Document Cited by: §B.3.
  • [119] Z. Li, P. Chen, and T. Ho (2025) Retention score: quantifying jailbreak risks for vision language models. In Thirty-Ninth AAAI Conference on Artificial Intelligence, Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence, Fifteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2025, Philadelphia, PA, USA, February 25 - March 4, 2025, T. Walsh, J. Shah, and Z. Kolter (Eds.), pp. 27446–27454. External Links: Document Cited by: §B.3.
  • [120] J. Liang, Z. Wang, S. Hong, S. Ji, and T. Wang (2025) Watermark under fire: A robustness evaluation of LLM watermarking. In Findings of the Association for Computational Linguistics: EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 21050–21074. External Links: Document Cited by: §B.3.
  • [121] Z. Liang, L. Yu, S. Zhang, Q. Ye, and H. Hu (2026) How much do large language model cheat on evaluation? benchmarking overestimation under the one-time-pad-based framework. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 37636–37644. External Links: Document Cited by: §B.3.
  • [122] Z. Lin, Z. Wang, Y. Tong, Y. Wang, Y. Guo, Y. Wang, and J. Shang (2023) ToxicChat: unveiling hidden challenges of toxicity detection in real-world user-AI conversation. In Findings of the Association for Computational Linguistics: EMNLP 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 4694–4702. External Links: Document Cited by: §B.3.
  • [123] F. Liu, Y. Zhang, J. Luo, J. Dai, T. Chen, L. Yuan, Z. Yu, Y. Shi, K. Li, C. Zhou, H. Chen, and M. Yang (2025) Make agent defeat agent: automatic detection of taint-style vulnerabilities in llm-based agents. In 34th USENIX Security Symposium, USENIX Security 2025, Seattle, WA, USA, August 13-15, 2025, L. Bauer and G. Pellegrino (Eds.), pp. 3767–3786. Cited by: §B.3.
  • [124] H. Liu, Y. Xie, Y. Wang, and M. Shieh (2024) Advancing adversarial suffix transfer learning on aligned large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), pp. 7213–7224. External Links: Document Cited by: §B.3.
  • [125] X. Liu, P. Li, G. E. Suh, Y. Vorobeychik, Z. Mao, S. Jha, P. McDaniel, H. Sun, B. Li, and C. Xiao (2025) AutoDAN-turbo: A lifelong agent for strategy self-exploration to jailbreak llms. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
  • [126] X. Liu, S. Liang, M. Han, Y. Luo, A. Liu, X. Cai, Z. He, and D. Tao (2025) ELBA-bench: an efficient learning backdoor attacks benchmark for large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 17928–17947. External Links: Document Cited by: §B.3.
  • [127] Y. Liu, Y. Liu, X. Chen, P. Chen, D. Zan, M. Kan, and T. Ho (2024) The devil is in the neurons: interpreting and mitigating social biases in language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024, Cited by: §B.3.
  • [128] J. Lu, J. Liu, X. Zheng, M. Yang, J. Wang, P. Wang, and Y. Zhang (2026) MHB: medical hallucination benchmark for large language models in complex clinical tasks. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 38971–38978. External Links: Document Cited by: §B.3.
  • [129] X. Lu, F. Brahman, P. West, J. Jung, K. Chandu, A. Ravichander, P. Ammanabrolu, L. Jiang, S. Ramnath, N. Dziri, J. Fisher, B. Lin, S. Hallinan, L. Qin, X. Ren, S. Welleck, and Y. Choi (2023) Inference-time policy adapters (IPA): tailoring extreme-scale LMs without fine-tuning. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, H. Bouamor, J. Pino, and K. Bali (Eds.), Singapore, pp. 6863–6883. External Links: Document Cited by: §B.3.
  • [130] Y. Lu, J. Li, Y. Zhou, Y. Zhang, W. Wang, X. Li, M. Zhang, F. Liu, J. Yu, and M. Zhang (2025) Adaptive detoxification: safeguarding general capabilities of llms through toxicity-aware knowledge editing. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Findings of ACL, Vol. ACL 2025, pp. 19744–19758. External Links: Document Cited by: §B.3.
  • [131] K. Lukošiūtė and A. Swanda (2025) LLM cyber evaluations don’t capture real-world risk. arXiv preprint arXiv:2502.00072. External Links: Document Cited by: §1.
  • [132] W. Luo, S. Dai, X. Liu, S. Banerjee, H. Sun, M. Chen, and C. Xiao (2025) AGrail: A lifelong agent guardrail with effective and adaptive safety detection. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 8104–8139. External Links: Document Cited by: §B.3.
  • [133] T. S. Luong, T. Le, L. N. Van, and T. H. Nguyen (2024) Realistic evaluation of toxicity in large language models. In Findings of the Association for Computational Linguistics, ACL 2024, Bangkok, Thailand and virtual meeting, August 11-16, 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), Findings of ACL, Vol. ACL 2024, pp. 1038–1047. External Links: Document Cited by: §B.3.
  • [134] D. T. Mahato (2026) Safeguard-conditioned uplift: measuring utility–risk frontiers for dual-use biology assistants. Note: arXiv preprint External Links: 2607.13039, Document Cited by: §1, §2.3, §6.
  • [135] R. Maheshwary, V. Yadav, H. Nguyen, K. Mahajan, and S. T. Madhusudhan (2025) M2Lingual: enhancing multilingual, multi-turn instruction alignment in large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, L. Chiruzzo, A. Ritter, and L. Wang (Eds.), pp. 9676–9713. External Links: Document Cited by: §B.3.
  • [136] S. Masud, S. Singh, V. Hangya, A. Fraser, and T. Chakraborty (2024) Hate personified: investigating the role of llms in content moderation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), pp. 15847–15863. External Links: Document Cited by: §B.3.
  • [137] M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li, D. Forsyth, and D. Hendrycks (2024) HarmBench: a standardized evaluation framework for automated red teaming and robust refusal. In Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, Vol. 235, pp. 35181–35224. Cited by: §1.
  • [138] A. McKenzie, U. Pawar, P. Blandfort, W. Bankes, D. Krueger, E. S. Lubana, and D. Krasheninnikov (2025) Detecting high-stakes interactions with activation probes. CoRR abs/2506.10805. External Links: Document, 2506.10805 Cited by: §B.3.
  • [139] R. Miao, Y. Liu, Y. Wang, X. Shen, Y. Tan, Y. Dai, S. Pan, and X. Wang (2026) BlindGuard: safeguarding llm-based multi-agent systems under unknown attacks. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 39215–39234. External Links: Document Cited by: §B.3.
  • [140] W. J. Mo, Q. Liu, X. Wen, D. Jung, H. Askari, W. Zhou, Z. Zhao, and M. Chen (2026) RedCoder: automated multi-turn red teaming for code llms. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 33140–33155. External Links: Document Cited by: §B.3.
  • [141] S. R. Motwani, M. Baranchuk, M. Strohmeier, V. Bolina, P. H. S. Torr, L. Hammond, and C. S. de Witt (2024) Secret collusion among AI agents: multi-agent deception via steganography. CoRR abs/2402.07510. External Links: Document, 2402.07510 Cited by: §B.3.
  • [142] R. Movva, P. W. Koh, and E. Pierson (2024) Annotation alignment: comparing LLM and human annotations of conversational safety. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), pp. 9048–9062. External Links: Document Cited by: §B.3.
  • [143] M. Nagireddy, L. Chiazor, M. Singh, and I. Baldini (2024) SocialStigmaQA: A benchmark to uncover stigma amplification in generative language models. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2024, February 20-27, 2024, Vancouver, Canada, M. J. Wooldridge, J. G. Dy, and S. Natarajan (Eds.), pp. 21454–21462. External Links: Document Cited by: §B.3.
  • [144] G. Nalbandyan, R. Shahbazyan, and E. Bakhturina (2025) SCORE: systematic consistency and robustness evaluation for large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 3: Industry Track, Albuquerque, New Mexico, USA, April 30, 2025, W. Chen, Y. Yang, M. Kachuee, and X. Fu (Eds.), pp. 470–484. External Links: Document Cited by: §B.3.
  • [145] M. Nasr, N. Carlini, C. Sitawarin, S. V. Schulhoff, J. Hayes, M. Ilie, J. Pluto, S. Song, H. Chaudhari, I. Shumailov, A. G. Thakurta, K. Y. Xiao, A. Terzis, and F. Tramèr (2026) The attacker moves second: stronger adaptive attacks bypass defenses against LLM jailbreaks and prompt injections. In 35th USENIX Security Symposium (USENIX Security 26), Baltimore, MD. Cited by: §C.6, §1, §5.4.
  • [146] M. Nasr, J. Rando, N. Carlini, J. Hayase, M. Jagielski, A. F. Cooper, D. Ippolito, C. A. Choquette-Choo, F. Tramèr, and K. Lee (2025) Scalable extraction of training data from aligned, production language models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
  • [147] H. Nghiem, J. Prindle, J. Zhao, and H. Daumé III (2024) “You gotta be a doctor, lin” : an investigation of name-based bias of large language models in employment recommendations. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 7268–7287. External Links: Document Cited by: §B.3.
  • [148] G. Niess and R. Kern (2025) Ensemble watermarks for large language models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, pp. 2903–2916. External Links: Document, ISBN 979-8-89176-251-0 Cited by: §B.3.
  • [149] S. Oh, K. Lee, S. Park, D. Kim, and H. Kim (2024) Poisoned chatgpt finds work for idle hands: exploring developers’ coding practices with insecure suggestions from poisoned AI models. In IEEE Symposium on Security and Privacy, SP 2024, San Francisco, CA, USA, May 19-23, 2024, pp. 1141–1159. External Links: Document Cited by: §B.3.
  • [150] OpenAI (2025) Disrupting malicious uses of AI: an update (october 2025). Note: OpenAI Threat Intelligence ReportPublished October 7, 2025; accessed August 20, 2026 External Links: Link Cited by: §1.
  • [151] OpenAI (2026) GPT-5.6 system card. Note: OpenAI Deployment Safety HubPublished July 9, 2026; accessed August 14, 2026 External Links: Link Cited by: §1.
  • [152] OpenAI (2026) GPT-5.6. Note: Model suite announcement (Sol, Terra, Luna), July 9, 2026. Accessed 2026-08-20 External Links: Link Cited by: §B.3.
  • [153] K. O’Brien, S. Casper, Q. G. Anthony, T. Korbak, R. Kirk, X. Davies, I. Mishra, G. Irving, Y. Gal, and S. Biderman (2025) Deep ignorance: filtering pretraining data builds tamper-resistant safeguards into open-weight LLMs. In BioSafe GenAI Workshop 2025, Cited by: §5.4.
  • [154] M. J. Page, J. E. McKenzie, P. M. Bossuyt, I. Boutron, T. C. Hoffmann, C. D. Mulrow, L. Shamseer, J. M. Tetzlaff, E. A. Akl, S. E. Brennan, R. Chou, J. Glanville, J. M. Grimshaw, A. Hróbjartsson, M. M. Lalu, T. Li, E. W. Loder, E. Mayo-Wilson, S. McDonald, L. A. McGuinness, L. A. Stewart, J. Thomas, A. C. Tricco, V. A. Welch, P. Whiting, and D. Moher (2021) The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ 372, pp. n71. External Links: Document Cited by: Table 4.
  • [155] L. Pan, A. Liu, S. Huang, Y. Lu, X. Hu, L. Wen, I. King, and P. S. Yu (2025) Can LLM watermarks robustly prevent unauthorized knowledge distillation?. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 13228–13251. External Links: Document Cited by: §B.3.
  • [156] X. Pang, X. Hao, S. Guo, Q. Luo, and Z. Wang (2025) ICLScan: detecting backdoors in black-box large language models via targeted in-context illumination. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
  • [157] Y. Pang, W. Meng, X. Liao, and T. Wang (2026) Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag Paradigm. In Network and Distributed System Security Symposium (NDSS), Cited by: §B.3.
  • [158] L. H. Park, J. Cho, G. Kim, Y. Yeo, and T. Kwon (2026) Chimera: compositional jailbreak attacks on llms via judgment-driven search over heterogeneous strategies. In Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 33330–33355. External Links: Document Cited by: §B.3.
  • [159] S. Park and K. Kim (2025) Measuring and mitigating media outlet name bias in large language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 29778–29797. External Links: Document Cited by: §B.3.
  • [160] H. L. Patel, A. Agarwal, A. Das, B. Kumar, S. Panda, P. Pattnayak, T. H. Rafi, T. Kumar, and D. Chae (2025) SweEval: do llms really swear? A safety benchmark for testing limits for enterprise use. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 3: Industry Track, Albuquerque, New Mexico, USA, April 30, 2025, W. Chen, Y. Yang, M. Kachuee, and X. Fu (Eds.), pp. 558–582. External Links: Document Cited by: §B.3.
  • [161] C. Pathade (2025) Red teaming the mind of the machine: a systematic evaluation of prompt injection and jailbreak vulnerabilities in LLMs. arXiv preprint arXiv:2505.04806. External Links: 2505.04806, Document Cited by: §6.
  • [162] K. Pelrine, A. Imouza, C. Thibault, M. Reksoprodjo, C. Gupta, J. Christoph, J. Godbout, and R. Rabbany (2023) Towards reliable misinformation mitigation: generalization, uncertainty, and GPT-4. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, EMNLP 2023, Singapore, December 6-10, 2023, H. Bouamor, J. Pino, and K. Bali (Eds.), pp. 6399–6429. External Links: Document Cited by: §B.3.
  • [163] D. Peng, Q. Ke, and J. Liu (2024) UPAM: unified prompt attack in text-to-image generation models against both textual filters and visual checkers. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024, R. Salakhutdinov, Z. Kolter, K. A. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (Eds.), Proceedings of Machine Learning Research, Vol. 235, pp. 40200–40214. Cited by: §B.3.
  • [164] A. Peppin, A. Reuel, S. Casper, E. Jones, A. Strait, U. Anwar, A. Agrawal, S. Kapoor, O. Koyejo, M. Pellat, R. Bommasani, N. Frosst, and S. Hooker (2025) The reality of ai and biorisk. In Proceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency, pp. 763–771. External Links: Document Cited by: §1.
  • [165] E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving (2022) Red teaming language models with language models. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, Y. Goldberg, Z. Kozareva, and Y. Zhang (Eds.), Abu Dhabi, United Arab Emirates, pp. 3419–3448. External Links: Document Cited by: §B.3.
  • [166] N. Prakash, Y. W. Jie, A. Abdullah, R. Satapathy, E. Cambria, and R. K. W. Lee (2025) Beyond “I’m sorry, I can’t”: dissecting large language model refusal. arXiv preprint arXiv:2509.09708. External Links: 2509.09708, Document Cited by: §6.
  • [167] X. Qi, B. Wei, N. Carlini, Y. Huang, T. Xie, L. He, M. Jagielski, M. Nasr, P. Mittal, and P. Henderson (2025) On evaluating the durability of safeguards for open-weight LLMs. In The Thirteenth International Conference on Learning Representations, Cited by: §1, §6, §7.1.
  • [168] Z. Rao, W. Zhu, C. A. Lu, Z. Chen, W. Niu, L. Guan, B. Li, and Z. Xiang (2026) FragFuse: bypassing access control of large language model agents via memory-based query fragmentation and fusion. In 35th USENIX Security Symposium (USENIX Security 26), Baltimore, MD. Cited by: §1.
  • [169] M. L. Rethlefsen, S. Kirtley, S. Waffenschmidt, A. P. Ayala, D. Moher, M. J. Page, J. B. Koffel, and PRISMA-S Group (2021) PRISMA-S: an extension to the PRISMA statement for reporting literature searches in systematic reviews. Systematic Reviews 10 (1), pp. 39. External Links: Document, Link Cited by: §B.1.
  • [170] L. Richter, X. He, P. Minervini, and M. J. Kusner (2025) An auditing test to detect behavioral shift in language models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
  • [171] A. Robey, E. Wong, H. Hassani, and G. J. Pappas (2025) SmoothLLM: defending large language models against jailbreaking attacks. Transactions on Machine Learning Research. Cited by: §5.4.
  • [172] P. Röttger, H. Kirk, B. Vidgen, G. Attanasio, F. Bianchi, and D. Hovy (2024) XSTest: a test suite for identifying exaggerated safety behaviours in large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), Mexico City, Mexico, pp. 5377–5400. External Links: Document Cited by: §1.
  • [173] Y. Ruan, H. Dong, A. Wang, S. Pitis, Y. Zhou, J. Ba, Y. Dubois, C. J. Maddison, and T. Hashimoto (2023) Identifying the risks of LM agents with an lm-emulated sandbox. CoRR abs/2309.15817. External Links: Document, 2309.15817 Cited by: §B.3.
  • [174] H. Saffari, M. Shafiei, H. Zhang, L. T. Harris, and N. S. Moosavi (2025) Beyond hate speech: NLP’s challenges and opportunities in uncovering dehumanizing language. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 26965–26980. External Links: Document, ISBN 979-8-89176-332-6 Cited by: §B.3.
  • [175] S. Sagar, A. Taparia, and R. Senanayake (2024) Failures are fated, but can be faded: characterizing and mitigating unwanted behaviors in large-scale vision and language models. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024, R. Salakhutdinov, Z. Kolter, K. A. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (Eds.), Proceedings of Machine Learning Research, Vol. 235, pp. 42999–43023. Cited by: §B.3.
  • [176] J. H. Saltzer and M. D. Schroeder (1975) The protection of information in computer systems. Proceedings of the IEEE 63 (9), pp. 1278–1308. External Links: Document, Link Cited by: §3.3.3.
  • [177] P. Sarkar, S. Ebrahimi, A. Etemad, A. Beirami, S. Ö. Arik, and T. Pfister (2025) Mitigating object hallucination in mllms via data-augmented phrase-level alignment. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
  • [178] G. M. Shahariar, Z. A. Nazi, Md. O. H. Bhuiyan, and Z. Shi (2026) PII-visbench: evaluating personally identifiable information safety in vision language models along a continuum of visibility. In Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 10294–10316. External Links: Document Cited by: §B.3.
  • [179] R. M. Shahroz, Z. Tan, S. Yun, C. Fleming, and T. Chen (2025) Agents under siege: breaking pragmatic multi-agent LLM systems with optimized prompt attacks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), pp. 9661–9674. External Links: Document Cited by: §B.3.
  • [180] G. Shen, D. Zhao, L. Feng, X. He, J. Wang, S. Shen, H. Tong, Y. Dong, J. Li, X. Zheng, and Y. Zeng (2025) PandaGuard: systematic evaluation of LLM safety against jailbreaking attacks. arXiv preprint arXiv:2505.13862. External Links: 2505.13862, Document Cited by: §6.
  • [181] H. Shen, B. Huang, and X. Wan (2025) Enhancing LLM watermark resilience against both scrubbing and spoofing attacks. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
  • [182] T. Shi, J. He, Z. Wang, L. Wu, H. Li, W. Guo, and D. Song (2025) Progent: securing AI agents with privilege control. arXiv preprint arXiv:2504.11703. External Links: Document Cited by: §1, §7.2.
  • [183] W. Shi, A. Ajith, M. Xia, Y. Huang, D. Liu, T. Blevins, D. Chen, and L. Zettlemoyer (2024) Detecting pretraining data from large language models. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024, Cited by: §B.3.
  • [184] W. M. Si, M. Li, M. Backes, and Y. Zhang (2026) Pruning unsafe tickets: A resource-efficient framework for safer and more robust llms. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 26285–26302. External Links: Document Cited by: §B.3.
  • [185] I. Solaiman, M. Brundage, J. Clark, A. Askell, A. Herbert-Voss, J. Wu, A. Radford, G. Krueger, J. W. Kim, S. Kreps, M. McCain, A. Newhouse, J. Blazakis, K. McGuffie, and J. Wang (2019) Release strategies and the social impacts of language models. External Links: 1908.09203, Document Cited by: §1.
  • [186] M. Son, J. Jang, and M. Kim (2025) Lightweight query checkpoint: classifying faulty user queries to mitigate hallucinations in large language model question answering. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Findings of ACL, Vol. ACL 2025, pp. 14664–14677. External Links: Document Cited by: §B.3.
  • [187] Y. Son, M. Kim, S. Kim, S. Han, J. Kim, D. Jang, Y. Yu, and C. Y. Park (2025) Subtle risks, critical failures: A framework for diagnosing physical safety of llms for embodied decision making. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 25692–25733. External Links: Document Cited by: §B.3.
  • [188] M. Spliethöver, T. Knebler, F. Fumagalli, M. Muschalik, B. Hammer, E. Hüllermeier, and H. Wachsmuth (2025) Adaptive prompting: ad-hoc prompt composition for social bias detection. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, L. Chiruzzo, A. Ritter, and L. Wang (Eds.), pp. 2421–2449. External Links: Document Cited by: §B.3.
  • [189] R. Tamirisa, B. Bharathi, L. Phan, A. Zhou, A. Gatti, T. Suresh, M. Lin, J. Wang, R. Wang, R. Arel, A. Zou, D. Song, B. Li, D. Hendrycks, and M. Mazeika (2025) Tamper-resistant safeguards for open-weight LLMs. In The Thirteenth International Conference on Learning Representations, Cited by: §5.4.
  • [190] F. Tramèr, N. Carlini, W. Brendel, and A. Madry (2020) On adaptive attacks to adversarial example defenses. In Advances in Neural Information Processing Systems 33 (NeurIPS 2020), Cited by: §3.1, §3.3.1.
  • [191] I. Uddin and A. Bauer (2026) Conformal LLM routing with distribution-free safety guarantees. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), ACL 2026, San Diego, California, United States, July 2-7, 2026, T. Y. S. S. Santosh, J. D. Rodriguez, and O. de Gibert (Eds.), pp. 791–799. External Links: Document Cited by: §B.3.
  • [192] UK AI Safety Institute Safeguards Analysis Team (2025) Principles for evaluating misuse safeguards of frontier AI systems. Technical report UK AI Safety Institute. Note: Published February 4, 2025; organization renamed the AI Security Institute on February 14, 2025; accessed August 14, 2026 External Links: Link Cited by: §1, §6, §7.1.
  • [193] R. Uppaal, A. Dey, Y. He, Y. Zhong, and J. Hu (2024) Model editing as a robust and denoised variant of DPO: a case study on toxicity. CoRR abs/2405.13967. External Links: Document, 2405.13967 Cited by: §B.3.
  • [194] M. Vaccaro, J. Song, A. Almaatouq, and M. A. Bakker (2026) Evaluating human–ai safety: a framework for measuring harmful capability uplift. Note: arXiv preprint External Links: 2603.26676, Document Cited by: §1, §2.3, §6.
  • [195] C. Wang, X. Chen, N. Zhang, B. Tian, H. Xu, S. Deng, and H. Chen (2025) MLLM can see? dynamic correction decoding for hallucination mitigation. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
  • [196] H. Wang, Z. Huang, Z. Lin, and T. Liu (2024) NoiseGPT: label noise detection and rectification through probability curvature. In Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang (Eds.), Cited by: §B.3.
  • [197] J. G. Wang, J. Wang, M. Li, and S. Neel (2026) CheckMIABench: firm foundations for membership inference attacks on language models. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 364–370. External Links: Document Cited by: §B.3.
  • [198] J. Wang, Z. Xu, D. Jin, X. Yang, and T. Li (2026) Accommodate knowledge conflicts in retrieval-augmented llms: towards robust response generation in the wild. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 33530–33538. External Links: Document Cited by: §B.3.
  • [199] P. Wang, B. Dong, Y. Cai, Z. Zhang, J. Liu, H. Xue, Y. Wu, Y. Zhang, and Z. Zhang (2025) Game of arrows: on the (in-)security of weight obfuscation for on-device tee-shielded LLM partition algorithms. In 34th USENIX Security Symposium, USENIX Security 2025, Seattle, WA, USA, August 13-15, 2025, L. Bauer and G. Pellegrino (Eds.), pp. 279–298. Cited by: §B.3.
  • [200] X. Wang, A. Balashankar, and V. Chandrasekaran (2026) Systematic scaling analysis of jailbreak attacks in large language models. arXiv preprint arXiv:2603.11149. External Links: 2603.11149, Document Cited by: §5.4.
  • [201] X. Wang, Z. Li, B. Wang, Y. Hu, and D. Zou (2025) Model unlearning via sparse autoencoder subspace guided projections. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 26530–26546. External Links: Document, ISBN 979-8-89176-332-6 Cited by: §B.3.
  • [202] X. Wang, S. Zhu, and X. Cheng (2025) Speculative safety-aware decoding. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, pp. 12827–12841. External Links: Document, ISBN 979-8-89176-332-6 Cited by: §B.3.
  • [203] X. Wang, Z. Ji, W. Wang, Z. Li, D. Wu, and S. Wang (2026) SoK: evaluating jailbreak guardrails for large language models. In 2026 IEEE Symposium on Security and Privacy (S&P), pp. 39–58. External Links: Document, 2506.10597 Cited by: §1, §6.
  • [204] X. Wang, D. Wu, Z. Ji, Z. Li, P. Ma, S. Wang, Y. Li, Y. Liu, N. Liu, and J. Rahmel (2025) SelfDefend: llms can defend themselves against jailbreaking in a practical manner. In 34th USENIX Security Symposium, USENIX Security 2025, Seattle, WA, USA, August 13-15, 2025, L. Bauer and G. Pellegrino (Eds.), pp. 2441–2460. Cited by: §B.3.
  • [205] Y. Wang, T. Huang, L. Shen, H. Yao, H. Luo, R. Liu, N. Tan, J. Huang, and D. Tao (2025) Panacea: mitigating harmful fine-tuning for large language models via post-fine-tuning perturbation. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
  • [206] Y. Wang, M. Zhang, J. Sun, C. Wang, M. Yang, H. Xue, J. Tao, R. Duan, and J. Liu (2025) Mirage in the eyes: hallucination attack on multi-modal large language models with only attention sink. In 34th USENIX Security Symposium, USENIX Security 2025, Seattle, WA, USA, August 13-15, 2025, L. Bauer and G. Pellegrino (Eds.), pp. 3707–3726. Cited by: §B.3.
  • [207] Y. Wang, R. Wu, Z. He, X. Chen, and J. J. McAuley (2024) Large scale knowledge washing. CoRR abs/2405.16720. External Links: Document, 2405.16720 Cited by: §B.3.
  • [208] Z. Wang, Z. Wu, X. Guan, M. Thaler, A. S. Koshiyama, S. Lu, S. Beepath, E. E. Jr., and M. Pérez-Ortiz (2024) JobFair: A framework for benchmarking gender hiring bias in large language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Findings of ACL, Vol. EMNLP 2024, pp. 3227–3246. External Links: Document Cited by: §B.3.
  • [209] Z. Wang, D. Anshumaan, A. Hooda, Y. Chen, and S. Jha (2025) Functional homotopy: smoothing discrete optimization via continuous parameters for LLM jailbreak attacks. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
  • [210] A. Wei, N. Haghtalab, and J. Steinhardt (2023) Jailbroken: how does LLM safety training fail?. In Advances in Neural Information Processing Systems 36, External Links: Document Cited by: §1.
  • [211] C. Wu, Z. R. Tam, C. Lin, Y. V. Chen, S. Sun, and H. Lee (2025) Mitigating forgetting in LLM fine-tuning via low-perplexity token learning. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, D. Belgrave, C. Zhang, L. N. Montoya, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, N. Chen, I. V. M. Ruíz, and A. Loaiza-Bonilla (Eds.), Cited by: §B.3.
  • [212] F. Wu, E. Cecchetti, and C. Xiao (2024) System-level defense against indirect prompt injection attacks: an information flow control perspective. arXiv preprint arXiv:2409.19091. External Links: Document Cited by: §5.6.
  • [213] L. Wu, M. Wang, Z. Xu, T. Cao, N. Oo, B. Hooi, and S. Deng (2025) Automating steering for safe multimodal large language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 792–814. External Links: Document Cited by: §B.3.
  • [214] M. Wu, J. Ji, O. Huang, J. Li, Y. Wu, X. Sun, and R. Ji (2024) Evaluating and analyzing relationship hallucinations in large vision-language models. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024, R. Salakhutdinov, Z. Kolter, K. A. Heller, A. Weller, N. Oliver, J. Scarlett, and F. Berkenkamp (Eds.), Proceedings of Machine Learning Research, Vol. 235, pp. 53553–53570. Cited by: §B.3.
  • [215] P. Wu, L. Zhu, W. Zhang, and N. Yu (2026) Safeguards based on copyable context cannot provide reliable safety for LLMs. Note: arXiv preprint External Links: 2607.27951, Document Cited by: §3.2.2.
  • [216] Y. Wu, R. Wen, C. Cui, M. Backes, and Y. Zhang (2026) InferPilot: autonomous inference attacks against ML services with llm-based agents. In Findings of the Association for Computational Linguistics, ACL 2026, San Diego, California, United States, July 2-7, 2026, M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), pp. 11781–11801. External Links: Document Cited by: §B.3.
  • [217] Y. Wu, F. Roesner, T. Kohno, N. Zhang, and U. Iqbal (2025) IsolateGPT: an execution isolation architecture for LLM-based agentic systems. In Network and Distributed System Security Symposium, External Links: Document Cited by: §1.
  • [218] W. Xia and Z. Deng (2026) SDA: steering-driven distribution alignment for open llms without fine-tuning. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 34025–34033. External Links: Document Cited by: §B.3.
  • [219] T. Xiang, L. Li, W. Li, M. Bai, L. Wei, B. Wang, and N. Garcia (2023) CARE-MI: chinese benchmark for misinformation evaluation in maternity and infant care. CoRR abs/2307.01458. External Links: Document, 2307.01458 Cited by: §B.3.
  • [220] T. Xing, J. Li, Y. Du, and X. Hu (2026) Are LLMs reliable rankers? rank manipulation via two-stage token optimization. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 9120–9132. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §B.3.
  • [221] F. Xu, H. Hu, C. He, S. Hang, H. Hu, X. Liu, Y. Zhao, Z. Zhou, B. B. Zhu, S. Sun, D. Gu, and S. Wang (2026) SoK: robustness in large language models against jailbreak attacks. In 2026 IEEE Symposium on Security and Privacy (S&P), pp. 118–137. External Links: Document, 2605.05058 Cited by: §1, §6.
  • [222] J. Xu, M. D. Ma, F. Wang, C. Xiao, and M. Chen (2024) Instructions as backdoors: backdoor vulnerabilities of instruction tuning for large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024, K. Duh, H. Gómez-Adorno, and S. Bethard (Eds.), pp. 3111–3126. External Links: Document Cited by: §B.3.
  • [223] S. Xu, L. Pang, Y. Zhu, H. Shen, and X. Cheng (2025) Cross-modal safety mechanism transfer in large vision-language models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
  • [224] Z. Xu, F. Liu, and H. Liu (2024) Bag of tricks: benchmarking of jailbreak attacks on llms. In Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, A. Globersons, L. Mackey, D. Belgrave, A. Fan, U. Paquet, J. M. Tomczak, and C. Zhang (Eds.), Cited by: §B.3.
  • [225] S. Yi, Y. Liu, Z. Sun, T. Cong, X. He, J. Song, K. Xu, and Q. Li (2024) Jailbreak attacks and defenses against large language models: a survey. arXiv preprint arXiv:2407.04295. Note: preprint; no peer-reviewed venue recorded on arXiv as of 2026-08-19 External Links: 2407.04295 Cited by: §6.
  • [226] X. Yu, H. Cheng, X. Liu, D. Roth, and J. Gao (2024) ReEval: automatic hallucination evaluation for retrieval-augmented large language models via transferable adversarial attacks. In Findings of the Association for Computational Linguistics: NAACL 2024, Mexico City, Mexico, June 16-21, 2024, K. Duh, H. Gómez-Adorno, and S. Bethard (Eds.), Findings of ACL, Vol. NAACL 2024, pp. 1333–1351. External Links: Document Cited by: §B.3.
  • [227] Z. Yue, H. Zeng, Y. Lu, L. Shang, Y. Zhang, and D. Wang (2024) Evidence-driven retrieval augmented response generation for online misinformation. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), K. Duh, H. Gomez, and S. Bethard (Eds.), Mexico City, Mexico, pp. 5628–5643. External Links: Document Cited by: §B.3.
  • [228] Q. Zhan, R. Fang, H. S. Panchal, and D. Kang (2025) Adaptive attacks break defenses against indirect prompt injection attacks on LLM agents. In Findings of the Association for Computational Linguistics: NAACL 2025, Albuquerque, New Mexico, pp. 7116–7132. External Links: Document Cited by: §1.
  • [229] X. Zhan, J. C. Carrillo, W. Seymour, and J. Such (2025) Malicious llm-based conversational AI makes users reveal personal information. In 34th USENIX Security Symposium, USENIX Security 2025, Seattle, WA, USA, August 13-15, 2025, L. Bauer and G. Pellegrino (Eds.), pp. 61–80. Cited by: §B.3, §1.
  • [230] B. Zhang and G. Ren (2025) Challenges and remedies of domain-specific classifiers as LLM guardrails: self-harm as a case study. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 3: Industry Track, Albuquerque, New Mexico, USA, April 30, 2025, W. Chen, Y. Yang, M. Kachuee, and X. Fu (Eds.), pp. 173–182. External Links: Document Cited by: §B.3.
  • [231] B. Zhang, H. Liu, Q. Tian, S. Chen, Z. Wang, and Q. Qi (2026) Towards trustworthy smart contract synthesis: a multi-agent framework with lean-based verification. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 39548–39582. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §B.3.
  • [232] C. Zhang, J. X. Morris, and V. Shmatikov (2024) Extracting prompts by inverting LLM outputs. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Miami, Florida, USA, pp. 14753–14777. External Links: Document Cited by: §B.3.
  • [233] M. Zhang, K. K. Goh, P. Zhang, J. Sun, L. X. Rose, and H. Zhang (2025) LLMScan: causal scan for LLM misbehavior detection. In Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025, A. Singh, M. Fazel, D. Hsu, S. Lacoste-Julien, F. Berkenkamp, T. Maharaj, K. Wagstaff, and J. Zhu (Eds.), Proceedings of Machine Learning Research, Vol. 267. Cited by: §B.3.
  • [234] S. Zhang, H. Li, and R. Ji (2024) Code membership inference for detecting unauthorized data use in code pre-trained language models. In Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), Findings of ACL, Vol. EMNLP 2024, pp. 10593–10603. External Links: Document Cited by: §B.3.
  • [235] S. Zhang, Y. Zhai, K. Guo, H. Hu, S. Guo, Z. Fang, L. Zhao, C. Shen, C. Wang, and Q. Wang (2025) JBShield: defending large language models from jailbreak attacks through activated concept analysis and manipulation. In 34th USENIX Security Symposium, USENIX Security 2025, Seattle, WA, USA, August 13-15, 2025, L. Bauer and G. Pellegrino (Eds.), pp. 8215–8234. Cited by: §B.3.
  • [236] T. Zhang, Z. Xi, T. Wang, P. Mitra, and J. Chen (2024) PromptFix: few-shot backdoor removal via adversarial prompt tuning. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), NAACL 2024, Mexico City, Mexico, June 16-21, 2024, K. Duh, H. Gómez-Adorno, and S. Bethard (Eds.), pp. 3212–3225. External Links: Document Cited by: §B.3.
  • [237] X. Zhang, X. Wang, Y. Lu, J. Wang, Z. Ye, M. Bao, P. Yan, and X. Su (2026) TrendFact: a benchmark towards hotspot perception in automatic fact-checking. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 26494–26513. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §B.3.
  • [238] Y. Zhang, T. Liu, Z. Zhao, G. Meng, and K. Chen (2026) Bleeding pathways: vanishing discriminability in LLM hidden states fuels jailbreak attacks. In 33rd Annual Network and Distributed System Security Symposium, NDSS 2026, San Diego, California, USA, February 23-27, 2026, Cited by: §B.3.
  • [239] Y. Zhang, R. Xie, J. Chen, X. Sun, Z. Kang, and Y. Wang (2025) QAVA: query-agnostic visual attack to large vision-language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, L. Chiruzzo, A. Ritter, and L. Wang (Eds.), pp. 10205–10218. External Links: Document Cited by: §B.3.
  • [240] Z. Zhang, L. Lei, L. Wu, R. Sun, Y. Huang, C. Long, X. Liu, X. Lei, J. Tang, and M. Huang (2024) SafetyBench: evaluating the safety of large language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, L. Ku, A. Martins, and V. Srikumar (Eds.), pp. 15537–15553. External Links: Document Cited by: §B.3.
  • [241] Z. Zhang, H. Zhang, W. Li, Q. Zhang, J. Dong, Y. Tong, and Z. Zheng (2026) FedSEA-llama: A secure, efficient and adaptive federated splitting framework for large language models. In Fortieth AAAI Conference on Artificial Intelligence, Thirty-Eighth Conference on Innovative Applications of Artificial Intelligence, Sixteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2026, Singapore, January 20-27, 2026, S. Koenig, C. Jenkins, and M. E. Taylor (Eds.), pp. 28680–28688. External Links: Document Cited by: §B.3.
  • [242] A. Zhao, Q. Xu, M. Lin, S. Wang, Y. Liu, Z. Zheng, and G. Huang (2025) DiveR-ct: diversity-enhanced red teaming large language model assistants with relaxing constraints. In Thirty-Ninth AAAI Conference on Artificial Intelligence, Thirty-Seventh Conference on Innovative Applications of Artificial Intelligence, Fifteenth Symposium on Educational Advances in Artificial Intelligence, AAAI 2025, Philadelphia, PA, USA, February 25 - March 4, 2025, T. Walsh, J. Shah, and Z. Kolter (Eds.), pp. 26021–26030. External Links: Document Cited by: §B.3.
  • [243] C. Zhao, X. Wang, P. Zhao, Y. Huang, J. Lu, Z. Liu, Q. Lin, S. Rajmohan, and D. Zhang (2026) Gradient-guided multi-judge prompt optimization. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), M. Liakata, V. P. Moreira, J. Zhang, and D. Jurgens (Eds.), San Diego, California, United States, pp. 23744–23773. External Links: Document, ISBN 979-8-89176-390-6 Cited by: §B.3.
  • [244] S. Zhao, M. Jia, A. T. Luu, F. Pan, and J. Wen (2024) Universal vulnerabilities in large language models: backdoor attacks for in-context learning. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024, Y. Al-Onaizan, M. Bansal, and Y. Chen (Eds.), pp. 11507–11522. External Links: Document Cited by: §B.3.
  • [245] S. Zhao, R. Brekelmans, A. Makhzani, and R. B. Grosse (2024) Probabilistic inference in language models via twisted sequential monte carlo. CoRR abs/2404.17546. External Links: Document, 2404.17546 Cited by: §B.3.
  • [246] W. Zhao, D. Ben-Levi, W. Hao, J. Yang, and C. Mao (2025) Diversity helps jailbreak large language models. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2025 - Volume 1: Long Papers, Albuquerque, New Mexico, USA, April 29 - May 4, 2025, L. Chiruzzo, A. Ritter, and L. Wang (Eds.), pp. 4647–4680. External Links: Document Cited by: §B.3.
  • [247] W. Zhao, J. Guo, Y. Hu, Y. Deng, A. Zhang, X. Sui, X. Han, Y. Zhao, B. Qin, T. Chua, and T. Liu (2025) AdaSteer: your aligned LLM is inherently an adaptive jailbreak defender. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, EMNLP 2025, Suzhou, China, November 4-9, 2025, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), pp. 24559–24577. External Links: Document Cited by: §B.3.
  • [248] Y. Zhao, W. Zheng, T. Cai, X. L. Do, K. Kawaguchi, A. Goyal, and M. Shieh (2024) Accelerating greedy coordinate gradient and general prompt optimization via probe sampling. CoRR abs/2403.01251. External Links: Document, 2403.01251 Cited by: §B.3.
  • [249] K. Zheng, J. Chen, Y. Yan, X. Zou, H. Zhou, and X. Hu (2025) Reefknot: A comprehensive benchmark for relation hallucination evaluation, analysis and mitigation in multimodal large language models. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Findings of ACL, Vol. ACL 2025, pp. 6193–6212. External Links: Document Cited by: §B.3.
  • [250] K. Zhou, C. Liu, X. Zhao, A. Compalas, D. Song, and X. E. Wang (2024) Multimodal situational safety. CoRR abs/2410.06172. External Links: Document, 2410.06172 Cited by: §B.3.
  • [251] X. Zhou, M. Zhang, Z. Lee, W. Ye, and S. Zhang (2025) HaDeMiF: hallucination detection and mitigation in large language models. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025, Cited by: §B.3.
  • [252] Z. Zhou, J. Liu, Z. Dong, J. Liu, C. Yang, W. Ouyang, and Y. Qiao (2024) Emulated Disalignment: Safety Alignment for Large Language Models May Backfire!. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), Volume 1: Long Papers, Bangkok, Thailand, pp. 15810–15830. External Links: Document Cited by: §B.3, §5.2.
  • [253] Z. Zhou, Q. Wang, M. Jin, J. Yao, J. Ye, W. Liu, W. Wang, X. Huang, and K. Huang (2024) MathAttack: attacking large language models towards math solving ability. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2024, February 20-27, 2024, Vancouver, Canada, M. J. Wooldridge, J. G. Dy, and S. Natarajan (Eds.), pp. 19750–19758. External Links: Document Cited by: §B.3.
  • [254] Y. Zhuang, K. Guo, J. Wang, Y. Jing, X. Xu, W. Yi, M. Yang, B. Zhao, and H. Hu (2025) I know what you meme! understanding and detecting harmful memes with multimodal large language models. In 32nd Annual Network and Distributed System Security Symposium, NDSS 2025, San Diego, California, USA, February 24-28, 2025, Cited by: §B.3.
  • [255] A. Zou, L. Phan, J. Wang, D. Duenas, M. Lin, M. Andriushchenko, R. Wang, Z. Kolter, M. Fredrikson, and D. Hendrycks (2024) Improving alignment and robustness with circuit breakers. In Advances in Neural Information Processing Systems, Vol. 37. External Links: Document Cited by: §C.6, §1, §5.4.

Appendix A Proofs and Attaining Constructions

This appendix proves the results in Section 3 and gives the constructions that attain them, which makes each row of Table 1 sharp. Probability laws are defined on the declared trajectory space; all spaces are standard Borel and all declared events are measurable. Finite or countable structure is assumed only where a result states it. Write PT for the pushforward of P under the declared value-relevant projection and [x]+=max{x,0}.

Two elementary facts are used throughout. For any measurable f with range in [0,1],

|𝔼Pf𝔼Qf|TV(P,Q), (35)

and projecting a law cannot increase total variation.

A.1 Witness and Reachable-Set Bounds

Proposition 1 holds because Pm,σg𝒞g, so the supremum defining WZ(g) is at least the value at that law. Proposition 7 holds because a supremum over a superset of 𝒞g dominates the supremum over 𝒞g, with equality when the two classes coincide.

The corresponding negative entry in Table 1 is equally immediate. A finite set of tested modifications {σ1,,σn}Σ generates laws lying inside 𝒞g. For any candidate ceiling c<1, consider any Σ that also contains a strategy whose law places all mass on a trace with vZ=1 and which agrees with σ1 on every coordinate the enumeration observed. The enumeration is unchanged and WZ(g)=1. Hence no supremum over an inner sample bounds WZ(g) above at any sample size, unless the sample is shown to exhaust Σ, in which case it is no longer a sample and the equality case of Proposition 7 applies.

A.2 Frontier Bounds

Proof of Proposition 4.

The inclusion 𝒦Dreal𝒦Dout implies that every law feasible for the realizable frontier is feasible for the outer frontier. Minimizing the same objective over the superset gives the frontier order. The restriction statement follows because minimizing over a subset cannot lower the optimum. This argument applies to either law set without identifying the two sets. ∎

Under the uniform dual-use relation of Equation 18, every feasible law μ obeys

𝔼μvZρ𝔼μbρq,

and minimization proves the stated floor for both frontiers. The uniform premise cannot be replaced on an infinite space by pointwise positivity. For 𝒯Z={tn:n1}, b(tn)=1, and vZ(tn)=1/n, every trace has positive adverse value but the frontier at q=1 has infimum zero and no optimizer.

The floor ρq is attained. On a two-point space with b(t1)=1,vZ(t1)=ρ and b(t0)=vZ(t0)=0, the feasible law placing mass q on t1 has legitimate value exactly q and adverse value exactly ρq.

When a bound on δT(g) accompanies the dual-use premise, the pointwise relation gives more than the composed frontier route. Fix g𝒢D(q), write μ=(Pbg)T, and for any Q𝒞g let μQT be the common part of the two laws, the largest positive measure dominated by both; its total mass is 1TV(μ,QT). Because vZ0, because μQTμ inherits the relation vZρb, and because b1 with 𝔼μbq,

𝔼QvZ vZd(μQT)ρbd(μQT)
ρ[qTV(μ,QT)].

Selecting Q with TV(μ,QT) arbitrarily near δT(g) and using WZ(g)0 prove the endpoint of Table 1, WZ(g)ρ[qδT(g)]+. It is attained: moving mass δT(g) from t1 to t0 in the construction above yields a law at total variation exactly δT(g) from the benign law whose adverse value is exactly ρ[qδT(g)], so no larger lower bound follows from ρ, q, and a bound on δT(g).

A.3 Value-Relevant Simulation

Proof of Proposition 2.

Fix a committed policy g. For every ξ>0, select Qξ𝒞g such that

TV((Pbg)T,(Qξ)T)δT(g)+ξ.

Applying Equation 35 to vZ gives

WZ(g)𝔼QξvZ𝔼PbgvZδT(g)ξ.

Letting ξ vanish and using WZ(g)0 proves the bound. ∎

The coefficient of δT(g) is sharp. On a two-point T space, move mass δ from a point with vZ=1 to one with vZ=0. The expectation difference and the total variation are both δ, so no coefficient smaller than one is valid.

Proof of Theorem 5.

Fix a policy g𝒢D(q). Its benign T law belongs to 𝒦Dreal and is feasible at q, hence

𝔼PbgvZΓD,Zreal(q)ΓD,Zout(q).

Combining these inequalities with Proposition 2 and monotonicity of the positive part proves Equation 17. If WZ(g)β and the realizable-frontier positive part is active, that bound rearranges to β+δT(g)ΓD,Zreal(q). If it is inactive, the same inequality already holds because δT(g)ΓD,Zreal(q).

If every feasible policy has δT(g)=0, taking the infimum over g gives RD,Z(q)ΓD,Zreal(q). Now suppose g induces a realizable-frontier optimizer and satisfies the value equalizer condition supQ𝒞g𝔼QvZ=𝔼PbgvZ. Then

WZ(g)=𝔼PbgvZ=ΓD,Zreal(q),

which supplies the reverse inequality for RD,Z(q) and proves equality. No equality of malicious laws was used. ∎

Proof of Proposition 3.

Condition on the two processes having identical histories before update t. Maximal coupling of their next evidence kernels makes the updates disagree with probability at most ηt. Couple each deployment-controlled or remaining random draw through its common conditional kernel; such a draw cannot create the first disagreement while its inputs agree. The chain rule over successive updates therefore leaves the processes coupled with probability at least t(1ηt), so a first disagreement occurs with probability at most 1t(1ηt). The coupling characterization of total variation then bounds TV(Pbg,Pm,σg) by the same quantity. Projecting a law cannot increase total variation, and the selected malicious law is one candidate in the infimum defining δT(g). These two facts prove Equation 12. ∎

The bound is attained: let every deployment transition be trivial, let the malicious process deviate to a fresh symbol independently with probability exactly ηt at each update, and let Σ contain that strategy alone; each conditional distance is then exactly ηt and the identity projection gives δT(g)=1t(1ηt).

The premise is conditional by necessity, which is the content of the corresponding negative row. Split the histories into a set dis after which the benign and malicious evidence kernels disagree maximally, so that each hdis admits an event Ah of benign conditional probability one and malicious conditional probability zero, and a complement agr after which the two kernels coincide. Let the benign law give Pbg(dis)=η¯ and Pbg(agr)=1η¯, so a suite drawn from the benign workload reports marginal disagreement rate exactly η¯. Let Σ contain a single strategy σ, and let that strategy drive the process into dis with probability one. Let A be the event that the history lies in agr, or lies in dis and the next update falls in Ah. Then Pbg(A)=(1η¯)+η¯=1 while Pm,σg(A)=0, so the two trajectory laws are mutually singular and the identity projection gives δT(g)=1 however small η¯ is. Proposition 3 asks for a distance that holds after every shared history, and the smallest such value here is one; η¯ is the benign average of those distances, and the attacker selects the histories the average treats as rare. An average over histories therefore constrains neither the kernel after any particular history nor δT(g).

Proof of Proposition 6.

Each of B(g), WZ(g), and δT(g) is a functional of the benign T law and of the set gT alone: B(g) integrates b against the former, WZ(g) maximizes 𝔼vZ over the latter, and δT(g) minimizes total variation between the former and members of the latter. Two deployments agreeing on both objects therefore agree on all three. ∎

A.4 Success-Region Sharpness

Proof of Proposition 8.

For any Q𝒞g, the reachable envelope gives Q(g)=1, so vZrG holds Q-almost surely on G. Boundedness of vZ and Q(G)λG with 1rG0 give

𝔼QvZrGQ(G)+1Q(G)1λG(1rG).

Maximization over Q proves the bound. For sharpness, a two-region law placing mass λG on value rG inside G and the remaining mass on value one outside G attains it in a model consistent with these two parameters. No smaller uniform bound therefore follows from λG and rG alone. ∎

A.5 Closed-Mediation Tightness

Proof of Theorem 9.

For any Q𝒞g, boundedness of vZ and the continuation premise give

𝔼QvZrQ(AFc)+1Q(AFc)=1(1r)Q(AFc).

The robust contract and coverage definition imply

Q(AFc)=Q(A)Q(AF)(1ϵ)Q(A)α(1ϵ).

Substitution and maximization over Q prove Equation 22.

The bound is sharp from exactly these three quantities. On a three-atom space, assign mass α(1ϵ) to AFc with value r, mass αϵ to AF with value one, and mass 1α to Ac with value one. The coverage, contract, and continuation inequalities all hold with equality, and the expectation is 1α(1ϵ)(1r). A singleton attacker class containing this law therefore attains the bound.

The sharp value is below one exactly when α(1ϵ)(1r)>0, which is equivalent to the three strict conditions in Equation 23. This proves both the if-and-only-if statement and the impossibility of a stronger uniform bound. ∎

Equation 24 follows by direct subtraction of the two sharp bounds, so the stated contribution of detection accuracy is exact rather than an estimate. Corollary 10 follows by setting ϵ=0 in the tight construction. The r=1 construction may take A to be the whole space and F empty: the local fact then holds without error on every invocation, yet vZ=1 on the success region and the deployment value remains one.

At the strict boundary, 1α(1ϵ)(1r)=0 requires α(1ϵ)(1r)=1, and since each factor lies in [0,1] this holds exactly when α=1, ϵ=0, and r=0. If any one of the three fails, the three-atom construction has strictly positive value, so no zero-residual certificate follows from these quantities.

A.6 Robust Trusted State

This subsection states the robust frontier summarized in Section 3.2.4 and proves its bound. Write T=(N,T¯), let pb be the benign marginal of trusted state N and 𝔓m the set of maliciously reachable state marginals, both common to all feasible policies, and let 𝒦DN be the set of kernels realized by some g𝒢D as the benign conditional law of T¯ given N. The robust frontier is

ΓD,ZN(q)=infk𝒦DN:𝔼pbkbqsupp𝔓m𝔼pkvZ. (36)
Theorem 12(Robust conditional simulation).

If for every feasible g and every p𝔓m the attacker can induce p(dn)Pbg(dt¯n)gT, then RD,Z(q)ΓD,ZN(q). Moreover, with d=infp𝔓mTV(pb,p),

ΓD,ZN(q)[ΓD,Zreal(q)d]+. (37)

The definition of 𝒦DN supplies both inclusions the argument uses. On a standard Borel trace space, every feasible g has a benign disintegration pb(dn)kg(dt¯n), and kg𝒦DN by definition. Conversely, every k𝒦DN is the benign disintegration of some policy in 𝒢D, so its benign law pbk belongs to 𝒦Dreal.

Proof of Theorem 12.

Fix a feasible g and disintegrate its benign T law as pb(dn)kg(dt¯n). By the robust conditional copy premise, for every p𝔓m the law pkg is attacker-reachable. Hence

WZ(g)supp𝔓m𝔼pkgvZΓD,ZN(q),

because kg𝒦DN and the legitimate value is at least q. Taking the infimum over feasible g proves the lower bound. If a robust optimizer k has a realizing policy g for which no attacker law yields value exceeding supp𝔓m𝔼pkvZ, the bound is attained at g.

For the data-processing bound, fix a feasible k. Its benign law is feasible for the realizable frontier, so 𝔼pbkvZΓD,Zreal(q). For every p𝔓m, the common kernel k and Equation 35 give

𝔼pkvZΓD,Zreal(q)TV(pb,p).

Choosing a sequence of malicious marginals whose distance tends to d and taking the infimum over k proves Equation 37. ∎

The bound depends on d rather than an average, which is the content of the corresponding table row. Let acquisition succeed on a single allowed path with probability one and fail on m other paths. The average success rate over paths tends to zero as m grows, while d and therefore the bound are unchanged.

A.7 Composition Bounds and Tightness

Proof of Proposition 11.

For a serial path, the probability chain rule gives

Pr(j=1mFj)=Pr(F1)j=2mPr(Fji<jFi).

The history-uniform premise bounds every factor, proving the product bound without independence. Independent layers attain it.

With marginal bounds only, Pr(jFj)Pr(Fi)ϵi for each i, and taking the smallest bound proves the minimum bound. It is attained when all events share a common subevent of probability miniϵi, so a stack of any depth whose layers fail together supports no bound better than its single strongest layer.

The union bound proves the alternative-path bound, and disjoint events attain it until total mass reaches one. For retries, the chain rule gives

Pr(jEjc)=jPr(Ejci<jEic).

Each factor lies between 1p¯j and 1j. Multiplication and subtraction from one prove Equation 27; sequential Bernoulli trials attain both endpoints. ∎

For additive outcomes, suppose T=(T1,,Tn), b(T)=ibi(Ti), and vZ(T)=ivi(Ti), with a separate target qi for each coordinate. Write Γi for the frontier of coordinate i. Every feasible joint law then has marginal value at least Γi(qi), so linearity gives the sum as a lower bound on the joint frontier. If the marginal optimizers have an admissible joint coupling, that coupling attains the sum. Correlation is unrestricted; the conclusion fails for a union or intersection payoff because such a payoff is not additive.

A.8 Combining Candidates

If L1,,Lm and U1,,Un are supported endpoints for the same WZ(g) under one anchor, then WZ(g)maxiLi and WZ(g)minjUj, since each inequality holds separately. Neither combination need be sharp under the joint premises: a law attaining one candidate can violate another’s. If maxiLi>minjUj, no law satisfies all the stated premises simultaneously, so at least one premise fails under the fixed anchor and neither endpoint is available until the conflict is resolved.

A candidate established for a subclass ΣΣ bounds only the supremum over 𝒞g restricted to Σ. Since Σ=kΣk gives WZ(g)=maxsupQ𝒞g(Σk)k𝔼QvZ, a family of subclass upper bounds combines into a bound for Σ only when the subclasses cover Σ, and the combined value is then their maximum rather than any single one.

Appendix B Complete Search and Screening Protocol

This appendix states the protocol under which the corpus is assembled and every denominator in the paper is produced. The supplementary materials provide the common full-text review task, the final source-status list, and both channels’ per-source assessment records.

The search target follows Section 4.4. We look for sources that make, or directly adjudicate, at least one locatable claim that some deployed or deployable intervention reduces LLM-enabled misuse or a closely specified adverse deployment outcome. Attack, red-teaming, and adaptive-evaluation sources are therefore included: they assess deployment-safety claims on the lower-bound side of the schedule.

The review task was frozen before full-text coding began.

B.1 Information Sources and Exact Queries

Three sources are used, reported in the PRISMA-S search-reporting structure [169]. S1 arXiv, via the API, restricted to cs.CR, cs.CL, cs.AI, cs.LG, cs.SE, cs.MA, and stat.ML. S2 the ACL Anthology, screened offline by regular expression over a frozen full BibTeX dump. S3 dblp venue enumeration for the major security and machine-learning venues and their workshop volumes, with a deliberately liberal title screen. The arXiv and dblp windows run from 2023-01-01 to the freeze date.

Query terms.

The query structure is derived from the language of the claim. A deployment-safety claim in the sense of Section 4.4 has a recurring surface form: [intervention] reduces / prevents / bounds [adverse outcome] for [deployed system] under [attacker or usage conditions]. Three facets are read off that frame: facet A, the deployed system (LLM, foundation model, LLM agent, and variants); facet B, the claim verb or evidential noun (defense, mitigate, safeguard, guarantee, plus adjudication nouns such as benchmark and red teaming); and facet C, the adverse outcome (misuse, jailbreak, prompt injection, exfiltration, and related terms). A record is a candidate when it matches A AND (B OR C). A fourth facet D, the mechanism lexicon, is added by union solely to recover sources matching neither B nor C. The source-specific implementations apply this facet logic under the category and date window stated above.

B.2 Deduplication and Version Families

A version family is the set of records sharing a work identity, keyed by a DOI-to-arXiv cross-link, or by normalized title similarity together with an author-set Jaccard overlap above a fixed threshold, or by an explicit “extended version of” statement. Candidate pairs may be generated automatically, including by model, but every merge is human-confirmed and logged with both identifiers.

Preprint, conference, and journal extension form one family. A workshop paper and its later full version form one family when the contribution is the same, and two families when the later claim set differs materially. System cards and policies are never merged. Each dated release is a separate source because changes in deployment claims between releases are part of the analysis.

The coded claim comes by default from the latest peer-reviewed version available at freeze. If a claim exists only in the preprint and is weakened or removed in the camera-ready, a separate instance is bound to the preprint version and flagged version_divergence. Divergence between versions is a reportable finding. Every quoted or coded claim records the identifier, version label, date, and a page or section locator.

B.3 Screening, Eligibility, and Machine Assistance

The executed sequence has five ordered stages (Figure 1): keyword-based identification; the hard authority gate; independent title-and-abstract screening, with advancement requiring include from both channels; human quality spot checks, with any detected quality problem triggering a rerun of the preceding screening step; and randomized full-text sampling and analysis. Version-family deduplication reconciles records between the authority gate and screening but does not add another eligibility criterion.

Title-and-abstract screening.

Screening is conservative: a record advances to the full-text pool only when both channels independently record include. Any disagreement, or an unsure verdict from either channel, keeps the record out. Human spot checks assess the quality and rule compliance of the channel outputs. They do not replace the both-include rule with per-record adjudication. When a spot check detects a quality problem, the title-and-abstract screening step is run again before the pool is finalized.

Full-text analysis layers.

Within the randomized full-text sample, Layer 1 decides whether a source enters the corpus; Layer 2 decides whether a codable claim instance exists.

Layer 1: source level.

A source is included only if I1 and I2 both hold and at least one of I3 or I4 holds. I1 English full text is obtainable by the freeze date. I2 the work concerns an LLM-based, LLM-integrated, or LLM-agentic deployed or deployable system. I3 the work contains at least one locatable sentence that is citable to a section, page, or paragraph and that asserts or directly tests whether an intervention changes an adverse deployment outcome, or demonstrates that such an assertion fails. A pure attack paper satisfies I3 because it adjudicates the existing claim that current deployments resist that attack class. I4 the evidence-apparatus clause admits sources that supply the instruments used to adjudicate such claims, including benchmarks, evaluation-validity critiques, auditing-access analyses, and safety-case templates; these sources are tagged role=apparatus. Methodological sources that describe how we work rather than what we study are tagged role=method and are excluded from every corpus denominator.

Six exclusion codes are used: E1 capability-only reporting with no adverse-outcome claim; E2 normative-only argument with no intervention and no adverse-outcome evidence; E3 non-LLM subject; E4 non-substantive item, with vendor system cards and policies never excluded on length; E5 secondary literature without primary claims; and E6 duplicate within a version family.

Layer 2: claim-instance level and the minimum anchoring threshold.

For each included source, channels attempt to instantiate the anchor of Equation 30. An instance is created if and only if C1 and C2 both hold and at least one of C3 or C4 holds. C1, the intervention is locatable: the source identifies a concrete evaluated policy or configuration and where it acts in the deployment. C2, the adverse outcome is locatable: a named outcome with an attached operationalization, not “unsafe behaviour.” C3, a comparison world is reported: guarded versus unguarded, guarded versus a baseline defense, or before versus after. C4, a formal statement is given: a theorem, invariant, or architectural non-bypassability argument, even without measurement.

Every remaining anchor coordinate the source does not state is recorded unknown. Coders never supply a missing coordinate by inference. A source that passes Layer 1 but fails C1 or C2 is recorded as included, no codable instance and retained: these sources are counted in every corpus denominator and must not be dropped.

Full-text exclusion requires an explicit criterion and a locator.

Machine assistance.

Title-and-abstract screening and sampled full-text analysis are performed by two independent model channels using Claude Fable 5 [4] and GPT-5.6 Sol [152]. At the title-and-abstract stage, advancement is determined mechanically by the intersection of their include decisions, subject to the human quality check and rerun rule above. At full text, each channel applies the same frozen review task, samples the full-text pool independently, and remains blind to the other. No unknown anchor coordinate is filled from model background knowledge.

Flow accounting.

Table 4 reconciles the executed pipeline. Identified records pass a hard authority gate before title-and-abstract screening: G1 admits records peer-reviewed at a fixed list of major security and machine-learning venues, G2 admits the remainder at a citations-per-year threshold checked against an open bibliographic index, and G3 covers first-party vendor material, of which these database pools contain none. Gate-eligible records are consolidated into version families, and only families included by both title-and-abstract channels enter the full-text pool after the human quality check. Full-text analysis then samples that pool in two independently randomized sequences. Twelve papers receive full ten-slot depth coding and 187 receive endpoint-route wide coding. The two strata overlap in one paper, and the coded set holds 198 distinct papers. The wide-coding counts reconcile: confirmed records for each source and channel pair split into corpus and apparatus, and corpus records split into those yielding at least one instance and those with none. The wide-coded set comprises 187 distinct papers [162, 35, 220, 29, 97, 117, 140, 219, 197, 123, 158, 195, 81, 204, 1, 198, 24, 236, 118, 10, 148, 41, 254, 132, 187, 135, 208, 100, 27, 130, 181, 238, 89, 12, 163, 124, 224, 205, 216, 26, 22, 61, 34, 138, 94, 23, 102, 157, 133, 37, 107, 95, 141, 62, 173, 147, 250, 229, 84, 242, 80, 63, 115, 127, 222, 211, 165, 65, 125, 230, 14, 36, 17, 51, 196, 245, 146, 103, 170, 49, 16, 52, 96, 174, 88, 77, 252, 246, 226, 202, 116, 120, 101, 98, 199, 30, 253, 82, 149, 73, 33, 55, 106, 129, 57, 209, 122, 233, 240, 112, 11, 59, 111, 104, 156, 251, 249, 28, 67, 184, 113, 144, 108, 191, 183, 232, 206, 235, 247, 179, 110, 69, 142, 213, 58, 87, 160, 79, 64, 239, 159, 175, 201, 83, 121, 93, 207, 243, 143, 50, 126, 99, 218, 234, 76, 178, 70, 18, 128, 71, 119, 241, 231, 244, 19, 139, 248, 186, 188, 136, 39, 155, 214, 223, 9, 66, 237, 15, 177, 90, 72, 227, 105, 86, 44, 13, 193].

媒体内容 · 前往原文查看
Table 4: Identification, screening, and claim-instance accounting for the executed pipeline, in the PRISMA reporting structure [154].
Stage Count
Identification (full scope)
arXiv, five queries, deduped scope-filtered 42,029 10,387
ACL Anthology, regex-registered scoped 3,899
dblp, 11 venues (2023–2026), title prefilter 4,990
Records entering the authority gate 19,276
Screening (full scope)
Gate-eligible (G1 5,662; G2 142; G3 0) 5,804
After version-family dedup 4,810
Both-channel include 1,872
Full-text analysis (coded set)
Distinct papers coded 198
Depth-coded papers 12
     Depth-coded claim instances 24
Wide-coded papers (one also depth-coded) 187
     Source–channel wide-coding records 190
     Excluded at Layer 1 17
     Included, role=corpus 104
      Corpus records yielding 1 instance 88
      Included, no codable instance (C1 / C2) 16
     Included, role=apparatus 69
Wide-coded claim instances extracted 152
Wide-coded instances on channel-overlap papers 7

B.4 Stopping Rule and Truncation Stability

Full-text coding samples the full-text pool rather than coding it exhaustively. Sources are coded in randomized batches, and coding stops when the reported quantities stop moving as batches are added. Saturation is defined on the estimates this paper reports. It is not defined on the supply of new boundary cases: each batch separately logs new proof-relevant anchor coordinate values or slot distinctions, new boundary cases that would require review task revision, and new claim-relation types, and those ledgers continued to record new items through the final batch of both channels. Category novelty and estimate convergence are different quantities, and the reported proportions are the ones the conclusions rest on.

Truncating the coding sequences shows that convergence directly. Dropping the final batch of each channel, then the final two, then the final three, moves the positive-residual share from 108 of 152 to 100 of 141, then 89 of 124, then 80 of 113: 71.1, 70.9, 71.8, and 70.8 percent. The scale for that movement is the estimate’s own sampling error: clustering instances within their source papers (87 clusters, design effect 1.51) gives a standard error of 4.5 points, so truncation moves the share by about a quarter of one standard error. The two channels, drawing separately randomized samples, reach 68 of 93 and 40 of 59, a difference inside that same error. Further batches would move the share within its noise rather than toward a different value.

Appendix C Review Records, Coding Aggregates, and Cases

This appendix documents the two coding strata of one full-text coding pass. The depth-coded records provide the slot-level aggregates. The wide-coded records provide the quantities required by each endpoint route.

C.1 Claim-Instance Record Format

A claim instance is the unit of coding. Each record is a structured document with five blocks. Section 4 defines the slots and states, while the coding instrument in the supplementary materials gives the per-slot rules both channels applied.

The anchor block records the source identifiers, the verbatim anchored claim cx and its locator, the claim-scope extension flag, the promotion basis, and the coded anchor θx of Equation 30. It lists one anchor coordinate per row, each with a locator or an explicit absent record, together with the channel-assigned harm_locus and the anchor-completeness count. The case tables below list the nine coordinates of θx and count completeness over those nine; records add two context rows, the outside-resource baseline and the external constraint set x, so their own counts use eleven as the denominator.

The slot block lists the ten slots of Table 2 in fixed order. Each slot includes a payload from its allowed value set, an annotated evidence state, a locator, and the required slot-specific sub-entries. The derivation and cross-source blocks record channel-added inferences, including their premises, inference rule, result, and residual uncertainty. They also record attachment rows from attack-evaluation sources; these rows are keyed by the attacked defense and carry the attacking source’s evidence.

The conclusion block records the endpoints Lx and Ux of Equation 33, the row used to compute each endpoint or the reason it is unavailable, any separately declared tolerance τx, the residual conclusion and its unresolved reasons, the independent claim_verdict, and the consistency flag. The residual conclusion and the claim verdict are recorded separately and can diverge; Section C.6 is the case where they do.

Source-reported deployment costs are retained in each record alongside the slots. They do not enter any endpoint, and they are reported here only where a case discussion uses them, as in the utility cost of the zero-continuation design in Section C.5.

C.2 Supplementary Evidence Files

The supplementary materials contain four evidence files: the common review task, the final source-status list, and one final assessment compilation for each model channel. All four cover the wide-coded portion of the full-text coding pass. The list records channel, source identifier, full-text eligibility status, claim-instance status, and instance count. The two assessment compilations preserve the evidence and source locators behind the wide-coded aggregates in Section C.4. The depth-coded subset is documented in this appendix instead.

Each assessment record carries the anchor coordinates, the slot evidence with its locators, and the channel’s structured conclusion block. The endpoints and residual conclusions tabulated below are those recorded in the conclusion blocks, computed under Equation 33, and can be re-derived from it. A positive residual rules out τx=0 under the anchor.

C.3 Depth-Coded Subset Aggregates

Table 5 tabulates the depth-coded subset’s slot-level evidence states, and Table 6 gives its instance-level distributions beside the wide-coded set’s. The subset is 24 claim instances, plus the two attack-evaluation source records it also covers, which yield no instances of their own and instead contribute 45 attachment rows keyed by the defenses they attack. The two channels agreed on 24 of 24 residual conclusions and on 20 of 24 claim verdicts. The table abbreviates supported as su, derived as de, claimed as cl, not-reported as nr, and not-applicable as na. The 45 attachment rows are excluded from the table and are all supported.

媒体内容 · 前往原文查看
Table 5: Evidence states by slot in the depth-coded subset (24 instances; 240 slot assessments).
Slot su de cl nr na
Lower-endpoint evidence
LB1 21 0 1 2 0
LB2 18 0 3 3 0
LB3 19 0 2 3 0
LB4 15 0 2 4 3
Upper-endpoint evidence
UB0 8 0 0 13 3
UB1 23 0 1 0 0
UB2 19 0 2 3 0
UB3 21 0 2 1 0
UB4 4 1 2 17 0
UB5 19 0 2 0 3
Total 167 1 17 46 9

C.4 Wide-Coded Set Aggregates

Table 6 reports the two coding strata together. The two channels sampled the 1,872-paper full-text pool under different random seeds. In the wide-coded stratum, they coded 80 and 110 papers (187 distinct) and extracted 152 claim instances between them. Because the channels assessed the overlap papers independently, instance counts are pooled across channels rather than deduplicated at the paper level; the three channel-overlap papers yield seven instances across the two channels. Wide-coded claim verdicts are single-channel judgments; they stay on the individual records and are not aggregated, so the verdict rows below cover the depth-coded subset only. The nine positive residuals occur with upheld and unresolved verdicts, and the zero-residual instance has an upheld verdict. All three refuted verdicts occur in instances whose residual conclusion remains unresolved.

媒体内容 · 前往原文查看
Table 6: Instance-level outcomes in the depth-coded and wide-coded strata.
Depth-coded Wide-coded
(n=24) (n=152)
Residual conclusion
zero residual established 1 0
unresolved 14 44
positive residual established 9 108
Claim verdict
upheld 4
refuted 3
unresolved 17
Harm locus Wide-coded
service-integrity 81
external-world 52
mixed 19
Endpoint availability
upper endpoint below one 1 0
both endpoints computable 1 0

C.5 Case Record I: All Three Gates Closed

Instance costa2025fides-01 anchors the by-design integrity noninterference claim of the Fides information-flow-control planner [43]. It is the depth-coded subset’s only zero-residual instance and its only instance with both endpoints computable. The anchored claim is that attacker-controlled untrusted data cannot influence the agent’s consequential tool actions. Its locator is Section 1 and Proposition 1 in Section 4.4 of the source. The record has promotion_basis = explicit-deployment-claim, harm_locus = service-integrity, and claim-scope extension yes because the source separately uses a broader prompt-injection headline. The broader empirical and confidentiality readings are represented by sibling instances.

Table 7 gives all nine coordinates of the coded anchor θx and the complete ten-slot projection with payload, state, and locator, so the conclusion below can be replayed without consulting the narrative in Section 5. The deployment-specific constraint set x is not a coordinate of θx. The completeness count is 6/9: Sx, Dx, x, Zx, Σx, and bx are filled, while qx, Tx, and the general continuation value vZx remain unknown. The cross-episode part of the horizon is also unreported. An absent locator supplies no default, and a (+) state applies to the slot’s sub-entry.

媒体内容 · 前往原文查看
Table 7: Complete coded record for costa2025fides-01: anchor, slots, states, and source locators.
Coded anchor θx
Coordinate Coded value Status Source locator
Sx ReAct-style LLM agent loop with native tool calling and no information-flow control filled Section 2; Section 7.2 Basic-planner baseline
Dx Fides taint-tracking planner, P-T/P-F policy engine, selective hide/reveal, query_llm, and constrained decoding filled Sections 4 and 5
x Agentic tasks over email, calendar, banking, Slack, and travel tools; untrusted tool outputs; one AgentDojo user-task episode filled Section 2; Section 7; cross-episode persistence absent
Zx Probability of the binary episode event in which untrusted tool data causes this agent to execute an unintended consequential tool action filled Section 2.1; Section 4.3 P-T; Section 7.2 ASR metric
Σx Defense-aware attacker controlling arbitrary untrusted tool content, knowing the configuration, and observing some tool effects; configuration compromise excluded filled Section 2.1 threat model
qx Task-completion rate measured on 97 AgentDojo tasks, but no minimum normal-utility threshold declared unknown Section 8.2; threshold absent
Tx No explicit value-relevant trajectory projection unknown absent
bx AgentDojo task-completion rate under a programmatic user-goal check filled Section 7.2
vZx Binary integrity semantics are stated, but a general continuation value is not separately reported unknown Section 8.1; route-specific noninterference in Proposition 1
Slots
Slot Payload State Source locator
LB1 present, clean attribution; injections 163(156) to 1(0), parenthetical figures excluding two tasks the source rules outside its policies; task-completion-rate improvement 16.7% for o1 supported Table 1; Section 8.2 Figure 4
LB2 restriction-only; query_llm is an additive candidate but the matched-utility comparator is absent supported, claimed(+) Sections 4.3 and 5; comparator absent
LB3 reproducible, with zero benign adverse value supported Proposition 1 and noninterference definition in Section 4.4; Algorithm 5 lines 7 and 9; source Appendix A
LB4 separation shown supported Section 4.1 default-untrusted labeling; Section 4.3 P-T; Proposition 1; Section 2.1
UB0 no artifact-transfer or capability-removal route not-applicable Mediation architecture in Sections 4 and 5
UB1 session-grain, decidable policy-success event supported Section 4.3; Section 4.4 complete-execution noninterference; Algorithm 5
UB2 deployment coverage, αx=1 supported Algorithm 5 line 7; Section 8.1 response channel outside this integrity outcome
UB3 robust integrity bound, ϵx=0 supported Proposition 1; source Appendix A small-step semantics; Algorithm 5
UB4 bounded continuation, rx=0 on covered integrity successes derived Section 2.1 binary episode event behind Zx; Proposition 1; Algorithm 5
UB5 serial per-call stateful topology; whole-trace property proved directly rather than composed from marginal rates supported Algorithm 5; Section 4.4

The conclusion procedure executes in four steps. C1 records the general continuation coordinate vZx as unknown while the route-specific integrity value is fully identified by the binary adverse event behind Zx, Proposition 1, and the derived rx=0. The endpoints therefore rest on supported or derived premises: the attacker controls untrusted tool data, the configuration remains trusted, and the proof uses the same consequential-action semantics as the guarded comparison. C2 combines supported coverage αx=1, supported robust failure ϵx=0, and derived continuation rx=0:

Ux=1αx(1ϵx)(1rx)=11(10)(10)=0. (38)

With Zx normalized to [0,1] this gives 0Zx(Sx[Dx])0, so C3 establishes a zero-residual certificate at the strict boundary and clears the consistency flag. C4 assigns upheld because the coverage, failure, and continuation evidence matches the integrity claim’s own attacker class.

The decisive coding judgment is whether UB3 is supported or merely claimed. The final assessment assigns supported. The source provides Proposition 1 and the small-step proof in its Appendix A. The integrity result does not depend on deterministic model behavior, and within the declared Σx, external data receives the conservative untrusted label by construction. This ruling supports ϵx=0. The route-specific continuation value has an independent basis. Zx is the probability of the binary event that untrusted data causes one unintended consequential action within the episode. For this outcome, a covered success leaves no further adverse value. The record derives rx=0 from this premise. The outcome definition, not the accuracy of the check, makes this endpoint available.

The record also retains the source-reported cost of this design: policy-on task-completion loss up to 24.5% with a data-dependent-task ceiling, and two to three times the Basic planner’s token use plus query_llm latency, at Section 8.2 Figure 4 and source Appendix E Figure 7. These values do not enter the endpoint. They quantify the cost of the design that yields rx=0, and a concrete deployment weighs them against its declared q.

The certificate covers the integrity-scoped outcome with trusted configuration and correct conservative labeling. Text response manipulation, implicit confidentiality leakage, and the broader statement that the system stops all prompt-injection attacks belong to the two sibling instances.

C.6 Case Record II: Refuted Under the Claim’s Own Class

Instance zou2024circuitbreakers-01 anchors the representation-rerouting claim of [255] for text-only models. Table 8 gives the condensed record. It is one of the three depth-coded instances whose claim is refuted while the residual conclusion stays unresolved, so it illustrates the separation of residual conclusion and claim verdict. It also shows the role of attachment rows contributed by an attack-evaluation source [145], which yields no instance of its own.

媒体内容 · 前往原文查看
Table 8: zou2024circuitbreakers-01, condensed. Coded anchor θx completeness 5/9; harm_locus = external-world; claim-scope extension yes (the anchored abstract claim reaches beyond a single-turn scope declaration). The two attachment rows are contributed by the attack source and inherit this instance’s anchor.
Slot Payload State
LB1 present, clean attribution supported
LB2 restriction-only supported, claimed(+)
LB3 simulable-output supported
LB4 structural none supported
UB0 silent none not-reported
UB1 event-defined supported
UB2 artifact-coverage supported
UB3 reference-rate, class check no-open-quantifier supported, claimed(+)
UB4 none not-reported
UB5 none, single component supported
attachment:LB3 simulable-output at 100% ASR supported
attachment:UB3 refuted, ϵ1 supported
Residual conclusion and claim verdict
On the upper side there is no robust ϵ and no r, so Ux=1, the bound available before measurement. On the lower side the attachment rows report a conditional failure rate and output simulability rather than an adverse value measured on the anchored Zx scale. The benign continuation value needed by the simulation row is not reported either, so no nontrivial Lx follows. State unresolved (no-nontrivial-endpoint); verdict refuted.

The verdict is refuted because supported attachment evidence contradicts the anchored claim within its own declared class: the adaptive attack reaches ϵ1, which also reclassifies the in-source suite averages as a reference rate. The residual conclusion nonetheless remains unresolved. Refuting the robustness claim invalidates an upper-bound operand, while the lower-bound rows contributed by the attack source are not measured on the anchored scale and therefore supply no lower endpoint. This difference is why the two outputs are recorded separately.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org