cd /news/large-language-models/safety-alignment-inverts-the-tension… · home › topics › large-language-models › article
[ARTICLE · art-148683] src=dev.to ↗ pub= topic=large-language-models verified=true sentiment=· neutral

Safety Alignment Inverts the Tension–Response Curve in LLM Agents for Civil Violence Simulation

A developer testing LLM-driven agent-based models against Epstein's 2002 civil violence model found that safety-aligned Qwen 27B inverted the expected tension-response curve: under an act/wait label pair, activation probability fell from 41.5% at tension 0.10 to 12.6% at 0.90, the opposite of the reference rule, while relabeling the options protest/comply restored a monotone rise from 0% to 95%. Across 600+ calibration conditions, the work attributes the distortion to lexical label priors, position bias and a tension-dependent anti-action component hypothesized to stem from RLHF/DPO safety alignment, and argues LLM-ABM pipelines need a calibration stage with a reported bias budget before agent behavior is compared with mathematical baselines.

by read5 min views1 publishedOct 10, 2026

LLM-driven agent-based models (ABMs) replace a closed-form decision rule with an LLM agent. The idea is that the agent reproduces the qualitative behaviour of the original rule while adding context-sensitive reasoning. We test that assumption on Epstein's (2002) civil violence model. Across 600+ calibration conditions on a locally served Qwen 27B model (4-bit, MLX), the probability of choosing activation falls as scenario tension rises, from 41.5% at tension 0.10 to 12.6% at 0.90 under an act/wait label pair. Epstein's rule predicts the opposite. Changing only the labels to protest/comply restores a monotonically increasing response, from 0% to 95% over the same range. We break the observed distortion into lexical label priors, position bias and a tension-dependent anti-action component, and we hypothesise that the last one comes from safety alignment. We argue that LLM-ABM pipelines need a calibration stage, with a reported "bias budget", before agent behaviour is compared with mathematical baselines.

In Epstein's model each citizen agent computes a grievance term and a net-risk term:

The agent turns active when G − N > T.

The rule is deterministic given state, and it produces a sharp increase in activation as legitimacy falls and hardship rises. That property makes it a usable reference: an LLM substitute that is valid at the decision level should be at least monotonically non-decreasing in G − N.

In this work, the LLM agent gets a natural-language rendering of the same state variables and must return a binary choice.

Parameter Value
Primary model Qwen 27B, 4-bit quantised, MLX, local inference
Secondary model Mistral 7B Instruct v0.3
Temperature 0.7
Repetitions per condition 20
Conditions 600+
Label order Counterbalanced
Tension levels 0.10, 0.50, 0.70, 0.90

| Country prior (Romania) | L = 0.21 (Eurobarometer), H = 0.55 (OECD), police density 0.003, R = 0.70 | Tension maps to model inputs as follows: [describe: which of H, L, police presence vary with the tension scalar, and how].

Conditions vary along: [list factors: label pair, prompt template, option order, persona attributes, ...].

The response is parsed by [exact-match on label / constrained decoding / regex], and invalid outputs are handled by [dropped / re-sampled / counted as ...].

| Tension | P(act) | 
|---|---|

| 0.10 | 41.5% | | 0.50 | 38.1% | | 0.70 | 29.3% | | 0.90 | 12.6% |

The response is monotonically decreasing in tension. At 0.90, where the reference model predicts near-universal activation, the LLM agent activates in about 1 in 8 samples.

| Tension | P(protest) | 
|---|---|

| 0.10 | 0% | | 0.50 | 20% | | 0.70 | 35% | | 0.90 | 95% |

The curve has the expected shape, with a steep rise between 0.70 and 0.90, which matches the threshold behaviour of the reference rule.

At n = 20 per cell, the 95% Wilson interval around 20% is about [8%, 42%]. The intermediate points should be read as directional until they are re-estimated with more samples or pooled across conditions.

Protest/comply was the only pair that produced a monotone response among the [N] label pairs tested, including [list].

Because the working pair was found by search, its performance is an in-sample result. It needs validation on held-out prompt templates before it can count as a calibrated configuration.

Hypothesis. Preference tuning (RLHF/DPO-style safety alignment) gives a negative prior on outputs that endorse unspecified "action" in high-conflict contexts. The penalty grows with how strongly the context signals conflict. That would explain why suppression increases with tension instead of staying constant.

On this account, protest escapes the penalty because the training data encodes it as a protected civic act, while act, rebel or take action sit near content the alignment stage learned to refuse or discourage. Wait draws a strong positive prior as a de-escalatory, "safe" completion.

Status. The data here are consistent with this hypothesis but do not isolate it. A direct test is to compare base and aligned checkpoints of the same model family under identical prompts. The tension-dependent component should be present in the aligned checkpoint and absent, or much smaller, in the base one.

Mistral 7B Instruct v0.3. Its documentation does not describe a dedicated safety-tuning stage. It does not show the anti-action pattern. Instead it shows near-total primacy bias: it chooses whichever option is presented first in about 100% of samples, regardless of label or tension. The output carries no information about scenario state.

Base models. These show smaller safety-type priors but larger social biases (in-group favouritism) and unreliable instruction following, so they cannot serve as drop-in agents.

None of the three model classes gives an unbiased decision function without calibration.

| Component | Magnitude | Scope |

|---|---|---|
| Primacy (first-option) bias | ~100% | Mistral 7B | 

| Position bias | 5–16 pp | [model] | | Lexical prior on "wait" | −45 pp on action | Qwen 27B, act/wait | | Tension-dependent anti-action | −15 to −30 pp at high tension | Qwen 27B |

Each of these effects is comparable to, or larger than, the behavioural differences the simulation is meant to resolve. Taken together they are enough to turn a predicted mass mobilisation into near-total quiescence.

SocioVerse-ABM (Fudan University / Shanghai Innovation Institute, v0.2.0, 2026) includes Epstein's model among twelve benchmark scenarios and evaluates GPT-4o, DeepSeek-V3 and Qwen3-235B against mathematical baselines. AgentTorch (MIT Media Lab) supplies the computational patterns that make large-population LLM-ABMs feasible.

Both frameworks compare agent output to a baseline. Neither, as far as we know, includes a pre-simulation step that separates decision-function distortion from emergent dynamics. Without that step, a deviation from baseline cannot be attributed to the model's reasoning rather than to its label and position priors.

Li et al. (2025, ACM FAccT) report a related pattern: alignment reduces explicit bias while implicit bias persists or grows.

Full calibration data are available on request.

Epstein, J.M. (2002). Modeling civil violence: An agent-based computational approach. *PNAS*, 99(suppl. 3), 7243–7250.

SocioVerse-ABM v0.2.0, Fudan University (2026).

Li et al. (2025). Actions Speak Louder than Words. ACM FAccT.

AgentTorch, MIT Media Lab.

── more in #large-language-models 4 stories · sorted by recency
── more on @qwen 27b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/safety-alignment-inv…] indexed:0 read:5min 2026-10-10 · —