# Safety Alignment Inverts the Tension–Response Curve in LLM Agents for Civil Violence Simulation

> Source: <https://dev.to/dracopol/safety-alignment-inverts-the-tension-response-curve-in-llm-agents-for-civil-violence-simulation-32pf>
> Published: 2026-10-10 08:47:12+00:00

LLM-driven agent-based models (ABMs) replace a closed-form decision rule with an LLM agent. The idea is that the agent reproduces the qualitative behaviour of the original rule while adding context-sensitive reasoning. We test that assumption on Epstein's (2002) civil violence model. Across 600+ calibration conditions on a locally served Qwen 27B model (4-bit, MLX), the probability of choosing activation falls as scenario tension rises, from 41.5% at tension 0.10 to 12.6% at 0.90 under an *act/wait* label pair. Epstein's rule predicts the opposite. Changing only the labels to *protest/comply* restores a monotonically increasing response, from 0% to 95% over the same range. We break the observed distortion into lexical label priors, position bias and a tension-dependent anti-action component, and we hypothesise that the last one comes from safety alignment. We argue that LLM-ABM pipelines need a calibration stage, with a reported "bias budget", before agent behaviour is compared with mathematical baselines.

In Epstein's model each citizen agent computes a grievance term and a net-risk term:

The agent turns active when G − N > T.

The rule is deterministic given state, and it produces a sharp increase in activation as legitimacy falls and hardship rises. That property makes it a usable reference: an LLM substitute that is valid at the decision level should be at least monotonically non-decreasing in G − N.

In this work, the LLM agent gets a natural-language rendering of the same state variables and must return a binary choice.

| Parameter | Value | 
|---|---|
| Primary model | Qwen 27B, 4-bit quantised, MLX, local inference | 
| Secondary model | Mistral 7B Instruct v0.3 | 
| Temperature | 0.7 | 
| Repetitions per condition | 20 | 
| Conditions | 600+ | 
| Label order | Counterbalanced | 
| Tension levels | 0.10, 0.50, 0.70, 0.90 | 
| Country prior (Romania) | L = 0.21 (Eurobarometer), H = 0.55 (OECD), police density 0.003, R = 0.70 | 

Tension maps to model inputs as follows: `[describe: which of H, L, police presence vary with the tension scalar, and how]`.

Conditions vary along: `[list factors: label pair, prompt template, option order, persona attributes, ...]`.

The response is parsed by `[exact-match on label / constrained decoding / regex]`, and invalid outputs are handled by `[dropped / re-sampled / counted as ...]`.

| Tension | P(act) | 
|---|---|
| 0.10 | 41.5% | 
| 0.50 | 38.1% | 
| 0.70 | 29.3% | 
| 0.90 | 12.6% | 

The response is monotonically decreasing in tension. At 0.90, where the reference model predicts near-universal activation, the LLM agent activates in about 1 in 8 samples.

| Tension | P(protest) | 
|---|---|
| 0.10 | 0% | 
| 0.50 | 20% | 
| 0.70 | 35% | 
| 0.90 | 95% | 

The curve has the expected shape, with a steep rise between 0.70 and 0.90, which matches the threshold behaviour of the reference rule.

At n = 20 per cell, the 95% Wilson interval around 20% is about [8%, 42%]. The intermediate points should be read as directional until they are re-estimated with more samples or pooled across conditions.

Protest/comply was the only pair that produced a monotone response among the `[N]` label pairs tested, including `[list]`.

Because the working pair was found by search, its performance is an in-sample result. It needs validation on held-out prompt templates before it can count as a calibrated configuration.

**Hypothesis.** Preference tuning (RLHF/DPO-style safety alignment) gives a negative prior on outputs that endorse unspecified "action" in high-conflict contexts. The penalty grows with how strongly the context signals conflict. That would explain why suppression increases with tension instead of staying constant.

On this account, *protest* escapes the penalty because the training data encodes it as a protected civic act, while *act*, *rebel* or *take action* sit near content the alignment stage learned to refuse or discourage. *Wait* draws a strong positive prior as a de-escalatory, "safe" completion.

**Status.** The data here are consistent with this hypothesis but do not isolate it. A direct test is to compare base and aligned checkpoints of the same model family under identical prompts. The tension-dependent component should be present in the aligned checkpoint and absent, or much smaller, in the base one.

**Mistral 7B Instruct v0.3.** Its documentation does not describe a dedicated safety-tuning stage. It does not show the anti-action pattern. Instead it shows near-total primacy bias: it chooses whichever option is presented first in about 100% of samples, regardless of label or tension. The output carries no information about scenario state.

**Base models.** These show smaller safety-type priors but larger social biases (in-group favouritism) and unreliable instruction following, so they cannot serve as drop-in agents.

None of the three model classes gives an unbiased decision function without calibration.

| Component | Magnitude | Scope | 
|---|---|---|
| Primacy (first-option) bias | ~100% | Mistral 7B | 
| Position bias | 5–16 pp | `[model]` | 
| Lexical prior on "wait" | −45 pp on action | Qwen 27B, act/wait | 
| Tension-dependent anti-action | −15 to −30 pp at high tension | Qwen 27B | 

Each of these effects is comparable to, or larger than, the behavioural differences the simulation is meant to resolve. Taken together they are enough to turn a predicted mass mobilisation into near-total quiescence.

**SocioVerse-ABM** (Fudan University / Shanghai Innovation Institute, v0.2.0, 2026) includes Epstein's model among twelve benchmark scenarios and evaluates GPT-4o, DeepSeek-V3 and Qwen3-235B against mathematical baselines. **AgentTorch** (MIT Media Lab) supplies the computational patterns that make large-population LLM-ABMs feasible.

Both frameworks compare agent output to a baseline. Neither, as far as we know, includes a pre-simulation step that separates decision-function distortion from emergent dynamics. Without that step, a deviation from baseline cannot be attributed to the model's reasoning rather than to its label and position priors.

Li et al. (2025, ACM FAccT) report a related pattern: alignment reduces explicit bias while implicit bias persists or grows.

Full calibration data are available on request.

Epstein, J.M. (2002). Modeling civil violence: An agent-based computational approach. *PNAS*, 99(suppl. 3), 7243–7250.

SocioVerse-ABM v0.2.0, Fudan University (2026).

Li et al. (2025). Actions Speak Louder than Words. *ACM FAccT*.

AgentTorch, MIT Media Lab.
