{"slug": "safety-alignment-inverts-the-tension-response-curve-in-llm-agents-for-civil", "title": "Safety Alignment Inverts the Tension–Response Curve in LLM Agents for Civil Violence Simulation", "summary": "A developer testing LLM-driven agent-based models against Epstein's 2002 civil violence model found that safety-aligned Qwen 27B inverted the expected tension-response curve: under an act/wait label pair, activation probability fell from 41.5% at tension 0.10 to 12.6% at 0.90, the opposite of the reference rule, while relabeling the options protest/comply restored a monotone rise from 0% to 95%. Across 600+ calibration conditions, the work attributes the distortion to lexical label priors, position bias and a tension-dependent anti-action component hypothesized to stem from RLHF/DPO safety alignment, and argues LLM-ABM pipelines need a calibration stage with a reported bias budget before agent behavior is compared with mathematical baselines.", "body_md": "LLM-driven agent-based models (ABMs) replace a closed-form decision rule with an LLM agent. The idea is that the agent reproduces the qualitative behaviour of the original rule while adding context-sensitive reasoning. We test that assumption on Epstein's (2002) civil violence model. Across 600+ calibration conditions on a locally served Qwen 27B model (4-bit, MLX), the probability of choosing activation falls as scenario tension rises, from 41.5% at tension 0.10 to 12.6% at 0.90 under an *act/wait* label pair. Epstein's rule predicts the opposite. Changing only the labels to *protest/comply* restores a monotonically increasing response, from 0% to 95% over the same range. We break the observed distortion into lexical label priors, position bias and a tension-dependent anti-action component, and we hypothesise that the last one comes from safety alignment. We argue that LLM-ABM pipelines need a calibration stage, with a reported \"bias budget\", before agent behaviour is compared with mathematical baselines.\n\nIn Epstein's model each citizen agent computes a grievance term and a net-risk term:\n\nThe agent turns active when G − N > T.\n\nThe rule is deterministic given state, and it produces a sharp increase in activation as legitimacy falls and hardship rises. That property makes it a usable reference: an LLM substitute that is valid at the decision level should be at least monotonically non-decreasing in G − N.\n\nIn this work, the LLM agent gets a natural-language rendering of the same state variables and must return a binary choice.\n\n| Parameter | Value | \n|---|---|\n| Primary model | Qwen 27B, 4-bit quantised, MLX, local inference | \n| Secondary model | Mistral 7B Instruct v0.3 | \n| Temperature | 0.7 | \n| Repetitions per condition | 20 | \n| Conditions | 600+ | \n| Label order | Counterbalanced | \n| Tension levels | 0.10, 0.50, 0.70, 0.90 | \n| Country prior (Romania) | L = 0.21 (Eurobarometer), H = 0.55 (OECD), police density 0.003, R = 0.70 | \n\nTension maps to model inputs as follows: `[describe: which of H, L, police presence vary with the tension scalar, and how]`.\n\nConditions vary along: `[list factors: label pair, prompt template, option order, persona attributes, ...]`.\n\nThe response is parsed by `[exact-match on label / constrained decoding / regex]`, and invalid outputs are handled by `[dropped / re-sampled / counted as ...]`.\n\n| Tension | P(act) | \n|---|---|\n| 0.10 | 41.5% | \n| 0.50 | 38.1% | \n| 0.70 | 29.3% | \n| 0.90 | 12.6% | \n\nThe response is monotonically decreasing in tension. At 0.90, where the reference model predicts near-universal activation, the LLM agent activates in about 1 in 8 samples.\n\n| Tension | P(protest) | \n|---|---|\n| 0.10 | 0% | \n| 0.50 | 20% | \n| 0.70 | 35% | \n| 0.90 | 95% | \n\nThe curve has the expected shape, with a steep rise between 0.70 and 0.90, which matches the threshold behaviour of the reference rule.\n\nAt n = 20 per cell, the 95% Wilson interval around 20% is about [8%, 42%]. The intermediate points should be read as directional until they are re-estimated with more samples or pooled across conditions.\n\nProtest/comply was the only pair that produced a monotone response among the `[N]` label pairs tested, including `[list]`.\n\nBecause the working pair was found by search, its performance is an in-sample result. It needs validation on held-out prompt templates before it can count as a calibrated configuration.\n\n**Hypothesis.** Preference tuning (RLHF/DPO-style safety alignment) gives a negative prior on outputs that endorse unspecified \"action\" in high-conflict contexts. The penalty grows with how strongly the context signals conflict. That would explain why suppression increases with tension instead of staying constant.\n\nOn this account, *protest* escapes the penalty because the training data encodes it as a protected civic act, while *act*, *rebel* or *take action* sit near content the alignment stage learned to refuse or discourage. *Wait* draws a strong positive prior as a de-escalatory, \"safe\" completion.\n\n**Status.** The data here are consistent with this hypothesis but do not isolate it. A direct test is to compare base and aligned checkpoints of the same model family under identical prompts. The tension-dependent component should be present in the aligned checkpoint and absent, or much smaller, in the base one.\n\n**Mistral 7B Instruct v0.3.** Its documentation does not describe a dedicated safety-tuning stage. It does not show the anti-action pattern. Instead it shows near-total primacy bias: it chooses whichever option is presented first in about 100% of samples, regardless of label or tension. The output carries no information about scenario state.\n\n**Base models.** These show smaller safety-type priors but larger social biases (in-group favouritism) and unreliable instruction following, so they cannot serve as drop-in agents.\n\nNone of the three model classes gives an unbiased decision function without calibration.\n\n| Component | Magnitude | Scope | \n|---|---|---|\n| Primacy (first-option) bias | ~100% | Mistral 7B | \n| Position bias | 5–16 pp | `[model]` | \n| Lexical prior on \"wait\" | −45 pp on action | Qwen 27B, act/wait | \n| Tension-dependent anti-action | −15 to −30 pp at high tension | Qwen 27B | \n\nEach of these effects is comparable to, or larger than, the behavioural differences the simulation is meant to resolve. Taken together they are enough to turn a predicted mass mobilisation into near-total quiescence.\n\n**SocioVerse-ABM** (Fudan University / Shanghai Innovation Institute, v0.2.0, 2026) includes Epstein's model among twelve benchmark scenarios and evaluates GPT-4o, DeepSeek-V3 and Qwen3-235B against mathematical baselines. **AgentTorch** (MIT Media Lab) supplies the computational patterns that make large-population LLM-ABMs feasible.\n\nBoth frameworks compare agent output to a baseline. Neither, as far as we know, includes a pre-simulation step that separates decision-function distortion from emergent dynamics. Without that step, a deviation from baseline cannot be attributed to the model's reasoning rather than to its label and position priors.\n\nLi et al. (2025, ACM FAccT) report a related pattern: alignment reduces explicit bias while implicit bias persists or grows.\n\nFull calibration data are available on request.\n\nEpstein, J.M. (2002). Modeling civil violence: An agent-based computational approach. *PNAS*, 99(suppl. 3), 7243–7250.\n\nSocioVerse-ABM v0.2.0, Fudan University (2026).\n\nLi et al. (2025). Actions Speak Louder than Words. *ACM FAccT*.\n\nAgentTorch, MIT Media Lab.", "url": "https://wpnews.pro/news/safety-alignment-inverts-the-tension-response-curve-in-llm-agents-for-civil", "canonical_source": "https://dev.to/dracopol/safety-alignment-inverts-the-tension-response-curve-in-llm-agents-for-civil-violence-simulation-32pf", "published_at": "2026-10-10 08:47:12+00:00", "updated_at": "2026-10-10 09:10:33.973515+00:00", "lang": "en", "topics": ["large-language-models", "ai-safety", "ai-agents", "ai-research", "machine-learning"], "entities": ["Qwen 27B", "Mistral 7B Instruct v0.3", "Epstein", "MLX", "Eurobarometer", "OECD"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/safety-alignment-inverts-the-tension-response-curve-in-llm-agents-for-civil", "markdown": "https://wpnews.pro/news/safety-alignment-inverts-the-tension-response-curve-in-llm-agents-for-civil.md", "text": "https://wpnews.pro/news/safety-alignment-inverts-the-tension-response-curve-in-llm-agents-for-civil.txt", "jsonld": "https://wpnews.pro/news/safety-alignment-inverts-the-tension-response-curve-in-llm-agents-for-civil.jsonld"}}