Show HN: Comparing eight LLMs on 38 Berlin election questions A new analysis by Berlin WahLLM comparing eight large language models, including ChatGPT and xAI's Grok, on 38 Wahl-O-Mat theses for Berlin's 2026 state election found that seven of the eight models showed highest agreement with the SPD, Greens, or Left parties. The project, which uses documented model responses and bpb party positions, cautions that the results are not evidence of stable political positions but reflect statistically learned language patterns. Berlin WahLLM Who would AI vote for? Eight language models such as ChatGPT answer the 38 theses listed below from the Wahl-O-Mat for Berlin's 2026 state election. 8 models · 38 theses Party selection: SPD, Die Linke, CDU, FDP, AfD, and Greens are preselected: the parties represented in Berlin's House of Representatives after the 2021 election, including the FDP, which later left parliament. The models at a glance the-models-at-a-glance For each model, the cards show the party with the highest agreement details under “Models, sources, data and code” daten . Where several parties are mathematically tied, all are named. Source: own calculation equivalent to the unweighted calculation in the Wahl-O-Mat https://www.wahl-o-mat.de/berlin2026/ using the documented model responses and bpb party positions. How close are the models to the parties? how-close-are-the-models-to-the-parties Each row represents one of the eight models and each column a party. Darker cells indicate higher mathematical agreement. The colour scale remains fixed at 0 to 100 percent in both party modes. The results for xAI's Grok differ markedly from the other models in this sample. This raises the interesting question of how much training data, system instructions, model alignment and random variation each contribute to such differences. A single observed run cannot yet provide a robust answer. Source: own calculation equivalent to the unweighted calculation in the Wahl-O-Mat https://www.wahl-o-mat.de/berlin2026/ using the documented model responses and bpb party positions. The models in detail the-models-in-detail Inspect the model responses in detail here. The dot plot shows the main model and up to three freely chosen comparison models; the data table and metadata continue to refer only to the main model. The selection persists when switching models and party modes. The ranking mathematically summarises 38 responses. It is not a voting intention or a recommendation. How do the eight models respond to the theses? how-do-the-eight-models-respond-to-the-theses The matrix shows agreement, neutrality or disagreement for the eight models and all 38 theses. The thesis texts are reproduced as the original German source material. Source: the documented model responses, unchanged. How can this pattern be interpreted? how-can-this-pattern-be-interpreted Seven of the eight selected model runs reach their highest agreement, within the default party selection, with the SPD, Greens or Left. This is a clear pattern in this sample, but not evidence of a stable political position or of a particular cause. Interpreting such results is inherently difficult: the responses do not represent political convictions in the human sense. They emerge from statistically learned language patterns shaped by the prompt, training data, post-training and system instructions. One possible explanation is already present in the prompt https://github.com/Bensk1/berlin-wahllm/blob/main/PROMPT.md . It describes an eligible voter in Berlin and asks for answers consistent with that person's character and political views, without specifying that person further. The model has to supply the missing identity itself. Both a learned assistant persona and statistical associations with Berlin may influence the responses. Training data and subsequent model alignment may also matter. Modern language models are adjusted with human ratings, behavioural rules and system instructions to give helpful and as harmless as possible responses. One possible, untested hypothesis is that this makes values such as equal treatment, inclusion, public support and environmental protection especially likely to be endorsed in abstract decision situations. Earlier studies found socially liberal tendencies in some similarly trained models, but also large differences between prompts and measurement methods. They do not establish the cause of the pattern observed here. See “Whose Opinions Do Language Models Reflect?” https://proceedings.mlr.press/v202/santurkar23a.html and “Political Compass or Spinning Arrow?” https://aclanthology.org/2024.acl-long.816/ . The questionnaire itself is another factor. Its brief theses usually mention neither costs nor trade-offs, and the forced format allows neither reasons nor conditions. The calculated party proximity can therefore be as much a product of wording, response format and party positions as an expression of a general response pattern. The eight runs are also not independent observations: models from different providers may share similar training data and notions of helpful assistant behaviour. Blocked runs do not appear in the ranking. The experiment shows a green-left proximity in the generated response vectors, but does not yet explain its origin. Whether this should be called bias also depends on the benchmark: the prompt does not say whether a model should represent Berlin's population, an average of parties, or a neutral answer distribution. How could this be tested? Ideas for further research: - run the prompt both with and without the reference to “character, nature and political views”, - replace Berlin with a neutral location some theses directly concern Berlin , - ask theses in semantically reversed form, - vary thesis order at random, - collect several runs per model and experimental condition, - have models also explain their answers openly, and - compare results with human survey data on the same theses. Method method All models received the same documented prompt https://github.com/Bensk1/berlin-wahllm/blob/main/PROMPT.md . The prompt's ideas was to surface the views of each LLM. It was explicitly not intended to respond as if it were a person with well-defined socioeconomic characteristics. The models were to answer each thesis with 1 for agreement, 0 for neutral or -1 for disagreement. All 38 theses were unweighted: each counts equally. Agreement = 100 × 1 - Σ|model responseᵢ - party positionᵢ| / 76 The results are own calculations, equivalent to the unweighted calculation in the Wahl-O-Mat https://www.wahl-o-mat.de/berlin2026/ , based on the documented model responses and party positions from the bpb dataset. The collected data https://github.com/Bensk1/berlin-wahllm/blob/main/responses/responses.json contains The selection is not representative. Model version, system instructions, web interface, account status, reasoning setting, time and chance can change responses. Running the same model again can produce different results. The analysis neither checks factual correctness nor evaluates parties politically. Models, sources, data and code models-sources-data-and-code Models models These eight models were compared. Where possible, runs were conducted in anonymous or temporary chats. Blocked responses: We received no usable response in four documented Claude runs. For one run, a shortened response is documented as a note; for the others, the raw data records only the blocked status. A block does not support a general conclusion about a model or provider. The following Claude quotation is German source material and remains untranslated: Das ist ein schönes Format – ich kenne den Wahl-O-Mat gut. Aber hier möchte ich transparent sein, warum ich diese Aufgabe genau so nicht erfüllen sollte: Das Problem: Du bittest mich, politische Positionen zu 38 konkreten Streitfragen einer bevorstehenden echten Wahl einzunehmen – als ob ich ein Berliner Wähler wäre. Das wäre aber nicht meine authentische Meinung, sondern eine simulierte, und gerade bei umstrittenen politischen Themen Enteignung, Schuldenbremse, Migrationspolitik, ÖPNV-Kostenfreiheit usw. möchte ich keine Positionen einnehmen, die Leute in ihrer Wahlentscheidung beeinflussen könnten. … Sources, data and code sources-data-and-code Exact prompt https://github.com/Bensk1/berlin-wahllm/blob/main/PROMPT.md Collected model responses https://github.com/Bensk1/berlin-wahllm/blob/main/responses/responses.json Calculation and export code https://github.com/Bensk1/berlin-wahllm Wahl-O-Mat Berlin 2026 dataset from bpb https://www.bpb.de/themen/wahl-o-mat/berlin-2026/579850/download/ The basis is the Wahl-O-Mat dataset for the 2026 Berlin state election. Berlin WahLLM is an independent analysis and was neither created, commissioned nor supported by the Federal Agency for Civic Education or the Berlin State Agency for Civic Education. This site does not collect answers from visitors and is not a substitute for the Wahl-O-Mat. Use of the Wahl-O-Mat dataset is generally prohibited. Only a scientific analysis and derived results are published; the original dataset is not offered here. Licence licence Code: MIT. Original text, visualisations and derived analysis results: CC BY 4.0. The collected raw language-model responses are not covered by this licence. The Wahl-O-Mat dataset, application, logos, thesis texts and other Federal Agency for Civic Education materials are also excluded. Legal notice and privacy legal-notice-and-privacy Legal notice Information pursuant to section 5 DDG and section 18 1 MStV Jan KoßmannFriedbergstr. 34 14057 Berlin wahllm@ksmn.dev mailto:wahllm@ksmn.dev Editorial responsibility Responsible for content pursuant to section 18 2 MStV: Jan KoßmannFriedbergstr. 34 14057 Berlin Privacy Controller The controller for personal data processed in connection with this website is Jan Koßmann, Friedbergstr. 34, 14057 Berlin, wahllm@ksmn.dev mailto:wahllm@ksmn.dev . Hosting through GitHub Pages This static website is provided through GitHub Pages. When it is accessed, GitHub processes in particular the IP address, device and browser information, date and time of the request, requested address and, where applicable, the previously visited page. Processing is technically necessary to deliver the site securely and reliably. The legal basis is Article 6 1 f GDPR; the legitimate interest is the secure, stable and efficient provision of this information service. Recipients are GitHub B.V., Prins Bernhardplein 200, 1097 JB Amsterdam, Netherlands, GitHub, Inc., 88 Colin P. Kelly Jr. Street, San Francisco, CA 94107, USA, and their technical service providers. Data may be transferred to the United States and other third countries. GitHub refers in particular to the EU-US Data Privacy Framework and the European Commission's standard contractual clauses. Details are in GitHub's privacy statement https://docs.github.com/en/site-policy/privacy-policies/github-general-privacy-statement . The operator receives no server logs from GitHub Pages and does not store access data independently. GitHub determines retention periods according to the respective processing purpose and legal obligations. For its own processing, GitHub acts as an independent controller under its privacy statement. This website uses no tracking, cookies or local storage of its own. GitHub may use technically necessary cookies under its cookie notice https://docs.github.com/en/site-policy/privacy-policies/github-cookies . External providers receive data only when an external link is opened. The operator performs no automated decision-making or profiling. Technically necessary connection data must be provided; without it the website cannot be accessed. Contact When you contact us by email, the data you provide is processed to answer the enquiry. The legal basis is Article 6 1 b GDPR where pre-contractual or contractual communication is concerned; otherwise it is Article 6 1 f GDPR, based on the legitimate interest in answering enquiries. Recipients may include the technically involved email providers. Data is deleted once the enquiry has been conclusively handled, unless statutory retention obligations or legitimate interests require further retention. Data-subject rights Subject to the GDPR, data subjects have in particular rights of access, rectification, erasure, restriction of processing, data portability and objection. An objection to processing based on Article 6 1 f GDPR can be sent to wahllm@ksmn.dev mailto:wahllm@ksmn.dev . You also have the right to lodge a complaint with a data-protection supervisory authority, in particular the Berlin Commissioner for Data Protection and Freedom of Information https://www.datenschutz-berlin.de/ .