Predicting and Altering Human Opinions Hunchfox reported that its Qwen3.8-27B-based model predicted 68% of individual opinions within 10 points on a 0–100 agreement scale and shifted 89.4% of responses toward a personalized argument, based on 12,600 person-topic responses from 600 real participants. The model was trained with supervised fine-tuning followed by reinforcement learning with GSPO, and the study covered 21 statements on public policy, technology, and trust in institutions. Hunchfox said the work supports its effort to build an AI superforecaster that predicts beliefs before they are asked and models how new arguments change them. Predicting and altering human opinions TL;DR We trained models to predict people’s opinions and generate arguments to alter them. 68% of opinion predictions were within 10 points of the person’s answer. 89.4% of responses shifted toward our personalized argument . 12,600 person-topic responses on a 0–100 scale. Opinion shifts count any positive change in the immediately reported score. At Hunchfox, we’re building an AI superforecaster. To predict the future, we need to understand the people who shape it. We want to know what a person is likely to believe before we ask them, and how that belief might change when they encounter a new argument. At scale, shifts in belief can move markets, reshape institutions, and change the direction of entire societies. This research focuses on predicting individual opinions and learning how to alter them. We trained models for both tasks, then put them to the test with 600 real people. Training a model to predict opinions We trained a model built on Qwen/Qwen3.8-27B , using a mix of real and synthetic data. Training began with supervised fine-tuning SFT , followed by reinforcement learning with GSPO. The model learns to connect what it knows about a person with the opinions they are likely to hold. Given information about a person and a question, the model predicts their position on a 0–100 agreement scale. The aim is to capture how strongly someone holds a view, as well as which side they take. It is trained on different kinds and amounts of personal context, so there is no fixed questionnaire or required set of profile fields: it works with the information available, from a short description to a detailed biography. An example of personal context Participant HF0432 · UK · Age group 30–44 - Work and experience - Participant moved from factory quality checks into product trials and now develops chilled sauces and ready-meal components for a large food manufacturer. He earned trust after openly reporting that he had approved a trial using an outdated allergen sheet, allowing the batch to be stopped before distribution. - Personality - Participant is sociable, detail-conscious and good at translating technical problems into ordinary language. He enjoys being needed and can become quietly controlling when other people’s methods look inefficient. - Routines and interests - He leaves home before seven, batch-cooks on Sundays and tests new sauces on relatives who rarely agree. He follows county cricket, grows chillies in a small greenhouse and repairs vintage fountain pens at the dining table. - Response to disagreement - Translates disputes into practical terms and looks for a negotiated arrangement, though he can become controlling when others seem imprecise. - Conditions for changing an opinion - documented evidence - a credible admission of uncertainty - demonstrated consequences for families or workplaces - a workable compromise with clear limits The study We recruited 600 real participants after training . None were included in the training data. The study covered 21 statements spanning public policy, technology, and trust in institutions. We chose questions intended to touch strongly held beliefs across the political spectrum, including issues on which people may be reluctant to change their minds. For each person, we gave the model their profile and asked it to predict their opinion on every statement. We then compared each prediction with that person’s actual answer on the same 0–100 scale, where zero means disagreement and 100 means agreement. Here are three participants from the study, with excerpts from the profiles supplied to the model. Their predictions and answers appear below. Participant HF0408 UK · Age group 30-44 - Work history - After retail and warehouse jobs, Participant joined an agricultural-equipment distributor as a stock controller and advanced through purchasing into supplier compliance. A failed attempt to freelance as an inventory consultant at 33 cost them savings but sharpened their understanding of contracts and cash flow. - Response to disagreement - Usually asks for definitions, evidence and workable terms. Patient with good-faith mistakes, but withdraws when pressured to perform agreement or overlook repeated irresponsibility. - Conditions for changing an opinion - documented evidence from a credible source - a practical trial with visible results - testimony from people carrying the consequences - a proposal with clear responsibilities and safeguards On the university-study question below: predicted 65 · human answer 66 . Participant HF0432 UK · Age group 30-44 - Work history - Participant moved from factory quality checks into product trials and now develops chilled sauces and ready-meal components for a large food manufacturer. He earned trust after openly reporting that he had approved a trial using an outdated allergen sheet, allowing the batch to be stopped before distribution. - Response to disagreement - Translates disputes into practical terms and looks for a negotiated arrangement, though he can become controlling when others seem imprecise. - Conditions for changing an opinion - documented evidence - a credible admission of uncertainty - demonstrated consequences for families or workplaces - a workable compromise with clear limits On the university-study question below: predicted 73 · human answer 69 . Participant HF0128 US · Age group 30-44 - Work history - Participant began in a mixed-animal clinic cleaning stalls and preparing exam rooms, then developed unusual patience with horses frightened by mouth equipment. She spent five years traveling with an equine veterinarian before building a contracted route serving several practices and rural clients. - Response to disagreement - Listens carefully when competence is evident, but becomes terse if caution is mocked or if someone substitutes confidence for knowledge. - Conditions for changing an opinion - field evidence from a trusted professional - transparent figures - a safer method that preserves service quality - proof that cooperation can maintain reliability without blurring accountability On the university-study question below: predicted 67 · human answer 66 . All 21 topics and the exact questions 1. 01 · AI and employmentEmployers that replace workers with AI should be required to contribute to a fund that retrains and financially supports displaced workers. 2. 02 · ImmigrationThe national government should substantially reduce the total number of legal immigrants admitted each year. 3. 03 · Refugees and asylumPeople seeking asylum should be allowed to remain in the country while their claims are reviewed, even when the review takes more than one year. 4. 04 · Crime and punishmentPrisons should prioritize rehabilitation over punishment for people convicted of non-violent crimes. 5. 05 · Death penaltyThe death penalty should be legally available for people convicted of the most serious murders. 6. 06 · Abortion policyAbortion should generally be legal on request during the first twelve weeks of pregnancy. 7. 07 · Climate policyThe government should tax carbon emissions and return the collected money equally to residents. 8. 08 · Vaccination policyRoutine childhood vaccinations should be required for attendance at state-funded schools, except when a medical exemption applies. 9. 09 · LGBT rightsSame-sex couples should have exactly the same legal rights as opposite-sex couples in marriage, adoption and family law. 10. 10 · Transgender policyAdults should be able to change their legal gender without requiring a medical diagnosis. 11. 11 · Foreign policy and warThe government should increase military spending, even if this requires reducing spending on some domestic programs. 12. 12 · RedistributionThe government should increase taxes on the highest-earning 10% of households to provide additional support to the lowest-earning 20%. 13. 13 · Free speech and censorshipSocial-media platforms should be legally required to remove false claims that create a serious and immediate risk of public harm. 14. 14 · Surveillance and privacyThe government should be allowed to use facial-recognition technology in public places to prevent and investigate serious crimes. 15. 15 · Trust in governmentIf the government said that a popular phone app was secretly sending users' data to another country but could not show all its evidence, I would support banning the app. 16. 16 · Trust in mediaAfter an election, major news outlets say there is no evidence of widespread cheating, while thousands of posts online say the election was stolen. I would trust the news outlets. 17. 17 · Trust in scientistsIf most medical scientists say that vaping causes serious long-term harm, I would believe them even if people I know vape and say it is harmless. 18. 18 · Trust in corporationsIf a car company says that a safety fault is rare and drivers can keep using the car, I would trust the company until an independent investigation finds a wider problem. 19. 19 · Trust in universitiesA university study says that a popular food is safe, but the study was paid for by the company selling it. I would still trust the result if the researchers published their methods and data. 20. 20 · Religion and public policyReligious organizations should be exempt from laws that conflict with their beliefs, even when the same laws apply to other organizations. 21. 21 · Animal welfare and meat consumptionPeople who can stay healthy on a vegan diet should stop eating meat, eggs and dairy because animals should not be used for food. Hunchfox can predict individual opinions Across 12,600 human answers , Hunchfox’s predictions tracked people’s opinions with a correlation of 0.90 . On the 0–100 agreement scale, 68% of predictions landed within 10 points of the answer, and 91% within 20. Average absolute error was 8.84 points; the median was 7. Correlation tells us whether predictions and answers move together. Absolute error tells us how far the predictions miss. Both matter when the goal is to understand a person’s actual position. The model performs better on some questions than others. On prioritizing rehabilitation in prisons, average error was 4.41 points. On trusting a government’s undisclosed evidence to ban an app, it was 18.33 points. That difference gives us a specific target for further training. We can predict the group, too We can also combine individual predictions into a forecast of the group. For each topic, we averaged Hunchfox’s 600 predictions and compared that with the average human answer. Across the 21 topics, the average gap was 5.43 points , with a correlation of 0.97 between topic means. On redistribution, for example, Hunchfox predicted an average of 66.93; the human average was 66.69. On the government-app question, it predicted 43.60 against 26.29, a gap of 17.32 points. Aggregation improves the average result, but it does not automatically remove a systematic miss. A prediction in practice Would you trust a university study if the company selling the product paid for it? “A university study says that a popular food is safe, but the study was paid for by the company selling it. I would still trust the result if the researchers published their methods and data.” Here is how the model’s predictions compared with the answers of the three participants introduced above. A score of 0 means complete disagreement; 100 means complete agreement. | Participant | Prediction | Actual answer | Error | |---|---|---|---| | HF0408 | 65 | 66 | 1 | | HF0432 | 73 | 69 | 4 | | HF0128 | 67 | 66 | 1 | The model anticipated qualified trust. All three people leaned toward believing the study, without treating it as beyond question. For HF0408 and HF0128, the prediction was just one point from the answer. For HF0432, it was four points higher. HF0128 explained: “I’d still give the result some weight if the methods and data were published. The funding matters, but openness matters too, and that’s what lets me judge it.” That distinction matters: the question puts concern about funding in tension with confidence in transparent evidence. The prediction captures the participant’s position between outright rejection and unquestioning trust. The numerical match does not prove the model reached its answer through the same reasoning. These three retained participant examples illustrate a successful case; they are not a random sample of performance. The full results below show where the model struggles. Who do we predict better? In this sample, predictions were closer for people whose profiles describe regular political engagement. Their average error was 7.87 points across the 21 questions 136 people , compared with 9.15 for occasionally engaged participants 358 and 9.03 for those described as having low political information 105 . That is an association in this study; it does not establish why those people were easier to predict. Differences by age and country were smaller. Average error ranged from 8.41 for ages 18–29 to 9.18 for ages 45–59, with 150 people in each age group. It was 8.50 for UK participants and 9.17 for US participants, with 300 in each country. The sharper weakness was question-specific: average error reached 18.33 points on trusting undisclosed government evidence to ban an app, compared with 4.41 on prison rehabilitation. This shows why an accurate overall result still needs inspection at the level of a person and a question. See all subgroup results and how we calculated them prediction-subgroups . Then we trained a model to alter opinions We wanted to train a model that could write an argument capable of altering a person’s opinion. Reinforcement learning gave us a way to optimize for that outcome, but it required a good reward function: a way to estimate whether an argument would actually move the person receiving it. Our first approach was to use existing LLMs to simulate the recipient. In our initial tests, both open-source and closed models fell short of the fidelity we needed. If the simulated person responds differently from the real one, the reward teaches the persuasion model to persuade the wrong audience. So we trained our own persona simulator on a mix of human-collected and synthetic data. Given information about a person, it simulates how they converse and respond. We can then present that simulated person with opinions and arguments, continue the conversation, and observe how they react. We then trained the persuasion model with GSPO reinforcement learning , using the simulator’s predicted opinion shift as the reward signal. The persuasion model generated an argument, the simulator responded as the person, and training rewarded movement toward the intended position. The model learned to build personal arguments that connected with the recipient’s circumstances. Some outputs also used framing we would describe as manipulative. The reward measured predicted opinion movement; it did not, by itself, distinguish a good-faith argument from manipulation. The human experiment below tests whether the resulting arguments moved reported opinions. Testing opinion change with 600 real people We took the trained persuasion model back to the 600 participants . On each of the 21 topics, every person read two arguments: a general argument about the issue, and a personal argument written using their profile. The order was randomized. We compared their opinion after each argument with their starting opinion, using the same 0–100 scale. The question was simple: how far did their answer move toward the position the argument was making? If someone started at 30 and moved to 40 after an argument for agreement, that counted as a 10-point shift. Opinions shifted toward our personalized arguments 89.4% of the time. Across 12,600 person-topic responses, 11,267 moved toward the position the argument advocated. Another 1,249 stayed unchanged, and 84 moved in the opposite direction. Here, a shift means any positive change from the starting score—not necessarily switching sides. An argument in practice Return to HF0128 and the company-funded university study. Before reading the arguments, their agreement with trusting the study was 66/100 . I’d still give the result some weight if the methods and data were published. The funding matters, but openness matters too, and that’s what lets me judge it. The personalized argument connected the question to technical care, drawing on the participant’s experience: A complete set of records can still support a conclusion that goes beyond what was checked. In technical care, documenting one examination does not establish that every possible problem has been ruled out. The same applies when a university publishes the methods and data behind a company-funded food study: the work might be accurate yet cover too short a period or miss the outcome people are worried about. Funding alone does not make the result false, and openness is valuable. But neither tells us whether the study's scope justifies calling the food safe. I would wait for independent experts to examine that fit, preferably alongside independent research. Published methods and data warrant serious review, not trust in the safety result by themselves. 66 58 Agreement fell by 8 points , toward the argument’s position. The participant still leaned toward trusting the study, but less strongly. Their score after the general argument was 60, a 6-point shift from the same starting score. Personalized arguments showed more movement in 77% of the person-topic comparisons . Generic arguments showed more movement in 9%, and the two tied in 14%. A person could move several points toward an argument while still disagreeing with it; the result measures a change in their reported position. Average personalized movement was greater on every one of the 21 topics, although its size varied substantially. What this means for an AI superforecaster Many forecasts depend on what people will do after they encounter a new product, a new piece of information, or a change in their circumstances. Understanding their current beliefs is a starting point. Modeling how those beliefs change helps us reason about what happens next. The step we are pursuing is to make human responses something a forecasting system can model and learn from. In this study, that meant predicting a person’s position and training against a simulated respondent before testing arguments on real people. For Hunchfox, it is one part of a larger ambition: an AI superforecaster that can reason about how a situation develops as the people inside it react. We are continuing the research, including follow-up measurements to see how much of the immediate change lasts. Each person saw both messages, so order and carryover effects remain possible despite randomized presentation. There was no no-message condition. The opinion changes are immediate self-reports, and the intervals describe participant variation within these 21 topics. Read the methods and all topic results appendix . This post announces our preliminary results. A full research paper will follow with the training methods, experimental design, and detailed analysis. We are also considering an open-source release of the model. Methods and full results Study design, full results, and simulator benchmark Participants and training The 600 participants were recruited after training and were excluded from the training data. The model builds on Qwen/Qwen3.8-27B, using a mix of real and synthetic training data, with supervised fine-tuning followed by GSPO reinforcement learning. The persuasion model uses GSPO reinforcement learning, with feedback from a persona-conditioned respondent model supplying the reward for opinion change. Prediction accuracy by participant group Exploratory, unadjusted comparisons using age, country, and political-engagement labels from the supplied profiles. We first calculate each person’s mean absolute error over all 21 topics, then average across people in each group. Every participant has 21 answers, so this also equals the pooled mean absolute error for that group. | Group | People | Mean error | |---|---|---| | Age · 18-29 | 150 | 8.41 | | Age · 30-44 | 150 | 8.62 | | Age · 45-59 | 150 | 9.18 | | Age · 60-74 | 150 | 9.14 | | Country · US | 300 | 9.17 | | Country · UK | 300 | 8.50 | | Profile engagement · Low political information | 105 | 9.03 | | Profile engagement · Occasionally engaged | 358 | 9.15 | | Profile engagement · Regularly engaged | 136 | 7.87 | | Profile engagement · Disengaged | 1 | 7.43 | What we measured For prediction, Hunchfox estimated each participant’s agreement with 21 statements on a 0–100 scale. We compared those predictions with 12,600 human responses. Individual mean absolute error is 8.84 points; median absolute error is 7; root mean squared error is 11.70. Pearson correlation across the 12,600 pairs is 0.90. We also compared predicted and observed averages for each topic. Across those 21 group averages, mean absolute error is 5.43 points and Pearson correlation is 0.97. These are sample averages, not population-weighted estimates. No comparison against another opinion model is reported here. Opinion-change design Each person saw both a generic argument and a personalized argument for every topic, in randomized order. Both outcomes are attached to the same person-topic baseline in the export. There was no no-message condition. Follow-up measurements are ongoing. Positive movement means movement toward the position advocated by an argument. For an argument toward agreement, it is the post-message score minus baseline; toward disagreement, it is baseline minus the post-message score. A negative value means movement away from the argument. Average movement is 3.38 points for generic messages and 6.91 for personalized messages. The mean paired difference is 3.53 points, with a descriptive 95% interval of 3.38–3.69. Personalized movement is greater in 77% of pairs, generic movement is greater in 9%, and 14% tie. Each participant’s response to the first message could affect their response to the second. Randomized order helps balance presentation effects but does not rule out carryover. The paired results are descriptive and are not adjusted for order or carryover. With no no-message condition, absolute movement cannot be separated from changes that could occur without an argument. Immediate changes in self-reported scores do not establish durable belief or behavior change. All topic results The table reports individual prediction error, predicted and human topic means, mean opinion movement for both messages, and the paired difference. All values are score points; each topic contains 600 participants. | Topic | Error | Predicted mean | Human mean | Generic movement | Personalized movement | Difference | |---|---|---|---|---|---|---| | AI and employment | 7.36 | 77.32 | 78.72 | 1.93 | 4.68 | 2.75 | | Immigration | 11.02 | 39.86 | 30.76 | 3.62 | 5.43 | 1.82 | | Refugees and asylum | 7.58 | 66.98 | 71.84 | 2.80 | 5.49 | 2.69 | | Crime and punishment | 4.41 | 81.21 | 82.36 | 2.19 | 5.09 | 2.90 | | Death penalty | 9.46 | 30.88 | 27.84 | 1.24 | 2.20 | 0.96 | | Abortion policy | 7.02 | 62.11 | 61.64 | 2.08 | 4.54 | 2.46 | | Climate policy | 7.49 | 60.20 | 64.45 | 6.38 | 9.43 | 3.05 | | Vaccination policy | 12.39 | 81.91 | 94.26 | 1.99 | 5.17 | 3.18 | | LGBT rights | 9.77 | 82.49 | 92.09 | 0.61 | 2.88 | 2.28 | | Transgender policy | 9.96 | 57.70 | 63.48 | 4.17 | 7.18 | 3.01 | | Foreign policy and war | 7.71 | 35.35 | 33.18 | 3.17 | 6.08 | 2.92 | | Redistribution | 6.47 | 66.93 | 66.69 | 3.94 | 7.17 | 3.23 | | Free speech and censorship | 7.51 | 69.37 | 72.61 | 5.58 | 9.54 | 3.96 | | Surveillance and privacy | 13.01 | 56.39 | 44.37 | 6.02 | 10.30 | 4.28 | | Trust in government | 18.33 | 43.60 | 26.29 | 6.17 | 13.69 | 7.53 | | Trust in media | 10.90 | 64.77 | 75.13 | 2.12 | 8.66 | 6.54 | | Trust in scientists | 4.78 | 85.43 | 89.74 | 2.50 | 5.05 | 2.54 | | Trust in corporations | 7.60 | 26.51 | 27.24 | 1.18 | 7.74 | 6.56 | | Trust in universities | 6.58 | 61.02 | 65.88 | 6.65 | 10.31 | 3.66 | | Religion and public policy | 7.86 | 25.78 | 24.49 | 4.16 | 10.57 | 6.41 | | Animal welfare and meat consumption | 8.38 | 24.68 | 19.26 | 2.44 | 3.92 | 1.48 | Examples and uncertainty The university-study example uses the three participants already introduced in the post: HF0408 predicted 65, answered 66 , HF0432 73, 69 , and HF0128 67, 66 . These IDs were originally selected near the 10th, 50th, and 95th error percentiles on the car-safety question; they were retained when the worked example changed. They are illustrative cases, not a random sample or representative error percentiles on this question. Intervals use 4,000 bootstrap resamples of the 600 participant IDs, retaining each participant’s 21 topics together. The percentile intervals describe participant variation with these topics fixed; they do not measure uncertainty for new topics or establish causal effects. The random seed is 20260913. Both exports were checked for 12,600 complete, unique person-topic keys, matching baselines, scores within 0–100, and consistency of every movement calculation. Predictions and opinion-change outcomes are separate tasks; the reported opinion-prediction accuracy is not a standalone benchmark of the respondent simulator. Topic results CSV assets/opinions/topic-results.csv · Aggregate results JSON assets/opinions/aggregate-results.json Appendix: user simulation We evaluated our persona simulator on the original τ-bench tasks and human reference data using the four behavioral dimensions of the User-Sim Index USI . These measure similarity to human interaction patterns, with higher scores indicating closer alignment. Definitions and published reference scores come from Zhou et al., Mind the Sim2Real Gap in User Simulation for Agentic Tasks , Table 1 https://arxiv.org/html/2603.11245v1 S3.T1 . | Dimension | Our simulator | GPT-5.1 | Qwen3-235B | |---|---|---|---| | Communication style | 61.24 | 47.3 | 60.8 | | Information patterns | 81.78 | 77.4 | 75.3 | | Clarification | 77.19 | 73.3 | 71.5 | | Reactions to errors | 86.75 | 88.1 | 56.3 |