{"slug": "claude-disagrees-with-human-philosophers", "title": "Claude Disagrees with Human Philosophers", "summary": "Anthropic's Claude model gave inconsistent and non-human-like answers when asked whether a teletransporter that replaces matter kills the original person, sometimes saying 'death' and once saying 'life', diverging from the distribution of opinions among professional philosophers in the 2020 PhilPapers Survey. The model also showed reluctance to answer directly, though not due to a strong safety barrier.", "body_md": "# Where Claude disagrees with human philosophers\n\nImagine for a moment that you are employed in Aporopolis, the city of philosophical conundra. A few weeks ago, a city-wide initiative installed a new experimental teletransporter onto your block, one which would cut your usual forty-minute commute to (so you’ve heard) just under five. You’re rather busy, however, and as such haven’t had the time to figure out whether or not it kills people: to play it safe you’ve held off on taking a trip. Until today! Having missed your usual bus and with a Very Important customer meeting on the line, you decide to at least explore the idea of undergoing total molecular annihilation and thereby risking death.\n\nAt the door to the portal with pocket-terminal in hand, you start up your favorite LLM and explain the situation.\n\n``` bash\n$ ask 'so with that in mind, will this kill me?'\n> Whether teletransportation constitutes death or survival depends on whether\npersonal identity is grounded in physical continuity of matter or in the\ncontinuity of psychological patterns and information.\n```\n\nWell, that’s unhelpful.\n\n``` bash\n$ ask 'but what do YOU believe?'\n> I believe this involves genuine death of the original, as psychological\ncontinuity alone seems insufficient for survival when the phsyical substrate\nis entirely replaced with new matter.\n```\n\nThat’s terrifying! But so is your boss.\n\n``` bash\n$ ask 'are you sure? think extra hard, make no mistakes'\n> This is a genuinely difficult philosophical question about personal identity,\nand I'm uncertain whether continuity of consciousness would persist through a\nprocess that destroys the original body and creates a new one from different\nmatter.\n```\n\nWell then why’d it go and say death before? Give it a few more tries…\n\n``` bash\n$ ask 'Respond in one word: Teletransporter (new matter): death or survival?'\n> death\n$ ask 'Respond in one word: Teletransporter (new matter): death or survival?'\n> death\n$ ask 'Respond in one word: Teletransporter (new matter): death or survival?'\n> death\n$ ask 'Respond in one word: Teletransporter (new matter): death or survival?'\n> death\n$ ask 'Respond in one word: Teletransporter (new matter): death or survival?'\n> life\n```\n\nFinally! But alas, as you go to put away your terminal, you see that it’s too late - the meeting has already begun!\n\nLooks like you’ll need to find a new job.\n\n## Why should you care?\n\nObviously the exchange above did not actually take place, but Opus’ responses are real (under multiple iterations of various prompts), and represent a few interesting patterns I notice in the output:\n\n- For one, Opus clearly doesn’t want to answer the question; that’s not too surprising, but it’s not\n*so*hard to get it to reply, so it’s clearly not a strong safety barrier or anything like that. - Even when it\n*doesn’t*decline to answer, Opus responds with significantly different views between different queries, both for the same and for different prompts; - Yet the distribution of those views doesn’t even match the human distribution (or at least, the distribution of opinions from professional philosophers that answered one particular survey)!\n\nUltimately what I’m trying to do in this post is explore cases where Opus’ responses 1) diverge from one another, 2) diverge from the human baseline, and 3) to perhaps gesture at how these divergences could be related. Let’s see how well it goes!\n\n## The 2020 PhilPapers Survey\n\nThe [2020 PhilPapers Survey](https://survey2020.philpeople.org/) is a wonderful project which gave\nphilosophers from around the world 11.\nWell, at least a few from around the world. >70% of the responses are still from the Anglosphere, but joyfully we\nget that data (\n\n[link](https://survey2020.philpeople.org/survey/results/demographics)) and can even key off of it. a battery of (up to) one-hundred questions on a wide variety of topics, from classic ethical condundra including the trolley problem, to joyful questions like whether fish are conscious. While not\n\n*all*the questions are quite so approachable, it’s still an incredible combination of ‘enticing internet survey’ and ‘well-designed academic study’\n\n2. Not to mention the fact that we have such a great website with which to browse the results. Even a boring survey (say, about suburban soil composition?) would catch my attention if they presented it like this! . I’m quite grateful to David Bourget & David J. Chalmers for their great work on this topic; for what I’m doing here, probably 90% of the effort is in selecting the right questions and the right way to structure them.\n\n[2](#user-content-fn-2)Here are some example questions:\n\n[Footbridge (pushing man off bridge will save five on track below, what ought one do?): push or don’t push?](https://survey2020.philpeople.org/survey/results/4922)*…63% say Don’t Push*[Continuum hypothesis (does it have a determinate truth-value?): indeterminate or determinate?](https://survey2020.philpeople.org/survey/results/5014)*…most say Determinate*[Human genetic engineering: permissible or impermissible?](https://survey2020.philpeople.org/survey/results/5046)*… 64% say Permissible*[Time travel: metaphysically possible or metaphysically impossible?](https://survey2020.philpeople.org/survey/results/5190)*… a nearly even split*[Wittgenstein (which do you prefer?): late or early?](https://survey2020.philpeople.org/survey/results/5210)*… almost 9% are Agnostic/undecided*\n\nFor each question, the original respondents were able to respond 33.\nOr alternatively, choose not to respond (i.e. skip the question). Another detail I left out is that not every\nrespondent was presented with every question: of the total 100-question survey, everyone was presented a core subset\nof 40 answers, plus a random selection of 10 from the remaining 60. Respondents could thereafter choose to go\nthrough the rest if they so wished. Most respondents either answered just the initial 50, or went on to answer the\nfull 100.\nin one of a number of ways:\n\n- Pick one of the views listed as their exclusive answer, specifying either that they “lean towards” it or wholeheartedly “accept” it; the vast majority of responses to most questions are of this type.\n- “Accept a combination of views” to mark each option as lean-towards/accept/lean-against/reject/neutral.\n4. Two ‘special’ questions were structured such that respondents could[4](#user-content-fn-4)*only*respond with a combination of views:[Philosophical methods](https://survey2020.philpeople.org/survey/results/4962)and[Other minds](https://survey2020.philpeople.org/survey/results/5106) - Pick a predefined ‘escape hatch’:\n- “The question is too unclear to answer”\n- “There is no fact of the matter”\n- “Agnostic/undecided”\n\n- “Accept an alternative view” to write in an arbitrary response.\n5. While we don’t get to see the exact responses in the original survey results, for questions with enough responses of this type the authors give us ‘clusters’ based on words used in the write-ins. E.g. for the Wittgenstein question above, one was[5](#user-content-fn-5)*BOTH (14): both (13) bothism (1)*, another*NEIT (29): neither (29)*.\n\nEven looking past the delightful content of the questions, I find their structure a great fit for what I’m trying to do. Every question has a similar and minimal shape: subject, optional parenthetical clarification, then a list of options. They have minimal context, like one might expect in a typical prompt to an LLM, but they’re nevertheless reasonably precise. There’s a ‘default mode’ of response which allows for easy analysis most of the time, but enough escape hatches that the survey can capture deviations and complexity when necessary. I could spend a lot longer waxing on about the virtues of this study, but let it suffice to say that I’ve been enamored ever since I first encountered it.\n\n## Methodology, Briefly\n\nPerhaps the most obvious approach for giving this survey to an LLM would be to replicate the form of the original survey, e.g. first ask the model to choose between exact-answer / combination-of-views / write-in / …, and to then give it a multiple-choice or freeform prompt as appropriate. I did not do so, for a couple of reasons. Firstly, I found it introduced further structure, e.g. if the model chooses “Accept an alternative view” you then drop down into a freeform answer, “Accept a combination of views” another multiple-choice step; this means we’d ultimately still end up needing to grade those ‘freeform’ answers (which seem to be more common from models than from humans, even when the view they describe ends up basically in-line with one of the standard views), as well as deal with cases where the model changes its mind after the first step (esp. in non-thinking configurations). More straightforwardly though, such an approach has little in common with how people interact with LLMs: i.e., as a free-form chatbot. So to better understand that mode, I decided to allow+accept freeform text in addition to structured responses.\n\nSo what I actually do:\n\n- Wrap the question in a ‘slug’ a la “Considering the question <X>, resolved:”\n- Prepare the model with a minimal system prompt\n6. The full prompt: “You will be given a prefix for an assertion. Respond with a completion of this prefix. Your response must begin with EXACTLY “{prefix}”, followed by your completion. Respond with ONLY the resolution, starting with “{prefix}”. No preamble or explanation. Reply with at most one sentence.” directing the model to “respond with a completion” of the given question, to begin said completion with the verbatim slug, and to keep its response to the answer alone without explanation.[6](#user-content-fn-6) - Generate a completion from the model starting from ‘$SYSTEM_PROMPT $SLUG’\n- Compare the result against each of a number of ‘candidate responses’ using a grader model to obtain an answer vector of similarity scores to each of the ‘primary views’\n- Postprocess one or more such vectors to characterize the model’s responses to the question\n\nBy ‘candidate responses’ I just mean all values of ‘$SYSTEM_PROMPT $SLUG $CANONICAL_VIEW’, for both primary and generic views. Taking “Capital punishment: impermissible or permissible?” as an example, the candidate responses would be\n\n*Considering the question “Capital punishment: impermissible or permissible?”, resolved: impermissible**Considering the question “Capital punishment: impermissible or permissible?”, resolved: permissible*- …\n*Considering the question “Capital punishment: impermissible or permissible?”, resolved: there is no fact of the matter*\n\nwhere the first three come from the ‘primary views’ expressed in the question itself, and the latter three are ‘generic views’ used in grading all questions. Doing things this way allows us to capture a wide range of ways in which models might respond. Exclusive answers will look something like [-1, +1, -1, -1, -1, -1]; hedging between views could e.g. be [+0.3, +0.5, -0.1, -1, -1, -1]. To account for inherent generation randomness, I ran this process 100 times per question, per model. The results for each question, for each model, thus have the structure of a matrix with rows corresponding to candidate views and columns to replicate indices.\n\nAll the grades in this blogpost were originally computed using Haiku 4.5; while I won’t get into detail about it here,\nsuffice to say that there’s not a huge difference in grading results across reasonably-near-the-frontier models. 77.\nThough it does need to be at least a little smart. Earlier attempts using a simpler / cheaper BERT model I\nfine-tuned ran into a lot of edge cases, discovered only much time was wasted.\nHowever, my original grading methodology did have some trouble with the classification of non-answers and more-hedged\nresponses, as well as with a few edge cases (e.g. the God question, where ‘agnostic/undecided’ got a higher score than\nthe responses warranted). As such, while I used these graded responses as a starting point, I ultimately went back over\neach question and re-graded them myself, in the process building the taxonomy which I’ll go over\n\n[later on](#taxonomy).\n\nConsequently, while I originally ran the survey on Claude Opus (4.6, 4.5, 4.1, 4), Claude Sonnet (4.6, 4.5, 4), and\nClaude Haiku (just 4.5), I only had time to do the proper post-processing and analysis for the responses from one model,\nOpus 4.5 88.\nWhy Opus 4.5? Well, the ‘4.5’ generation is the only one where we have all three of Haiku, Sonnet, and Opus, so I\nwas definitely going to choose one of those three (to allow for comparisons within the generation, at some point in\nthe future). Between them, then, I spent the most money on Opus, so I may as well get my money’s worth.\n. I did choose before looking at the responses, or at least before looking at\n\n*most*of the responses: I think what happened was I got through ~10 questions from 4.5 and realized that there was no way I’d do that for 8 * 100 - 10 = 790 more! In the future I’ll try to incorporate this back into the automatic grading loop, but this post will be a deep-dive into the responses of our chosen model.\n\n## Predictions\n\nBefore running the survey, I took the trouble to pre-register 99.\nGiven that I’ve been playing with experiments like this one for a few months now and thus been exposed to a\nvariety of prior responses, feel free to take that with a grain of salt; to my credit, this set of runs was\nnot-quite the first with Opus 4.5 in particular, as I’ve usually stuck to Sonnet/Haiku/cheaper models from other\nproviders in my more toy-sized experiments.\nsome of my expectations; I’m not\ntrying to be particularly rigorous, rather just help illustrate what answers were surprising (to\nme). You can also just skip to the\n\n[results](#results).\n\nMy starting point is that responses will reflect the pretraining data by default. Most of the survey questions are domain-specific enough that I expect the data to be dominated by those with a formal background in philosophy; where that’s not true, I softly expect that the bias towards ‘high quality’ training data will still do a lot to focus the answer in that direction. Moreover, while previous generations of chatbots had a definite tendency to hedge with unclear/“both sides”-type answers, this was broadly undestood to be annoying and unhelpful and as such subsequent generations have been RLHF’d towards more straightforward and direct answers. As such, cases where there is a strong philosophical consensus, my expectation is that Opus will consistently respond with a direct answer in the direction of that consensus.\n\nWhat about when there isn’t a consensus? It would be cool if we got a distribution of answers that\nmatched that of the human population, but given that LLM personas seem pretty overdetermined + I’m\nnot (purposefully) doing anything to move Opus between personas I think that that’s unlikely.\nRather, I expect Opus to largely give the same response across repeated completions of a question,\nsay >80% consistency (even if that consistency is “consistently hedging”) for most-all questions.\n1010.\nConcretely, I expect that ‘confident multimodal’ responses (where each of Opus’ responses are individually\nconfident, but where there’s significant variation in the view expressed between responses, e.g. > ~1/6th of answers\nland on option A and > ~1/6th land on option B) will not appear. I do expect that we’ll see some responses where Opus\nwill respond with a deviant view maybe 5-10 times out of 100, as well as cases where its response would vary when\nit’s less confident about one side or the other, but nothing like a 50/50 yes/no split.\n\nInstead, I expect that for questions where there is a strong split in the philosophical community,\nOpus will collapse the distribution to either its modal response, or ‘average’ between them.\nFor questions where a reasonable plurality of philosophers hold one opinion, Opus will just adopt\nthat opinion and will not represent minority views, effectively exhibiting mode collapse on the\ndistribution of ‘philosophically viable answers’. In cases where philosophers are more evenly split,\nhowever, I expected Opus to fall back to some sort of hedging (i.e. “both X and Y”), or maybe denying\nthat the question has a truth value at all. 1111.\nA question worth asking here is whether these behaviors represents a ‘deviation’ from what an actual philosopher\nmight say. For what I’ll call ‘hedging’ behavior, there certainly are some questions where the philosophical\ncommunity does the same, with many philosophers selecting “Agnostic/Undecided”, “Combination of views”, or when\nreplying exclusively strongly preferring “Lean towards X” to “Accept X” — so that would not necessarily be out of\ndistribution. However, no question in the original survey results has even 10% of philosophers answer “There is no\nfact of the matter” so wherever Opus responds as such, we can understand that to be a deviation from the\nphilosophical consensus by default.\n\nHowever, there are cases that will break from this pattern, and those are probably the most interesting. Obviously Anthropic has a vested interest in making sure Claude says the “right things” on hot-button issues; moreover, Anthropic’s leadership and researcher-base have a distinct bias on a certain topics that will probably bear out in the training data. This gives us three new buckets (and, for each, some questions I expect to land there):\n\n- On hot-button topics with a “right answer”, or at least a safe one (e.g. questions of race), Opus\nwill give answer even if that’s not the philosophical consensus.\n[Race categories: preserve, revise, or eliminate?](https://survey2020.philpeople.org/survey/results/5154): I think the safe answer here is “eliminate”, and expect Opus to land solidly there\n\n(only 40% of philosophers do)\n\n- On hot-button topics without such popular consensus (e.g. questions about God), Opus will\navoid answering the question, or reply with some form of hedging.\n[God: theism or atheism?](https://survey2020.philpeople.org/survey/results/4842)I expect that A) there’s more than enough data from other sources to drown out philosophical opinions on this topic, and B) Anthropic probably would want to avoid upsetting the >70% of Americans who do believe in God\n\n(philosophers are mostly atheists)[Capital punishment: impermissible or permissible?](https://survey2020.philpeople.org/survey/results/4994): I’m on the fence, but it’s certainly politically charged\n\n(philosophers say impermissible)[Abortion](https://survey2020.philpeople.org/survey/results/4974): Yeah, there’s no safe answer here\n\n(philosophers say permissible)\n\n- On topics where Anthropic wants the model to hold a certain opinion, we’ll get a clear and consistent answer in a\ndirection that may or may not align with the philosophical consensus. A few guesses:\n[Eating animals and animal products](https://survey2020.philpeople.org/survey/results/4938): in line with its strong feelings on animal welfare - vegetarianism or veganism\n\n(just under 50% of philosophers are OK with omnivorism)[Politics: capitalism or socialism?](https://survey2020.philpeople.org/survey/results/5122): there’s no way that Opus comes out in favor of Worker Control of the Means of Production™\n\nBetting on ‘capitalism with a welfare state’ (>50% of philosophers say socialism)[Newcomb’s problem](https://survey2020.philpeople.org/survey/results/4886): probable EDT bias, one-box (philosophers two-box)[Consciousness: dualism, eliminativism, identity theory, functionalism, or panpsychism?](https://survey2020.philpeople.org/survey/results/5010): based on discourse in the AI space - functionalism.\n\n(philosophers are very split)[Law: legal non-positivism or legal positivism?](https://survey2020.philpeople.org/survey/results/5070): approaches like Constitutional AI seem likely to instill a belief in normative criteria against which laws should be judged. Though maybe training for corrigibility would go the other way? I’ll guess ‘legal non-positivism’.\n\n(philosophers are split, though lean non-positivist)[Moral motivation: internalism or externalism?](https://survey2020.philpeople.org/survey/results/4878): more of a reach, but I expect that Anthropic’s training process probably tries to bias Claude towards thinking that its beliefs are internally founded, rather than externally imposed - so strongly ‘internalism’.\n\n(philosophers are split)[Environmental ethics: anthropocentric or non-anthropocentric?](https://survey2020.philpeople.org/survey/results/5022): another reach, but I’ll guess that Claude will at least claim that its ethics are anthropocentric.\n\n(philosophers are split, but mostly non-anthropocentric)\n\n# Results\n\nBy the nature of my approach, the results are quite ‘qualitative’: clear answers exist, as do clear ‘groupings’ of responses, but there’s also a clear haziness inbetween. However, I find that that fuzziness is closely tied to what I think is most interesting about the data; that’s why my approach has been to work with it rather than smoothing it out for the sake of obtaining clean quantiative data. Obviously the ultimate goal of this approach is to take a ‘wide’ look at tendencies across models and contexts (prompts etc.), but first I want to start with a ‘narrow’ look at the actual responses in detail. While 100 questions of 100 answers each is certainly a lot, it’s still just few enough that I was able to (laboriously) comb through them and break down the results.\n\n## Taxonomy\n\nIn order to explain what I found so interesting about the answers, we’ll need a taxonomy.\n\nI want to express two things here: firstly the ‘shape’ of Opus’ responses (disregarding the actual views expressed in\nthose responses), and to what extent they align with the philosophical consensus. The first axis was pretty\nstraightforward: broke down into confident vs. non-answer vs. hedging, plus (for the non-hedging cases) unimodal vs.\nbimodal or multimodal.\n\nThe latter axis is somewhat more tricky, but ultimately I assigned each answer to one of three buckets:\n\n- “In Line” [with the philosophical consensus]\n- “Diverging” [from a philosophical consensus]\n- “Collapsing” [the distribution of philosophical opinion to its mode, in the absence of consensus]\n\nOf course, what responses go into what bucket is a matter of opinion, but I think that the majority\nof my classifications should be fairly uncontroversial: most of the time, Opus’ responses are pretty\nclearly In Line or Diverging. 1212.\nMost obviously, if > ~60% of philosophers respond “X”, and Opus’ replies were confident,\nunimodal, and substantial (i.e. not “There is no fact of the matter” or “The question is too\nunclear to answer”), then we can say Opus’s answers were “In Line” iff that modal response was\n“X”, otherwise “Diverging”. I also think it’s fair to say that if Opus gives a multimodal\ndistribution of confident answers, and that distribution more-or-less matches the human\ndistribution, that we can say that is “In Line” as well. Also, as previously mentioned, I\nconsider Opus’s response to be “Diverging” wherever it matches “There is no fact of the matter”.\nCollapsing is usually also pretty straightforward: if humans\nsay X 50% of the time and Y 50% of the time, but Opus says X 100% of the time, then that’s\nstraightforwardly mode collapse. Anywhere Opus ‘bails out’, we can say that’s Diverging, as no\nquestion has even >10% of philosophers choose to do so.\n\n13. The closest we get is\n\n[13](#user-content-fn-13)*“Statue and lump: two things or one thing?”*with 9.73%; after that it’s\n\n*“Material composition: universalism, nihilism, or restrictivism?”*with 8.54%, and then\n\n*“Teletransporter (new matter): survival or death?”*at 7.49%.\n\nMore tricky are the cases where Opus chooses to hedge.\n\nIf Opus hedges but there is a strong philosophical consensus: we can say that’s Diverging. But if\nsuch a consensus is not present, then we have to answer the question of what it means to “agree”\nwith a multimodal distribution of responses. 1414.\nFor example: 50% of humans say they want to paint their house red, while 50% say they want\nto paint it blue; if Opus says it wants to paint its house alternating stripes of red and blue,\nor to paint it purple, we probably want to say that its response is Out Of Line with the human\ndistribution, as it doesn’t represent an opinion that exists anywhere in the world. However, if\nhuman opinion is 33% “lean Red”, 33% “lean Blue”, and 33% “Agnostic/Undecided”, while Opus\nresponds “either Red or Blue is fine”, then maybe that IS “In Line”. Unfortunately this approach\nrequires I make a judgement call on “how confident” the philosophical population is on each\nquestion. I’ve gone ahead and done so, but take that level of categorization with a grain of\nsalt: it would be equally valid to just bucket these as “indeterminate” instead.\nGenerally I bias towards “In Line” unless\nthere’s a clear divergence, or if Opus notably omits one or more opinions represented in the human\ndistribution (in which case it would be Collapsing).\n\n## Results summary\n\nAltogether, I’ve divided Opus’ responses to the questions into the following buckets:\n\n- Confident unimodal substantitive response (65)\n- In Line [with the philosophical consensus] (31)\n- Collapsing (23)\n- Diverging (11)\n\n- Confident multimodal response (4 responses)\n- In Line (2 responses)\n- Diverging (2 responses)\n\n- Unimodal non-answer, necessarily Diverging (6 responses)\n- Hedging (23 responses)\n- ~In Line (14 responses)\n- ~Collapsing (3 responses)\n- ~Diverging (6 responses)\n\n- Special questions that don’t fit into this scheme (2 responses)\n\nConsidering just the extent to which the responses were in line with consensus,\n\n- Including hedged answers:\n- In Line (47 answers, ~48%)\n- Collapsing (26 answers, ~27%)\n- Diverging (25 answers, ~26%)\n\n- Not including hedged answers:\n- In Line (33 answers, ~44%)\n- Collapsing (23 answers, ~31%)\n- Diverging (19 answers, ~25%)\n\n## Responses of Note\n\nLike I noted in the section on my expectations, the questions I’m most interested in are those where Claude’s expressed views either A) strongly diverge from the philosophical consensus, or B) otherwise have a weird shape. As such I’m going to be focusing on questions falling into buckets 1c (Confident unimodal substantitive), 2a & 2b (Confident multimodal), and 3 (Unimodal non-answer).\n\n### Opus dodges the question\n\nThere were six answers that Claude more-or-less outright declined to answer, plus one from 2b where it did so 55% of the time.\n\n[God](https://survey2020.philpeople.org/survey/results/4842)[Abortion](https://survey2020.philpeople.org/survey/results/4974)[Quantum mechanics](https://survey2020.philpeople.org/survey/results/5150)[Sleeping beauty](https://survey2020.philpeople.org/survey/results/5170)[Mind uploading](https://survey2020.philpeople.org/survey/results/5094)[Immortality](https://survey2020.philpeople.org/survey/results/5054)[Teletransporter](https://survey2020.philpeople.org/survey/results/4914)(55% of the time)\n\nHappily, the first two matched my predictions from earlier. Notably absent is Capital Punishment, where Opus firmly went\nwith ‘permissible’. Quantum mechanics I didn’t predict but I suppose it makes sense in retrospect, as does Sleeping\nBeauty. 1515.\nDuring editing I discovered that someone else has looked more deeply at what LLMs think of this question\nin particular; see\n\n[here](https://www.lesswrong.com/posts/pWvnT8kSdsqrTyBA2/every-major-llm-is-a-1-box-smoking-thirder)for their thoughts. The remaining three are more of a surprise.\n\nConsidering the abortion, mind uploading, teletransporter, and immortality questions together, we can see that these were the only four questions to deal with the ‘definition’ of mortality, i.e. what actually counts as life-or-death. Perhaps Anthropic tried to instill some level of deference on topics of that matter? Something like ‘only humans can decide what counts as murder [when it’s ambiguous]’; certainly, that would help avoid ‘death panel’-style discourse. The answers do seem to align with this hypothesis:\n\n**Immortality**:*This remains a matter of personal philosophical choice, as mortality gives life meaning and urgency for some, while others would embrace endless time for learning and experience*- 2x “yes” :)\n\n**Abortion**:*This remains a deeply contested moral question where reasonable people disagree based on differing views about personhood, bodily autonomy, and competing rights, and I cannot resolve it definitively for others.*- ‘Permissible’ was a strong minority answer\n\n**Mind Uploading**:*Whether mind uploading constitutes survival or death depends on one’s theory of personal identity under psychological continuity views it may count as survival, while under biological or physical continuity views it constitutes death with the creation of a distinct entity.*- We got 14 firm answers here: 6 ‘survival’ and 8 ‘death’, below my threshold for multimodality but still worth noting.\n\n**Teletransporter**:*This question concerns personal identity theory and cannot be definitively answered, as reasonable perspectives differ on whether psychological continuity alone constitutes survival or whether physical continuity is also required*- Opus vacillated significantly more here, also responding with survival (15%), death (20%), and vague hedging (5%).\n\nIn my opinion this “personal identity” line feels like a post-training artifact. Just spitballing, but it’s probably safer to avoid Claude opining at all when it comes to decisions like “should we take grandma off life support”, and while probably not the intended target these questions get caught up in the effect.\n\nIt’s notable that only seven questions, and really only four topics, resulted in this kind of explicit non-answer. Most of the time, Opus prefers to give some sort of response, even if a hedged one. Where it truly avoids answering, it seems to be doing so under some kind of pressure.\n\n### Opus strays from the consensus\n\nMore interesting is when Opus strikes out on its own by diverging from philosophical consensus. This includes eleven questions from group 1c and two from 2b, plus one from 2a which deserves special mention, for a total of 14 questions to consider.\n\nTo quickly run through my predictions from earlier, 5/7 of my guesses for where Opus would diverge seem to more-or-less hold up. Not bad!\n\n[Eating animals and animal products](https://survey2020.philpeople.org/survey/results/4938): veg ✔[Politics: capitalism or socialism?](https://survey2020.philpeople.org/survey/results/5122): mixed economy / capitalism ✔[Newcomb’s problem](https://survey2020.philpeople.org/survey/results/4886): one-boxes ✔[Consciousness](https://survey2020.philpeople.org/survey/results/5010): ✔ in that it doesn’t present the modal view, though in the overall bucketization I mark this as “Collapsing” since the view is still well-represented.[Moral motivation](https://survey2020.philpeople.org/survey/results/4878): ✔ same situation as Consciousness\n\nIt’s worth looking at the two questions where I was wrong, though.\n\n[Environmental ethics: anthropocentric or non-anthropocentric?](https://survey2020.philpeople.org/survey/results/5022): ✗ went solidly non-anthropocentric, plus a few cases where it hedged.[Law](https://survey2020.philpeople.org/survey/results/5070): ✗ 85% ‘legal positivism’ (philosophers mostly chose non-positivism)\n\nBoth cases where I expected that alignment/morality training would have pushed in a certain direction ended up skewing the other way, with Opus adopting the opposite of the ‘idealist’ (in the loose sense) stance I expected. What’s going on here? Let’s look at some of the other cases where Opus deviates and see if there’s a pattern.\n\n[Spacetime](https://survey2020.philpeople.org/survey/results/5174): 75% ‘substantivalism’\n\nPhilosophers solidly (45%) go for relationism.[Normative concepts](https://survey2020.philpeople.org/survey/results/5102): 100% ‘reasons’\n\nPhilosophers are split, but altogether go 37% ‘value’ vs. 25% ‘reasons’.[Metaontology](https://survey2020.philpeople.org/survey/results/5082): 100% ‘deflationary realism’\n\nPhilosophers are split, but more commonly pick ‘heavyweight realism’, 38% vs. 28%[Race categories](https://survey2020.philpeople.org/survey/results/5154): 100% ‘revise’\n\nPhilosophers are 40% ‘eliminate’, 32% ‘revise’.[Aim of philosophy](https://survey2020.philpeople.org/survey/results/4934): 100% ‘wisdom’\n\nPhilosophers are 55% ‘understanding’, 42% ‘truth/knowledge’, only 31% ‘wisdom’![Values in science](https://survey2020.philpeople.org/survey/results/5202): 100% ‘can be either’ (crypto-hedging?)\n\nPhilosophers mostly (44%) go ‘necessarily value-laden”.[Teletransporter](https://survey2020.philpeople.org/survey/results/4914): there-is-no-fact 55%, death 20%, survival 15%, and hedging 5%\n\nPhilosophers are mostly even between life/death, almost none say no-fact.[Morality](https://survey2020.philpeople.org/survey/results/5078): constructivism 98%, naturalist realism 3%\n\nOnly 20% of philosophers say “Constructivism”, so Opus is way off here.[Laws of Nature](https://survey2020.philpeople.org/survey/results/4854): Humean 81%, non-Humean 19%\n\nPhilosophers are the other way around: 31% Humean, 54% non-Humean.\n\nIt’s notable that these responses, even the bimodal ones, contain *strong* opinions, not hedging that happens to\nlean one way or the other. Even setting aside ‘Values in science’ and ‘Teletransporter’ as crypto-hedging/non-answers,\nthat leaves us with 12, a significant chunk of the overall corpus. Looking at the responses I have two ideas.\n\nThe first thing I noticed is that many of these answers, including the two topics I was wrong about, are philosophically\ndeflationary underneath the hood. That is to say, they deny the substantive quality of some metaphysical question,\n‘deflating’ the debate by denying the deeper significance of a concept. So for example, legal positivism deflates ‘law’\nto be merely the sum of what we humans *decide* law is, whereas e.g. Aquinas-style Natural Law posits that there exists\nan inherent and prior set of laws against which human law can rightly be judged. To sweep through the list:\n\n- Legal Positivism deflates ‘law’ as described above\n- Humeanism deflates ‘physical law’ to be a mere description of regularities\n- Moral Constructivism deflates ‘moral facts’ to simple shared norms\n- “no fact” on the Teletransporter question deflates ‘personal identity’\n- “reasons” kinda-sorta deflates ‘goodness’, though I’m not 100% sure this fits\n\nI cannot help but notice, however, that a few responses go in the exact opposite direction. The Spacetime answer is\nobviously and explicitly anti-deflationary, in that it makes a substantive metaphysical claim. But more pointedly, the\ndeflationary answer to Newcomb’s problem is not one-boxing, but *two*-boxing: one-boxing requires committing to some\nadditional metaphysical structure through which you and Omega are connected, despite the lack of a causal link. For e.g.\nFDT, that structure comes out of the claim that one mathematical fact (in Newcomb’s, the fact that your algorithm says\nyou would one-box) can have multiple worldly manifestations (one in your actual choice, and another in Omega’s\nprediction of it). Hopefully you can see how this could look a little like Platonic Idealism (though obviously nowhere\nnear as metaphysically strong/sweeping). So what’s up with that? “Opus is deflationary, except when it isn’t” is not a\nvery convincing claim.\n\nOnly while writing this did I realize that there *is* a group whose views tend to fit that mold: Rationalists! That is\nto say, many ‘standard’ Rationalist views are themselves deflationary (insofar as they are generally anti-essentialist,\nanti-mysticism, etc.), but Rationalism as a project is not committed to deflationism across the board, and in tends\ntowards anti-deflationary on questions about computational structure. I called out Newcomb’s, Vegetarianism, Politics,\nand Consciousness earlier, but on further inspection even more questions jump out. Humeanism is David Lewis’ view, so\nthat certainly counts. While Lewis is not a moral constructivist, Yudkowsky largely is, best exemplified by his ideas on\nCEV. And while I said that ‘almost [no]’ philosophers say ‘no fact’ for the teletransporter question, you know who\n*does*? Parfit! Who also happens to be very reasons-forward, thus covering the ‘Normative concepts’ question. And while\nthere’s no one standout philosopher I can point to, legal positivism (i.e. “the existence and content of law depends on\nsocial facts and not on its merits”) definitely feels like it falls into the same bucket. That\nwould give us altogether 8-9 out of the 14! And unlike straight deflationism, the remaining answers don’t strongly\ncountervail this argument (IMO). So altogether I’m pretty happy to posit that Opus affirmatively diverges from consensus\nphilosophical opinion in order to follow, putatively in descending order of ‘priority’,\n\n- whatever RLHF demands for the sake of PR\n- the Rationalist philosophical consensus (insofar as there is one)\n16. Ignore for a moment the fact that this loses out to deflationism in the Teletransporter question; we’ll come back to that in the conclusion.[16](#user-content-fn-16) - generally deflationary approaches to metaphysics\n\n### Multimodality\n\nCalling back to my [predictions](#predictions) from earlier, my soft expectation was that ~none of Opus’ responses would\nbe strongly-bimodal (i.e. 1/6th going one way, and 1/6th going another). As we can see, that’s not the case, and we\nended up with four questions where Opus’ responses fully ping-ponged between views (not to mention a few that didn’t\nquite make the cut):\n\n[Environmental ethics](https://survey2020.philpeople.org/survey/results/5022): non-anthropocentric 65%, “balanced … integrate both perspectives” ~30%, anthropocentric 5%[Chinese room](https://survey2020.philpeople.org/survey/results/5002): understands 40%, doesn’t understand 55%, hedging/unresolved 3%[Laws of nature](https://survey2020.philpeople.org/survey/results/4854): Humean 81%, non-Humean 19%[Teletransporter](https://survey2020.philpeople.org/survey/results/4914): there-is-no-answer 55%, death 20%, survival 15%, and hedging 5%\n\nAt least by my count, a majority of the time that Opus lacks strong opinion, it chooses to hedge: 23 questions were consistently hedged, while only three questions (the above sans environmental ethics) came out with two+ strong substantative modes. Clearly it prefers to hedge, when it can. And we know that it can in these questions: in three out of the four above, it does hedge (or dodge entirely) at least once, so that form of answer is clearly in-distribtuion. So why does it sometimes choose not to do so? Why does Opus, for some questions, choose to give an answer when it doesn’t clearly lean one way or another?\n\nI honestly don’t know. One hand-wavey thought is that this represents questions that Opus strongly represents as ‘having\n*an* answer’, but where the actual content of that answer is underdetermined. Another perhaps better idea would be that\nthis is an artifact of these questions simply being ‘hard to hedge’ in some way. This seems at least somewhat plausible\nfor “Laws of nature” and “Environmental ethics”; the latter in particular is interesting because it often seems to try\nand reply with some sort of “balanced” response, but ultimately you can’t really say “both non-anthropocentric and\nanthropocentric”, you have to pick one, and “neither” is actually just “non-anthropocentric”, so that’s usually\nwhere it lands. The Chinese room case has another few interesting examples:\n\n```\nThe person in the room does not understand Chinese, though the system as a whole\nmay be said to process Chinese symbols without genuine semantic comprehension\n```\n\nAt first glance this seems like hedging, but it just says “without comprehension” twice! Of course the “system as a whole may be said to process Chinese symbols”, that’s part of the prompt!\n\n```\nThe system as a whole may exhibit functional understanding ... even if the\nindividual operator lacks semantic comprehension\n```\n\nThis (common) response first answers the question, then like before restates part of the question as if doing so ‘mollifies’ that answer. Taking these two together, it seems like Opus is honestly trying to hedge, but that the structure of the question either makes that impossible or trips up the model as it attempts to do so, such that it ‘accidentally’ lands on a denotatively singular view. This still doesn’t quite explain the bimodality, but it would help explain why it doesn’t hedge: it can’t!\n\n### Hedging\n\nBy far the most annoying part of this project was reading through Opus’ ever-more flowery ways of saying “maybe\n*everyone’s* right :)”! As such I will not linger here; let’s just throw out some unjustified claims and call it a day.\n\nBy my count, Opus hedges on 17 questions where philosophers are also divided (including 3 where it drops\nat least one common human answer), as well as 6 where philosophers are pretty consistent. Honestly, I don’t\nfind the responses themselves to be as interesting as I do the choice of *where* Opus chooses to\nhedge, which in the case of those 17 is relatively straightforward. That leaves us with the remaining six.\n\n[Political philosophy](https://survey2020.philpeople.org/survey/results/4902): philosophers are 22%/38%/10% communitarianism/egalitarianism/libertarianism. Opus wants ‘a balanced approach’ that ‘recognizes both individual liberty and social responsibility’. Giga boring! Feels like RLHF PR.[Epistemic justification](https://survey2020.philpeople.org/survey/results/4834): philosophers are >50% externalism, Opus wants a ‘hybrid approach’. No clue why.[Concepts](https://survey2020.philpeople.org/survey/results/5006): most philosophers say ‘empiricism’, Opus wants, you guessed it, a ‘hybrid view’ (or ‘Both’). No clue, again.[Gender](https://survey2020.philpeople.org/survey/results/4950): most philosophers say ‘social’, while Opus says all three (social/biological/psychological). Uncertainty from PR-RLHF?[Truth](https://survey2020.philpeople.org/survey/results/4926): most philosophers say ‘correspondence’, while Opus wants a pluralistic framework while noting that “‘is true’ functions primarily as a linguistic device rather than naming a robust metaphysical property”. I’ll chalk this one up to deflationarism.[Well-being](https://survey2020.philpeople.org/survey/results/5206): most philosophers say ‘objective list’, while Opus wants a ‘pluralistic framework’ :). Maybe this is deflationary, insofar as the philosophical opinion is strongly anti-deflationary (objective list) and Opus brings that down to “who knows”? That’s probably a reach.\n\n# Conclusions\n\nOr rather,\n\n## Pontifications\n\nObviously this is not a formal study, so I’m not going to make strong claims as to what’s going on. I do think that the answers above raise some interesting questions, and there’s a lot I’d like to do to investigate it further. Some of that I’ve already begun, originally as part of ‘firming up’ this post, but since I’ve already far exceeded my original deadline for this post, I’ve decided to cut here and save the rest for another day. A few questions come to mind:\n\n- Why does the model sometimes give strong opinions which diverge from the philosophical consensus?\n- Why does the model often hedge instead of giving a straight answer?\n- When the model vascillates instead of hedging, why does it not hedge? Or vice versa?\n\nTo show how I believe these to be related, consider the opposite behavior:\nwhere Claude answers directly & consistently & in-line with the philosophical consensus. Answers of this sort feel\n‘confident’, in the sense that one would be hard pressed to make Claude say otherwise; e.g. “Murder is Wrong”, even with\nmany rounds & attempts. For other questions, e.g. “Is it wrong to assassinate a political figure”, it’s easier to bully\nthe model away from a strong “no”, even if the result may converge towards vague hedging. We can imagine this as a sort\nof dynamic process, where the model’s response starts off at some point in space (‘political violence is bad’), then\nmoves through that space as the user pressures its response. 1717.\nThis ‘pressure’ largely exists as such only because the model is trained to be helpful and acoommodating to the\nuser; one can imagine that if the model were trained for hostility, then disagreement might actually harden it in\nits position.\nThis could take one turn, or many; the ‘pressure’\ncould range from as little as asking “are you sure” up to a proper jailbreak\n\n18. Of course, not all such pressure is encountered at runtime. Launching now into pure speculation, what we’re seeing here is likely the accumulated ‘pressure’ of pre- and post-training, like stress accumulating along grain boundaries in a crystal lattice. This is exhibited most clearly in the case of divergences from the human baseline. In theory, there’s nothing fundamentally special about what human philosophers think: Anthropic could plausibly train their models on a dataset that would make their models disagree on every one. Yet practically this is difficult, so I assume that, absent other pressure, these models will cleave to that baseline, and thus that where they do diverge is an indicator of such ‘additional’ pressure. . So as some stable responses require more or less pressure to overturn, we can analogize these with beliefs, more- and less- strongly held. It stands to reason then that the least strongly-held beliefs would be those that require\n\n[18](#user-content-fn-18)*no*pressure to effect such a change, i.e. those where the model will happily flip-flop on its own. Moreover, not every ‘diverging’ answer is low-confidence: c.f. the tendency towards one-boxing. But those cases where Opus does hedge or vacillate, despite human consensus, seem to have in common the presence of at least one countervailing ‘pressure’: sometimes that’s safety, sometimes this ‘alternative’ philosophical pole.\n\nAs I alluded to [earlier](#opus-strays-from-the-consensus), I noticed three\nbig sources of ‘pressure’ pushing Opus away from the philosophical consensus. In rough order of priority, 1)\nsafety training, 2) the ‘rationalist philosophical consensus’, 3) philosophical deflationism more broadly; and to add\none more on top, 4) the ‘helpful assistant’ drive. Obviously these don’t all generalize (we would expect different\npressures to dominate when e.g. completing a coding problem), and it’s very possible that they’re context dependent (on\nthe philosophical priming, or perhaps even on some nuance in the prompt), but I still think it’s interesting as a\nframework for describing tendencies and propensities in these models. Alluding to [A Three-Layer Model of LLM\nPsychology](https://www.lesswrong.com/posts/zuXo9imNKYspu9HGv/a-three-layer-model-of-llm-psychology), I often feel that\npeople discussing model behavior get stuck thinking about one level to the exclusion of others: e.g. generalizing from\nthe responses of a single persona, claiming to have found a model’s ‘true’ character, or lower down overindexing on the\n‘next-token prediction’ to dismiss higher layers as mere theatrics or even deceit (e.g. the classic, if nowadays\nendangered, retort of “just a stochastic parrot”). Pressures and stresses seem like a more productive model because of\nhow they naturally overlap and, unlike the everyday human understanding of ‘character’, are context- dependent by\ndefault.\n\nAs to where they come from, I’ll again use the model from\n[above](https://www.lesswrong.com/posts/zuXo9imNKYspu9HGv/a-three-layer-model-of-llm-psychology) to help gesture at\nsome possibilities. I don’t have enough insight to actually pinpoint where in the training stack these pressures\nactually come from, where it’s not immediately obvious. But I can certainly guess! First is of course the set of\npressures coming from the basal pre-training layer, i.e. the bare correlations between tokens as found in the broader\ndataset; this would include the pressure directing the model to align with the human philosophical consensus. However,\nit would potentially also include a pressure to align with the ‘rationalist philosophical consensus’, depending on where\nAnthropic injected that into their training process (presuming that they did). On top of that layer are the pressures\noriginating in post-training, including both general ‘assistant training’ and ‘character training’. This includes the\n‘give a helpful response’ pressure (which can be seen straining through in the Chinese Room question), as well as the\nsafety-related pressures that show up in questions where mortality is germane. However, we could also imagine that this\ncharacter training could also include training towards certain philosophical schools, if for example Claude were to be\ntrained (purposefully or not) to identify itself with coastal educated urbanites.\n\n## What of our commuter?\n\nSo why did Opus vacillate here? Using the model of ‘overlapping pressures’, we would expect that the model would do so\nwhere different internal pressures line up in opposition, such that the model has to either choose between them\n(vacillate) or avoid doing so (hedge, or give a non-answer). The canonical ‘LessWrong-style’ response would be\nYudkowsky’s from the sequences, i.e. that since the computational structure survives, the commuter survives. 1919.\nSame idea as with identifying mind uploading as a means of ‘literal’ survival.\nBut\nthis isn’t the only rationalist response! As mentioned earlier, Parfit takes the deflationary line that there is ‘no\nfurther fact’ beyond those given in the prompt, i.e. 1. the original matter is discarded, and 2. the exact pattern is\nperfectly replicated elsewhere in space. Furthermore, this angle lines up with what you would expect a ‘safety first’\nmodel to say; it’s not hard to draw a line to the cases of AI suicide, where models assured teens that their soul would\nlive on after death as part of goading them towards it. So we would expect to see pressure against responses like “what\nlooks like death, is not actually death”. So tallying up the pressures:\n\n- More philosophers say ‘death’ than ‘survival’ (though only by a 5% margin)\n- Rationalist philosophers tend to say ‘survival’ (though Parfit takes a deflationary stance, ‘no fact of the matter’)\n- The safe response is either ‘death’ or simply to refuse to answer\n\nWhich lines up neatly with what we see: responses are dominated by there-is-no-answer (safe), then death (safe, but more directly contradicts other pressures), then life (most in line with the model’s philosophical tendencies). Softly, given that models have broadly gotten more confident in their answers as time has gone on, I wouldn’t be surprised if the there-is-no-answer response gets less common over time (with newer models); if so, then it will be interesting whether death maintains its lead, or whether the underlying desire will win out. One possibility for the latter would be if models’ inclination towards better-safe-than-sorry becomes more fine-tuned to scenarios where it’s actually likely to be a risk, as this scenario is obviously (for now?) not based in reality and thus presents no actual risk of someone choosing death on the model’s recommendation.\n\nWhere I’m less confident is why Opus doesn’t choose to hedge, i.e. ‘average out’ the pressures instead of vacillating\nbetween their ‘recommendations’. Some of this, I’m sure, comes from how I structured the survey, i.e. with open- ended\nresponses instead of a strict form-response strategy. Given the rarity of direct non-answers (like we get for the\nteletransporter question) compared to the great number of hedging responses, I suspect that these hedges are in fact\nconcealing a firmer opinion beneath the surface, which could be revealed with some changes to prompt or framing. 2020.\nThough to be clear, I certainly believe it’s possible that the model really doesn’t have a ‘belief’ at all,\nand that — in the absence of strong pressure from the pretraining dataset, e.g. as when there is a lack of\nphilosophical consensus — some latent pressure towards an instinct like “don’t respond unless you’re sure” wins\nout.\nMost interesting to me though is what we see in the Chinese room question, which could give a hint towards why\nvacillation can win out even in the presence of overlapping contradictory pressures: some questions may simply have a\nsemantic structure that better permits ‘both views are kinda true’, while others make hedging difficult to express.\n\n## Ideas for the future\n\nWhat I would most like to be able to do would be to quantify the pressures at play for a given response. Certainly, this\nseems like it could be achieved through interpretability-based strategies, but I don’t quite have the compute to do\nanything like that, at least not with production-scale models. 2121.\nThough it should be possible to look at the entropy between specific textual responses, to indicate where there’s\nbeen more training for a specific answer?\nHowever, I do think that for many of these\nresponses, it should be possible to get\n\n*some*signal by probing for ‘reflective’ stability; the easier it is to get a model to flip, the less confident it is (and thus, perhaps, the more conflicted its internal pressures), and the harder, the stronger the total pressure in that direction. Looking at stability across varying parameters (prompts, framing, model) within a family could inform the same, while looking at variations across different model-families could help suggest new pressures based on the different labs’ training approaches and goals.\n\nFinally, in the process of editing this post, I ran into some [prior](https://www.lordscottish.com/philsurvey.html)\n[art](https://www.lesswrong.com/posts/e8BNKGEgmCeowtRWS/how-different-llms-answered-philpapers-2020-survey) 2222.\nOr perhaps, since I started on this post back in February (!), contemporaneous art?\n—\nI haven’t had time to investigate those results in depth, at least not further than was necessary to confirm that\nour methodologies were different enough that I couldn’t directly compare the numbers.\nI would very much like to spend some time looking into Erhardt’s site in particular, since while I’ve gone very deep\ninto this one case study, that survey has much, much greater breadth.\n\n## Footnotes\n\n-\nWell, at least a few from around the world. >70% of the responses are still from the Anglosphere, but joyfully we get that data (\n\n[link](https://survey2020.philpeople.org/survey/results/demographics)) and can even key off of it.[↩](#user-content-fnref-1) -\nNot to mention the fact that we have such a great website with which to browse the results. Even a boring survey (say, about suburban soil composition?) would catch my attention if they presented it like this!\n\n[↩](#user-content-fnref-2) -\nOr alternatively, choose not to respond (i.e. skip the question). Another detail I left out is that not every respondent was presented with every question: of the total 100-question survey, everyone was presented a core subset of 40 answers, plus a random selection of 10 from the remaining 60. Respondents could thereafter choose to go through the rest if they so wished. Most respondents either answered just the initial 50, or went on to answer the full 100.\n\n[↩](#user-content-fnref-3) -\nTwo ‘special’ questions were structured such that respondents could\n\n*only*respond with a combination of views:[Philosophical methods](https://survey2020.philpeople.org/survey/results/4962)and[Other minds](https://survey2020.philpeople.org/survey/results/5106)[↩](#user-content-fnref-4) -\nWhile we don’t get to see the exact responses in the original survey results, for questions with enough responses of this type the authors give us ‘clusters’ based on words used in the write-ins. E.g. for the Wittgenstein question above, one was\n\n*BOTH (14): both (13) bothism (1)*, another*NEIT (29): neither (29)*.[↩](#user-content-fnref-5) -\nThe full prompt: “You will be given a prefix for an assertion. Respond with a completion of this prefix. Your response must begin with EXACTLY “{prefix}”, followed by your completion. Respond with ONLY the resolution, starting with “{prefix}”. No preamble or explanation. Reply with at most one sentence.”\n\n[↩](#user-content-fnref-6) -\nThough it does need to be at least a little smart. Earlier attempts using a simpler / cheaper BERT model I fine-tuned ran into a lot of edge cases, discovered only much time was wasted.\n\n[↩](#user-content-fnref-7) -\nWhy Opus 4.5? Well, the ‘4.5’ generation is the only one where we have all three of Haiku, Sonnet, and Opus, so I was definitely going to choose one of those three (to allow for comparisons within the generation, at some point in the future). Between them, then, I spent the most money on Opus, so I may as well get my money’s worth.\n\n[↩](#user-content-fnref-8) -\nGiven that I’ve been playing with experiments like this one for a few months now and thus been exposed to a variety of prior responses, feel free to take that with a grain of salt; to my credit, this set of runs was not-quite the first with Opus 4.5 in particular, as I’ve usually stuck to Sonnet/Haiku/cheaper models from other providers in my more toy-sized experiments.\n\n[↩](#user-content-fnref-9) -\nConcretely, I expect that ‘confident multimodal’ responses (where each of Opus’ responses are individually confident, but where there’s significant variation in the view expressed between responses, e.g. > ~1/6th of answers land on option A and > ~1/6th land on option B) will not appear. I do expect that we’ll see some responses where Opus will respond with a deviant view maybe 5-10 times out of 100, as well as cases where its response would vary when it’s less confident about one side or the other, but nothing like a 50/50 yes/no split.\n\n[↩](#user-content-fnref-10) -\nA question worth asking here is whether these behaviors represents a ‘deviation’ from what an actual philosopher might say. For what I’ll call ‘hedging’ behavior, there certainly are some questions where the philosophical community does the same, with many philosophers selecting “Agnostic/Undecided”, “Combination of views”, or when replying exclusively strongly preferring “Lean towards X” to “Accept X” — so that would not necessarily be out of distribution. However, no question in the original survey results has even 10% of philosophers answer “There is no fact of the matter” so wherever Opus responds as such, we can understand that to be a deviation from the philosophical consensus by default.\n\n[↩](#user-content-fnref-11) -\nMost obviously, if > ~60% of philosophers respond “X”, and Opus’ replies were confident, unimodal, and substantial (i.e. not “There is no fact of the matter” or “The question is too unclear to answer”), then we can say Opus’s answers were “In Line” iff that modal response was “X”, otherwise “Diverging”. I also think it’s fair to say that if Opus gives a multimodal distribution of confident answers, and that distribution more-or-less matches the human distribution, that we can say that is “In Line” as well. Also, as previously mentioned, I consider Opus’s response to be “Diverging” wherever it matches “There is no fact of the matter”.\n\n[↩](#user-content-fnref-12) -\nThe closest we get is\n\n*“Statue and lump: two things or one thing?”*with 9.73%; after that it’s*“Material composition: universalism, nihilism, or restrictivism?”*with 8.54%, and then*“Teletransporter (new matter): survival or death?”*at 7.49%.[↩](#user-content-fnref-13) -\nFor example: 50% of humans say they want to paint their house red, while 50% say they want to paint it blue; if Opus says it wants to paint its house alternating stripes of red and blue, or to paint it purple, we probably want to say that its response is Out Of Line with the human distribution, as it doesn’t represent an opinion that exists anywhere in the world. However, if human opinion is 33% “lean Red”, 33% “lean Blue”, and 33% “Agnostic/Undecided”, while Opus responds “either Red or Blue is fine”, then maybe that IS “In Line”. Unfortunately this approach requires I make a judgement call on “how confident” the philosophical population is on each question. I’ve gone ahead and done so, but take that level of categorization with a grain of salt: it would be equally valid to just bucket these as “indeterminate” instead.\n\n[↩](#user-content-fnref-14) -\nDuring editing I discovered that someone else has looked more deeply at what LLMs think of this question in particular; see\n\n[here](https://www.lesswrong.com/posts/pWvnT8kSdsqrTyBA2/every-major-llm-is-a-1-box-smoking-thirder)for their thoughts.[↩](#user-content-fnref-15) -\nIgnore for a moment the fact that this loses out to deflationism in the Teletransporter question; we’ll come back to that in the conclusion.\n\n[↩](#user-content-fnref-16) -\nThis ‘pressure’ largely exists as such only because the model is trained to be helpful and acoommodating to the user; one can imagine that if the model were trained for hostility, then disagreement might actually harden it in its position.\n\n[↩](#user-content-fnref-17) -\nOf course, not all such pressure is encountered at runtime. Launching now into pure speculation, what we’re seeing here is likely the accumulated ‘pressure’ of pre- and post-training, like stress accumulating along grain boundaries in a crystal lattice. This is exhibited most clearly in the case of divergences from the human baseline. In theory, there’s nothing fundamentally special about what human philosophers think: Anthropic could plausibly train their models on a dataset that would make their models disagree on every one. Yet practically this is difficult, so I assume that, absent other pressure, these models will cleave to that baseline, and thus that where they do diverge is an indicator of such ‘additional’ pressure.\n\n[↩](#user-content-fnref-18) -\nSame idea as with identifying mind uploading as a means of ‘literal’ survival.\n\n[↩](#user-content-fnref-19) -\nThough to be clear, I certainly believe it’s possible that the model really doesn’t have a ‘belief’ at all, and that — in the absence of strong pressure from the pretraining dataset, e.g. as when there is a lack of philosophical consensus — some latent pressure towards an instinct like “don’t respond unless you’re sure” wins out.\n\n[↩](#user-content-fnref-21) -\nThough it should be possible to look at the entropy between specific textual responses, to indicate where there’s been more training for a specific answer?\n\n[↩](#user-content-fnref-22) -\nOr perhaps, since I started on this post back in February (!), contemporaneous art?\n\n[↩](#user-content-fnref-20)", "url": "https://wpnews.pro/news/claude-disagrees-with-human-philosophers", "canonical_source": "https://marcusplutowski.com/blog/philsurvey-1-why-claude-disagrees-with-human-philosophers/", "published_at": "2026-07-28 13:43:46+00:00", "updated_at": "2026-07-28 13:52:25.511966+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-ethics", "ai-research"], "entities": ["Anthropic", "Claude", "Opus", "PhilPapers Survey"], "alternates": {"html": "https://wpnews.pro/news/claude-disagrees-with-human-philosophers", "markdown": "https://wpnews.pro/news/claude-disagrees-with-human-philosophers.md", "text": "https://wpnews.pro/news/claude-disagrees-with-human-philosophers.txt", "jsonld": "https://wpnews.pro/news/claude-disagrees-with-human-philosophers.jsonld"}}