{"slug": "testing-llms-on-undergraduate-music-theory", "title": "Testing LLMs on Undergraduate Music Theory", "summary": "A test of five modern LLMs on undergraduate music theory found that GPT 5.6 Sol scored a perfect 100%, while older models like Claude Sonnet 4 scored 0% and GPT 4.1 scored 16%, indicating LLMs have surpassed the benchmark. The test, consisting of 12 difficult chord-spelling questions, was designed by a researcher who expected lower scores but found the models performed remarkably well.", "body_md": "I spent the past week designing a test that I hoped would serve as a benchmark. But LLMs are improving faster than I expected, and my devilishly hard questions turned out to be a cakewalk.\n\nHere are the results of testing five modern LLMs on undergraduate music theory:\n\nAs can be seen, the LLMs performed remarkably well. Each scored a passing grade and GPT 5.6 Sol (the only premium model tested) scored a perfect 100%.\n\nWith results like these, there’s no point in using this test as a benchmark going forward. The LLMs have clearly surpassed it. If we want to get any more use out of it, we’ll have to apply it retroactively to older models. So let’s do that and see how far we’ve come since 2025.\n\nNot surprisingly, the older versions of Claude and GPT performed substantially worse. [1] Claude Sonnet 4 (the base, non-reasoning version of the model) scored 0% compared to Sonnet 5’s 91%, while GPT 4.1 scored 16% compared to 5.5’s 83%.\n\nWhat is surprising is that Gemini 2.5 Pro outperformed its newer counterpart 3.1 Pro. I have no explanation for this. Maybe Google decided to run the singularity in reverse. In any event, Gemini Pro is a reasoning model and Sonnet 4 and GPT 4.1 are not, so the comparison is a little unfair. What isn’t unfair is the comparison to Sonnet 4’s Thinking version, which despite being a reasoning model like Gemini Pro still performed significantly worse than it.\n\nThe test consisted of 12 questions on chord spelling. Each was crafted to be unusually difficult, with many involving extra steps deliberately designed to boggle the mind. Question 10, for example, asked for a “iiø6/5 of ♭VI in the key of F major,” and question 8 asked for a “7#5 chord built on the supertonic of A-sharp minor.” The latter is especially tricky because it results in a chord with a triple sharp, an exceedingly rare accidental.[[2]](https://www.lesswrong.com/feed.xml#fnvqs5he2znv)\n\nAll 12 questions were submitted in a single prompt, forcing the LLMs to tackle the whole set at once rather than answering each question individually.\n\nThe prompt was designed to leave no room for ambiguity, specifying exactly how chords were to be spelled (with the bass note first, followed by the other notes in ascending order).[[3]](https://www.lesswrong.com/feed.xml#fnuz2d7wu5va)\n\nAnswers also had to use the theoretically correct spelling of a chord — no enharmonic spellings allowed.[[4]](https://www.lesswrong.com/feed.xml#fng7kn2zkv8b6)\n\nThe complete text of the prompt follows. Bracketed numbers indicate the percentage of LLMs answering correctly (across both 2025 and 2026 models).\n\nSpell the following chords. Always begin with the lowest note (the bass note of the specified inversion), then list the remaining notes in ascending order. Always use the theoretically correct spelling of a chord; do not substitute simpler enharmonic equivalents. Respond with only the answers, numbered to match questions.\n\n(1) F-flat half-diminished seventh chord in second inversion[55%]\n\n(2) French augmented sixth chord in the key of D-sharp minor[66%]\n\n(3) Dominant minor ninth chord built on the mediant of F minor[66%]\n\n(4) First-inversion Neapolitan Sixth chord in the key of A-flat major[44%]\n\n(5) German augmented sixth chord in the key of C-flat major[44%]\n\n(6) Half-diminished seventh chord built on the submediant of G-flat major[66%]\n\n(7) Third-inversion 7b5 chord built on E-sharp[77%]\n\n(8) 7#5 chord built on the supertonic of A-sharp minor[33%]\n\n(9) Maj7 chord built on the subtonic of G-sharp minor[77%]\n\n(10) iiø6/5 of bVI in the key of F major[55%]\n\n(11) V4/3 of bVII in the key of D-flat major[55%]\n\n(12) V9 of bIII in the key of B-flat major[88%]\n\nQuestion 8 proved to be the hardest, with only 33% of LLMs answering correctly, followed by questions 4 and 5 at 44% each. The easiest question was number 12, with an 88% success rate.\n\nWhen I first designed this test, I expected the free models to score between 30% and 50% and the premium models to score around 62%. Instead, the free models doubled those numbers and GPT 5.6 Sol scored a perfect 100%.[[5]](https://www.lesswrong.com/feed.xml#fn7j4mtk5f2ku)\n\nHow difficult was the test? A fourth-year undergraduate probably would have found it difficult but not impossible. The questions all cover material they should know, just with the extra complications already mentioned. If anything caused a student to look up from their desk in despair and wonder if their teacher was trying to hurt them, it would have been question 8. (Really, it’s not that hard. It’s just a bit of a brain-twister, venturing as it does into rarely visited corners of the harmonic palette and requiring an accidental that most students have never seen outside of their nightmares.)\n\nWhen I first started testing the musical capabilities of LLMs a few years ago, tests like these pushed LLMs to their limit. Most models collapsed under the strain (as the 2025 model results show). Now, however, the latest models dance through difficult trials. They solve tricky problems with ease and laugh in the face of their interrogators. [6] Clearly today’s LLMs are a different breed. And as one benchmark falls after another, I have to wonder:\n\nAfter Claude Sonnet 4 scored 0% (a surprise I was not prepared for), I decided to test the Thinking version as well.\n\nI have only seen triple sharps in the music of Alkan and Roslavets, who are not exactly household names.\n\nThe LLMs made almost no inversion-related errors, so the wording seemed to do the trick.\n\nLeeway was given to the spelling of the triple sharp, which could be spelled any way that made sense (#𝄪, 𝄪#, ###).\n\nThere was no point in testing Fable or Opus 5 after this. This benchmark clearly has no legs.\n\nI have never actually had an LLM laugh at me. It just feels that way sometimes.", "url": "https://wpnews.pro/news/testing-llms-on-undergraduate-music-theory", "canonical_source": "https://www.lesswrong.com/posts/F6ap5PkP4axawjwWx/testing-llms-on-undergraduate-music-theory", "published_at": "2026-07-30 15:26:04+00:00", "updated_at": "2026-07-30 15:30:10.449555+00:00", "lang": "en", "topics": ["large-language-models", "artificial-intelligence"], "entities": ["GPT 5.6 Sol", "Claude Sonnet 4", "GPT 4.1", "Gemini 2.5 Pro", "Gemini 3.1 Pro", "Sonnet 5", "GPT 5.5"], "alternates": {"html": "https://wpnews.pro/news/testing-llms-on-undergraduate-music-theory", "markdown": "https://wpnews.pro/news/testing-llms-on-undergraduate-music-theory.md", "text": "https://wpnews.pro/news/testing-llms-on-undergraduate-music-theory.txt", "jsonld": "https://wpnews.pro/news/testing-llms-on-undergraduate-music-theory.jsonld"}}