AI-Generated Multiple-Choice Quizzes: I Found a Bias in Answer Placement (and How to Check Before You Ship) A developer who generated 300 AI-written multiple-choice practice questions for an AI certification exam found that the correct answer's position advanced sequentially (1→2→3→4→1) in 58% of adjacent questions, far above the 25% expected from random placement, allowing the pattern to be guessed without understanding the content. After scrambling option order and enabling the platform's randomize setting, the rate dropped to 32% for the 120 questions published. The developer also reported that an auto-fill feature truncated the published course subtitle at a 120-character limit, prompting a weekly automated check comparing live page text against the approved version. When you ask an AI to generate multiple-choice questions for a certification exam, it churns out a lot of plausible-looking questions in an instant. They read naturally, without obvious weirdness. But when I actually counted, I found a significant bias. This post will share three checks you can perform before releasing AI-generated multiple-choice questions for study quizzes, internal assessments, or practice tests to other people. No programming required. I had an AI generate 300 practice questions for a specific AI certification exam. They were 4-option multiple-choice, each with an explanation. My goal was to publish them as "Practice Tests" on a platform like Udemy. I fed the questions without answers or explanations to a different AI than the one that generated them. The result: 300 out of 300 questions were answered correctly by the second AI. This made me feel pretty confident at first. However, as the next check revealed, this perfect match needed to be interpreted with caution. The prompt I used was something like this: "Solve the following 4-option multiple-choice questions. In addition to the answers, mark any questions you found difficult or confusing, and any where the factual information seemed old or questionable." Out of 300 questions, 23 were flagged. These included descriptions of outdated product generations or statements with ambiguous attribution. I removed all of them from consideration. I listed the correct answer's position 1st, 2nd, 3rd, or 4th option for each question in order. What I found was a clear pattern: the correct answer often cycled sequentially, like 1 → 2 → 3 → 4 → 1... For adjacent questions, the correct answer's position "advanced by one" in 58% of cases. In a truly random distribution, this would be around 25%. In some sections, it was as high as 83% This meant that even without understanding the content, someone could guess the answer by recognizing the positional pattern. The 300-question match from Check 1 might not entirely reflect the solving AI's true ability if it caught onto this pattern for certain sections. The fix is simple: scramble the order of the options while keeping the question and explanation paired. For the 120 questions I actually published, I got this down to 32%. I also enabled the platform's "randomize option order" setting. How to test: In a spreadsheet, list "Question Number" and "Correct Answer Position." Then, count how many rows show a "+1" progression from the previous question with 4 wrapping around to 1 . If it's significantly over 25%, you should definitely reorder. Platform policies often state that AI-drafted questions are permissible, provided they are reviewed and edited by a human, and that AI usage is disclosed in the description. Mass-producing low-quality courses with AI can lead to account suspension. Therefore, all 120 questions I published were manually reviewed before release. I published the course on September 28th. However, after publication, I noticed that the course description subtitle at the top of the page was cut off mid-sentence. Half-width English characters like "AI" were missing, truncated at a 120-character limit. The cause was an auto-fill feature, where text intended for another field had been inserted into this one. I had checked the character count for the description, but I hadn't looked at the adjacent field that I hadn't explicitly typed into. I only noticed it when I viewed the live course page with my own eyes. This field is used not only at the top of the page but also in search result descriptions. Now, I have a system that automatically reads the published page text weekly and alerts me if it differs from the approved version. How to test: When you save anything that involves AI generation or auto-filling, don't just check the fields you explicitly typed into. Compare all fields on the screen against your expectations, focusing on content, not just character count. Sales-wise, I had 0 enrollments in the first 10 days. I suspect the main reason is that it's not appearing high in search results, though I haven't confirmed this. I'll check the numbers again at the beginning of November and write an update. Before releasing AI-generated multiple-choice questions to others: What I built: Practice Exam for the Generative AI Passport Udemy, 2 x 60 questions https://www.udemy.com/course/seisei-ai-passport-mogi-shiken/?referralCode=6469391F79A2107CC456 https://www.udemy.com/course/seisei-ai-passport-mogi-shiken/?referralCode=6469391F79A2107CC456 I'll update this post with the November numbers. I build and run small Python systems — trading bots, RAG APIs, scheduled automation — and write up whatever breaks along the way. If a provider-agnostic RAG Q&A API is useful to you, mine is MIT-licensed on GitHub: rag-faq-api https://github.com/masaoshimadaOpen/rag-faq-api . It runs and passes its full test suite with no API key offline stub LLM + hashing embedder , swaps to Claude / Gemini / OpenAI via one env var, and ships a retrieval-quality harness Hit@k / MRR / Recall@k with a chunking sweep.