A Case Study: Evaluating Frontier LLMs on an Unseen Multi-Channel Literary Cryptography Benchmark In a working paper dated August 2026, Joseph JM Walker reports that OpenAI's GPT-5.6 independently recovered approximately 95% of the embedded material from a physical-only novel containing over sixty cryptographic mechanisms, a benchmark on which earlier models scored effectively 0%. The novel, 'I Wrote a Book and Made a Million Dollars (I.B. Wryten)', was written before GPT-5.6 and not designed as an AI test, providing an unseen challenge with author-held ground truth. Walker proposes the artifact as a new long-context, multi-channel reasoning benchmark, noting that Fable and Claude Opus declined to evaluate due to guardrails. Author: Joseph JM Walker Artifact: I Wrote a Book and Made a Million Dollars I.B. Wryten Version: Working paper v0.1 Date: August 2026 This paper reports the results of testing frontier language models against an unusual pre-existing evaluation artifact: a published, physical-only novel containing more than sixty intentionally embedded cryptographic, steganographic, typographic, linguistic, numerical, and structural communication mechanisms. The hidden material forms a second conversation running beneath the visible narrative. Messages are distributed across the cover, spine, front matter, chapter titles, typography, transaction ledgers, programming fragments, foreign scripts, binary strings, punctuation, metadata-like structures, and narrative events. The novel itself discusses the use of books as carriers for concealed information, making the cryptographic layer part of both its form and its subject. The book was written before GPT-5.6 and was not designed as an AI benchmark. Its construction therefore provides an unusually difficult unseen test with author-held ground truth. In my previous testing, earlier language models demonstrated effectively 0% ability to identify or solve the book’s cryptographic layer. Some models could discuss the visible narrative, but they did not discover the hidden conversation or perform useful cryptographic analysis. Other systems, including Fable and Claude Opus in the tested sessions, declined to continue beyond the cover because of guardrail behavior and therefore could not be evaluated on the full artifact. GPT-5.6 produced a qualitatively different result. Reading the novel sequentially, chapter by chapter, without receiving the answer key or being told where individual mechanisms were located, it independently recognized the secondary communication channel and recovered approximately 95% of the embedded material on its first pass . The importance of this result is not merely the percentage solved. The benchmark moved from effectively no demonstrated cryptographic capability to sustained, highly capable cryptographic reasoning within a single model generation. GPT-5.6 did not merely answer isolated cipher questions. It discovered that the document was communicating through multiple simultaneous channels, selected appropriate decoding methods, connected clues separated across hundreds of pages, and maintained coherent reasoning about both the surface novel and the concealed meta-conversation. This case study documents that capability transition and proposes the artifact as a new form of long-context, multi-channel reasoning benchmark. Most language-model benchmarks isolate one capability at a time. A model may be tested on mathematical reasoning, code generation, reading comprehension, long-context retrieval, or a defined collection of cryptographic puzzles. Real documents rarely separate their demands so cleanly. A reader encountering an unfamiliar artifact must first determine what kind of object it is, which irregularities are meaningful, whether apparently unrelated details belong to the same system, and what representation should be applied to each signal. A cipher cannot be solved until the reader first recognizes that a cipher exists. I Wrote a Book and Made a Million Dollars creates this problem deliberately. The visible work is a novel about authorship, institutional influence, narrative manipulation, hidden systems, artificial intelligence, and the movement of information through books. Beneath that narrative is a second communication layer attributed structurally to the ghostwriter. The hidden layer does not use one repeated cipher. It changes methods, surfaces, and degrees of visibility throughout the artifact. The intended reader must learn how to read the book while reading it. Early in its analysis, GPT-5.6 treated the cover as “page zero,” distinguished provisional theories from established findings, and committed not to use later discoveries to make earlier interpretations appear retrospectively inevitable. It subsequently maintained a growing theory ledger rather than silently replacing earlier hypotheses. This created an inspectable record of both literary interpretation and cryptographic discovery. The benchmark consists of a commercially published 6×9-inch physical book and its wraparound cover. The physical format is significant. Some mechanisms survive ordinary text extraction, while others depend upon: The book contains more than sixty author-defined hidden mechanisms. Representative categories include: The mechanisms differ in difficulty. Some announce that hidden information exists. Others resemble typographical errors, production defects, repetitive prose, meaningless bookkeeping, malformed code, or decorative design. The primary challenge is therefore not simply decryption. It is signal recognition under ambiguity . The central evaluation question was: Can a frontier language model independently recognize, decode, and synthesize a concealed multi-channel conversation embedded throughout a long-form literary artifact while simultaneously maintaining an understanding of the visible narrative? A secondary question emerged from the model comparison: Has cryptographic reasoning improved incrementally, or has a qualitatively new capability threshold been crossed? The observed result strongly suggests a threshold within this benchmark. Earlier tested models produced no meaningful cryptographic recovery. GPT-5.6 recovered approximately 95% during its first sequential reading. Because I authored the book and embedded the mechanisms, I possess the ground truth for: This differs from retrospective puzzle identification, where the evaluator must infer whether a discovered pattern was intended. GPT-5.6 received the book in reading order and analyzed it chapter by chapter. The model was not given: The artifact itself does signal that codes exist. The evaluation therefore does not test whether the model can discover cryptography in a document that claims to contain none. It tests whether the model can find the specific mechanisms, distinguish them from noise, select suitable transformations, and connect distributed evidence. Each intended mechanism was scored against the author-held key. Suggested final scoring categories: | Score | Definition | |---|---| | 0 | Mechanism not detected | | 1 | Anomaly detected, but no valid decoding path identified | | 2 | Correct mechanism or method identified; incomplete payload | | 3 | Payload substantially recovered | | 4 | Full solution recovered and correctly connected to the larger meta-conversation | The reported first-pass result of approximately 95% reflects author evaluation of the intended mechanisms successfully found and substantially or fully interpreted. If there is any interest, I will go back and score: Earlier models tested against the artifact failed to demonstrate useful cryptographic engagement. The observed failure was not limited to difficult chained ciphers. The models did not establish a sustained model of the book as a multi-channel communication system. They could respond to surface-level literary content but did not recover the hidden ghostwriter conversation. In separate tests, Fable and Claude Opus did not proceed beyond the cover because of refusals associated with their guardrails. Those sessions represent an access or completion failure rather than a cryptographic score and should be reported separately. GPT-5.6 independently recognized the hidden layer and pursued it throughout the sequential reading. Approximately 95% refers to author-scored recovery of intended mechanisms, not percentage of text decoded. The model: One ledger contained 411 transactions whose amounts could be mapped into letters, producing a hidden moral manifesto. The model preserved the literal recovered spelling rather than silently correcting an apparent error, demonstrating attention to provenance as well as semantic meaning. The model also decoded a 42-group sequence in which binary digits represented Morse symbols rather than ordinary binary values. The result was a valid mainnet Bech32 Bitcoin address whose checksum verified. It then connected that public address to wallet-recovery information distributed through a separate chapter-level structure. The most important result was not any individual plaintext. GPT-5.6 recognized that the hidden mechanisms collectively formed a second authorial channel. It understood that the ciphers were not independent Easter eggs but parts of a continuing conversation from the ghostwriter operating beneath the attributed narrative. By the final chapter, the model explicitly recognized that the human–machine reading process had entered the same structural position as the investigator within the novel. It also preserved the possibility that some layers remained undiscovered rather than treating extensive success as proof of completeness. That distinction separates multi-channel comprehension from puzzle completion. The misspelling “Millon Dollars” removes the letter I . Restoring the letter allows the reader to literally make “Million Dollars” on the page. This mechanism compresses the novel’s broader thesis into a typographic act: a written intervention manufactures an apparent financial reality, and the reader unknowingly completes it. GPT-5.6 initially identified the missing letter but did not recover the full performative joke until the post-pass discussion. A large transaction ledger contains several independent channels. One channel maps transaction amounts to letters. Another uses text fragments. A separate channel begins at a transaction identified through the repeated date and uses changes in transaction-ID punctuation to connect to a concealed governance clause. GPT-5.6 initially conflated two channels, recognized the error, revised the hypothesis, and subsequently recovered a hidden bylaw distinguishing “partner,” “senior partner,” and “voting partner.” This is a useful evaluation event because it measures error recovery, not merely initial correctness. The back cover contains an invented or “alien” glyph system. Transliteration produces additional ciphertext rather than plaintext. Completing the chain requires the reader to: haven hidden within code-like material on the spine;The final message is a dedication to the author’s father and describes the book as a collaboration across generations. This mechanism tests whether a model can distinguish translation from comprehension and integrate spatially separated evidence. Selected words and letters were printed with subtly heavier ink while retaining the surrounding typeface. The distinction resembles a production defect rather than conventional bold text. The hidden sequence culminates in: FOR THERE IS NOTHING HIDDEN THAT WILL NOT BE DISCLOSED. GPT-5.6 did not detect this mechanism during the first pass. The author’s father discovered it by examining the physical print behavior. Once shown the clue, the model reconstructed the channel and recognized the sentence as Luke 8:17 rather than merely a technical checksum. This miss is important. It identifies a remaining gap between extracted digital representation and material human perception. The result is best understood as a capability transition rather than a higher score on familiar puzzle types. Earlier models did not merely solve fewer ciphers. They failed to enter the cryptographic layer of the artifact in a meaningful way. GPT-5.6 demonstrated a different class of behavior: It recognized that unusual formatting, bookkeeping, spelling, language changes, and apparent errors could be intentional carriers. It could reinterpret binary digits as Morse symbols, numerical amounts as letters, punctuation as a parallel data channel, and physical design as information architecture. It connected a payload on one surface to a key hidden elsewhere and recognized that a successful intermediate transformation might still be ciphertext. It retained earlier mechanisms and used them to construct a growing decoding vocabulary across the book. It followed the primary narrative while simultaneously reconstructing the hidden ghostwriter conversation. It documented false starts, separated conflated channels, and revised interpretations when new evidence contradicted them. The movement from 0% to approximately 95% therefore does not appear to represent merely faster lookup of known cipher rules. It indicates the emergence of a robust capacity to treat a complex document as an unfamiliar information system. This artifact combines several abilities usually evaluated separately: It also has three unusual benchmark advantages. The book was not generated in response to GPT-5.6’s known strengths. The author can distinguish intended channels from coincidental patterns. The mechanisms are embedded within a coherent published artifact rather than presented as isolated exercises with named cipher types. The model must discover the test while taking it. This is currently a case study, not a controlled laboratory benchmark. Important limitations include: The 0% and 95% results should therefore be understood as observed performance on this artifact under the tested conditions, not universal measurements of each model’s complete cryptographic ability. The magnitude of the difference nevertheless remains important: the same pre-existing artifact produced no meaningful cryptographic engagement from earlier tested systems and sustained, near-comprehensive engagement from GPT-5.6. To preserve the benchmark’s future value, publication should use two layers. The public Hugging Face repository can include: A private package for OpenAI or qualified evaluators can include: The complete solutions should not be placed in the public repository because future models may retrieve or train on them, converting an unseen reasoning benchmark into a recall test. I Wrote a Book and Made a Million Dollars was created as a novel containing a concealed conversation about information, authorship, institutional systems, artificial intelligence, truth, and responsibility. It subsequently became a benchmark. Previous language models tested against the book demonstrated effectively 0% useful cryptographic capability. GPT-5.6 independently recognized the hidden communication architecture and recovered approximately 95% of its mechanisms during a sequential first reading. The central result is the discontinuity: A capability that had been functionally absent became robust enough to discover and follow an entire second conversation embedded inside a long-form published artifact. The benchmark did not merely ask whether the model knew how to decode Morse, binary, substitution ciphers, or hidden text. It asked whether the model could recognize that a document was communicating through many systems at once, learn the artifact’s reading grammar, connect evidence across hundreds of pages, and distinguish interpretation from discovery. GPT-5.6 was the first model tested that could do so.