Grand Coder: bounded results from internal AI engineering evaluations An internal evaluation of the AI engineering tool Grand Coder found that three configurations passed 16/18, 16/18, and 15/18 isolated executions versus a 13/18 baseline after a post-scoring audit excluded two task families with invalid hidden requirements, leaving 72 of 96 runs. The tool's creator, who has more than twenty years in systems and infrastructure engineering, reported the results as bounded and not independently reproduced, noting that other evaluations showed saturated tests with no difference, apparent wins that disappeared after scoring corrections, and long-duration failures. The implementation and executable benchmark fixtures remain private, limiting independent verification. I’ve spent more than twenty years in systems and infrastructure engineering. Grand Coder grew out of trying to make that experience useful to the AI agents I work with. I don’t primarily think of it as making a model smarter; the aim is to help a capable model reason and work more diligently. I’m its creator, and our team uses it. What follows is a report on internal historical evaluations—not an independently reproduced public benchmark. One bounded positive result A tiered engineering study ran 96 isolated executions: a matched baseline and three Grand Coder configurations. A documented post-scoring audit found invalid hidden requirements in two task families. Those entire families were excluded across all four conditions, leaving 72 runs, or 18 per condition: - Baseline: 13/18 passed. - Grand Coder configurations: 16/18, 16/18, and 15/18 passed. The comparison held the shared baseline setup constant and measured Grand Coder’s additional contribution—not our whole setup against a bare model. The gains were concentrated in two retained task families; the other four were already passing in every condition. The exclusions were post-scoring, not preregistered. This is not a universal 16.7-percentage-point improvement. What did not support a broad win Other evaluations included saturated tests with no difference, apparent wins that disappeared after scoring corrections, and long-duration failures. In one build study, blinded review scores improved but corrected behavioral tests showed no advantage, at higher cost. Work on an evolving project and a structured decision task found narrower improvements, not a general solution to autonomous engineering. The practical lesson for us has been to separate three questions: did the artifact look better to reviewers, did it actually pass the behavioral checks, and what did it cost? An improvement on one is not automatically an improvement on the others. The evaluator also needs auditing: a hidden test that imposes an invalid requirement can produce a misleading treatment comparison. Disclosure limits I’m keeping the implementation private for now. The field note shares findings and limitations, not the instructions or executable benchmark fixtures. That limits what others can independently verify, and the historical results do not automatically validate every later revision. My conclusion is that Grand Coder has produced specific measurable benefits and is useful in our ongoing work. It doesn’t need perfect marks for those improvements to matter, but it does need honest reporting about where they occurred and where they didn’t. Field note: Grand Coder and more diligent AI engineering https://www.sheilastudios.com/field-notes/grand-coder-more-diligent-ai-engineering For others evaluating engineering agents: how do you report post-scoring test corrections and localized gains without either overstating the result or throwing away a useful finding?