3JSBench: A benchmark for LLM-generated 3D objects General Context Labs released 3JSBench, described as the first open-source benchmark for evaluating LLMs on generating coherent Three.js 3D assets, using 697 deterministic geometric checks across 100 tasks. The company, which runs the AI game-generation platform instaplay.ai with over 40,000 games, said models mostly build the right parts but fail to connect them, with almost half of attempts missing when one piece must rest on, hang from, or lean against another. In side-by-side comparisons, 1,000+ game creators cast 9,500+ votes and preferred Opus 5.5 slightly above Astra even though Astra passed more correctness checks. General Context Labs on X: "Frontier models are solving Millenium Problems but still struggle to build complex 3D objects. How do we even measure what they’re getting wrong? Introducing 3JSBench, the first open-source benchmark to evaluate LLMs on the ability to generate coherent @threejs assets." / X General Context Labs on X: "Frontier models are solving Millenium Problems but still struggle to build complex 3D objects. How do we even measure what they’re getting wrong? Introducing 3JSBench, the first open-source benchmark to evaluate LLMs on the ability to generate coherent @threejs assets." Frontier models are solving Millenium Problems but still struggle to build complex 3D objects. How do we even measure what they’re getting wrong? Introducing 3JSBench, the first open-source benchmark to evaluate LLMs on the ability to generate coherent @threejs assets. Frontier models are solving Millenium Problems but still struggle to build complex 3D objects. How do we even measure what they’re getting wrong? Introducing 3JSBench, the first open-source benchmark to evaluate LLMs on the ability to generate coherent @threejs assets. Having built instaplay.ai, an AI game generation platform home to over 40,000 games, we noticed a recurring problem: models can generate game mechanics for increasingly complex worlds, but 3D objects fail to meet the bar. We built 3JSBench to measure the commonShow more Unlike multimodal evaluations that rely on vLLM-as-judge screenshots, 3JSBench measures geometric correctness deterministically. We built a custom linker that identifies objects within generated scenes, allowing 697 deterministic checks across 100 tasks. One failed requiredShow more Models mostly build the right parts, but struggle to connect them properly. When one piece has to rest on, hang from, or lean against another, almost half of attempts miss. Intersecting geometry and flickering textures are also common. Failures like these make assets unusable inShow more Deterministic checks alone can’t tell you what looks good, so we collected 9500+ votes from 1000+ game creators who compare assets side by side. Users prefer Opus 5.5 slightly above Astra, despite Astra passing more of the correctness checks. You can explore every task, inspect the verifiers, and interact with the generated objects yourself. Reach out if you’re interested in learning more or collaborating: email protected