# AI 3D Character Creation — No GPU Required

> Source: <https://dev.to/plastikelectrik/ai-3d-character-creation-no-gpu-required-44mj>
> Published: 2026-09-27 22:53:29+00:00

**Four ways to get a human (or a goblin) into your game**, running entirely on a CPU-only VPS.

A while back I wrote about **CHARFORGE** and **3D GAME ENGINE** as two apps. 

As of this week, that framing is a little outdated — they're one system now, and the interesting part isn't "generate a face," it's that there are now four completely different ways to get an actor into the game, and the engine doesn't care which one you used — and none of it touches a GPU.

This post is about that unification, and about the theory + plumbing behind each of the four paths.

All of it runs on a 6-vCPU, 11 GB RAM VPS. No CUDA, no rented GPU instance, no cloud inference bill.

That constraint is the whole story, so it's worth saying up front instead of burying it in section 6.

The core idea: an "actor" is an interface, not a mesh

The single most important architectural shift in this update is that the game engine no longer thinks in terms of "CHARFORGE characters."

It thinks in terms of actors, and an actor is anything that can answer a small set of questions: how fast do you move, how do you play an animation, how do you die, how do you aim, where do you look, where does a weapon attach to your hand.

Once that interface exists, it stops mattering whether the thing behind it is:

A hand-sculpted CHARFORGE human

An imported GLB/VRM/FBX/OBJ model someone else made

A procedurally grown creature (goblin, robot, alien, four-legged beast)

A CPU-generated human built by Blender + MPFB2, or reconstructed from a single photo

All four get baked once, rigged, retargeted if needed, and dropped into the shared character library, where the game engine treats them identically.

That's the theoretical payoff of building around an interface instead of a data format: your roadmap stops being "support format X" and becomes "satisfy the actor contract," which is a much shorter list of problems.

Let's walk through the four paths.

**Path 1: Sculpt a human from morphs (the original CHARFORGE loop)**

This is the one I covered before, so briefly: a 19,158-vertex MakeHuman base mesh, 163 bones, deformed via morph weights for gender/age/muscle/weight/ancestry macros plus 136 extra body and facial targets.

Engine.weights(P) turns a parameter object into a weight vector; the skeleton uses Euler ZXY rotations solved at each bone head:

M = M_parent · T(h) · R · T(−h)

All of this — sculpting, skinning, the shader that adds pores and freckles — runs in JavaScript in the browser's CPU. No GPU compute shaders, no WebGPU requirement, just WebGL for display.

What's new here isn't the math, it's what happens after you finish sculpting: every character now lands in MySQL automatically, not just in browser localStorage.

That single change removes an entire category of "oh no I cleared my cache" disasters, and it unlocked something more interesting — versioning.

Every character keeps up to 15 historical versions (throttled to one new snapshot per 5 minutes), restorable, with a 30-day trash instead of instant deletion.

A character generator turned into a character database, which is a very different product even though the sculpting code barely changed.

**Path 2: Import someone else's model**

This is the "stop reinventing everyone else's work" path. You can drop in:

GLB / VRM files as-is — VRM keeps its own material pipeline (via @pixiv/three-vrm), because VRM materials are a whole ecosystem of their own and re-deriving them would be pointless.

FBX / OBJ, converted to GLB either in-browser or, when that fails, server-side via Blender running headless on CPU.

CC0 KayKit asset packs, added once per pack, contributing roughly 95 shared animation clips to the whole library.

The theoretically interesting bit is rig detection and retargeting. Nobody agrees on bone names.

KayKit, Mixamo, VRM/VRoid, Unreal/MPFB, MakeHuman, and raw Blender exports all name their skeletons differently, so the system does name-based humanoid bone detection across all of them, with a structural fallback (walking the hierarchy by shape rather than by label) for anything unrecognized.

Once a rig is identified, KayKit's ~95 animation clips get baked once per model in world space, with bone directions aligned, so both T-pose and A-pose rigs animate correctly without per-model manual retargeting.

There's a small, very "we hit this in production" fix worth mentioning: non-VRM imports were rendering semi-transparent because of how their materials were flagged.

The fix was blunt and effective — force solid materials on anything non-VRM, with alpha-test at 0.5 for textures that actually need cutouts.

Sometimes the correct architectural decision is "stop being clever about transparency for models we don't control."

**Path 3: Grow a creature procedurally**

The creature lab doesn't import anything — it generates robots, toon people, goblins, aliens, golems, mushroom folk, beast-men, and four-legged animals from type + seed + sliders, rebuilt live in the browser, no server round-trip at all. This is architecturally the cheapest of the four paths: no file I/O, no conversion, no rig detection, just deterministic procedural generation from a small parameter set — the same "seed in, character out" philosophy the world generator already uses for environments.

**Path 4: Ask a CPU to generate a human for you**

This is the one I'm most proud of, mostly because it sounds like it shouldn't work without a GPU, and yet:

Realistic human generation via Blender + MPFB2. A job queues on charforge-3d, a dedicated worker running Blender 4.5.14 headless with the MPFB2 add-on (23 skins, 20 outfits) — pure CPU rendering, no CUDA — and produces a fully rigged human mesh. This takes minutes, not seconds; CPU-rendered character generation always will. But it lands automatically in the library when it's done, whether or not anyone's still watching the job.

Photo → 3D reconstruction. Upload a photo, rembg strips the background, TripoSR reconstructs a 3D mesh from that single image running on torch in CPU mode, and an auto-rigger adds a skeleton. This is genuinely the most "we're living in the future" feature in the whole system, and also the most honest about its limits: single-image reconstruction gives you a rough body and an approximate face and hands. It's not going to fool anyone at close range. But as a "turn a photo into an actor you can drop into a scene in a couple of minutes, with no dedicated hardware" pipeline, it's a small miracle that it works at all.

Both jobs share the same discipline as the image-generation jobs from the original CHARFORGE pipeline: queue it, poll it, never hold an HTTP connection open for it, and the result gets committed to the library automatically instead of requiring someone to babysit the tab.

The local AI roster — three services, one CPU, zero GPUs

Service What it does    Hardware

Text — Ollama, llama3.2:3b    Character identity: name, age, occupation, personality, backstory, quote    CPU

Image — PyTorch + diffusers, SD-Turbo Garment fabric generation, AI portrait photos   CPU

3D — Blender 4.5.14 + MPFB2, TripoSR, rembg   Realistic human generation, photo-to-3D, format conversion  CPU

Three genuinely different AI modalities — language, diffusion, 3D reconstruction — sharing 11 GB of RAM and zero accelerators. Nothing here runs concurrently by accident: text, image, and 3D generation are explicitly taking turns, with Ollama unloaded whenever the 3D or image jobs need the headroom. That's the actual engineering answer to "how do you run three kinds of AI without a GPU" — you don't make each one faster, you make sure they never fight each other for the same memory at the same time.

**Why the library matters more than any single generator**

Here's the theoretical point I actually want to leave you with: the interesting engineering problem stopped being "how do I generate a 3D character" and became "how do I make four completely different generation methods — none of them GPU-accelerated — produce the same kind of citizen in the same database."

Every path — sculpted, imported, procedural, or AI-reconstructed — ends up as a row with a kind (human, model, creature), an author, a thumbnail, tags, folders, favorites, and now version history. The game engine's actor interface and the library's shared schema are really the same idea applied twice: define the contract once, and let every producer satisfy it independently. That's what let a photo-to-3D reconstruction — a wildly different pipeline from morph-based sculpting — show up in a game's cast list "like any other character," with zero special-casing in the game engine.

What's still rough (staying honest)

Photo-to-3D quality is genuinely rough — approximate faces and hands, this is not a scanning solution.

MPFB2 human generation and TripoSR reconstruction both take minutes and compete for the same RAM as the image worker.

Unusual imported rigs fall back to structural detection and can animate imperfectly.

The browser-side editor still does morph/skinning math in JavaScript on the CPU (the game engine already moved to GPU-skinned rendering for playback; the editor hasn't yet).

None of these are foundational problems — they're the normal tax on ambition when your hardware budget doesn't include an accelerator, and they're all on the roadmap.

**The takeaway**

If you're building anything that ingests content from multiple sources — user-generated, imported, procedurally generated, AI-reconstructed — the lesson generalizes past 3D characters entirely: don't build four separate pipelines that each know how to be "the" content type. Build one interface that any pipeline can satisfy, and let the diversity of sources become a strength instead of a maintenance burden. And don't assume any of this needs a GPU until you've actually checked — a 3B-parameter LLM, a Turbo-distilled diffusion model, and a single-image 3D reconstructor are all, deliberately, chosen because they tolerate CPU inference. A goblin from the creature lab and a human reconstructed from a photo have almost nothing in common under the hood — and in this system, they can stand in the same lineup, take the same bullet, and neither the engine, the database, nor a single GPU ever had to know they were made differently.
