Experiment

I Let a 27B Local LLM Build a 3D Game for an Hour. No Hands On.

A local Qwen 27B model at Q4 on one RTX 5090 drove the Pi coding agent for an hour and shipped a playable 3D game, tests and deploy included.

By JimmySep 30, 2026 4 min read
Illustration of a local AI coding agent building a floating-island platform game
AI generated editorial illustration

TL;DR: Qwen3.8-27B-GGUF at Q4 quant, running locally on an RTX 5090 through Unsloth, drove the Pi coding agent for about an hour with zero intervention and shipped a complete, playable 3D platformer — physics, art direction, UI, mobile controls, its own test suite, and a Cloudflare deploy. The whole thing. Here's how I set it up, and what surprised me.

The experiment

I wanted to answer one simple question: how good are local LLMs at real agentic coding right now? Not benchmarks, not "write me a fibonacci function" — a full multi-file project where the model has to plan, write, test, debug, and ship on its own.

So I framed it as an experiment with three parts: find a good prompt, build the local inference stack, hand it to an agent — then walk away and see what comes back.

Step 1 — finding the prompt

First I asked ChatGPT to search for the best prompt for something that's a simple yet nice 3D game. It turned up this one — the Jelly Jungle brief: a pink jelly character crossing 13 floating islands using triple jumps and spring mushrooms.

Honestly, it's a perfect test case. Simple enough that a small model can hold it in its head, but rich enough to exercise everything: game physics, procedural 3D world building, camera work, UI, mobile touch controls, even an asset-replacement pipeline. And the brief itself is beautifully written — it reads like a design doc, not a vibe.

Step 2 — the local stack (Unsloth)

I set up Unsloth, pulled Qwen3.8-27B-GGUF in the UD-Q4_K_XL quant, and then fiddled with the settings until I found what I think is the best balance:

Setting Value
Context length ~128K
KV cache dtype q8_0
Speculative decoding MTP (multi-token prediction)
VRAM budget 95% — capped so it won't crash my other running apps
GPU memory ≈ 25.5 GB on an RTX 5090

One takeaway from the fiddling: ~128K context is the sweet spot for small agentic coding projects like this — big enough to hold the whole project in its head, small enough to stay fast.

Step 3 — the agent (Pi)

Then I set up Pi, the coding agent harness, pointed it at the local model, dropped the Jelly Jungle brief in front of it, and hit start.

And then… I walked away.

The run: ~1 hour, zero intervention

This is the part that surprised me. For an hour plus, the model worked without a single nudge from me:

  • Scaffolded the whole project (Vite + Three.js + vanilla JS — no game engine)
  • Wrote fixed-timestep game physics: triple jump, spring mushrooms, spinning candy-bar obstacles, crumbling islands, checkpoints
  • Built the entire 3D world procedurally — 13 islands, cloud sea, palm trees, crystals, a finish portal
  • Matched the reference UI: title screen, HUD, jump meter, pause/help dialogs, mobile joystick + bounce button
  • Wrote its own headless test suite (Playwright + software GL) and verified every mechanic before calling it done
  • Deployed the whole thing to a Cloudflare Worker

The end result:

And if you want to compare, the original Jelly Jungle was built with Astra — see this post. Side by side, it's honestly impressive how close a local 27B gets.

The verdict: this is just a 27B model!

Let that sink in for a second. This isn't a frontier API model with a hundred billion parameters and unlimited context. It's:

  • 27B parameters
  • Q4 quantization (UD-Q4_K_XL)
  • q8_0 KV cache
  • Running on a single RTX 5090, ≈ 25.5 GB VRAM

And it produced a complete, polished, playable game in about an hour of autonomous work. Very capable. Local LLMs are not where they were even six months ago.

What's next

Two things:

  1. Next, I'm going to run this same model through our Cowork module — our own agentic harness that we're building and releasing soon. It handled Pi surprisingly well on its own, so I'm really curious how it performs inside our proper harness.
  2. For production, I don't think Unsloth is the right fit — it's great for dev and experimentation, but for a real serving setup I'd rather stand up a dedicated machine running SGLang or vLLM. That's the plan.

Stay tuned!