Coderefercoderefer
CoursesWebinarsBlogAbout
Coderefercoderefer

Discover. Learn. Automate. Grow.

Learn

  • Courses
  • Webinars
  • Blog
  • Search

Company

  • About
  • Contact
  • FAQ

Legal

  • Privacy Policy
  • Terms of Service
  • Refund Policy
© 2026 Coderefer. All rights reserved.
HomeBlogHow to Build a Full Game With GPT-6 Astra (39-Hour Runbook)
How to Build a Full Game With GPT-6 Astra (39-Hour Runbook)
ai-toolsSeptember 20, 20266 min read

How to Build a Full Game With GPT-6 Astra (39-Hour Runbook)

A solo dev built PaperRoute, a finished Paperboy-style browser game, with GPT-6 Astra, Blender and Meshy in 39 tracked hours and 1.56B tokens. Here's the 6-step process he used.

V

Vamsi Tallapudi

Manager, Architect Technology at Cognizant

ai-tools openai gpt-6-astra game-development 2026
Share:

You can build a finished game with GPT-6 Astra, but not in one prompt. Solo developer Emm Tee (@builtbysketch) shipped PaperRoute, a Paperboy-style browser game, using Astra for code, Blender for models and Meshy for characters. It took 39 tracked hours, 1.56 billion tokens and a six-step process he's now published.

My take: the token count is what got the clicks, but the process is the part worth stealing.

What Is PaperRoute?

PaperRoute is a delivery game you play in a browser tab. You ride a bike through an American neighbourhood, throw newspapers at mailboxes, keep dogs off your back wheel and try to finish the week high on a league table. It lives at paperroute.lol, and even the landing page and leaderboard are styled as newspapers.

A demo video went viral on X first. Then he posted a long article with the whole runbook, and that's what I'm working from here.

What Did It Actually Cost?

He tracked everything with DevClocked, from the first planning commit on July 1 to September 12. The numbers match the public PaperRoute devlog:

MetricNumber
Tracked time39 hours (25.2 human, 13.8 agent-only)
Tokens1.56 billion (1.53 billion cached reads)
API value$2,175 (not his actual bill)
Commits90 across 11 days

Read that table carefully. 1.53 billion of those tokens are cached reads, which are cheap. And $2,175 is what the usage would cost at API rates, not what he paid. So "$2K to build a game" is the wrong headline.

He also admits presence tracking only started on September 6, so earlier hours count him as at the keyboard by default. Fair play for saying it.

How Do You Build a Full Game With Astra?

Here's his order, condensed:

  1. Write your own brief. He rewrote his July brief instead of reusing it, and picked one reference close to what he wanted: the original Paperboy. Clone the mechanics first, spin it later.
  2. Mechanics before art. Brief plus two or three rough images, and art direction held back. Six commits before midnight got him a browser build with aimed paper throws, mailbox scoring, a park course and tests.
  3. Split engine and style into separate streams. When a house looks wrong, you don't want to argue about throw physics in the same conversation.
  4. Script Blender from Python. Every prop is a script that builds the mesh and exports a GLB. Seven house families by 10am on day two.
  5. Use Meshy when a character needs a face. More on that below.
  6. Review renders on every change. Then save a day for flourishes.

The part I like: nothing here is magic. It's a normal production order, with an agent doing the typing.

Why Did He Bring In Meshy?

Astra is fine at houses and simple props. Characters broke it. His rider looked like a wooden puppet, Astra couldn't build a face, and more tokens got him nowhere.

The fix was a small pipeline. Concept images in ChatGPT, then Meshy for about $8 (roughly 300 generations, and he's still under quota), then Astra to simplify, rig and fit the mesh to the bike.

That's the honest lesson for me. Knowing when to stop prompting the model and reach for a different tool is a skill, and burning tokens on a problem the model can't solve is the most common mistake I see with coding agents.

The catch: the rider is 88,550 triangles because he kept the face, and he hasn't tested that on a real phone yet.

Why Are Review Renders the Actual Job?

This is the idea I'd steal first. Early on he built a harness where the agent renders its own work, spots what's wrong and sends the renders back. He says 60-70% of the improvement between render one and render two came from the agent doing this itself.

Astra set up capture scripts that put the Three.js rider through steering, throw, sprint and fall poses, saving front, side, rear and clay turnarounds each time. When something's off, you point at the frame ("the cap doesn't cover the hair") and that's your prompt.

It works for any visual output from a coding agent, not just games. Our post on Higgsfield's hell grind landed on the same shape: generate a lot, review hard, keep little.

Is One-Shot Game Generation Enough?

No. And he says so himself: one-shot games lack detail, and taste is still the moat. About 80% of his session time went into fixing bad meshes and small details, not building the game.

That matches what you'd expect. The window-smash camera got its own branch. So did a full rainy day with puddle splashes and tyre marks. He lengthened the skate park so the route ends on something fun instead of just stopping. None of that is the model. That's a person deciding what's worth a second play.

Test risky ideas in their own branch before you fold them in. Boring advice, and it's right.

What Should You Be Skeptical Of?

I'd hold a few things loosely:

  • The name. His tweet says "GPT 5.6 Astra". The article, the devlog coverage and OpenAI's launch all say GPT-6 Astra, so that looks like a typo.
  • The dollar figure. It's API value, not a bill. It says nothing about what a Plus or Pro subscriber would pay.
  • Performance. Sustained 60fps on a phone isn't proven. One run held it, another averaged 58.
  • Sample size. One developer, one game, a strong eye for art direction. I haven't rebuilt anything with this workflow, so this is his process and his numbers, not my test results.

What's the Takeaway?

Astra launched in early September with "welcome to the AGI era" energy. This build is a useful counterweight. The model did the typing, and the human still did the briefing, the reviewing and the throwing-away. That's not a knock on Astra. It's just what shipping something looks like.

If you're picking a model for this kind of work, compare notes with Claude Opus 5 vs Fable 5. And for OpenAI's earlier push into agents, see ChatGPT Work.

I'm going to try the stream-splitting and render-review loop on a small project before I believe any of it works for me. If it does, that loop is what I'd keep.

Official Post

The full article and devlog are linked from paperroute.lol if you want every checkpoint and render.

Get AI tricks that save you hours every week

New AI tools, automation workflows, and course drops — straight to your inbox. Join 2,400+ builders.

Read Next

Claude Cowork and Chat Are Merging Into One App
ai-tools

Claude Cowork and Chat Are Merging Into One App

Anthropic is merging Claude Cowork and Claude chat into one app, rolling out to Pro and Max subscribers over the next few weeks.

September 17, 20264 min read
Gemini 3.8 Live Wants You to Stop Typing and Just Talk
ai-tools

Gemini 3.8 Live Wants You to Stop Typing and Just Talk

Gemini 3.8 Live and 3.8 Live Extended Thinking add real-time voice, camera vision, and live narration to Google's AI. Here's what's different.

September 16, 20264 min read
August 2026 AI Model Releases: Every Launch That Actually Mattered
ai-tools

August 2026 AI Model Releases: Every Launch That Actually Mattered

August 2026 brought over 20 AI model launches, from Grok 4.6 to GLM-5.3-Flash. Here's what shipped, what got corrected, and what's actually worth using.

September 1, 20267 min read
View all posts →