assetyardby Binary Helix

MAKING OF KNELLDEEP · EVERY NUMBER FROM A FILE IN THE REPOSITORY

Knelldeep, start to store, in one day

How one owner and an agent fleet built a voiced dungeon crawler in about a day. Every asset came from the Asset Yard command line, and every number below names the file it comes from.

Knelldeep key art: three heroes stand on a stone ledge across a sinkhole from an armoured warden, with a cracked bronze bell hanging in the dark.
The key art. One assetyard generate image request at seed 6001, 25.7 seconds, from the record image:key_art_title in assets/ledger.json.

This is a tutorial, not a brag sheet. The section on what went wrong is the longest, and it is the section worth your time.

1. What we built, in one screen

Knelldeep is a turn-based dungeon crawler with timed checks. Three heroes stand against three monsters. You press as a bar crosses a gold zone to strike, and hold a ring inside a band to parry. A last parry at the edge of death is the hook the design was built around.

The honest clock. Commit times come from git log --date=iso-strict. File times come from knelldeep-raw/, the download folder that git ignores.

MilestoneSourceWhenElapsed
Design document, nothing else659c3336c2026-09-15 03:35:18T+0:00:00
Voice cast chosenaudition file times03:53:26 to 03:59:12T+0:18:08
Content and tools committed8d273d77704:57:28T+1:22:10
583 asset files landedraw file times04:10:09 to 06:15:06T+0:34:51 to T+2:39:48
Playable start to finish970e66dd012:51:51T+9:16:33
Version code 2 on Playb765b33702026-09-16 00:56:35T+21:21:17

git log --oneline 659c3336c~1..b765b3370 counts 17 commits from the design to that release, and 13 with the four Knelldeep folders as a path filter. Section 10 of docs/knelldeep/DESIGN.md is headed "Week plan".

The asset run went overnight. The ledger at src/knelldeep/assets/ledger.json holds 562 records and 18,230.9 seconds of request time, which is 5.06 hours. File times put 583 raw files inside one window of 2 hours 4 minutes 57 seconds: 270 voice lines, 199 sound files, 102 pictures, 6 music tracks and 6 rooms. The tool runs three requests at once, so 5 hours of request time packs into 2 hours of wall clock. Nobody watched it.

Every record carries "attempt": 1, and the tool sets ATTEMPTS = 3, so a retry was there and no job used one. That is not the same as no retakes, and section 9 is honest about the difference.

Three heroes with a lantern, a rope and a bell shield face a Rattle Sentry and two Wick Worms in a candle-lit crypt.
Every room starts as a painted picture. One command rebuilt it as a walkable 3D scene, and the party stands inside it.
The same crypt in a fight. Health bars float over each unit and labelled intent cards show the next enemy move.
The monsters say what they will do before they do it. Drag hits Back for 15. Then you answer with a tap on a closing ring.

2. Why it followed two smaller games

Knelldeep is the third game on this site, and each earlier one answered a question that would have sunk the big build.

The 2D crawler at /demo/ proved the timing hook. Press as the needle crosses the gold. Hold while the rune ring is inside the band. That is the input model Knelldeep uses for every attack, heal and defence. Playing it answered what no design document can: whether a timed press is fun when a still picture is the only thing moving. It also proved that one catalog voice gives two performances by changing one number.

assetyard generate tts --voice VOICE_KEY --expressiveness 0.05 --text "Three floors down, the water stops moving."
assetyard generate tts --voice VOICE_KEY --expressiveness 0.95 --text "You should not have come down here!"

The 3D build at /demo3d/ proved rooms from pictures. Each of its three rooms is one generated picture, rebuilt as a GLB with a walkable floor and the camera that took the picture. That is the whole of Knelldeep's stage technology, tested on three rooms instead of six.

Both demos failed where we could still afford it. The 2D demo needed a per-asset key threshold for one figure, whose armour had near-black gaps against a black field. On three figures that is a line of config. On 74 pictures it is a broken pipeline. Prove the hook in a toy. Prove the stage in a toy. Then spend a night of generation.

3. What you need before you start

The commands below need three things on the machine. The tool names none of them in its help text.

The CLI, installed and signed in. run_job in sites/assetyard/tools/knelldeep_assets.py builds [sys.executable, "-m", "uncen_cli", ...]. The assetyard command in every ledger record is that same program under its friendly name. The quickstart has the full page.

python -m pip install https://uncen.ai/static/downloads/uncen_cli-0.11.1-py3-none-any.whl
assetyard login
assetyard balance

Three Python packages, for the pack step only. key_figure, key_glyph, key_frame and recompress_glb import numpy, PIL and scipy inside their own bodies. Generation runs without them. The pack step does not, and the pack step writes every file the game loads.

FFmpeg, at the path the tool hard-codes. The module constant reads FFMPEG = Path(r"C:\Program Files\ImageMagick-7.1.0-Q16-HDRI\ffmpeg.exe"). That is one Windows machine, not a setting.

Two more Windows habits are baked in. portable() strips a Windows repository prefix and nothing else, so on another operating system a machine path lands in the ledger. The Android scripts in section 6 are PowerShell.

4. The plan before any asset

The first commit is DESIGN.md and nothing else: 625 lines, one file, no code and no pictures, by git show --stat 659c3336c. It names four competing concepts, picks one, fixes the hook in a sentence, and sizes the content before anything is generated. Its estimates were not all right. Its own section 8 budgeted "about 80 x 2 takes = 160" pictures and "about 28 MB". The real run made 102 raw pictures, kept 74, and ships 41.5 MiB. Being wrong in writing beats being vague.

The content files are the single source of truth. After the design came a folder of JSON at src/knelldeep/content/.

FileHolds
prompts.jsonEvery picture: id, kind, prompt, size, seed, human review notes.
script.jsonAll 275 voice lines: id, speaker, text, expressiveness.
casting.json, cast.jsonWhat each speaker should sound like, and the voice the audition chose.
sfx.json135 cue ids, with prompt, duration and variant count.
music.json7 tracks, with prompt, duration, kind and lyrics.
rooms.jsonWhere figures, lamps and hotspots sit in each room.
facing.jsonThe figures the model painted facing the wrong way.

Commit 8d273d777 put 4,438 lines of that JSON into the repository in one go, before a single asset existed.

A manifest beats a pile of ad hoc requests. The game reads the plan, not the folder, so a file nothing names cannot ship. A diff is a review: when the room prompts were wrong, one git diff showed all ten changing the same way. And tools/knelldeep-content-check.mjs exits 1 on a missing line, a duplicate id or a banned word.

5. Step by step, with the real commands

Four commands drive everything. The docstring of sites/assetyard/tools/knelldeep_assets.py holds the first three. The fourth is a separate tool.

python sites/assetyard/tools/knelldeep_assets.py plan
python sites/assetyard/tools/knelldeep_assets.py generate --kind image --only room_nave_steps
python sites/assetyard/tools/knelldeep_assets.py pack --strict
python sites/assetyard/tools/knelldeep_casting.py

Read --only before you use it. It matches a job id or a kind:id key exactly, and it matches a prefix when the value ends in *, as in --only room_*. A value that matches nothing selects nothing, prints 0 jobs and stops, with no error. The docstring's own example is that case: --only nave_steps under --kind image selects zero jobs, because the image job id is room_nave_steps. Read the count that plan prints before you spend anything.

plan prints 527 jobs: 275 voice lines, 165 sound files, 74 pictures, 7 music tracks and 6 rooms. 521 of them ship. The six room pictures do not, because the room GLB carries its picture as a texture. The ledger stores each command the way a person types it from the repository root, and the tool adds --out, --overwrite and --json itself.

The ledger also records the image model behind each picture. The picture commands below leave the model out, so they run as printed on the server's current image model. A new run can paint a different picture.

5.1 Rooms first

Rooms came first because everything else has to match their light. The style block in prompts.json says so: "Every room puts its warm source at left and a cool teal source at right. One set of figures serves every room, so this layout matches the shared figure light."

Each room prompt is built from a shared template: a view sentence, a scene sentence, a finish sentence. The slowest of the six took 198.5 seconds.

assetyard generate image --prompt "Wide lens view from raised eye level, looking slightly down. Broad flat open floor across the lower half, clear on both sides. Empty foreground, deep background. Sunken cathedral nave, a wide flagstone landing. Shallow steps rise far off to a cracked rose window glowing teal. Tall iron candle stands burn at left. Cold teal moonlight falls through a tall window at right. Toppled pews line the walls. Painterly, cold teal and bell-bronze palette, soft edges, no ink outlines. No people, no text, no logo, no watermark." --width 1664 --height 928 --seed 1102

Read the first two sentences again. They are not scenery. They are the requirements of the 3D rebuild, written as a picture. A raised viewpoint keeps floor in frame. A broad flat floor gives the walkable grid something to cover. An empty foreground keeps a pillar out of the camera's lap.

5.2 Rooms to 3D

One command per room. The picture is the input.

assetyard generate environment --reference-image sites/assetyard/knelldeep-raw/images/room_nave_steps.png

All six GLBs landed between 05:13:15 and 05:13:50 by their file times: 35 seconds of downloads for six walkable rooms. The requests took 542.7 seconds, median 88.4 seconds. One function, source_shas, records the checksum of the source picture in the room's ledger record, so a new picture makes the room stale and it runs again.

5.3 Figures on a green screen

Every hero, monster, portrait and NPC is a still picture with a transparent background. The model paints it on a flat key colour and the pack step keys it out. The key sentence is not written by hand: the tool strips any black-field wording out of the stored prompt with a regular expression, then appends its own.

assetyard generate image --prompt "Full body game character art, head to feet: a broad grey-bearded older bell-founder, scorched leather apron over dented bronze scale armor, bronze-headed hammer, a cracked bronze chapel bell strapped to his arm as a shield, standing guard, bell shield raised, hammer low, facing three-quarter toward the right. Warm key light from upper left, teal rim light from right. Painterly, teal and bell-bronze palette, soft edges, no ink outlines.  No text, no logo. On a solid flat pure green screen background, #00FF00, evenly lit, no shadow, no floor, no gradient." --width 1024 --height 1536 --seed 2101

That is one figure standing still. At 246.9 seconds it is the slowest asset of the build. Icons and UI frames kept a black field on purpose, because for a glowing glyph you want brightness to become alpha.

assetyard generate image --prompt "Single game icon: a notched short sword pointing down on a slant. Bold simple shape readable at small size, centered, bell-bronze and pale teal with a soft glow, painterly. On a plain pure black background. No text, no letters, no border." --width 512 --height 512 --seed 4001
The party screen. Three heroes stand in the line, Front, Middle and Back, and a fourth waits on the bench.
Four hero portraits and twelve hero poses, all keyed off a flat green field. assets/manifest.json lists them as portrait_* and hero_*.
The Knelldeep title screen: a camp at the edge of the sinkhole, with a fire, an easel and a lantern-carrying balladeer.
The title screen is a room like any other. The camp picture became a GLB, and the menu sits in front of it.

5.4 The cast, by audition

knelldeep_casting.py ran for 5 minutes 46 seconds, from 03:53:26 to 03:59:12 by the audition file times. Per speaker it searched the catalog, dropped anything a filter forbids, generated one audition line with each of the four best candidates, measured the median speaking pitch, and scored each one.

Read fit() before you copy it, because it scores and it does not reject. A pitch outside the band in PITCH_RANGE, 80 to 150 Hz male and 165 to 260 Hz female, adds 100.0 plus the distance to the band. A party voice closer than MIN_PITCH_GAP, 12.0 Hz, to a party voice already cast adds 40.0. Then min(heard, key=fit) takes the lowest score, whatever it is. Both rules are preferences, and the shipped cast breaks the second: cast.json gives Oswin 119.4 Hz and Corvin 113.5 Hz, both male, both in PARTY, 5.9 Hz apart. Corvin's other three candidates measured 166.7, 152.4 and 200.0 Hz, all outside the male band, so the closest voice still won on score.

The filters do reject, and they run before any audition. CHILD_PATTERN drops child, youth, kid, teen, girl and boy, with the comment "No role in the game is a child". HUMAN_ONLY_EXCLUDE drops aliens, monsters, robots and mechs for every non-monster role. TAKEN blocks the three voices other Asset Yard demos use.

Nine roles were filled from 33 auditions, and 27 survive in cast.json. The tenth voice, the Warden, carries "method": "owner choice". The kept auditions run from 85.1 Hz to 326.5 Hz, and the nine cast voices from 102.6 Hz, Pell, to 207.8 Hz, the Gloom Stalker. A casting decision you can re-hear is one you can argue with.

The camp menu, titled The Balladeer's Fire, with 95 Tolls to spend and cards for healing, Charms and hero conversations.
At the Balladeer's Fire you spend Tolls on healing, on a Charm, or on nothing at all. Talking to your heroes is free, and every line is generated.

5.5 275 voice lines

One command per line. The voice key comes from cast.json, so no line names a voice by hand.

assetyard generate tts --voice ayhf2_anchor_veteran_c1 --expressiveness 0.90 --text "Not today. Not this bell."

The ledger records 7,269.3 seconds for the 275 lines, which is 2.02 hours, median 17.5 seconds. The manifest voice durations sum to 678,480 ms, which is 11 minutes 18 seconds, over ten speakers and 1,638 words.

5.6 Sound effects and music

199 sound requests, 6,196.0 seconds, median 17.7 seconds. Short cues are two seconds. Room ambience is thirty seconds and has to loop.

assetyard generate sfx --prompt "One bright heavy bronze bell struck hard with a ringing shimmer, dry close recording, no background music." --duration 2
assetyard generate sfx --prompt "Empty stone cathedral room tone, slow echoing drips, faint candle flutter, distant hollow wind, steady and loopable, no background music, no voices, no sudden sounds." --duration 30

The negative half of that second prompt does real work. "No background music, no voices, no sudden sounds" is what stops a bed from turning into a scene.

plan asks for 165 sound files today, not 199. The other 34 ledger records are cues the design later dropped, such as the crit impact variants. The ledger keeps them and the plan does not name them, which is the same silent orphaning that section 9 counts for pictures.

The ledger holds 8 music records and 688.2 seconds. The seven still in the plan cost 583.1 seconds, and six of those are instrumental.

assetyard generate music --prompt "Dark fantasy exploration ambience in a drowned cathedral, distant muffled bells, slow low strings, soft organ swells, sparse water percussion, wandering and unresolved" --duration 120 --music-kind ambient --instrumental

The seventh is sung. Its command takes the same shape with three more flags: --music-kind song, --vocal, and --lyrics, whose value is the whole lyric sheet as one argument, with [verse] and [chorus] markers and an escaped newline between lines. That value runs to two verses and two choruses, so read the command in the ledger record music:ballad_of_knelldeep rather than here. The 150 second track took 74.8 seconds to make.

5.7 The pack step

pack turns raw downloads into the files the game loads. It deletes every packed folder first and rebuilds it, so a file belonging to a deleted asset cannot survive into a release. It keys the figures off green, keeps the largest shape, removes the spill and crops to the alpha box. It saves WebP, and it shrinks a figure only when the crop is taller than its cap: FIGURE_MAX_HEIGHT is 1024 px for a hero or a monster, PORTRAIT_SIZE is 512 px for a portrait. It keys icons and frames off black. It normalises the audio to MP3: voice at 64 kbps and -18 LUFS, sound effects at 96 kbps and -16 with leading silence trimmed, music at 128 kbps and -20. It rewrites each room GLB with a JPEG texture. It flips no pixels: for each figure named in facing.json it writes "mirror": true into the manifest, and the stage draws that figure flipped at run time.

Then it writes manifest.json with every shipped file, its size and its duration. With --strict it exits 1 on any problem. The packed game holds 521 files: 68 pictures, 6 room GLBs, 275 voice lines, 165 sound files across 135 cues, and 7 music tracks. Those files hold 43,538,422 bytes, which is 41.5 MiB.

6. From the packed folder to the store

pack gives you assets. It does not give you a game you can install. The route to Google Play is a second project, sites/knelldeep-app/, and docs/knelldeep/ANDROID.md is the whole of it. Section 2 holds the build. Sections 3, 4 and 5 hold the keystore, the upload, and the steps only the account holder can do.

Knelldeep ships as a Capacitor app, and sites/assetyard/src/knelldeep/ is the single source. npm run build copies app/, core/, data/, content/ and assets/ into dist/, then adds the vendored three.js and one standalone full screen page. Nothing is fetched at run time, so the game plays in aeroplane mode.

cd sites\knelldeep-app
npm install
powershell -ExecutionPolicy Bypass -File tools\install-android-toolchain.ps1
npm run build
npx cap sync android
npm run android:build

The toolchain lives under ~/android-tools and stays off the machine's PATH, so each command points at it itself. npm run android:apk gives a debug APK for a phone over USB, and npm run android:build gives the release .aab for Play. tools/android-build.ps1 runs the web build, the sync and Gradle in one go, then runs jarsigner -verify and throws when the bundle is not signed. Look at the packed game first: node tools\serve.mjs serves dist/ from the root of an origin, which is what the webview gives the game.

7. The two tool ideas worth copying

Regenerate when the command changes. is_current is true only when the raw file exists, the stored command equals the command the tool would run now, and every input checksum matches. Edit one word and that asset runs again. Edit nothing and a re-run costs nothing.

A dependency graph, not a script. A room job declares depends_on: ["image:room_nave_steps"]. A stale picture marks its room stale, and the runner holds a job back until its inputs are done. That is what lets one generate call rebuild a picture and its room in the right order.

8. What went wrong

This is the useful part. Every item is in the history.

A black field ate dark armour. Figures were first painted on black and keyed by brightness. Dark armour touches black, so the flood fill ran into the figure, and no threshold keeps dented bronze scale armour while it drops a dark field. The fix is a flat pure green screen and a colour difference key, g - max(r, b). Icons kept black. The stale wording still sits in the review notes of prompts.json: "Keep a clear gap of black around the figure". Comments rot.

The green key needed three guards. A model does not always paint the field you asked for. The key can split a raised weapon off a body. A floor shadow survives as a speck. So key_figure measures the median keyness of a 12 pixel border and raises "has no clear green field" below 20, instead of shipping a grey box. It keeps every connected part at or above 2% of the largest, so a raised hammer stays and specks go.

The rooms came back full of people. The shared template said "room for three figures each side", and the model painted three figures. The trailing "No people" did not undo it. The phrase became "clear on both sides", and the intent moved into the review notes, which the generator never sees. git show cbb524398 changed ten room prompts, and that phrase is the only change in each.

The Candle Leech came back wrong. The commit body says only "The Candle Leech is a slug.", so the exact fault is not on record. The change is. The first prompt described the creature by what it lacks, "a long limbless segmented worm body and no legs", in a tall 1024x1536 frame. The fix changed three things at once: landscape 1536x1024, a positive posture, "a long fat slug-like body lying low along the ground in an S curve", and a new seed, 3114 to 3131.

The Gloom Stalker read as a film alien. It was fixed in the working tree before the first content commit, so the offending prompt was never recorded. The prompt that works names the silhouette and the material in full, down to "a bald cracked grey head with a sewn-shut eyeless face".

Eight voice lines broke at high expressiveness. The model lost the cast pitch on short, emphatic and dying lines. Six of the eight sat at 0.85 or 0.90, such as "The bell holds." The other two sat at 0.70 and 0.60. Every one dropped by exactly 0.30, same voice, same text. The script still holds 40 lines at 0.85 and 12 at 0.90, so this was a repair, not a retreat.

The ledger key collided. status_stun is both an icon and a cue, and keyed on the id alone one silently overwrote the other. The key became kind plus id. All 28 orphan pictures in the raw folder carry an icon_ prefix that no id uses now, which looks like an earlier defence against the same collision. No commit shows that prefix, because the ids are already bare in the first content commit. Read it off the filenames, not off the history.

Four bugs were game code, not assets. Section 12 of docs/knelldeep/ARCHITECTURE.md holds all four. A fallen hero left a gap in the line, so a lone healer in Front healed the monster's bites for hundreds of rounds: one rule, closedRanks(slots), took 42 stalls in 8,640 bot fights down to 17, and 15 in 7,200 two-hero fights down to 0. That rank change broke every saved fight log, so 99 of 102 bot logs threw part way through and every log now carries a version stamp. Every 3D room showed dark cracks along its depth edges until one fill plane replaced a dark backdrop copy. And the caster's timing check crossed in 400 ms, against 1,200 ms for the tank, so she was impossible at 30 fps.

Figures faced the wrong way, and nothing automated caught it. The prompts say "facing three-quarter toward the right" and the model obeys about half the time. Nothing in the pipeline can read a picture's facing, so a person looks at every keyed figure and lists the wrong ones in facing.json. Six of the twelve hero poses are on that list. One hero's three poses were added last, in af2bb12ef at 23:48:32 on the build day, T+20:13, after the owner saw him fighting away from the enemy line.

The store listing showed an interface the game no longer had. The party screen's swap button became drag and drop in the commit that packed the Android app, so the screenshots were older than the build. docs/knelldeep/ANDROID.md now carries the rule in bold: take the screenshots again after any change to a screen the listing shows.

Two upload rules cost a release each. Google Play keeps a used version code forever, across a deleted app, and the first upload went to an app that was then deleted and recreated, so the release went out as code 2. Gradle also emits an unsigned .aab without complaining, and Play is the first thing that objects, so tools/android-build.ps1 now runs jarsigner -verify and throws.

Two browsers refused the game. The preview server served .mjs without a JavaScript media type, so Chrome rejected every module. That fix was one commit, 70 seconds after the commit that added the game code. On Android a WebView suspends the audio context on a call or an alarm, so the game now resumes a suspended context on every gesture, on visibilitychange, and on pagehide and pageshow.

The casting tool auditioned children for adult roles. A catalog search on descriptive terms matches words, not suitability, and pitch alone cannot tell a child from an adult. Three filters now run ahead of the audition. The evidence survives on disk: 33 audition files, including child voices under a hero role and a giant under an old ferryman, against 27 kept.

What we would do differently. Write the ledger append-only, with the take number in the key. Start figures on a green screen, and carry the reason for a technique when you carry the technique. Fix the ledger key before you generate. Measure the timing checks on a slow machine on day one. Take the store screenshots last, always. Record the model bake-off.

9. What it cost, and what we cannot tell you

Time, measured. The ledger records 18,230.9 seconds across 562 records, which is 5.06 hours. The 521 records that ship hold 16,910.3 seconds, or 4.70 hours. The median request was 18.0 seconds. Read these as request time, not GPU time: the tool starts a clock, runs the whole CLI subprocess and stops after the download, so queue wait and transfer sit inside the number.

KindRequestsTotalMedianSlowest
Voice2752.02 h17.5 s207.8 s
Sound effects1991.72 h17.7 s180.5 s
Pictures740.98 h27.7 s246.9 s
Music80.19 h82.8 s134.0 s
3D rooms60.15 h88.4 s158.0 s

Credits, as far as the repository states them. A sound effect is 1 credit per request, and so is a music track at every supported duration. A 3D environment is 0 credits during the free beta. So the 199 sound requests cost 199 credits, the 8 music tracks cost 8, and the 6 rooms cost nothing. Pictures and voice are priced from database rows, not from code. A voice line is quoted as round(max(1, len(text) / 12.5) * credits_per_second). At 1 credit per second the 275 lines come to 678 credits, and at 2 they come to 1,362. We did not read the database for this article, so we will not tell you which.

The one cost figure in the repository, with three caveats. docs/knelldeep/ROADMAP.md section 4 says: 560 generations for the shipped game, every one on attempt 1, 5.02 GPU hours in all, and "At the 1000 credit pack price of 5 USD, about 2.80 USD of generation for a fully voiced game." That count is already behind the ledger, which holds 562 records and 5.06 hours. The arithmetic assumes one credit for every chargeable generation, and the repository's own quote code prices voice per second. The $5 pack is also not on sale: the pricing page calls it planned and says credit-pack checkout is not open. Treat 2.80 USD as the floor of an estimate, not a receipt.

What we honestly cannot say. No GPU-side timing exists for these requests, and no record says which host served each one. The music model's licence is listed only as a per-provider record. And the bake-off that chose the image model left one comment, which says the winner "won the speed and look test on 2026-09-15", with no timings, no scores and no sample pictures.

The ledger flatters the build. run_job calls the CLI with --overwrite, and save_record replaces the record under the same key, so a second take destroys the first take's record and file. The ledger cannot show a retake, and "every one on attempt 1" means no CLI retry, not no second try. The honest numbers come from elsewhere:

Anyone copying this pipeline should write the ledger append-only, with the take number in the key.