From intent to image

Building the visual system behind interactive storytelling.

FableVerse lets users shape their own stories. Changing relationships, emotions, actions, characters, and narrative direction in real time.

My challenge was helping the image-generation system understand what those moments were supposed to feel like, not simply what objects needed to appear.Our approach was to develop a way to translate user language into cinematic direction.
Base reference character 1, unmodified
Base reference character 2, unmodified
Base reference character 3, unmodified

Reference images

Three fixed base characters, unmodified, used as the control across every test below.

What we were solving for: FableVerse stories can change with every user interaction. The image system needed to visualize those moments while preserving the user's emotional intent. Gen AI could often generate the correct subject, but miss the intended feeling. By translating intent into camera, pose, expression, composition, and lighting direction, we could give the system clearer visual instructions.

What the solve meant: We could translate a user’s emotional intent into repeatable visual direction, rather than relying on the model to interpret it on its own. 

User intent

"you're too close omg"

Meaning

Playful intimacy

Cinematic direction

Close-up ¾, blush. POV, high angle, forced perspective

Pose + expression

Slight recoil. Blushed cheeks, flustered smile

Generate → test

Same direction, three base characters

Validate

PASS

Generated result — character 1, playful-intimacy direction
Generated result — character 2, playful-intimacy direction
Generated result — character 3, playful-intimacy direction

Not "write prompt, get picture." Interpret → structure → generate → compare → evaluate → refine.

The result: a reusable visual language that could be applied across characters

"{{ row.phrase }}"

{{ row.meaning }}

{{ row.direction }}

Generated output
Generated output
Generated output
{{ row.result }}

{{ row.note }}

Failure is data.

Target — "wait I didn't expect that"

Expression ✓ — medium close, wide eyes, mild shock

Full override of the base pose ✕

Failing generation — hand from the base image persisted

Shown after the fifth regeneration attempt.

Observed behavior

Five generations, same result each time: the expression landed correctly, but a hand pose inherited from the base image never changed.

Diagnosis

This was no longer a wording problem. Repeated identical failures meant the system was under-weighting a pose override in favor of the base image — a model-adherence limit, not a prompt limit.

Next iteration

Isolate the variable → strengthen the override → regenerate → compare → document.

Resolved on the next character after a single regeneration once the override was strengthened.

A shared visual language

Using my skillset in webtoon/cinematography, I created a visual language that could be applied and repeated.

Close-up framing example

Close-up

Face fills the frame. Intensity, intimacy, emotional escalation.

Mid-shot framing example

Mid shot

Torso up, hands visible. Gesture clarity, relational tension.

Wide-shot framing example

Wide shot

Full body, environment visible. Space, movement, safety resets.

What this system enabled

01 — Translate

Natural user language became structured visual direction.

02 — Standardize

Cinematic concepts became a shared vocabulary for artists, prompts, and AI systems.

03 — Test

The same behavior could be evaluated across multiple characters and repeated generations.

04 — Diagnose

Failures could be separated into camera, pose, expression, or model-adherence problems.

05 — Improve

Testing created actionable feedback for prompt refinement and the broader generation pipeline.

Creative quality shouldn't depend on a lucky generation.

The goal was to make visual intent measurable, teachable, and repeatable.

AI Lab · Methods, Systems & Experiments

Creative judgment, translated into systems.

I use AI as a creative production tool. So my work focuses on turning visual direction into repeatable language, testing frameworks, and human-led workflows that improve quality without removing artists from the process.

Prompt Architecture  ·  Visual Evaluation  ·  Model Testing  ·  Human-Led AI  ·  Product Experiments

Anime illustration of a dark-skinned girl with braided black hair adorned with colorful beads, wearing a black shiny sleeveless unitard, gold hoop earrings, and gold cuffs, throwing a roundhouse kick, camera is looking up from below, action scene

Prompt sheet
Reference
Camera diagram Camera-angle / shot-composition diagram

I treat prompting as visual systems design.

Production prompt engineering is not about finding impressive adjectives. It's the process of translating creative intent into structured instructions, isolating failure variables, testing repeated outputs, and refining the system until the result becomes reliable.

{{ step.n }}

{{ step.label }}

{{ step.detail }}

Treating prompt as a hierarchy

"A lone swordsman crosses a ruined temple courtyard during a storm."

{{ layer.label }}

{{ layer.value }}

{{ layer.why }}

Validation criteria

{{ c }}

Covers built from this hierarchy

Married the Cartel Prince — cover MWS — cover Le Festival de Lumière — cover OG Idea — cover Quan Millz — cover Anime Gym Bro — cover LUMIA — cover Fourth Wing — cover Harry Styles fan cover God Game — cover

I brought cinematic shot language into prompt language.

My art-direction background taught me to think in shots rather than descriptions. Instead of telling a model to make an image "more dramatic," I define the camera, staging, perspective, movement, subject hierarchy, and light behavior responsible for creating drama.

"{{ row.subjective }}"

{{ row.structured }}

I convert abstract creative feedback into instructions a generative system can execute and a creative team can evaluate.

Dark elf — camera authority, mood lighting, environmental story

"Fairy standing in a forest" leaves the camera and light undefined, so the model defaults to a flat, centered shot. Naming a low cinematic angle gives the model a camera position instead of a guess, and "mist swirling" plus "glowing mushrooms casting violet light" turns the setting into a light source that models the mood rather than just describing it.

Sexy fireman — mid-stride motion, implied lighting, emotional posture

"Standing confidently" is a static pose with no information about force or intent. Capturing the figure mid-stride gives the pose a direction and momentum, "helmet glinting" implies a specific light source without naming one, and "jaw set with determination" gives the model a legible emotional target instead of a vague mood word.

Elemental mage — force of nature, active verbs, focal hierarchy

"Elemental powers around them" is ambiguous about what those powers are doing. Naming molten earth, coiling flames, and arcing water gives each element an active verb, and placing the mage "at center of roaring vortex" establishes a clear focal hierarchy so the composition reads correctly instead of scattering attention across the effects.

Dark elf before/after prompt translation Sexy fireman before/after prompt translation Elemental mage before/after prompt translation

Quality needs a shared definition.

{{ layer.label }}

{{ item }}

Seven problems, worked through.

Click a row to expand the full diagnosis, change, and validation behind each outcome.

{{ row.problem }}

{{ row.value }}

⌄

What I identified

{{ row.identified }}

What I changed

{{ row.changed }}

How I validated it

{{ row.validated }}

The experiment is only useful when it becomes a system.

My AI work sits between creative direction and technical implementation. I identify what strong visual output requires, translate those requirements into structured language, test them against real model behavior, and document what the wider team can reuse.

Explore selected work Discuss an AI creative system