Devlog #2 — Meet Aric, Mira and… Whoever These People Are

In our first Devlog, we discovered that giving an AI the ability to improvise a world is relatively easy.

Getting it to remember that world is considerably harder.

Apparently, pictures have the same problem.

Meet Aric, Mira and Dawn

Every Dungeon Master needs some characters to make their life unnecessarily difficult.

Ours currently has these three.

Aric is our human knight and player character.

Mira is an elven scout travelling with him. She has already demonstrated the useful ability to notice things Aric doesn’t — which also means our Dungeon Master has had to learn that sometimes she should roll the dice, not him.

And Dawn is Aric’s horse.

Dawn should theoretically be the easiest of the three for an AI to understand.

She is a horse.

This assumption turned out to be optimistic.


Aric

Aric. Human knight. This is what he is supposed to look like. Please remember his face. It will become relevant shortly.

Mira

Mira. Elven scout, archer and companion. Also, definitely this person.

Dawn

Dawn. Horse. So far, arguably the easiest member of the party to recognise consistently.

Then we asked ComfyUI to illustrate the adventure

This is where things become interesting.

AERIS can take what is happening during the adventure and ask our local ComfyUI setup to turn the scene into an image.

The results can be surprisingly good.

They can also be surprisingly confident about things that never happened.

Aric and Mira during the adventure. Probably. At least there are two people.

The problem isn’t really generating a nice-looking fantasy image.

Modern image models are already rather good at that.

The difficult part is something much less spectacular:

Making Aric still look like Aric.

And Mira like Mira.

And keeping their equipment, clothing, position and even the number of people in the scene reasonably connected to what is actually happening in the game.

Apparently, “Aric standing next to Dawn” and “Aric riding Dawn” are two entirely different concepts.

Who knew?

Character consistency is more of a suggestion

And if we let the process wander far enough, eventually we get something like this:

Aric, Mira and… we’re going to need to check the guest list.

It’s a perfectly respectable fantasy image.

There is only one minor problem.

We have absolutely no idea who these people are.

Somewhere between the game state, the scene description, the image prompt and the diffusion model, our cast appears to have been quietly replaced.

Nobody informed the Dungeon Master.

Pictures need a memory too

And this brings us back to exactly the same problem we encountered with the Dungeon Master itself.

Something being plausible isn’t enough.

A new image can be beautiful, atmospheric and completely believable as fantasy art while still being wrong for this particular adventure.

So we’re experimenting with more consistent character descriptions, explicit scene information and a small quality-control step that can check whether the generated picture actually resembles the world AERIS intended to illustrate.

We’re not trying to make every image identical.

Variation is part of the fun.

But slowly, we’d like Aric, Mira and Dawn to stop feeling like three prompts and start feeling like the same three characters travelling through the same world.

The world can be improvised. It shouldn’t be randomly replaced.

Which, now that we think about it, is pretty much what we’re trying to teach AERIS in the first place.

Dawn, meanwhile, would presumably just like everyone to stop using her for software testing.


All images in this post were generated locally with ComfyUI.

RTX 2070 status: still alive.

Leave a comment