Replay a fight with different starting conditions. Check that a purchase wasn’t charged twice after a write error. Find places where a Russian label doesn’t fit on a small screen. Development has plenty of tasks that benefit from being done carefully and the same way many times. AI can help our workflow substantially with this kind of work.
We use it to prepare code and scripts, analyze documents and reports, adapt text, and iterate on art. But every result still raises a question: what confirms that it’s right for this particular game?
Define the check first
If we ask an AI to “check the boss,” we might get a convincing report that doesn’t address the situation that matters. A specific scenario is more useful: choose a spore, survive its charge, take cover, and restore the fight from a save. Then we can see exactly what was checked and repeat the result.
In Guzzle, these scenarios help us trace a sequence of rules. In Last Road, they help us check a deal calculation and make sure state stays unchanged after an invalid command. In Smoke and Flame, they help us check that effect points line up with the car and that wheels are restored after repairs. The tasks differ, but the principle is the same: a check needs an observable result.
You have to see the image
Generation helps us create options and materials, but a file can’t tell us on its own whether it works in the game. A character can lose important features, a car’s construction can be wrong, or an interface can have the wrong proportions. That’s why we review the result in the context where it will be used.
The new pencil sheets in this workshop were also created with generative tools. We label them as illustrations for the retrospective. We don’t paint over real gameplay screenshots to make the story more convincing: they need to show the version being discussed.
The decision stays with the creators
AI can suggest an approach, write a draft, and help us check it. It doesn’t remove the need to choose a direction, judge the game’s character, and personally review the result. An automated fight doesn’t confirm that a person understood the prompt, and a successful export doesn’t prove the car looks good.
We don’t promise a measured percentage of time saved without our own measurements. The practical benefit is elsewhere: repeatable work becomes more accessible, and we can describe problems more precisely. We want to spend the attention we’ve freed up on decisions that actually change the game. For example, figuring out why a friendly spore looks suspicious and deciding how much of that suspicion to leave in for the player to enjoy their next bite.

