Academy / Characters that stay themselves / Lesson 1

Free lesson · 8 min

Building a dataset from one face

Twenty-four pictures of a person who does not exist, consistent enough to train on, made from a single generated selfie in about twenty minutes. The exact recipe behind every Gmanski character.

Every character on this site started as one picture. Not a photo of anyone — a Krea 2 render of a described person — and from that one picture we build a training set of twenty-four images that agree on the face. This is the free lesson of the course because it is the step people get wrong most, and the step the trainer cannot fix for you later.

Step 1: one cover, one clean face

Write the person as a sentence you will reuse everywhere: "a stunning 26-year-old woman with long sun-bleached wavy blonde hair, bright blue eyes, a light tan, smooth clear glowing skin, a slim athletic figure and a soft easy smile". That sentence is her "look", and it goes into every prompt from now on. Render a cover with it: a candid selfie, front camera, natural light, face fully visible.

Then crop her face out of the cover. Not the whole picture — a tight head crop: roughly 0.85 face-widths either side of the centre, from the hairline down to about 2.4 face-heights (chin, neck, the top of the shoulders). We measured this: the tight crop improved likeness in every identity test compared with feeding the full selfie. The model should see a face, not a background.

Two things must not be in that crop, because the training set will inherit them: props and clothing with character. A sun hat ended up on a third of one character's dataset; a hotel bathrobe on half of another's. We re-rendered both covers without them. If the cover has an accessory, generate a plain studio portrait of the same person (black tank top, grey background) and crop the face from that instead.

Step 2: twenty-four scenes, three distances

A LoRA generalises when the dataset varies everything except the person. Our fixed scene list:

That gives varied lighting (window, golden hour, overcast, neon, candle, studio), varied expression (laughing, serious, neutral, smiling), both profiles, hair up and down, and clothing from swimwear to coats. The person is the only constant.

Step 3: scene first, then identity

Each dataset image is made in two passes, both on serverless:

  1. Text-to-image of the scene with her look sentence: A stunning 26-year-old woman … , close-up portrait, head turned three-quarters to the left, golden hour sunlight outdoors, blurred trees. adult, mature face, natural photograph, clear smooth skin, sharp focus on the face. Sizes: 896×1024 for close-ups, 832×1040 half-body, 768×1152 full-length.
  2. Identity edit of that picture with the face crop: the Krea 2 Identity Edit LoRA at 0.8, 12 steps euler / beta, scene at 0.85 MP and reference at 1.0 MP, instruction "Replace the person in image 1 with the person from image 2. Use the face, head shape, hairline, skin tone, age and expression of the person in image 2 exactly. Keep the pose, framing, clothing, lighting and setting of image 1."

Why two passes instead of one text-to-image? Because the first pass draws a woman matching the sentence and the second pins the woman. One pass alone drifts: same description, twenty-four slightly different people. The edit pass is what makes them the same person.

Resize the result so the long side is 1024 and save as JPEG at 90. Each image gets a caption file: trigger, a woman, close-up portrait, three-quarter view, golden hour light outdoors — the trigger word, then a short description of what is not her (framing, light, place). Never describe her face in captions; the trigger has to carry that.

Step 4: look at the sheet before you train

Make a contact sheet of the 24 and look for three things:

Eighteen good images train fine; twenty-four is our target; more than thirty from a single source face adds little.

What you have now

A folder of 24 JPEGs with captions, all of one fictional person, in enough situations that the next lesson's training can tell her apart from her clothes, her backgrounds and her lighting. Total cost on a serverless RTX 5090: about 24 minutes and under a dollar.

Takeaway: one tight, prop-free face crop; scenes first, identity second; vary everything except her; and read the contact sheet before you spend ninety minutes training.