Every character on this site started as one picture. Not a photo of anyone — a Krea 2 render of a described person — and from that one picture we build a training set of twenty-four images that agree on the face. This is the free lesson of the course because it is the step people get wrong most, and the step the trainer cannot fix for you later.
Step 1: one cover, one clean face
Write the person as a sentence you will reuse everywhere: "a stunning 26-year-old woman with long sun-bleached wavy blonde hair, bright blue eyes, a light tan, smooth clear glowing skin, a slim athletic figure and a soft easy smile". That sentence is her "look", and it goes into every prompt from now on. Render a cover with it: a candid selfie, front camera, natural light, face fully visible.
Then crop her face out of the cover. Not the whole picture — a tight head crop: roughly 0.85 face-widths either side of the centre, from the hairline down to about 2.4 face-heights (chin, neck, the top of the shoulders). We measured this: the tight crop improved likeness in every identity test compared with feeding the full selfie. The model should see a face, not a background.
Two things must not be in that crop, because the training set will inherit them: props and clothing with character. A sun hat ended up on a third of one character's dataset; a hotel bathrobe on half of another's. We re-rendered both covers without them. If the cover has an accessory, generate a plain studio portrait of the same person (black tank top, grey background) and crop the face from that instead.
Step 2: twenty-four scenes, three distances
A LoRA generalises when the dataset varies everything except the person. Our fixed scene list:
- 8 close-ups — facing camera, three-quarter left, three-quarter right, profile, selfie from above, serious low-key studio, neon at night, hair tied back in flat bathroom light.
- 10 half-body — café window, gym mirror, rooftop at sunset, beach, kitchen morning, car seat, restaurant at night, city street in a blazer, sofa by a fire, hotel pool.
- 6 full-length — sidewalk in jeans, bedroom mirror in a dress, beach at sunset, yoga studio, hotel lobby in an evening dress, park bench in autumn.
That gives varied lighting (window, golden hour, overcast, neon, candle, studio), varied expression (laughing, serious, neutral, smiling), both profiles, hair up and down, and clothing from swimwear to coats. The person is the only constant.
Step 3: scene first, then identity
Each dataset image is made in two passes, both on serverless:
- Text-to-image of the scene with her look sentence:
A stunning 26-year-old woman … , close-up portrait, head turned three-quarters to the left, golden hour sunlight outdoors, blurred trees. adult, mature face, natural photograph, clear smooth skin, sharp focus on the face.Sizes: 896×1024 for close-ups, 832×1040 half-body, 768×1152 full-length. - Identity edit of that picture with the face crop: the Krea 2 Identity Edit LoRA at 0.8, 12 steps euler / beta, scene at 0.85 MP and reference at 1.0 MP, instruction "Replace the person in image 1 with the person from image 2. Use the face, head shape, hairline, skin tone, age and expression of the person in image 2 exactly. Keep the pose, framing, clothing, lighting and setting of image 1."
Why two passes instead of one text-to-image? Because the first pass draws a woman matching the sentence and the second pins the woman. One pass alone drifts: same description, twenty-four slightly different people. The edit pass is what makes them the same person.
Resize the result so the long side is 1024 and save as JPEG at 90. Each image gets a caption file: trigger, a woman, close-up portrait, three-quarter view, golden hour light outdoors — the trigger word, then a short description of what is not her (framing, light, place). Never describe her face in captions; the trigger has to carry that.
Step 4: look at the sheet before you train
Make a contact sheet of the 24 and look for three things:
- Any black-and-white image. Remove it. One monochrome frame in a set of 24 taught a LoRA to render roughly one in six pictures in black and white, for scenes that had nothing to do with it. We fixed it afterwards by swapping scenes, but it costs an hour.
- A repeated garment or prop. If the same white halter top shows up in eight frames, it becomes part of her. Re-render those scenes or add the garment to their captions so the model can separate it.
- Blank or failed frames. The serverless worker occasionally returns a black image; the scripts detect brightness under a threshold and retry, but check.
Eighteen good images train fine; twenty-four is our target; more than thirty from a single source face adds little.
What you have now
A folder of 24 JPEGs with captions, all of one fictional person, in enough situations that the next lesson's training can tell her apart from her clothes, her backgrounds and her lighting. Total cost on a serverless RTX 5090: about 24 minutes and under a dollar.
Takeaway: one tight, prop-free face crop; scenes first, identity second; vary everything except her; and read the contact sheet before you spend ninety minutes training.