The Character Portrait H3 workflow takes one reference photo and returns a still of that person in a described scene, with the identity lock H3 is known for in video. Our first version of it produced soft, slightly smeared faces, and the fix had nothing to do with the model's quality. It had to do with asking a video model for a video and then taking a frame out of the middle.
What was wrong
The template shipped as a straightforward reference-to-video run: 2.5 MP, 22 frames, sigma shift 12, then ImageFromBatch grabbing frame 11. Three problems, stacked:
- Frame 11 is mid-motion. Even with "static camera, no motion" in the prompt, the model moves. Twenty-two frames in, the head has turned a few degrees and the hair has drifted; the frame carries motion blur you cannot see at thumbnail size and cannot miss at 100%.
- 2.5 MP spreads the model thin. H3's detail budget at that size goes into the whole frame; the face gets a small share.
- Shift 12 is tuned for motion. Higher sigma shift favours temporal coherence over per-frame sharpness. For a still, coherence is irrelevant.
The recipe that works
Tested side by side on the same reference and seed, judged at 100% crops of the face:
- Resolution: 1 MP — 832×1248 portrait (or 1216×832 landscape). The face gets more of the model's attention and the render is roughly four times faster.
- Length: 5 frames. The minimum the sampler is happy with. There is no motion to speak of in five frames.
- Take frame 0. The first frame is the one the model conditions most tightly on the reference and the prompt; every later frame is a prediction from it. Set
ImageFromBatchto index 0. - Sigma shift 6, not 12. Sharper per-frame detail; nothing lost, because there is no sequence to keep coherent.
- Reference image size: max. Feed the reference at its full size. One caveat in the next lesson: at
maxthe model also copies the reference's background, so the reference should be a head crop against something neutral.
Everything else stays: the ref2va int8 checkpoint, the 8-step turbo LoRA, res_multistep.
What you get
On the fictional characters we tested with, the frame-0 / 1 MP / shift-6 version put the identity through cleanly: face, hairline, even an earring from the reference, in scenes the reference never saw (café, rooftop at dusk, beach at golden hour, studio). Render time on a warm RTX 5090 is 12 to 16 seconds per still on serverless, which is faster than the Krea 2 identity edit for the same job.
The prompt grammar H3 expects for reference work is specific: start with <Picture 1> is a woman. (or the description you want), then A single photorealistic still: she is <scene>., and end with Static camera, no motion. The angle-bracket token is how the model binds the reference to "she".
When to use this instead of Krea 2
- You have one good photo and need a portrait, fast, with no LoRA.
- You want the H3 look (its skin and light rendering is distinctly camera-like) in a still.
- You are going to make video of the same person next: the still and the clip come from the same model with the same identity lock, so they match.
Use Krea 2 (course 2) when you need hundreds of images, text-to-image freedom, or the LoRA ecosystem.
The DaSiWa and Remix variants
The same template exists for the two community H3 checkpoints. The settings carry over unchanged; the difference is taste. Remix ships its own text encoder and must be loaded with it; DaSiWa drops in. Try all three on one reference once and pick by the skin.
Takeaway: 1 MP, 5 frames, frame 0, shift 6, full-size head-crop reference. Ask a video model for a photograph and it will give you one.