Academy / Video on one GPU / Lesson 1

Free lesson · 6 min

H3 stills: five frames, keep the first

MiniMax H3 is a video model, and it makes a better identity-locked portrait than most image models if you stop asking it for a video. The settings that turned soft, mid-motion stills into sharp ones.

The Character Portrait H3 workflow takes one reference photo and returns a still of that person in a described scene, with the identity lock H3 is known for in video. Our first version of it produced soft, slightly smeared faces, and the fix had nothing to do with the model's quality. It had to do with asking a video model for a video and then taking a frame out of the middle.

What was wrong

The template shipped as a straightforward reference-to-video run: 2.5 MP, 22 frames, sigma shift 12, then ImageFromBatch grabbing frame 11. Three problems, stacked:

  1. Frame 11 is mid-motion. Even with "static camera, no motion" in the prompt, the model moves. Twenty-two frames in, the head has turned a few degrees and the hair has drifted; the frame carries motion blur you cannot see at thumbnail size and cannot miss at 100%.
  2. 2.5 MP spreads the model thin. H3's detail budget at that size goes into the whole frame; the face gets a small share.
  3. Shift 12 is tuned for motion. Higher sigma shift favours temporal coherence over per-frame sharpness. For a still, coherence is irrelevant.

The recipe that works

Tested side by side on the same reference and seed, judged at 100% crops of the face:

Everything else stays: the ref2va int8 checkpoint, the 8-step turbo LoRA, res_multistep.

What you get

On the fictional characters we tested with, the frame-0 / 1 MP / shift-6 version put the identity through cleanly: face, hairline, even an earring from the reference, in scenes the reference never saw (café, rooftop at dusk, beach at golden hour, studio). Render time on a warm RTX 5090 is 12 to 16 seconds per still on serverless, which is faster than the Krea 2 identity edit for the same job.

The prompt grammar H3 expects for reference work is specific: start with <Picture 1> is a woman. (or the description you want), then A single photorealistic still: she is <scene>., and end with Static camera, no motion. The angle-bracket token is how the model binds the reference to "she".

When to use this instead of Krea 2

Use Krea 2 (course 2) when you need hundreds of images, text-to-image freedom, or the LoRA ecosystem.

The DaSiWa and Remix variants

The same template exists for the two community H3 checkpoints. The settings carry over unchanged; the difference is taste. Remix ships its own text encoder and must be loaded with it; DaSiWa drops in. Try all three on one reference once and pick by the skin.

Takeaway: 1 MP, 5 frames, frame 0, shift 6, full-size head-crop reference. Ask a video model for a photograph and it will give you one.