{"id": "6f3c1c2a-7b9e-4c3a-9d2e-gmanski-minimax3", "revision": 0, "last_node_id": 22, "last_link_id": 25, "nodes": [{"id": 1, "type": "MarkdownNote", "pos": [-1480, -340], "size": [470, 640], "flags": {}, "order": 0, "mode": 0, "inputs": [], "outputs": [], "properties": {"Node name for S&R": "MarkdownNote"}, "widgets_values": ["## MiniMax Music 3 (Gmanski)\n\n**Idea in, finished song out.** The Song Planner writes the structured caption and tagged lyrics that MiniMax Music 3 needs, then the music model generates the track.\n\n### Run it\n1. Describe the song in **Song Planner (Gmanski)** → *idea*. Pick vocals, language and length.\n2. Queue. The planner runs first (30–60 s on the local Ollama model), then the music model.\n3. Listen in **Preview Audio**. Files are saved under `output/audio/Gmanski/`.\n\n### Providers\n- **ollama** (default, free): uses `gemma3:27b` on this pod. Pulled automatically at boot; check `ollama list`.\nSet the `GMANSKI_API_KEY` environment variable on the pod (recommended) or paste the key in the *api_key* widget (it is then saved inside this workflow file). Each plan costs credits.\n\n### Seeds\n- The **planner seed** is fixed on purpose: once you like the caption and lyrics, change only the **Music seed** to get new takes of the same song.\n- Change the planner seed (or the idea) to get a new plan.\n\n### Hand-editing the text\nFlip **use manual text** on either switch and paste your own caption / lyrics into the *Manual caption* / *Manual lyrics* boxes. The planner is skipped while a switch is on.\n\n### Models (auto-downloaded to `/workspace/models/...` on first boot, ~22 GB)\n- `diffusion_models/minimax_music3_dit_fp16.safetensors`\n- `text_encoders/minimax_music3_text_encoder_pruned_bf16.safetensors`\n- `vae/minimax_music3_dav.safetensors`\n\nCheck `/workspace/logs/minimax-music3-download.log` if the loaders show the files as missing.\n\n### Tips\n- `max_duration` is fed from the planner (planned length × 1.4, max 300 s) so a song is never cut off mid-lyric.\n- Raise `batch_size` on the empty latent to get several takes from one planning pass.\n- Song length above ~240 s needs a 300 s cap and takes noticeably longer.\n"], "title": "Read me: MiniMax Music 3 (Gmanski)", "color": "#432", "bgcolor": "#653"}, {"id": 2, "type": "PrimitiveFloat", "pos": [-1480, 340], "size": [330, 60], "flags": {}, "order": 1, "mode": 0, "inputs": [{"name": "value", "type": "FLOAT", "link": null, "widget": {"name": "value"}}], "outputs": [{"name": "FLOAT", "type": "FLOAT", "links": [1]}], "properties": {"Node name for S&R": "PrimitiveFloat"}, "widgets_values": [120.0], "title": "Song length (seconds)"}, {"id": 3, "type": "GmanskiSongPlanner", "pos": [-960, -120], "size": [470, 600], "flags": {}, "order": 2, "mode": 0, "inputs": [{"name": "provider", "type": "COMBO", "link": null, "widget": {"name": "provider"}}, {"name": "idea", "type": "STRING", "link": null, "widget": {"name": "idea"}}, {"name": "genre_hint", "type": "STRING", "link": null, "widget": {"name": "genre_hint"}}, {"name": "vocal_config", "type": "COMBO", "link": null, "widget": {"name": "vocal_config"}}, {"name": "language", "type": "COMBO", "link": null, "widget": {"name": "language"}}, {"name": "duration_seconds", "type": "FLOAT", "link": 1, "widget": {"name": "duration_seconds"}}, {"name": "seed", "type": "INT", "link": null, "widget": {"name": "seed"}}, {"name": "temperature", "type": "FLOAT", "link": null, "widget": {"name": "temperature"}}, {"name": "ollama_model", "type": "STRING", "link": null, "widget": {"name": "ollama_model"}}, {"name": "ollama_url", "type": "STRING", "link": null, "widget": {"name": "ollama_url"}}, {"name": "unload_after", "type": "BOOLEAN", "link": null, "widget": {"name": "unload_after"}}, {"name": "cloud_model", "type": "COMBO", "link": null, "widget": {"name": "cloud_model"}}, {"name": "api_base_url", "type": "STRING", "link": null, "widget": {"name": "api_base_url"}}, {"name": "api_key", "type": "STRING", "link": null, "widget": {"name": "api_key"}}, {"name": "age_band", "type": "COMBO", "link": null, "widget": {"name": "age_band"}}, {"name": "character_id", "type": "STRING", "link": null, "widget": {"name": "character_id"}}, {"name": "character_library_path", "type": "STRING", "link": null, "widget": {"name": "character_library_path"}}], "outputs": [{"name": "caption", "type": "STRING", "links": [2]}, {"name": "lyrics", "type": "STRING", "links": [4]}, {"name": "max_duration", "type": "FLOAT", "links": [14]}, {"name": "debug", "type": "STRING", "links": []}], "properties": {"Node name for S&R": "GmanskiSongPlanner"}, "widgets_values": ["ollama", "An upbeat hard-rock / J-pop rock song about eating far too much at a festival and the tummy ache that follows. One female singer, two guitarists, bass and drums. At least one quick, shredding guitar solo. Big singalong chorus.", "hard rock / j-pop rock", "female vocals", "English", 120.0, 0, "fixed", 0.8, "gemma3:27b", "http://localhost:11434", true, "claude", "", "", "general", "", "/workspace/gmanski/characters.json"], "title": "Song Planner (Gmanski)"}, {"id": 4, "type": "PrimitiveStringMultiline", "pos": [-420, -340], "size": [380, 220], "flags": {}, "order": 3, "mode": 0, "inputs": [{"name": "value", "type": "STRING", "link": null, "widget": {"name": "value"}}], "outputs": [{"name": "STRING", "type": "STRING", "links": [3]}], "properties": {"Node name for S&R": "PrimitiveStringMultiline"}, "widgets_values": [""], "title": "Manual caption (optional)"}, {"id": 5, "type": "PrimitiveStringMultiline", "pos": [-420, -80], "size": [380, 280], "flags": {}, "order": 4, "mode": 0, "inputs": [{"name": "value", "type": "STRING", "link": null, "widget": {"name": "value"}}], "outputs": [{"name": "STRING", "type": "STRING", "links": [5]}], "properties": {"Node name for S&R": "PrimitiveStringMultiline"}, "widgets_values": [""], "title": "Manual lyrics (optional)"}, {"id": 6, "type": "ComfySwitchNode", "pos": [-420, 240], "size": [330, 110], "flags": {}, "order": 5, "mode": 0, "inputs": [{"name": "on_false", "type": "STRING", "link": 2}, {"name": "on_true", "type": "STRING", "link": 3}, {"name": "switch", "type": "BOOLEAN", "link": null, "widget": {"name": "switch"}}], "outputs": [{"name": "output", "type": "STRING", "links": [6]}], "properties": {"Node name for S&R": "ComfySwitchNode"}, "widgets_values": [false], "title": "Caption: use manual text?"}, {"id": 7, "type": "ComfySwitchNode", "pos": [-420, 390], "size": [330, 110], "flags": {}, "order": 6, "mode": 0, "inputs": [{"name": "on_false", "type": "STRING", "link": 4}, {"name": "on_true", "type": "STRING", "link": 5}, {"name": "switch", "type": "BOOLEAN", "link": null, "widget": {"name": "switch"}}], "outputs": [{"name": "output", "type": "STRING", "links": [7]}], "properties": {"Node name for S&R": "ComfySwitchNode"}, "widgets_values": [false], "title": "Lyrics: use manual text?"}, {"id": 8, "type": "ShowText|pysssss", "pos": [30, -340], "size": [420, 330], "flags": {}, "order": 7, "mode": 0, "inputs": [{"name": "text", "type": "STRING", "link": 6}], "outputs": [{"name": "STRING", "type": "STRING", "links": [8, 11]}], "properties": {"Node name for S&R": "ShowText|pysssss"}, "widgets_values": [], "title": "Caption (what the model gets)"}, {"id": 9, "type": "ShowText|pysssss", "pos": [30, 30], "size": [420, 470], "flags": {}, "order": 8, "mode": 0, "inputs": [{"name": "text", "type": "STRING", "link": 7}], "outputs": [{"name": "STRING", "type": "STRING", "links": [12]}], "properties": {"Node name for S&R": "ShowText|pysssss"}, "widgets_values": [], "title": "Lyrics (what the model gets)"}, {"id": 10, "type": "GmanskiExtractVocalDetails", "pos": [30, 540], "size": [420, 90], "flags": {}, "order": 9, "mode": 0, "inputs": [{"name": "caption", "type": "STRING", "link": 8}], "outputs": [{"name": "vocal_details", "type": "STRING", "links": [9]}], "properties": {"Node name for S&R": "GmanskiExtractVocalDetails"}, "widgets_values": [], "title": "Extract Vocal Details (copy into Save Character)"}, {"id": 11, "type": "ShowText|pysssss", "pos": [500, 540], "size": [420, 150], "flags": {}, "order": 10, "mode": 0, "inputs": [{"name": "text", "type": "STRING", "link": 9}], "outputs": [{"name": "STRING", "type": "STRING", "links": []}], "properties": {"Node name for S&R": "ShowText|pysssss"}, "widgets_values": [], "title": "Voice you liked? Copy this into Save Character"}, {"id": 12, "type": "UNETLoader", "pos": [500, -340], "size": [440, 90], "flags": {}, "order": 11, "mode": 0, "inputs": [{"name": "unet_name", "type": "COMBO", "link": null, "widget": {"name": "unet_name"}}, {"name": "weight_dtype", "type": "COMBO", "link": null, "widget": {"name": "weight_dtype"}}], "outputs": [{"name": "MODEL", "type": "MODEL", "links": [17]}], "properties": {"Node name for S&R": "UNETLoader", "models": [{"name": "minimax_music3_dit_fp16.safetensors", "url": "https://huggingface.co/Comfy-Org/MiniMax-Music-3/resolve/main/diffusion_models/minimax_music3_dit_fp16.safetensors", "directory": "diffusion_models"}]}, "widgets_values": ["minimax_music3_dit_fp16.safetensors", "default"]}, {"id": 13, "type": "CLIPLoader", "pos": [500, -210], "size": [440, 110], "flags": {}, "order": 12, "mode": 0, "inputs": [{"name": "clip_name", "type": "COMBO", "link": null, "widget": {"name": "clip_name"}}, {"name": "type", "type": "COMBO", "link": null, "widget": {"name": "type"}}, {"name": "device", "type": "COMBO", "link": null, "widget": {"name": "device"}}], "outputs": [{"name": "CLIP", "type": "CLIP", "links": [10]}], "properties": {"Node name for S&R": "CLIPLoader", "models": [{"name": "minimax_music3_text_encoder_pruned_bf16.safetensors", "url": "https://huggingface.co/Comfy-Org/MiniMax-Music-3/resolve/main/text_encoders/minimax_music3_text_encoder_pruned_bf16.safetensors", "directory": "text_encoders"}]}, "widgets_values": ["minimax_music3_text_encoder_pruned_bf16.safetensors", "minimax", "default"]}, {"id": 14, "type": "VAELoader", "pos": [500, -60], "size": [440, 60], "flags": {}, "order": 13, "mode": 0, "inputs": [{"name": "vae_name", "type": "COMBO", "link": null, "widget": {"name": "vae_name"}}], "outputs": [{"name": "VAE", "type": "VAE", "links": [23]}], "properties": {"Node name for S&R": "VAELoader", "models": [{"name": "minimax_music3_dav.safetensors", "url": "https://huggingface.co/Comfy-Org/MiniMax-Music-3/resolve/main/vae/minimax_music3_dav.safetensors", "directory": "vae"}]}, "widgets_values": ["minimax_music3_dav.safetensors"]}, {"id": 15, "type": "SeedNode", "pos": [500, 40], "size": [300, 90], "flags": {}, "order": 14, "mode": 0, "inputs": [{"name": "seed", "type": "INT", "link": null, "widget": {"name": "seed"}}], "outputs": [{"name": "seed", "type": "INT", "links": [13, 21]}], "properties": {"Node name for S&R": "SeedNode"}, "widgets_values": [12345, "randomize"], "title": "Music seed"}, {"id": 16, "type": "MiniMaxMusic3TextEncode", "pos": [500, 180], "size": [470, 330], "flags": {}, "order": 15, "mode": 0, "inputs": [{"name": "clip", "type": "CLIP", "link": 10}, {"name": "caption", "type": "STRING", "link": 11, "widget": {"name": "caption"}}, {"name": "lyrics", "type": "STRING", "link": 12, "widget": {"name": "lyrics"}}, {"name": "seed", "type": "INT", "link": 13, "widget": {"name": "seed"}}, {"name": "max_duration", "type": "FLOAT", "link": 14, "widget": {"name": "max_duration"}}, {"name": "cfg_scale", "type": "FLOAT", "link": null, "widget": {"name": "cfg_scale"}}, {"name": "top_k", "type": "INT", "link": null, "widget": {"name": "top_k"}}], "outputs": [{"name": "CONDITIONING", "type": "CONDITIONING", "links": [16, 18]}, {"name": "seconds", "type": "FLOAT", "links": [15]}], "properties": {"Node name for S&R": "MiniMaxMusic3TextEncode"}, "widgets_values": ["", "", 0, "fixed", 168.0, 1.7, 50], "title": "MiniMax Music 3 Text Encode"}, {"id": 17, "type": "EmptyMiniMaxMusic3LatentAudio", "pos": [1040, 180], "size": [360, 90], "flags": {}, "order": 16, "mode": 0, "inputs": [{"name": "seconds", "type": "FLOAT", "link": 15, "widget": {"name": "seconds"}}, {"name": "batch_size", "type": "INT", "link": null, "widget": {"name": "batch_size"}}], "outputs": [{"name": "LATENT", "type": "LATENT", "links": [20]}], "properties": {"Node name for S&R": "EmptyMiniMaxMusic3LatentAudio"}, "widgets_values": [120.0, 1], "title": "Empty MiniMax Latent (batch_size = takes per run)"}, {"id": 18, "type": "ConditioningZeroOut", "pos": [1040, 310], "size": [300, 40], "flags": {}, "order": 17, "mode": 0, "inputs": [{"name": "conditioning", "type": "CONDITIONING", "link": 16}], "outputs": [{"name": "CONDITIONING", "type": "CONDITIONING", "links": [19]}], "properties": {"Node name for S&R": "ConditioningZeroOut"}, "widgets_values": []}, {"id": 19, "type": "KSampler", "pos": [1040, 400], "size": [360, 270], "flags": {}, "order": 18, "mode": 0, "inputs": [{"name": "model", "type": "MODEL", "link": 17}, {"name": "positive", "type": "CONDITIONING", "link": 18}, {"name": "negative", "type": "CONDITIONING", "link": 19}, {"name": "latent_image", "type": "LATENT", "link": 20}, {"name": "seed", "type": "INT", "link": 21, "widget": {"name": "seed"}}, {"name": "steps", "type": "INT", "link": null, "widget": {"name": "steps"}}, {"name": "cfg", "type": "FLOAT", "link": null, "widget": {"name": "cfg"}}, {"name": "sampler_name", "type": "COMBO", "link": null, "widget": {"name": "sampler_name"}}, {"name": "scheduler", "type": "COMBO", "link": null, "widget": {"name": "scheduler"}}, {"name": "denoise", "type": "FLOAT", "link": null, "widget": {"name": "denoise"}}], "outputs": [{"name": "LATENT", "type": "LATENT", "links": [22]}], "properties": {"Node name for S&R": "KSampler"}, "widgets_values": [0, "fixed", 30, 1.7, "euler", "simple", 1.0]}, {"id": 20, "type": "VAEDecodeAudio", "pos": [1460, 400], "size": [250, 60], "flags": {}, "order": 19, "mode": 0, "inputs": [{"name": "samples", "type": "LATENT", "link": 22}, {"name": "vae", "type": "VAE", "link": 23}], "outputs": [{"name": "AUDIO", "type": "AUDIO", "links": [24]}], "properties": {"Node name for S&R": "VAEDecodeAudio"}, "widgets_values": []}, {"id": 21, "type": "GmanskiFilenamePrefix", "pos": [1460, 510], "size": [330, 130], "flags": {}, "order": 20, "mode": 0, "inputs": [{"name": "title", "type": "STRING", "link": null, "widget": {"name": "title"}}, {"name": "folder", "type": "STRING", "link": null, "widget": {"name": "folder"}}, {"name": "fallback", "type": "STRING", "link": null, "widget": {"name": "fallback"}}], "outputs": [{"name": "filename_prefix", "type": "STRING", "links": [25]}], "properties": {"Node name for S&R": "GmanskiFilenamePrefix"}, "widgets_values": ["", "audio/Gmanski", "minimax_music3"], "title": "Song title (used as the saved filename)"}, {"id": 22, "type": "SaveAudioAdvanced", "pos": [1460, 680], "size": [330, 150], "flags": {}, "order": 21, "mode": 0, "inputs": [{"name": "audio", "type": "AUDIO", "link": 24}, {"name": "filename_prefix", "type": "STRING", "link": 25, "widget": {"name": "filename_prefix"}}, {"name": "format", "type": "COMFY_DYNAMICCOMBO_V3", "link": null, "widget": {"name": "format"}}, {"name": "audioUI", "type": "AUDIO_UI", "link": null, "widget": {"name": "audioUI"}}], "outputs": [{"name": "audio", "type": "AUDIO", "links": []}], "properties": {"Node name for S&R": "SaveAudioAdvanced"}, "widgets_values": ["audio/Gmanski/minimax_music3", "flac"], "title": "Save Audio (flac)"}], "links": [[1, 2, 0, 3, 5, "FLOAT"], [2, 3, 0, 6, 0, "STRING"], [3, 4, 0, 6, 1, "STRING"], [4, 3, 1, 7, 0, "STRING"], [5, 5, 0, 7, 1, "STRING"], [6, 6, 0, 8, 0, "STRING"], [7, 7, 0, 9, 0, "STRING"], [8, 8, 0, 10, 0, "STRING"], [9, 10, 0, 11, 0, "STRING"], [10, 13, 0, 16, 0, "CLIP"], [11, 8, 0, 16, 1, "STRING"], [12, 9, 0, 16, 2, "STRING"], [13, 15, 0, 16, 3, "INT"], [14, 3, 2, 16, 4, "FLOAT"], [15, 16, 1, 17, 0, "FLOAT"], [16, 16, 0, 18, 0, "CONDITIONING"], [17, 12, 0, 19, 0, "MODEL"], [18, 16, 0, 19, 1, "CONDITIONING"], [19, 18, 0, 19, 2, "CONDITIONING"], [20, 17, 0, 19, 3, "LATENT"], [21, 15, 0, 19, 4, "INT"], [22, 19, 0, 20, 0, "LATENT"], [23, 14, 0, 20, 1, "VAE"], [24, 20, 0, 22, 0, "AUDIO"], [25, 21, 0, 22, 1, "STRING"]], "groups": [], "config": {}, "extra": {"ds": {"scale": 0.62, "offset": [1480, 420]}, "gmanski_template": "minimax-music-3", "generator": "scripts/build_minimax_music_workflow.py", "source": "https://gmanski.com/workflows/minimax-music-3"}, "version": 0.4}