Documentation
Help & Options Guide
What each option does for you — faster or nicer, portrait or landscape, keep the face or replace it.
Produce polished, sharp images for finals and hero frames.
- Quality
- Best detail and realism — preferred for delivery assets.
- Speed
- Slower than Fast mode, in exchange for a more finished result.
- Best for
- Once you've locked the idea and need a finished image.
Replace the person in a video you have the rights to use with your character, keeping the original motion and background.
- Quality
- Preserves pose and background well; preview a short segment first on complex scenes.
- Speed
- Slow; long videos are processed in segments and stitched together.
- Best for
- When you want to replace a person in footage you have the rights to use.
Keep the exact face of the person in your reference photo when generating new images.
- Quality
- Very strong likeness lock; can slightly reduce creative freedom elsewhere.
- Speed
- Adds a little time over a standard image.
- Best for
- When your character must look consistent across many images.
Split long videos into segments so they run reliably and can resume if interrupted.
- Quality
- Seams between segments are worth a check; most scenes join smoothly.
- Speed
- Total time is longer but each segment finishes sooner.
- Best for
- Turn on for long clips for a more reliable run.
Work quickly at the same resolution to preview first and save time.
- Quality
- Same resolution as every tier — only trimmed to run faster.
- Speed
- Fastest tier.
- Best for
- When you want a quick look before the final.
Use maximum duration and resolution for the final output.
- Quality
- Highest quality for the tier you choose.
- Speed
- Slowest; may queue behind other work.
- Best for
- When you export the final for delivery.
Generate sound effects and short audio beds from a text description.
- Quality
- Great for foley, ambience, and stylized effects. Not for full songs or vocals.
- Speed
- Usually tens of seconds depending on requested length.
- Best for
- Describe concrete nouns and actions ("metal door creak", "rain on window"). Include mood, intensity, and desired length.
- Limitations
- Length is capped by the duration you set. Complex multi-layer mixes are best built from several short clips.
Generate short video from a text description or from an existing image.
- Quality
- Picture and sound are generated together — spoken lines get a matching voice; up to 15 seconds per clip.
- Speed
- Moderate for clips of a few seconds.
- Best for
- When you need a short video (with sound) from an idea, an image, or a set of reference images.
Make images and video sharper, rebuilding detail rather than just stretching the size.
- Quality
- Best on clean sources; noisy or heavily compressed input can reveal more flaws.
- Speed
- Images: seconds. Video: longer with duration.
- Best for
- For long clips, try a lighter setting first to check quality.
Increase the size of an image or video to a larger resolution.
- Quality
- Larger enlargement sharpens detail but also reveals more flaws — try a lighter setting first on noisy sources.
- Speed
- The larger the enlargement, the longer it takes.
- Best for
- Use a lighter setting for drafts, a larger one for finals.
Detect and isolate the person in each frame so only the character region is changed, not the background.
- Quality
- More precise isolation means cleaner hair and hand edges; poor isolation can bleed into the background.
- Speed
- Adds a little isolation time before generating.
- Best for
- Needed when replacing a character in video; check edges on complex scenes.
Detect body pose to keep the correct stance and motion when replacing a character.
- Quality
- Strong pose lock; fast motion or occlusion can reduce accuracy.
- Speed
- Light, quick per-frame processing.
- Best for
- Use on full-body shots; less important for face-only framing.