VisorSTUDIO Sign in

Documentation

Help & Options Guide

What each option does for you — faster or nicer, portrait or landscape, keep the face or replace it.

High-quality image

Produce polished, sharp images for finals and hero frames.

Quality
Best detail and realism — preferred for delivery assets.
Speed
Slower than Fast mode, in exchange for a more finished result.
Best for
Once you've locked the idea and need a finished image.

Put your character into a video

Replace the person in a video you have the rights to use with your character, keeping the original motion and background.

Quality
Preserves pose and background well; preview a short segment first on complex scenes.
Speed
Slow; long videos are processed in segments and stitched together.
Best for
When you want to replace a person in footage you have the rights to use.

Lock the face to a reference

Keep the exact face of the person in your reference photo when generating new images.

Quality
Very strong likeness lock; can slightly reduce creative freedom elsewhere.
Speed
Adds a little time over a standard image.
Best for
When your character must look consistent across many images.

Process long video in segments

Split long videos into segments so they run reliably and can resume if interrupted.

Quality
Seams between segments are worth a check; most scenes join smoothly.
Speed
Total time is longer but each segment finishes sooner.
Best for
Turn on for long clips for a more reliable run.

Fast mode

Work quickly at the same resolution to preview first and save time.

Quality
Same resolution as every tier — only trimmed to run faster.
Speed
Fastest tier.
Best for
When you want a quick look before the final.

High-quality mode

Use maximum duration and resolution for the final output.

Quality
Highest quality for the tier you choose.
Speed
Slowest; may queue behind other work.
Best for
When you export the final for delivery.

Create sound & effects

Generate sound effects and short audio beds from a text description.

Quality
Great for foley, ambience, and stylized effects. Not for full songs or vocals.
Speed
Usually tens of seconds depending on requested length.
Best for
Describe concrete nouns and actions ("metal door creak", "rain on window"). Include mood, intensity, and desired length.
Limitations
Length is capped by the duration you set. Complex multi-layer mixes are best built from several short clips.

Create video from text or image

Generate short video from a text description or from an existing image.

Quality
Picture and sound are generated together — spoken lines get a matching voice; up to 15 seconds per clip.
Speed
Moderate for clips of a few seconds.
Best for
When you need a short video (with sound) from an idea, an image, or a set of reference images.

Sharpen & upscale

Make images and video sharper, rebuilding detail rather than just stretching the size.

Quality
Best on clean sources; noisy or heavily compressed input can reveal more flaws.
Speed
Images: seconds. Video: longer with duration.
Best for
For long clips, try a lighter setting first to check quality.

Enlarge resolution

Increase the size of an image or video to a larger resolution.

Quality
Larger enlargement sharpens detail but also reveals more flaws — try a lighter setting first on noisy sources.
Speed
The larger the enlargement, the longer it takes.
Best for
Use a lighter setting for drafts, a larger one for finals.

Isolate the subject from the background

Detect and isolate the person in each frame so only the character region is changed, not the background.

Quality
More precise isolation means cleaner hair and hand edges; poor isolation can bleed into the background.
Speed
Adds a little isolation time before generating.
Best for
Needed when replacing a character in video; check edges on complex scenes.

Keep pose and motion

Detect body pose to keep the correct stance and motion when replacing a character.

Quality
Strong pose lock; fast motion or occlusion can reduce accuracy.
Speed
Light, quick per-frame processing.
Best for
Use on full-body shots; less important for face-only framing.