News · StarSinger

How to build a custom virtual singer for vertical video platforms

A step-by-step workflow for cloning your voice, writing tracks, and generating lip-synced vertical music videos.

By Sylvie Fontaine·August 11, 2026·3 min read
Key points
  • Voice cloning requires a 30-second initial audio sample or 30 minutes for higher audio fidelity.
  • Vertical video generation maps per-beat storyboard scenes directly to isolated song stems.
  • Creators retain full commercial ownership over original generated tracks paired with cloned vocals.

Building a Virtual Artist Workflow

Launching a music channel used to require an ensemble: a vocalist, a producer, a video editor, and a director. Modern synthetic media workflows compress this production pipeline into software. StarSinger combines prompt-driven audio synthesis, vocal cloning, and lip-synced vertical video rendering inside a unified interface.

If you want to publish consistent short-form music content, you need a predictable pipeline. Relying on piecemeal tools for audio generation, vocal pitch shifting, and external video assembly slows down output. Here is how to build a virtual artist persona from scratch using a cloned vocal profile and automated video generation.

Step 1: Train Your Vocal Model

Your virtual persona needs a distinct sound. You can select an existing artist persona from the platform library or create a custom vocal clone. Vocal cloning isolates your specific pitch, tone, timbre, and accent patterns to build a private voice model.

To clone your voice, navigate to the studio interface and record your audio sample directly through your microphone. A short 30-second recording builds a functional voice model. If you plan to produce complex tracks across wider vocal ranges, record a 30-minute training sample instead. The longer sample gives the model more data points for pitch transitions and subtle timbre changes. The vocal clone feature requires a $1.99 one-time unlock fee.

Once processed, the vocal model remains private to your profile. You can apply it to any song style in the platform catalog, from disco funk and hip-hop to drill syntax.

Step 2: Generate Audio Stems From Prompts

With a vocal model configured, you can start composing music tracks. The system handles song structure through direct text prompts or explicit verse and chorus lyrics. You can generate music across 11 languages, including English, Chinese, Japanese, Korean, Spanish, Portuguese, French, German, Italian, Russian, and English.

Enter a vibe description or type custom text into the prompt field. The engine does not run as a closed black box. It processes your request through four distinct operational stages:

  • Prompt Parsing: Translates your written lyrics or vibe description into structural audio parameters.
  • Stem Separation: Divides the generated audio track into four distinct stems: drums, bass, melody, and vocals.
  • Scene Storyboarding: Analyzes dynamic changes across the separated audio stems to map visual shot transitions.
  • Video Assembly: Formats and renders a lip-synced vertical video frame sequence matched to the active vocal stem.

Platform accounts include one free AI song generation per day. Additional track creation operates on a pay-as-you-go structure rather than a required monthly subscription.

Step 3: Render Beat-Synced Vertical Video

Audio alone rarely holds user attention on short-form feeds. You need vertical video matched directly to the underlying instrumentation. Once your audio track generates, you can append vertical video assets directly to the audio timeline.

The rendering pipeline uses separated audio stems to construct a per-beat storyboard. Snare hits, drum drops, and vocal entry points trigger dynamic scene changes optimized for vertical displays. Adding a standard beat-synced vertical video to a completed audio track costs $0.99. Higher-end cinematic music videos start at $2.99 for a 15-second clip, or $0.20 per second when purchased in volume bundles.

Processing times depend on clip duration and queue load. A basic 15-second clip renders faster than a full song, but standard video generation usually finishes in 10 to 15 minutes. The platform delivers an automated notification as soon as your finished video file finishes rendering.

Step 4: Verify Rights and Export for Distribution

Before publishing generated media to external networks, confirm your asset ownership status. Licensing terms depend directly on how your audio and visuals were constructed:

  1. Original Track Assets: Tracks generated from original text prompts or user-written lyrics, paired with your private cloned voice on original beats, belong entirely to you. You maintain full commercial ownership rights over these files.
  2. Copyrighted Song Covers: Re-creations or vocal covers of existing copyrighted songs are restricted to personal, non-commercial distribution.

Export your completed vertical video files directly to local storage. Maintaining a regular publishing cadence is easier when you decouple audio creation from video rendering. Use your daily free generation to test track prompts, unlock vocal models, queue video renders during production downtime, and export finished vertical video files straight to your publishing workflow.

More from StarSinger News
Published via Stork Wire — independent coverage for AI tool makers, in partnership with this site.