Skip to content

How to Make an AI Influencer Video for a Trending Sound

A voiceover and a trending sound cannot share one audio track. Here is the silent route: generate the clip without speech, cut it to the sound, then add the audio when you upload.

HexUGC is an AI influencer generator. You design one character once, then generate vertical videos of that character whenever you want them. Most of those videos are voiced: a script, a voice, a lip-synced read. A trending sound post is the exception, because the sound has to own the audio track, and that changes what you generate before you ever open the upload screen.

A voiceover and a trending sound cannot share a track

There is only one audio track on the file you upload, and the platform's sound sits on top of it. If your influencer is also talking, one of the two has to lose. Ducking the music under the voice kills the thing the sound was doing for you, and leaving both at full volume reads as a mistake within a second. Creators who post to trending audio solve this by never recording speech in the first place. The video carries motion and on-screen text, and the sound carries everything else.

So the useful question is not how to add music to a talking video. It is how to generate a clip that was never meant to talk.

Generate the scene with no voice

Every scene in the create wizard has a Voice toggle with two settings, With voice and No voice. No voice skips the script and the voiceover entirely and animates your chosen still directly, so what you are directing is movement rather than a read.

Two ways to drive that movement. Describe it in the scene direction box, which is fine for something simple like turning to the camera and laughing. Or hand over a motion reference: paste a link to a clip, or upload one from your phone. Your influencer keeps its own face and its own pose and performs that clip's motion, which is the closest thing to copying the format of a post you already know works. We set out how that step behaves in turning a TikTok into an AI avatar video. The practical rule is to pick a reference that opens in roughly the pose your still is already in.

Cut the video to the sound, not to a script

Silent clips come out at five or ten seconds. Anything you ask for up to seven snaps to five, and anything above it snaps to ten, so the length is a choice between two numbers rather than a slider.

That is worth planning around, because trending sounds have structure. A sound with a beat drop at eight seconds wants a scene change at eight seconds. Add scenes and they stitch into one file in order, so three ten-second scenes give you thirty seconds with two cuts you placed deliberately. The mechanics are in making an AI video longer than ten seconds.

What comes out, and what you add on upload

You download a 9:16 H.264 MP4 with a silent audio track on it. Nothing to mute, nothing to duck, and the editor treats it as a clean slate when you pick the sound.

Two things you add there rather than here. The sound itself, from the platform's own library, which is also the only version of it that counts as the trending sound. And the on-screen text, because our burned-in captions are generated from the spoken script and a silent clip has no script to read from. If you are posting AI video at all, read what TikTok and Instagram ask of AI content before you get into a cadence.

What it costs

A five second silent clip is 26 credits and a ten second one is 51. A motion reference changes which engine runs, not the price, so it costs the same as describing the movement yourself.

Worth knowing before you plan a week of these: silent is dearer per second than voiced. A ten second voiced video is 28 credits, because lip-syncing a still to audio is cheaper to run than driving a still through a reference clip. The opening pack is $5 for 50 credits.

What this does not do

We do not detect the beat of your sound or cut to it. You choose the scene lengths, and matching them to the music is your judgement, not an automatic step.

We have no music library and never will have the trending one, because that catalogue lives inside the platform. On TikTok, accounts registered as businesses draw from a commercial music library rather than the full trending list, and no generator can work around that.

We generate one video at a time. Multi-variant and batch generation are on our roadmap rather than shipped. We do not publish or schedule to TikTok, Instagram or YouTube either, so you upload the file yourself. And there is no free tier: you pay from the first generation.

Closing

Silent is not a lesser mode here. It is the one that leaves room for the sound.

Create your AI influencer and generate the first video.