The Higgsfield setup and the copy-and-paste prompt to swap two characters into the viral rap performance.
You've seen the AI versions of Quavo and Takeoff's COLORS performance with different characters doing the exact same moves. Same camera, same background, same audio, new faces. This is how it's made, step by step, with the prompt that keeps each character glued to the right performer.
Clip length, resolution and credit cost are from our own generation. Higgsfield's pricing and options change, so check the cost shown before you hit Generate.
Drop your email and the whole guide unlocks right here, with a copy button on the prompt and a Markdown download. You'll also get one practical AI system most weeks.
This guide makes one video: the Hotel Lobby performance with two characters of your choice swapped in. It uses Higgsfield's Genjutsu tool, one motion reference video and two reference images.
The original Quavo and Takeoff COLORS clip. Every gesture, head turn and mouth movement comes from here.
One image per performer. Face, hair, body proportions and clothing come from these.
Higgsfield replaces each performer with its assigned character and keeps everything else as it was.
The reason this trend works is that nothing is re-animated from scratch. The camera move, the framing, the background, the props, the lighting and the audio are all the original. Only the two people change. That is why the prompt spends so many words on what to preserve and only a few on what to replace.
The one thing that goes wrong most. The two characters swap sides partway through, or blend into each other. The prompt handles this by assigning each image to the performer who starts on that side of the frame, and telling the model to keep the replacement attached as they move.
Tick them off as you go. Your progress saves in this browser. 0 / 4
Face fully visible, no sunglasses, no hand over the mouth. The outfit you want them wearing in the video. Plain background helps.
Two people in frame, a heavy filter, a tiny face in a wide shot, or a crop that hides the body when the clip shows full bodies.
If the clip shows the performers head to toe, the model has to invent whatever your image doesn't show. Give it a full-body photo and it has nothing to guess.
The order matters. Motion reference first, then Image 1, then Image 2.
At higgsfield.ai, open Video in the top nav and select Genjutsu.
Inside Genjutsu you'll see a toggle for Motion transfer and Objects swap. Choose Objects swap. Higgsfield moves its labels around; if your version shows Motion Control instead, that's the same upload flow.
Upload the original video as the movement reference. This becomes @Video1 in the prompt.
Upload Image 1 for the performer who starts on the left of the frame and Image 2 for the performer who starts on the right. These become @Image1 and @Image2.
Switch on the Prompt field, paste the prompt from the next section, and link each @ reference to the file you uploaded.
Check the order before you spend credits. If you accidentally upload the right-hand character first, the whole clip comes back mirrored. Swap the images, don't rewrite the prompt.
One prompt. It tells Genjutsu who replaces whom, what to take from each image, and everything to leave alone.
How it's built: paragraph one assigns and anchors. Paragraph two says what to take from each image and what to copy from the video. Paragraph three is the preserve list. If you edit it, edit the first paragraph only.
Look at the settings and the credit cost next to Generate, then click it. Our 24 second test at 720p, 16:9 showed 330 credits on a 55% off promotion (480 at the standard rate). Yours will differ with length, resolution and whatever Higgsfield is running that week.
24.0 seconds, 720p, 16:9. One reference video, two reference images, the prompt above, no other settings changed.
Cost scales with clip length. If you're testing a new pair of images, trim to a few seconds first and only run the full clip once the faces hold.
If it passes, download it. If one thing fails, fix that one thing (usually the reference image) and run it again rather than rewriting the prompt.
The Hotel Lobby clip is the trend, but the method doesn't care about the song. Anything with two performers who are clearly separated in the frame will work the same way.
Keep it to two. The prompt is written for a two-person composition. Adding a third character means rewriting the assignment paragraph and uploading a third image, and consistency drops with each extra person.
| Problem | What to try |
|---|---|
| The characters switched sides | Check the upload order. Image 1 must be the performer who starts on the left. Swap the images and run again; leave the prompt alone. |
| A face drifts or changes partway through | Replace that reference image with a clearer, front-on, well-lit photo of one person. Trim the clip shorter for the retest. |
| Wrong outfit or body shape | The image is probably cropped. Use a full-body photo when the clip shows full bodies, and make sure the outfit is fully visible. |
| Hands or mic look wrong | Regenerate. Hands near the face are the hardest part of any clip. A shorter section that avoids the worst gesture is a fair trade. |
| The audio is missing or different | Make sure the prompt still ends with "Preserve the original audio." If it does, regenerate; if it persists, re-attach the original track in your editor. |
| The credit cost is too high | Trim the clip before uploading. Test a few seconds first, then run the section you actually want to post. |
Higgsfield's labels, modes and pricing change. If your screen looks different from this page, follow the app over the page.
This is one build. AI Systems Lab is where the rest live: practical AI systems you can run in your own business.