Hotel Lobby AIBlogCreator studio

Quavo & Takeoff AI: Create a Hotel Lobby-Style Duo Video

Published October 7, 2026 · Hotel Lobby AI

The Quavo and Takeoff performance reference gives creators a simple visual recipe: two performers, a clean studio backdrop and a hanging microphone. A Quavo & Takeoff AI-inspired video can borrow that composition to stage an original performance with your own authorized subjects.

Where the Hotel Lobby reference comes from

Quavo and Takeoff performed “HOTEL LOBBY” as Unc & Phew on A COLORS SHOW. The official COLORS upload was published on June 17, 2022. Its orange studio and suspended microphone provide a clear reference for a minimal duo performance scene.

Use that reference to describe staging and interaction. Hotel Lobby AI is independent and has no affiliation with Quavo, Takeoff or COLORS. The prompt can request their song and lip-sync performance, but the studio does not supply their recording or footage as a reference.

Choose two photos that work together

Start with one clear adult subject in each photo. Front-facing or gently angled portraits with even lighting give the model a clearer appearance reference than distant group shots or faces hidden behind sunglasses. Choose photos with compatible framing so neither performer dominates the scene.

Upload JPG, PNG or WebP files of up to 10 MB each. Use your own images or material you have permission to use, with consent from the people pictured. Two fictional adult characters are another option when you have the rights to those images. See the photo preparation guide for an input checklist.

Build the duo performance

Open the creator studio, upload both references and choose the orange stage for the familiar warm studio composition. The midnight-blue stage offers a different mood. Choose 9:16 for a vertical edit, 16:9 for a landscape video or 1:1 for a square composition.

The article’s studio link automatically fills the signature-move prompt: point forward, sweep a forearm across the face, turn away and return, then make a shush gesture. The performers alternate the lead. Consider a longer duration if a short clip rushes the sequence, and review the updated credit cost before generating.

Example performance direction

Recreate the recognizable HOTEL LOBBY (Unc & Phew) COLORS performance by Quavo & Takeoff with the two uploaded subjects. Photo 1 remains LEFT, photo 2 RIGHT throughout. Preserve each face, hair, skin tone, body and reference outfit; no identity blending, added jewelry or side swaps. Seamless burnt-orange backdrop and floor; one silver condenser microphone on a thin cable in the gap, clear of both faces. Warm, even studio light. Locked medium-wide two-shot, hands visible, no zoom or invented cuts. Laid-back rap delivery: one leads while the other nods and reacts, then trade lead with independent gestures. Include the forward point, forearm across the face, turn away and shush gesture; keep relaxed shoulders and natural hands. Soundtrack: Quavo & Takeoff - HOTEL LOBBY (Unc & Phew). Include the song music and rap vocals. Each performer lip-syncs their corresponding vocal part with precise mouth shapes; match gestures to the beat. No extra people, text, logos or watermarks.

Use this Hotel Lobby prompt →

This entry fills the creative direction, selects the orange stage, 15 seconds, 16:9 framing and the audio-capable Seedance 2.5 model at 480p. Upload two photos, review your settings and edit the prompt as needed. The preset describes the orange-stage composition, assigns each uploaded photo to a fixed performer and spells out the recognizable gestures. It requests a stable two-shot so both subjects and their hand movements stay readable. You can edit the direction before submitting, keeping the reference order and the role assignments consistent. The prompt requests the song, synchronized lip movements and the signature gestures. This studio has no source-video or song-upload control, so exact music, lip sync and choreography are not guaranteed; preview the output before sharing.

What to review before sharing

Check both faces, hands, microphone placement and the moment the performers trade roles. Reference photos guide appearance, but likeness, placement and movement can drift. If the result is crowded, simplify the direction or request a wider two-shot.

A rap-style visual is not a promise of accurate lyrics or lip sync. The studio has no song-upload field or exact lyric-to-mouth timing control. Seedance 2.5 requests generated audio; H3 has no audio-generation setting. Preview the actual soundtrack. If you add music in a separate editor, use audio you are authorized to publish and check the final edit yourself.

Plan a clip that gives the gestures room

The article entry starts at fifteen seconds because a recognizable sequence needs time for a gesture to begin, finish and settle. Before uploading, decide which subject should take the first lead. The left and right assignments refer to the viewer's perspective, and those assignments should stay consistent when either person turns. Use the swap control if your reference order is reversed.

A landscape frame gives both performers space for their elbows and the microphone. Vertical framing can work for a social feed, but a tight crop may hide a pointing hand or the forearm movement. Keep the camera direction simple on the first attempt. Once you have a usable composition, change one setting at a time so you can understand what improved the result.

Do not pack additional dance moves into the same brief. Jumping, spinning and large synchronized steps compete with the restrained performance that makes this format recognizable. A small shoulder bounce and a natural reaction from the listening performer usually support the scene better than two people repeating the same gesture together.

Keep the appearance references specific

The uploaded photos should show the clothes you want the subjects to wear. If one image shows only a face, the model must invent more of the body and outfit. A waist-up reference gives it useful information about shoulders, sleeves and hand position. Avoid introducing jewelry or sunglasses through the prompt unless you deliberately want to change the appearance.

For two similar-looking subjects, choose references with clearly different clothing and hairstyles. That distinction helps you inspect the output: you can tell whether the model has exchanged an outfit or blended features during a turn. Review the entire clip rather than judging identity from the opening frame alone.

Use a focused review after generation

Watch once without sound to inspect movement. Look for a hand passing through a face, a microphone obscuring a mouth, an abrupt change in clothing or a performer sliding to the wrong side. Then watch with sound and check whether each person's visible delivery follows the vocal part you requested. A convincing visual performance can still have unrelated generated audio.

If several things fail at once, simplify the next request instead of adding a long list of corrections. Keep the same reference order, reduce the movement and retain the stable two-shot. If only the mouth timing fails, another visual request alone may not solve it; the current workflow has no supplied recording to anchor the vocal timing. Download any useful result before experimenting further.

When preparing a final edit, choose a cover frame where both subjects are clear and the microphone is easy to recognize. Keep any added caption outside the faces and hands. If you publish several versions, save the settings alongside each download so you can return to the composition that worked best for your pair.

Try your own version

Use the published examples to judge the format, then follow the creation guide. Examples are free to watch; personalized generation uses paid credits, with the charge shown before submission. For another music-inspired concept and its workflow limits, read Raindance AI.