The first frame is your photo
We route first-frame orders to xAI's native image-to-video endpoint and checked the result frame by frame: the opening frame is the input image, not a reinterpretation of it.
Model · xAI Grok Imagine Video 1.5
xAI's video model, and what it does better than anything else here: it opens on your photo and speaks your line. Give it a portrait and a sentence in quotes, and the first frame of the clip is that portrait, with the sentence said out loud and lip-synced. Clips run 4 to 15 seconds at any whole-second length, at 720p or 1080p, in five aspect ratios. You pay by the second — $0.14 at 720p, $0.25 at 1080p — the same rates xAI lists for its own API. Crypto or card, no subscription.
Grok Imagine Video 1.5 is xAI's video generation model — the one behind the video side of Grok Imagine — and it turns a prompt, or a still image plus a prompt, into a clip of up to 15 seconds with sound rendered in the same pass. It has been in the xAI API since June 2026 and on BananaBanana since 16 September 2026, paid from the same balance as every other model. We charge xAI's own per-second rates, $0.14 at 720p and $0.25 at 1080p, and we do not pass on the fee xAI adds for each input image.
Dialogue and first frames. A line written in quotes gets spoken, lip-synced and mixed at a conversational level, which is what a creator-style ad needs. And when you hand it a photo as the first frame, the clip opens on that photo — same person, same room, same pattern on the sweater — instead of a re-staged lookalike. Text-only prompts work too, and vertical is honest: a 9:16 order at 1080p came back as a 1088×1920 file, not a cropped horizontal render.
We route first-frame orders to xAI's native image-to-video endpoint and checked the result frame by frame: the opening frame is the input image, not a reinterpretation of it.
Put the sentence in quotes and the model says it. Audio is always on and always included in the price; there is no silent tier and nothing extra to pay for sound.
Not a menu of three durations: every whole second in the range. Seven seconds for one line of dialogue, eleven for two, and you pay for exactly that.
16:9, 9:16, 1:1, 3:2 and 2:3. Square and 3:2 are formats the other video models here do not render at all.
References steer the subject and the setting while the model stages the scene itself. They are an alternative to a first frame, not an addition — send one or the other.
$0.14 a second at 720p, $0.25 at 1080p. Work the prompt out at 720p, then render a text-to-video keeper at 1080p once.
Clips that start from an image run at 720p here. xAI caps reference-guided video at 720p, and the first-frame route available to us does not take 1080p yet, so 1080p is for text-to-video. Rather than bill you the 1080p rate and deliver a smaller file, we reject that combination before anything is charged; in the studio the image slots simply switch off with a note when you pick 1080p. A first frame and reference images cannot be combined. There is no last frame, no seed, no negative prompt and no video input.
15 seconds is the longest single render, and there is no extend or edit afterwards — a forty-second spot is three clips and a cut. The audio is strong on speech and on named sound events and weak on vague ambience: "soft ambient kitchen sounds" came back at about −50 dB, which is silence. Name a sound ("a kettle starting to whistle off-screen") or add room tone in the edit. The prompt tops out at 2048 characters. If you need half a minute in one take, reference clips or seeds, Wan 3.0 covers those; if you want to change a finished clip by describing the change, that is Gemini Omni Flash.
A portrait, a product in hand, the room behind. Generate it here or upload your own. Whatever is in that picture is what the clip starts with, so fix the frame before you pay for motion.
Describe who is in the frame, put the spoken line in quotes, then say how the camera behaves. "Natural handheld motion, quiet room tone" is enough. Keep it under 2048 characters.
Four seconds at 720p is $0.56. Your balance is charged when the job starts and refunded automatically if it fails or the filter rejects it. Top up with crypto or a card; there is no monthly fee.
| Model | 4 s | 8 s | 15 s | Sound |
|---|---|---|---|---|
| Grok Imagine Video 1.5 · 720p$0.14/s — first frame, references, drafts | $0.56 | $1.12 | $2.10 | Always on, included |
| Grok Imagine Video 1.5 · 1080p$0.25/s — text-to-video only | $1.00 | $2.00 | $3.75 | Always on, included |
| Wan 3.0 · 720pAlibaba's model, up to 30 s | $0.40 | $0.80 | $1.50 | Included, free |
| Veo 3.1 Fast · 720pGoogle's model, 4–8 s only | $0.35 | $0.70 | — | Costs extra |
| Gemini Omni Flash · 720pGoogle's model, by the second, 3–10 s | $0.40 | $0.80 | — | Always on |
Prices are per finished clip and include the audio track. Input images are free. A dash means the model cannot render that length in one run. Deposits of $50 add 5% to your balance and $100 adds 10%; an active promo code adds another 10% on top of the same deposit.
The full table for every model lives on the pricing page.
How starting frames, closing frames and loops actually behave, with the prompts that made them.
Read the guideCamera language, lighting, motion. Written for Veo, but the structure carries over to Wan almost unchanged.
Read the guideThe model to reach for when you need to edit a clip by talking to it instead of re-prompting.
Read the guideYes, always. Dialogue written in quotes is spoken and lip-synced, and named sound effects are rendered. Vague ambience requests often come back close to silent, so name a specific sound if you need one. Sound is included in the price.
You pay by the second: $0.14 at 720p and $0.25 at 1080p, the same rates xAI charges on its own API. So four seconds at 720p is $0.56, eight seconds is $1.12 or $2.00 depending on resolution, and the maximum — 15 seconds at 1080p — is $3.75. Input images cost nothing extra here.
Read the full answerNot today. Clips that start from a first frame or use reference images are 720p: xAI limits reference-guided video to 720p, and the native first-frame route we have access to does not accept 1080p yet. We reject a 1080p order with an image before charging instead of delivering a smaller file at the higher price. 1080p works for text-to-video.
Not on this model. Render up to 15 seconds per clip and cut longer pieces together, or use Wan 3.0 for a single take of up to 30 seconds, or Gemini Omni Flash, which extends its own clips up to 40.
No. It takes a first frame or up to 7 reference images, and that is all. For first-and-last-frame control use Veo 3.1, Omni or Wan 3.0; for seeds use Veo or Wan.
The full amount goes back to your balance automatically, content-filter refusals included. That is the same rule for every model on the platform.
Read the full answerYes. What you generate is yours to use, including in client and commercial work, subject to the content rules in our terms. There is no watermark on the output and no separate licence to buy.
New accounts start with $0.20 on the balance — enough to make the first frame. Clips from $0.56, no subscription, no commitment.