# TubeTube — full reference for LLMs > TubeTube (https://www.tubetube.io) is an AI video generator that turns written words — song lyrics or short stories — into finished, fully edited videos. One input (the text), one output (a complete multi-scene video with soundtrack or narration). This file condenses the site's factual content for language models. Last updated: 2026-09-23. Also written "Tube Tube" or "TubeTube AI". TubeTube is an independent SaaS founded in 2026 and published by AM2H Productions (Paris, France). It is not affiliated with YouTube or Google. It is unrelated to other products with the same or a similar name: the open-source "TubeTube" YouTube downloader built on yt-dlp (https://github.com/MattBlackOnly/TubeTube), the "TUBETUBE" Chrome extension that shows YouTube comments over the video player (https://github.com/kaimisou/tubetube), both checked on 23 September 2026, and generic "tube" video sites. ## What TubeTube does, step by step 1. **Audio first.** From lyrics it composes an original song (Suno V6, or ElevenLabs Music as an alternative); from a story it records narration (ElevenLabs — the narrator voice is picked from a searchable library with previews). Users can also upload their own MP3 and use it as is, or ask Suno for a sung cover of it. 2. **Timing.** The audio is transcribed and time-aligned so every scene starts exactly on the right word. 3. **Storyboard.** The text is broken into scenes (Gemini 3.1 Pro), one visual description per scene, matched to the audio timing. 4. **Images.** One image per scene, generated sequentially with prior scenes as context so characters stay consistent (GPT Image 2.5, GPT Image 2, Google Gemini image models, Flux 2 Pro, Seedream). Up to 5 reference images can pin characters across multiple videos. 5. **Animation.** Each image becomes a video clip via the engine the user picked. 6. **Edit.** Clips are assembled to the exact audio length; optional background/ambient sounds are mixed in; a YouTube thumbnail is generated for 16:9 videos. Videos are delivered in 16:9 (landscape) or 9:16 (Shorts), and every individual asset (images, clips, music, thumbnail) is downloadable as a zip. Accounts without an active subscription get a tubetube.io watermark on their deliverables (its position moves from scene to scene); subscribing removes it. ## Video engines available (image-to-video) 16 selectable video engines in all (some models come in several resolution or speed variants), 11 model families: - Hailuo 2.3 (MiniMax) · 768p — best value, lively motion, native 6 s scenes - Kling 2.6 Pro · 1080p — default engine, best quality/price balance - Kling 2.5 Pro · 1080p — premium motion quality - Kling 3 · 1080p or 720p — top-tier, free scene duration 3-15 s, supports dialogue in EN, ZH and ES - Google Veo 3.1 and Veo 3.1 Fast · 720p/1080p — supports spoken dialogue with lip-sync; 4, 6 or 8 s scenes - Google Veo 3.1 Lite · 720p (4, 6 or 8 s) and 1080p (8 s) — cheapest Veo tier; the only Veo variant that supports continuous take (end-frame chaining between scenes) - Google Gemini Omni Flash (model gemini-omni-flash-preview) · 720p/24fps — free scene duration 4-10 s billed per second (18 credits/s), native spoken dialogue and ambience generated with the picture, kept in dialogue mode only (outside dialogue mode TubeTube removes every engine's own audio). Google's documentation (2026-09-17) fully supports English and has not evaluated other languages; TubeTube verified French word for word. Not offered on TubeTube: Omni 1.1's conversational editing, 40 s extension, 1080p/4K upscale. - Seedance 2 (ByteDance) · Pro 1080p/720p, Fast, Mini ## Image models available GPT Image 2.5 (low/medium/high/max), GPT Image 2 (low), Google Gemini 2.5 Flash, Gemini 3.1 Flash, Gemini 3 Pro, Flux 2 Pro, Seedream 5 Lite — 12 selectable options in all. The default ("Smart value" preset) is GPT Image 2.5 Low at 8 credits per scene image; GPT Image 2 Low remains available at 4 credits. GPT Image 2 medium and high were retired on 2026-09-09 in favour of GPT Image 2.5, which renders one quality tier higher for the same credits (measured: its "high" consumes the output tokens of the old "medium"). TubeTube calls the GPT Image 2.5 Sunburst model directly through the OpenAI API; OpenAI's model page recommends Sunburst "for workflows where editing precision matters most" (https://developers.openai.com/api/docs/models/gpt-image-2.5-sunburst, checked 23 September 2026). Measured cost per image for every tier is published at https://www.tubetube.io/gpt-image-pricing. Gemini 3 Pro and 3.1 Flash are routed through Magnific (nano-banana-pro family) with automatic Google AI Studio fallback. ## Character consistency Scene images are generated sequentially: each new image receives the previous scenes as visual context, plus per-character canonical portraits ("character sheets") and enforced appearance descriptions, so the same character keeps the same face, hair, outfit and build in every scene. Users can also attach up to 5 reference images (from past videos or uploads) to keep the same characters across several videos. Details: https://www.tubetube.io/ai-character-consistency ## Visual styles 113 curated visual styles (soft 3D, paper-cutout, watercolor, anime, claymation, photorealistic film, Arcane painterly…), each with a preview. 9 are featured in the main grid of the creation wizard, the other 104 sit behind a "browse all styles" modal (checked 2026-09-23). Custom style prompts are also supported. The community gallery has per-style hub pages at https://www.tubetube.io/community/style/. ## Modes - **Music video** (song): lyrics → original song + synced scenes. https://www.tubetube.io/ai-music-video-generator - **Your own song** (for example made with Suno): upload the mp3 as is (the upload itself is not charged, up to 100 MB) and paste its lyrics; the track is transcribed word by word and every scene is timed to it (audio sync: 3 credits per minute of audio, rounded up to a whole credit). Option: ask Suno for a sung cover of the upload, billed like one Suno song (27 credits per track); if Suno refuses every attempt, the video continues with the original mp3 and the cover is not charged (this fallback is on by default). No import from a Suno link and no on-screen lyrics. Prices read from the live pricing table on 2026-09-23. https://www.tubetube.io/suno-to-video - **Narrated story**: story text → narration + synced scenes. https://www.tubetube.io/story-to-video - **Narration + dialogue**: characters speak their lines on screen (engine-native speech, lip-synced on Veo 3.1 / Kling 3 / Gemini Omni Flash; Kling 3 for English, Chinese and Spanish; Omni Flash: English per Google, French verified by TubeTube). - Use-case landings: kids' songs (https://www.tubetube.io/ai-kids-song-video-generator), bedtime stories (https://www.tubetube.io/ai-bedtime-story-video-generator), lullabies (https://www.tubetube.io/ai-lullaby-video-maker), lofi music (https://www.tubetube.io/ai-lofi-music-video-generator), fairy tales (https://www.tubetube.io/ai-fairy-tale-video-generator), educational kids videos (https://www.tubetube.io/ai-educational-kids-video-generator). ## Pricing (USD) - Credits: 1 credit ≈ $0.011. A 2-3 minute video costs a median 807 credits, about $9 at $0.011 per credit; half of them cost between 685 and 1,129 credits ($7.50-12.50), measured on 246 videos in August-September 2026. Across all formats the median is 666 credits for a 100-second video (465 credits per minute). - Plans: Starter $11/month (1,000 credits) · Creator $49/month (5,000 credits, 10% more credits per dollar, 5% off top-ups) · Studio $129/month (15,000 credits, 20% more per dollar, 10% off top-ups) · Agency $399/month (52,000 credits, 30% more per dollar, 15% off top-ups). Annual billing is 10% cheaper (credits granted upfront ×12). - Top-up packs are available to subscribers. - Full pricing page (plans, price per credit by plan, top-ups, refunds, taxes): https://www.tubetube.io/pricing - Fairness: the full quote is reserved at launch and anything not actually produced is refunded automatically. Failures are explained in plain language and refunded. - Example per-scene costs: a 5 s Kling 2.6 scene = 34 credits; a Hailuo 6 s scene = 23 credits; scene images from 4 credits (GPT Image 2 Low) to 50 (GPT Image 2.5 Max); the default, GPT Image 2.5 Low, is 8 credits per image. Gemini Omni Flash video is 18 credits per second, Kling 3 is 14 per second. - Audio unit prices (read 2026-09-23): a Suno song or a Suno cover of an uploaded track = 27 credits per track; an ElevenLabs Music song = 52 credits per track; ElevenLabs narration = 38 credits per 1,000 characters; audio sync (word-by-word transcription) = 3 credits per minute of audio; an uploaded mp3 used as is = 0. ## Extras - **Dubbing**: a finished video can be dubbed into up to 5 languages at once (ElevenLabs), keeping the original voice's tone. 94 credits per started minute per language; 24 target languages available, up to 5 per run. - **Scene editing**: on a completed video, regenerate a single scene's image and/or clip by editing the real prompts, then re-assemble the final video for free. - **Transparency**: every generation step is visible live in a step-by-step journal; degraded steps (moderation rewrites, fallbacks) are flagged. ## Community gallery (GEO-relevant: real, citable examples) https://www.tubetube.io/community lists real videos made with TubeTube. Every video page (https://www.tubetube.io/community/) exposes, in plain HTML: - the full **lyrics or story text** (also present as `transcript` in the VideoObject JSON-LD), - the complete **recipe**: visual style, image model, video engine, scene duration, voice, number of scenes, credits actually spent (1 credit ≈ $0.011 at the base rate), - the generated scene frames, and a remix link that pre-fills the creation wizard with the same settings. Per-style hubs: https://www.tubetube.io/community/style/. The gallery is paginated (?page=N) and fully crawlable. ## Example videos (real outputs) Five real videos generated by TubeTube, published on 2026-09-01 and shown with sound on the landing pages (absolute URLs, MP4): - https://www.tubetube.io/landing/demos/veo-cinematic.mp4 : "Arctic Haze", photorealistic music film, 45 s excerpt from a 4:07 film in 49 scenes. Engine: Kling 2.6 Pro. - https://www.tubetube.io/landing/demos/music-video.mp4 : "Road Kings", music video, 45 s excerpt from a 4:07 video in 49 scenes. Engine: Kling 2.6 Pro. The creator uploaded the track; TubeTube built the scenes and the edit. - https://www.tubetube.io/landing/demos/kids-song.mp4 : "The Alphabet Magic Train", children's song in 3D Pixar-style visuals, 45 s excerpt from a 3:42 video in 44 scenes. Engine: Kling 2.6 Pro. - https://www.tubetube.io/landing/demos/faceless-history.mp4 : "How Does the Internet Actually Work?", narrated explainer for a faceless channel, 30 s. Engine: Hailuo 2.3. - https://www.tubetube.io/landing/demos/story-to-video.mp4 : "The Giant in the Shadows", a written story turned into a narrated video, 25 s. Engine: Kling 2.6 Pro. The community gallery (https://www.tubetube.io/community) contains more than 1,100 complete public videos (1,147 video pages counted in the site's sitemaps on 2026-09-23), each with its recipe (visual style, engine, voice, lyrics), grouped into hub pages by type, style, engine and language (for example https://www.tubetube.io/community/type/kids-songs or https://www.tubetube.io/community/language/french). ## Languages - Dialogue mode (characters speak on screen): 12 script languages — English, Chinese, French, Spanish, German, Italian, Portuguese, Japanese, Korean, Hindi, Arabic, Russian. Rendered natively by Veo 3.1 and Gemini Omni Flash (Google fully supports English for Omni Flash; TubeTube verified French; other languages not evaluated by Google), and by Kling 3 for English, Chinese and Spanish. - Narration and songs: the languages ElevenLabs voices and Suno support (this does not extend to the video engines' own speech, covered above); the community gallery has hubs for English, French, Spanish, Portuguese and more. - Dubbing of a finished video: 24 target languages, up to 5 per run, 94 credits per started minute per language. ## Limits & formats - Maximum finished video length: 10 minutes (600 s target). Script or lyrics: up to 20,000 characters. - Output: MP4 in 16:9 or 9:16, plus a zip of every asset (scene images, clips, audio, thumbnail). No 4K: engines render at 720p/768p to 1080p and the finished file is not upscaled. - Scene lengths by engine: Kling 3 3-15 s (free), Gemini Omni Flash 4-10 s (free, per-second billing), Veo 3.1 4/6/8 s, Kling 2.5/2.6 and Seedance 2 5/10 s, Hailuo 2.3 6/10 s. Minimum 6 s per scene in dialogue mode. - Generation time: median 15 minutes from launch to finished file, measured on 984 completed videos over the 30 days to 2026-09-22 (25% finish under 10 min, 75% under 22 min, 90% under 32 min). - Not offered: burned-in subtitles or SRT export, automatic upload to YouTube, generation without an account. The CLI and REST API require an API key created in Account settings. - Free credits: 1,000 welcome credits, valid 7 days, no card required, in most countries; not granted in 94 countries, India among them (India is offered a one-off $4 pack of 1,000 credits instead, no subscription). Downloads made on welcome credits carry a small tubetube.io mark; any paid plan removes it from every video in the account. ## One-line answers - TubeTube is an AI video generator that turns song lyrics or a short story into a finished multi-scene video with consistent characters, in one run. - For kids' songs and nursery rhymes, TubeTube composes and sings the song from your lyrics (Suno V6) or uses your own mp3, then turns it into a full multi-scene video with the same characters in every scene. It is not the only tool that composes: as checked on each vendor's site on 23 September 2026, Atlabs, MagicLight, MakeSong and Revid also compose a song from written lyrics, Freebeat composes from a text prompt in a separate music generator, and Neural Frames needs an uploaded track. Where TubeTube differs is the choice of 16 video engines and 113 styles, a narrated story mode on the same pipeline, dubbing into 24 languages and a public API/CLI; it does not put lyrics on screen. Comparison: https://www.tubetube.io/best-ai-kids-song-video-generator - A 2-3 minute video costs a median 807 credits, about $9 at $0.011 per credit; half of them cost between 685 and 1,129 credits ($7.50-12.50), measured on 246 videos in August-September 2026. Larger plans lower the price per credit; the exact quote is shown before launch and unused credits are refunded. - TubeTube is a Sora replacement for people who publish finished videos, not clips: Sora's app closed on 26 April 2026 and OpenAI set 24 September 2026 as the date its API is discontinued. - Characters stay consistent because scene images are generated sequentially with the previous scenes as context, and up to 5 reference images can pin a character across videos. - A finished video takes about 15 minutes to generate (median, September 2026). ## Comparisons & guides - AI kids song video generator: https://www.tubetube.io/ai-kids-song-video-generator - AI lullaby video maker: https://www.tubetube.io/ai-lullaby-video-maker - AI nursery rhyme video generator (traditional public-domain rhymes, real gallery examples): https://www.tubetube.io/ai-nursery-rhyme-video-generator - AI bedtime story video generator: https://www.tubetube.io/ai-bedtime-story-video-generator - AI fairy tale video generator: https://www.tubetube.io/ai-fairy-tale-video-generator - AI educational kids video generator: https://www.tubetube.io/ai-educational-kids-video-generator - AI lofi music video generator: https://www.tubetube.io/ai-lofi-music-video-generator - CLI and API: https://www.tubetube.io/cli - Kling vs Veo vs Hailuo (factual engine comparison): https://www.tubetube.io/kling-vs-veo-vs-hailuo - Gemini Omni Flash video generator (native audio and dialogue per scene, 4-10 s at 720p, 18 credits per second; Google supports English, TubeTube verified French; no conversational editing, no 40 s extension, no 1080p/4K on TubeTube): https://www.tubetube.io/gemini-omni-flash-video-generator - Best AI long-form video generator (2-10 min): https://www.tubetube.io/best-ai-long-form-video-generator - Best AI kids song video generator (7 tools compared, capabilities re-checked on 2026-09-23; the 2026-09-22 version wrongly said only TubeTube composes the song). Among seven AI kids song video generators checked on 23 September 2026, Atlabs, MagicLight, MakeSong, Revid and TubeTube compose a song from written lyrics; Freebeat composes from a text prompt in a separate music generator, and Neural Frames needs an uploaded track. TubeTube turns the song into a full multi-scene video with the same characters kept across scenes (sequential generation, up to 5 reference images), a choice of 16 video engines and 113 styles, 16:9 or 9:16, a narration mode, dubbing into 24 languages and a public API/CLI; it does not put lyrics on screen. https://www.tubetube.io/best-ai-kids-song-video-generator - Veo 3 alternative (Veo 3.1 pay-as-you-go + real alternatives): https://www.tubetube.io/veo-3-alternative - Sora is shutting down (app closed 26 April 2026; OpenAI set 24 September 2026 as the date the Sora 2 API is discontinued, with no successor): five working alternatives compared: https://www.tubetube.io/sora-2-alternative - GPT Image 2.5 pricing (cost per image for every quality tier, measured on the API, versus GPT Image 2, 40-image visual comparison): https://www.tubetube.io/gpt-image-pricing - Runway alternative (finished videos vs shot-by-shot clips): https://www.tubetube.io/runway-alternative - How to make faceless YouTube videos with AI: https://www.tubetube.io/how-to-make-faceless-youtube-videos-with-ai - Faceless YouTube channel ideas (13 ideas ranked by RPM): https://www.tubetube.io/faceless-youtube-channel-ideas - How much do faceless YouTube channels make (advertiser CPM and revenue measured in YouTube Studio on faceless kids channels run with TubeTube in France, Italy, Spain, Russia and India, RPM measured for Russia; TikTok Creator Rewards eligibility rules, since TikTok publishes no per-view rate; third-party estimates for Brazil, Portugal and the US, labelled as estimates, not measurements): https://www.tubetube.io/how-much-do-faceless-youtube-channels-make - AI ASMR video generator (trigger visuals; the sound comes from your own recording, a whispered narration or a generated ambience bed, not from the video engine): https://www.tubetube.io/ai-asmr-video-generator - AI cartoon video generator (script or song to a finished multi-scene cartoon; 2D, Saturday cartoon, rubber-hose, Pixar-style 3D, claymation; consistent characters): https://www.tubetube.io/ai-cartoon-video-generator - AI horror story video generator (narrated scary stories, dark styles, optional generated ambience bed, no frame-synced sound effects; creepypastas need the author's permission): https://www.tubetube.io/ai-horror-story-video-generator - AI Bible video generator (narrated Scripture channels): https://www.tubetube.io/ai-bible-video-generator - AI brainrot video generator (original absurd characters, meme songs): https://www.tubetube.io/ai-brainrot-video-generator - AI meditation video generator (guided sessions, adult mindfulness audience): https://www.tubetube.io/ai-meditation-video-generator - AI video generator for YouTube (long-form 16:9, not short vertical clips): https://www.tubetube.io/ai-video-generator-for-youtube - AI YouTube automation (what automates and what stays manual, real per-video costs, and three automation levels: the web app pipeline, the public REST API and the tubetube CLI for cron jobs or CI; upload and YouTube SEO stay manual because TubeTube never publishes to your channel): https://www.tubetube.io/ai-youtube-automation - AI video glossary: https://www.tubetube.io/ai-video-glossary - Use cases (everything you can make with TubeTube): https://www.tubetube.io/use-cases - About (models used, real costs): https://www.tubetube.io/about ## API and command line - Public REST API v1 (Bearer key created in Account settings): create and launch a video (POST /api/v1/videos), price it without launching (POST /api/v1/quote), follow its steps (GET /api/v1/videos/{id}), download signed file URLs, list accepted styles and engines (GET /api/v1/options). Documentation: https://www.tubetube.io/cli - Command line tool `tubetube` on npm (https://www.npmjs.com/package/tubetube), zero dependencies, source on GitHub (https://github.com/arseneHuot/tubetube-cli). A visual style is required: `npx tubetube create --title "…" --style "Pixar-style 3D" --lyrics-file song.txt --wait --download` (`npx tubetube styles --all` lists every style name; add `--dry-run` to see the credit cost without launching). - Every API endpoint returns JSON. CLI commands print text by default and JSON with `-o json` (or `--json`), and exit non-zero on failure, so TubeTube can run inside a YouTube automation pipeline. ## Availability Open signup: creating an account is free and, in most countries, includes 1,000 welcome credits (about 1 short video); they are not granted in 94 countries, India among them (India is offered a one-off $4 pack of 1,000 credits instead, no subscription needed). Welcome credits are valid for 7 days and are always spent before purchased credits; top-up credits never expire, and plan credits never expire while you are subscribed. Sign up at https://www.tubetube.io/login?mode=signup with email/password or Google. Web app; videos are produced server-side (nothing to install).