Build an AI animated video with Canva and free AI tools: a step-by-step Canva animation workflow, model limits, and export settings.
Build an AI animated video with Canva in five stages
An AI animated video with Canva is assembled from five stages: a written script, still images, narration, image-to-video animation, and timeline editing. Canva covers the still-image stage only. The movement comes from a separate image-to-video model, and the final assembly happens in an editor like Cap Cut, so the workflow succeeds or fails on how well those stages hand data to each other.
| Stage | Tool in the tutorial | What it produces | Free-tier limit to watch |
|---|---|---|---|
| Script | ChatGPT | Title, story broken into scenes, narration and image prompts | Prompt structure matters more than the model |
| Stills | Canva AI image generator | Four candidate images per prompt | No character memory between prompts |
| Narration | ElevenLabs | One audio file per scene | Monthly character allowance caps script length |
| Animation | Grok | A moving clip from a still image | Automatic mode plus prompted reanimation |
| Editing and export | Cap Cut | Timeline, captions, transitions, 4K export | Manual sync work per scene |
The pipeline is not invented for this article. It is the exact sequence in MR.AI.JO's Canva animation tutorial, published in October 2025, which walks from a children's story script about a boy named Ben and a golden puppy called Buddy to a finished export. Each stage below names the free constraint that decides how far you can push it.
A companion walkthrough of the same method is embedded at the end of this article as the source video.
Write the script as scenes, not as paragraphs
A scene-by-scene script is the single decision that makes the rest of the workflow possible. Asking a writing model for a story plus, per scene, a scene description, a narrator line, an image prompt, and a repeated character description produces output that every later tool can consume without manual reformatting.
The prompt works in two passes:
- Generate a list of title ideas, pick the one you like, and copy it.
- Paste the story prompt, replace the bracket placeholder with your chosen title, and press enter. The model then expands it into a full story divided into scenes.
The consistency trick sits inside the script itself. Each scene's image prompt repeats the same character description, which is what stops a character from changing appearance between scenes. When a scene's still does come back with the wrong face or fur colour, the usual cause is a prompt that dropped the character description, not a fault in the image model.
ChatGPT, OpenAI's assistant, is one option for this step, and the tutorial uses it for both the title and the full story expansion. Any capable writing model works. What matters is the prompt structure, not the brand: title ideas first, then a story expanded into scenes with narration and image prompts attached to each one.
Keep narration lines short. Text-to-speech timing, caption placement, and scene length are all easier to control when a narrator line fits in one breath. A children's story that reads well aloud usually lands near 15 to 25 words per scene.
Generate consistent characters with Canva's image AI
Canva's AI image generator returns four variations per prompt, which lets you pick the strongest frame before spending time on animation. It is a still-image tool: it does not animate, and it does not remember your character between prompts.
In Canva, open the Canva AI feature and select "create an image". Paste the first image prompt and leave the style setting alone, because the style is already described inside the prompt text. Canva then produces the four variations, and you can tweak the prompt and regenerate any frame before moving on.
That lack of character memory is the whole game. Canva has no built-in character-lock feature, so consistency comes from pasting the same character description into every scene prompt and rejecting any variation where the character drifts. The tutorial's demonstration of a 3D-style puppy story shows the failure mode directly: the first prompt carried the full character description and the second one did not, so the character changed between scenes. When that happens, tweak the prompt so the full character description appears correctly, and regenerate that scene only.
A short style Claude pays for itself. Fixing a rendering style and sticking to it across all prompts keeps lighting, lens feel, and colour treatment consistent even when the character description is long.
Once every still is approved, download the ones you want to use. The images become the input for the animation stage.
Narrate with text-to-speech and match voice to story
Text-to-speech with voice filters replaces a recording session, and filtering by language, category, quality, gender, and age narrows thousands of voices down to a handful worth auditioning. ElevenLabs is the text-to-speech service used in the tutorial, and it offers both scene-by-scene and full-script generation.
ElevenLabs' published plans and pricing page is the source to check before committing, because the free allowance changes and it sets a hard ceiling on script length. As of September 2026 the company lists a Free plan with a monthly character allowance and a Starter tier priced at $5 per month, with the specific credit figures shown on that page rather than assumed here.
Generate audio one scene at a time. It costs more clicks but gives you a clean file per scene, which makes syncing in the editor far easier and lets you regenerate a single bad line without re-rendering the whole narration.
Voice quality is worth the extra auditioning. The tutorial treats voice choice as the step that decides how the finished video feels, and it is the one stage where a bad pick cannot be fixed later by editing.
Animate stills: what image-to-video does and does not do
Image-to-video models turn a still into a moving clip without a written prompt, and a second pass with a custom prompt can add a specific action. That two-mode design, automatic animation plus prompted animation, is what makes free image-to-video usable for storytelling rather than just for looping backgrounds.
Grok, xAI's assistant, is the tool used for this stage in the tutorial, where an uploaded image is switched to video and animated through an edit surface. Grok also chats like ChatGPT, creates images, and keeps separate projects, but for this workflow only the image editor matters.
The animation process is short:
- Create an account or log in.
- Upload or drag in one of the downloaded stills.
- Click edit image. The editor can change style, swap the background or add objects, but the tutorial leaves those alone.
- Click make video. Grok animates the image with no prompt required.
- Download the clip if you like it, or type a custom prompt and press enter to reanimate it.
The automatic result includes camera movement and a sense of flow at no cost to the user, and the tutorial notes that the first animation even arrived with background music. That is the reason the stage is viable on a free workflow at all.
Prompted reanimation is where the model earns its place. Asking for a character to say a specific line in a specific tone produced matching lip sync, facial expression, and voice in the tutorial's demonstration, so a scene that needs dialogue does not require a separate animation pass. For a scene where the character stays silent, a prompt that describes only the movement, such as walking, is enough. Grok keeps a history of your animated images, so uploading the next still auto-animates it while you review the previous result.
Treat the generated clip as a starting point. Some clips read as ambient drift rather than acting. Review each clip before it enters the timeline and reanimate the ones where the camera movement contradicts the narration.
Edit the timeline and export in 4K with Cap Cut
Cap Cut is where the five stages become one video. Import the narration audio, the animated clips, and any extra assets, then assemble in this order:
- Drag the audio track onto the timeline.
- Add the animated scenes on top of it.
- Rename each scene so the order stays trackable.
- Change the aspect ratio to match where the video will be published.
- Match each animation's length to its narration, trimming the clip rather than the line.
- Add smooth transitions between scenes.
- Run the auto caption tool, pick a template, then adjust font, colour, and size to match the video's tone.
- Add sound effects, background music, and playful effects for personality.
- Play the video end to end and fix rough spots.
- Export at 4K for a crisp finish.
Two of those steps are easy to underestimate. Renaming scenes sounds cosmetic, but a timeline with 11 or more unlabeled clips is where sync errors hide. And matching animation to narration is the step that decides whether a clip reads as acting or as random movement.
The aspect ratio change matters as much as the export setting. A vertical cut for social platforms and a 16:9 cut for YouTube are different edits of the same timeline, not different exports.
What the free constraints actually cost you
None of the five stages require paid software, but free tiers trade money for time and control. Here is where each one bites:
- Canva's image generator gives you four variations per prompt but no character memory, so consistency rests on your prompt discipline.
- ElevenLabs' free allowance caps how long a script you can narrate before the $5 Starter tier becomes the cheaper option.
- Image-to-video generation is free in the tutorial's workflow, but each clip still needs a human review pass before it enters the timeline.
- Cap Cut handles captions, transitions and 4K export without a subscription, but every sync and rename is manual.
The tutorial's own claim is that the whole animation took a few hours, with no animation skills and no paid software. The honest version of that claim is that the hours go into the review-and-regenerate loop, not into the generation itself.
FAQ
How long does an AI animated video take to make with Canva? The tutorial reports a few hours for a complete short story, start to finish, working only with AI tools and Canva. Most of that time goes into reviewing generated images and clips and regenerating the ones that drifted, not into the generation itself.
Why does my AI character look different in every scene? Canva's image generator does not remember your character between prompts, and there is no character-lock feature. The fix is to keep the full character description inside every scene's image prompt. In the tutorial, the drift appeared exactly when the second prompt omitted the description the first one carried.
Is the whole AI animation workflow really free? The pipeline the tutorial demonstrates uses free tiers throughout. Canva's image generator, Grok's image-to-video, and Cap Cut's editor and 4K export carry no cost, while ElevenLabs offers a Free plan with a monthly character allowance alongside a $5 per month Starter tier. Longer narration is where the free ceiling shows up first.
Can I make a character talk in an AI animated video? Yes. Upload the still to Grok, click make video for the automatic pass, then type a custom prompt that gives the character a specific line and tone. The tutorial's test produced matching lip sync, facial expression, and voice from that single prompted reanimation.
What export settings should I use for a Canva AI animation? Set the aspect ratio in Cap Cut before you fine-tune timing, because a vertical and a widescreen version are different edits. For the final file, the tutorial recommends 4K to keep the result crisp.
Turn this workflow into written content
The part of this pipeline that is genuinely hard is not the generation, it is the structure: a script split into scenes, each scene carrying its own prompt, description and narration line. Get that skeleton right and the rest is review work. That same skeleton is what makes a video easy to turn into an article, because the scenes already carry the narrative order, the explanation, and the voice.
That is the idea behind Skala Blog. If you have knowledge, lessons, interviews or opinions sitting inside a YouTube video, paste the video URL, let it transcribe the recording, and generate a written article from it. The video keeps the audience you already have; the article gives that same material a second life where people search for it.
Gustavo dev doido
Fork this article
Start a new branch from the same video, shaped your way. You keep the credit; the original keeps the attribution.
A fork in another language is filed as a translation of this article, so the two pages point at each other. You can unlink it later from the editor.
0/240
You are creating
- Format
- For
- Language
- Source
- Your angle
You will be asked to sign in before it is generated.
Buy credits