InVideo AI V4.0 – Complete Guide: Text-to-Video
InVideo AI V4.0 – The Complete Guide: From Text to Finished Video
Traditional video editors require hours of learning, expensive equipment, and often an entire team. InVideo AI V4.0 turns this on its head: you write a single sentence or upload an article, and the platform creates a finished video – complete with script, voiceover, cuts, and music.
This guide will cover what you can achieve with the platform, how to use it step-by-step, and the best tips for maximizing quality.
Part 1 – What can InVideo AI V4.0 do?
InVideo AI V4.0 is not just a video editor – it's an entire video studio controlled by text commands. Version 4.0 brought a significant leap forward: a shift from traditional stock clips to pure generative AI, which creates every image and movement from scratch.
The platform is suitable for industrial-level production and scaling of various video content. Here are four main use cases:
You can create digital characters (AI Actors) who speak your script directly to the camera with dynamic cuts and text animations. The characters react naturally to speech – without you needing to appear on camera.
You can maintain YouTube or email marketing channels where AI handles the visualization and speaks with your cloned voice. Your face is never seen – but your voice is heard in every video.
By feeding the platform articles or PDF files, it can construct structured Hook → Problem → Solution explainer videos without the AI making up facts. The Notebook knowledge base keeps the content fact-based.
Thanks to the advanced Wan 2.7 video model, you can create fluid motion, light leaks, and depth of field for product shots. The result looks like professional advertising studio production.
🔍 What sets V4.0 apart from previous versions? Previous versions primarily used pre-made stock clips arranged by AI. V4.0 generates images and motion from scratch with the Wan 2.7 model – this means you can request very specific visuals that are not found in any stock library.
Part 2 – Step-by-step User Guide
Successful video production requires the correct sequence of steps to avoid wasting the platform's monthly video credits on unnecessary experiments. Follow these five steps in order – each step builds upon the previous one.
-
Locking Facts into the Knowledge Base (Notebook)
Don't start by writing a prompt in an empty field – this is the most common mistake. First, open the Notebook section from the left menu and upload the video's background material there: affiliate product specs, a blog article, product description, or your company's information.
The Notebook acts as the AI's fact bank. When the AI generates a script, it first retrieves information from the Notebook and does not invent it itself. This prevents so-called hallucinations – situations where the AI invents incorrect facts or prices.
✅ Tip: Also upload competitor analysis or customer reviews to the Notebook – the AI can use them to build more convincing arguments. -
Constructing a Two-Level Prompt
Write the main prompt in the workspace and clearly divide it into two parts. The first part deals with visual instructions: camera angles (close-up, wide shot), lighting (fluorescent, sunlight, studio light), cutting speed (fast, slow, cinematic), and color scheme.
The second part defines speech and voice: tone of voice (enthusiastic, calm, authoritative), speech speed, and rhythm. Use square brackets to separate the instructions:
[Create a fast-paced vertical video. Style: Cinematic 35mm, moody lighting. Voice: calm and authoritative, slow pace.]✅ Tip: The more precise the prompt, the less manual adjustment you'll need afterwards. Spend 5 minutes refining the prompt – you'll save 30 minutes in editing. -
Selecting Voice and Generation Engine
Before generation, you have two important choices. Voice: Choose Voice Cloning if you have uploaded at least a 3-minute voice sample to the system. A cloned voice builds trust with viewers and makes the video instantly recognizable. If voice cloning is not available, select one of the platform's pre-made voices that best suits the tone of your video.
Video Engine: Select Wan 2.7 when you want purely generated cinematic clips. If your topic is general (e.g., "person at a computer"), you can also use stock footage – it's faster. For specific topics, Wan 2.7 is the only option.
⚠️ Note: Voice Cloning requires the voice sample to be recorded clearly without background noise. A poor sample will produce an unnatural clone. -
Checking Visual Continuity (Boards)
Once the video is generated (usually takes 5–15 minutes), don't download it immediately. Open the Boards tool, which divides the video into individual scenes on a timeline.
Go through the scenes one by one. Look for places where the clip is too generic, looks wrong, or doesn't fit the context. Edit only the instruction for that specific scene – the other scenes will remain unchanged:
[Replace with a close-up of hands typing on a laptop, warm office lighting]✅ Tip: Pay particular attention to the first 3 seconds – it determines whether the viewer stays or scrolls past. -
Fine-tuning with Text Commands (Edit with AI)
Once the visuals are in order, finalize the video with the Edit with AI text box. This tool modifies general settings for the entire video without you needing to re-render from scratch. For example, you can give these commands:
"Change background music to more dynamic hip-hop"
"Add dynamic captions for trigger words"
"Shorten the video to 60 seconds by removing the slowest scenes"Once you are satisfied, download the video in your desired format (MP4, 1080p, or 4K with the Plus package).
⚠️ Note: Edit with AI commands also consume credits, but significantly less than generating an entirely new video.
💡 Rule of thumb: Spend 80% of your time preparing the prompt and Notebook, 20% on post-editing. A good initial setup produces a better video than even the best editing of a poor starting point.
⚡ Part 3 – Pro Tips to Simplify Use and Improve Quality
- Never regenerate, but iterate If you want to change the video's cuts or music, always use the Edit with AI box or the Boards view. Requesting an entirely new video will quickly deplete your credit balance and also change parts of the video that were already good. 💡 In practice: If the music doesn't fit, type "Change background music to upbeat lo-fi" into the Edit with AI field – done in 30 seconds.
- Use the Video Reference feature to lock characters If you are making a video series where the same character or product needs to appear from different angles, upload a reference image to the system before generation. This forces the AI to keep facial features and product shapes identical throughout the video. 💡 In practice: Upload an image of your product as a reference – the AI will use it as a base for all scenes where the product appears.
- Optimize for mobile platforms (Muted Viewing) Since over 70% of people watch short social media videos without sound, ask the AI to add Kinetic typography overlays (dynamic text animations) for key keywords directly in the prompt. This way, the video works even without sound. 💡 In practice: Add "Add bold kinetic text overlays for key words, optimized for silent viewing" to the prompt.
- Clean up stock footage immediately after generation By default, AI often offers easy and somewhat artificial stock videos (such as overly cheerful people in an office). Spend 10 minutes after video creation replacing these sections with either Wan 2.7 model generations or your own material via the Edit media menu. 💡 In practice: In the Boards view, find scenes with smiling people in an office – these are almost always generic and should be changed.
Ready to get started? Try InVideo AI for free and create your first video today!
🚀 Try InVideo AI for Free →