Blog

The ChatGPT prompt for a YouTube script that actually works, and the six pieces it leaves on your desk

Almost everyone typing a script prompt into a chat box gets something readable back in twenty seconds, and readable is the floor, not the target. This is the prompt structure that returns usable material, the word count that decides whether your video runs six minutes or twelve, and the honest accounting of everything that still has to happen after the text exists.

Ricardo AlmeidaFounder14 min read
A single glowing golden sheet of paper facing a production line of six golden slabs where only the first one is lit.

The prompt has seven parts, and six of them are things only you know

The prompt most people type is some version of write me a YouTube script about this topic. It returns clean paragraphs fast, and that speed is exactly what hides the problem: a language model will always produce something plausible, whether or not you gave it enough to work with. The quality of what comes back is decided almost entirely before you hit enter.

A prompt that returns material you can actually narrate carries seven parts. The model supplies one of them, the prose. You supply the other six, and every one you skip gets filled with the average of everything the model has ever read on the subject.

Write those six into the prompt itself instead of correcting them afterwards. Correcting a script that was built on the wrong assumption takes longer than writing the assumption down in the first place, because the wrong angle is spread across every paragraph rather than sitting in one line you can delete.

  • The audience and their level: someone who has never opened YouTube Studio needs a different first minute than someone with 40 videos published.
  • The runtime in minutes and the word budget that follows from it, stated as a number, because the model has no sense of screen time.
  • The angle in one sentence, meaning the specific claim this video makes that the other ten videos on the topic do not make.
  • The proof you already hold: your numbers, your sources, your screenshots, your experience. Without this the script argues from generalities.
  • The structure you want: hook, then how many blocks, then what the close asks for.
  • The voice and its constraints: first person or narrator, formal or blunt, the words you never want to hear in your channel, no invented statistics, no unsourced claims, and short sentences written to be spoken rather than read.

The word count that nobody checks: 1,680 words, not 900

Comfortable narration lands between 130 and 160 words per minute, and 140 is the safe center for explanatory content. That single ratio turns a runtime into a word budget: 5 minutes is 700 words, 8 minutes is 1,120, 10 minutes is 1,400, 12 minutes is 1,680, and 20 minutes is 2,800. If you want the full ruler and the five reasons it misses in practice, the word count math of a video script breaks it down by duration.

A chat box returns 700 to 900 words for a script request, because chat answers are shaped to be read on a screen rather than performed for twelve minutes. So the script that felt complete produces a six minute video, and the gap gets filled at the worst possible moment, during the edit, by stretching shots and slowing the read.

The fix is mechanical. Ask for the script block by block instead of whole: a hook of roughly 120 words, then five or six blocks of about 280 words each, then a close of 100. You get the runtime you planned, and you get something better than length, which is a structure you can cut without the whole piece collapsing.

One warning about the ruler: it counts speech, and a finished video contains runtime that no word occupies. Pauses between sentences, a beat of silence after a strong line, a music sting between blocks, the intro, the outro, and b-roll that plays with nothing over it. Budget the words at 140 per minute and expect the timeline to run slightly longer, never shorter.

Why ten channels in your niche just published the same script

Ask a model for a script on a popular topic and it converges on the shape that appears most often in its training material: a rhetorical question, a promise of what you will learn, three or five numbered points of roughly equal weight, a generic summary, a request to subscribe. It is competent, it is coherent, and it is the same skeleton your competitors received last week from the same prompt.

YouTube does not have a rule against that. Your viewer does. The first video of that shape gets watched, the second one gets recognised at the twenty second mark, and recognition is what a skip feels like from the inside. The retention graph shows it as a clean drop right after the hook, before any of your content has been heard.

Four moves break the pattern, and none of them need a better model. Open on the number instead of the question, because a specific figure is a promise a rhetorical question cannot make. Put your unequal point first rather than treating five points as equals. Include the counter example, the case where your own advice fails, which no generic script ever contains. And name the thing the topic usually avoids.

The deeper structural choice, whether the video is a list or a narrative, changes retention more than any sentence you rewrite. A list gives the viewer a natural exit at every item, a narrative gives one thread and no clean stopping point.

The one thing the model cannot look up: your own channel

A script prompt can reach for everything published about a topic and nothing at all about you. It does not know that your click through rate from search is 7 percent while your rate from the home feed is 3 percent, that your retention collapses at minute two on tutorials but holds to minute nine on teardowns, or that one of your 40 videos carried last month while the rest flattened.

That is why the prompt asking for a viral script disappoints so consistently. Virality is not a property of prose. It is the result of a specific audience, a specific offer and a specific track record, and the model has access to none of the three.

The workaround is to feed it. Paste your last three titles with their click through rates, the point where retention dropped, and the comment that repeated itself. Now the model has data instead of averages, and its suggestions stop being generic advice.

The problem is that gathering that context by hand, every week, for every channel, is the part people quietly stop doing after a month. Inside FalconVid it is not a manual step: the Senior AI Analyst reads your channels and your videos and writes to you every 2 days, in your language, saying what went up, what went down, which video carried the month and what the next move is. Plans from Pro upwards include that person permanently, and Starter tests it for 7 days from the panel.

The confident sentence that happens to be wrong

Every model writes wrong facts in exactly the same tone it writes right ones. There is no hesitation in the prose to warn you, which is precisely why script errors survive all the way to publication: the sentence reads well, the edit is about pacing, and nobody goes back to check a date that sounded reasonable.

On YouTube that has a cost beyond embarrassment. A corrected fact in a published video means either a pinned comment that most viewers never see, or a reupload that throws away the watch hours the video already earned. Neither is a repair.

Run a verification pass with a fixed list before the script becomes audio: every number, every date, every proper name, every price, every rule or policy, and every superlative. Superlatives deserve their own line on that list, because first, only and best are claims a model produces casually and a viewer disputes immediately in the comments.

For the anatomy of the errors that slip through most often, and the five categories worth checking first, what an AI script gets factually wrong has the pattern. The rule of thumb: if a sentence would embarrass you in a comment thread, it needs a source before it needs a voice.

The script is one of six pieces, and the other five are the expensive ones

Here is the accounting that turns a free script into an honest decision. A publishable video is six pieces, and the text is the cheapest of them in both money and hours.

Research comes first, because a script written without it is confident about things nobody verified. Then the script itself, 1,680 words for twelve minutes. Then narration, which is that text performed rather than read, with the pronunciation of names and numbers fixed. Then the visuals: 12 minutes is 720 seconds, and at honest pacing that is 72 to 120 cuts, since the first 30 seconds want cuts of 2 to 4 seconds while the calm stretches want 12 to 20. If you want that rhythm laid out properly, the cut rhythm of a 12 minute video has the arithmetic.

Then the cover, which decides whether any of the previous work is ever seen. Then publication itself: title, description, chapters, tags, end screen, schedule, and the same package again for every other network you post to.

Assembled by hand, even with AI helping at every single step, a 12 minute video costs 9.5 to 13.5 hours of work, or US$ 165 to US$ 570 outsourced piece by piece. The free script saved you perhaps 90 minutes of that. This is the number that decides whether a channel survives its third month, and the reason the search for a better script prompt is usually a search for the wrong thing. The same confusion drives people to clip generators, which is why a prompt is not a full video walks the same gap from the visual side.

Six golden panels in a row, only the first complete and the other five hollow outlines.

What the pipeline does with the same topic

FalconVid takes the topic and returns the finished video, and the difference from a chat box is not the writing. It is that research, script, narration, editing and sound design are handled by AI specialists working in parallel, on a calendar you approve once, with a finished video in up to 30 minutes.

Research plus script costs 307 credits, which is the number that matters most for anyone who has been rewriting prompts: changing the script costs 307, and regenerating the whole video costs 1,008 to 26,760. A complete 12 minute video is 1,008 credits in economy mode, 8,676 in balanced and 26,760 in premium, and the mode is a switch you flip per project, on every plan. Narration is 3 credits per minute in economy and 24 in premium, a cover is 479 in economy and 214 in premium.

You approve the calendar, not each script: the topics, the dates and the times. From there the videos build themselves before their slot and publish on schedule to YouTube, Instagram, TikTok, Rumble and Facebook, in up to 63 languages. Several channels run at the same time, each with its own calendar, identity and language, from 1 channel on Starter to 50 on Scale, and from 2 simultaneous generations on Starter to 50 on Scale.

For the fear behind every automation question, what if it comes out wrong, the answer is the Studio. You watch version one before it publishes and adjust: shorten the intro, swap a piece of media, change the music, or open the timeline. Those repairs cost 5 to 320 credits against 1,008 to 26,760 to regenerate, so being picky is cheap and being wrong is not expensive. The full stage by stage picture is on the long form video generator page.

You already have a script. Here is the honest decision

Keep using the chat box for what it is genuinely good at: breaking a topic into blocks, arguing with your angle, listing the objections your video should answer, generating twenty title variants so you can pick the one that is not a question. That is real value and it costs nothing.

The decision is not whether the model writes well. It is where your hours go. One video a week, written and produced by hand, is 9.5 to 13.5 hours per video and roughly two working days a week gone. Four videos a week is not that number times four, it is the point where a person stops sleeping and the calendar starts slipping, which is the actual reason most channels stop publishing in month three.

That ceiling belongs to the manual route, not to the format. Automated, the same four videos a week are generations running in parallel while you do something else, and 100 videos in economy mode cost 100,800 credits, roughly US$ 316, against 950 to 1,350 hours of your own time for the same 100. The relevant number for 2027 is that new channels entering the Partner Program need 8,000 qualified public watch hours in 365 days, and at 1,000 views per 12 minute video with 40 percent watched, that is about 100 videos.

So the honest sequence is this: get the angle from your own head, use the model to pressure test it, and stop trying to solve production with a prompt. The script was never the bottleneck. The other five pieces were, and they are the ones a pipeline was built to carry. Everything else on this site starts from the FalconVid home page.

FAQ

Got questions? We've got answers.

What is the best ChatGPT prompt for a YouTube script?

The one that carries seven parts. The model supplies one of them, the prose, and you supply the other six: the audience and their level, the runtime in minutes with the word budget spelled out, the angle in one sentence, the proof you already hold, the structure you want, and the voice with its constraints such as no invented statistics and short spoken sentences. Ask for the script block by block instead of whole, so a 12 minute video arrives at 1,680 words rather than 900.

How many words does ChatGPT need to write for a 10 minute video?

1,400 words at 140 words per minute, which is the safe center for explanatory narration. The comfortable band is 130 to 160 words per minute, so the same 1,400 words can land anywhere between roughly 9 and 11 minutes depending on the pace of the read.

Does YouTube penalise scripts written by AI?

No. YouTube judges the finished video, not the tool that produced the text. What gets punished is content that adds nothing, repeats an existing upload or misleads, and the required AI disclosure applies to realistic synthetic footage rather than to a written script. Weak scripts fail with viewers before they ever attract a policy problem.

Why does my AI script sound the same as other channels in my niche?

Because the same prompt returns the same skeleton: rhetorical question, promise, three or five equal points, generic summary, subscribe request. Break it by opening on a specific number, putting your unequal point first, including the counter example where your own advice fails, and naming what the topic usually avoids.

Do I still need to fact check the script?

Yes, every time. Models write wrong facts in the same confident tone as right ones. Check every number, date, proper name, price, rule and superlative before the text becomes narration, because correcting a published video means a pinned comment nobody reads or a reupload that throws away the watch hours already earned.

I have the script. What is left to do before I can publish?

Five pieces, and they are the expensive ones: research to verify what the script asserts, narration with pronunciation fixed, 72 to 120 cuts of visuals for a 12 minute video, a cover, and the publication package of title, description, chapters and schedule for each network. By hand that is 9.5 to 13.5 hours per video, or US$ 165 to US$ 570 outsourced.

Do I have to approve every script inside FalconVid?

No. You approve the calendar once, the topics, the dates and the times, and the videos build themselves before their slot and publish on schedule. Research plus script costs 307 credits, so if you want to change direction on a video the change is cheap, while regenerating the finished video costs 1,008 to 26,760 depending on the mode.

Is it worth paying for this if ChatGPT writes scripts for free?

The script was never the expensive part. A 12 minute video is 1,008 credits in economy mode, 8,676 in balanced and 26,760 in premium, against 9.5 to 13.5 hours of your own time to assemble the same thing by hand. Starter at US$ 47 with 15,000 credits covers about 10 to 12 videos a month mixing economy with one in balanced, and there is a 7 day trial with 2,000 credits plus a 7 day guarantee.

A prompt gives you text. A calendar gives you a channel.

FalconVid researches, writes, narrates, edits, captions and publishes to YouTube, Instagram, TikTok, Rumble and Facebook in up to 63 languages, from a calendar you approve once, with AI specialists working in parallel and a finished video in up to 30 minutes. Research plus script costs 307 credits and a full 12 minute video costs 1,008 credits in economy mode, 8,676 in balanced and 26,760 in premium, with engines like Veo 3 with audio and Seedance Pro, premium voices, 4K and captions included. Starter at US$ 47 with 15,000 credits, 1 channel and 2 simultaneous generations covers about 10 to 12 videos a month mixing economy with one in balanced. Pro at US$ 97 with 30,000 credits, 5 channels and 5 simultaneous generations adds the permanent Senior AI Analyst and a dedicated server, up to Scale at US$ 997 with 320,000 credits, 50 channels and 50 simultaneous generations. Every creation feature is on every plan: what changes by plan is volume, channels, simultaneous generations, dedicated server, the Senior AI Analyst and support, with a 7 day trial including 2,000 credits and a 7 day guarantee.

Create my channel now

Charged today · 7-day guarantee · Cancel anytime

Keep reading