Four out of ten cited videos have fewer than 1,000 views
Otterly.ai published a study on 2 March 2026 that looked at more than 100 million AI citation instances across a 30 day window, spanning ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, Microsoft Copilot and Gemini. One number in it should stop anyone who has ever felt too small to compete: 40.83% of the YouTube citations went to videos with under 1,000 views.
Not 4%. Forty point eight three percent. Four in every ten times an AI system quoted a YouTube video as its source, that video had an audience you could fit in a school assembly. And 35% of the channels being cited had fewer than 10,000 subscribers. The median cited channel had 41 videos published in total.
This is not a rounding error or a quirk of one platform. It is a structural difference between how a recommendation engine works and how a retrieval engine works. The YouTube homepage is a popularity machine: it shows you what people like you watched to the end. An AI answer is a reference machine: it needs a source that states the fact cleanly, and it does not care who else has watched it.
The practical consequence is uncomfortable for anyone who spent two years chasing the algorithm, and very good news for anyone starting now. There is one distribution channel in 2026 where a brand new channel and a 500,000 subscriber channel enter on genuinely equal footing, and almost nobody is optimising for it.
Where the citations actually come from, platform by platform
The second study is Surfer's, published in May 2026, built on 46 million citations pulled from 36 million AI Overviews between March and August 2025. It answers a different question: of everything Google's AI Overview cites, what share is YouTube? The answer is 23.3%. Wikipedia is second at 18.4% and Google.com third at 16.4%. Every legacy news brand sits below those three.
So YouTube is the single most cited domain in Google's AI answers. That is worth reading twice, because it means video is not a side channel to the written web any more. It is the primary reference layer that the written answer is built on top of.
The Otterly numbers then split YouTube citations by which assistant is doing the citing, and the spread is brutal:
- Perplexity: 38.7% of all YouTube citations
- Google AI Overviews: 36.6%
- Google AI Mode: 19.6%
- ChatGPT: 4.4%
- Microsoft Copilot: 0.5%
- Gemini: 0.2%
Read that table before you build a strategy around ChatGPT
Three quarters of the opportunity sits inside Google's own surfaces once you add AI Overviews and AI Mode together, and Perplexity punches far above its user count because it is built citation first. ChatGPT, which is the assistant most people picture when they say AI search, accounts for 4.4% of YouTube citations. That does not make it worthless, it makes it a different game with different mechanics.
There is also a measurable payoff attached to being cited at all. Analyses of AI Overview behaviour in 2026 found that brands cited inside an AI Overview earn roughly 120% more organic clicks per impression than uncited brands appearing on the same queries. Meanwhile zero click searches on Google reached 68% in early 2026. Both things are true at once: fewer people click, and the ones who do click go disproportionately to whoever got quoted.
None of these are official YouTube or Google figures. Neither company publishes citation share. These are third party measurements from tools that watch AI answers at scale, and they should be read as strong directional evidence, not as a published rate card. That distinction matters more here than in most posts, because the whole field is eighteen months old.
If you want the grounding on how impossible it is to get a real number out of the platform itself, we went through that in detail when we looked at what YouTube search volume data actually exists. The short version: the platform never published volume, and every tool that shows you one is modelling.
What correlates with being cited, and what does not
Otterly ran Pearson correlations between citation frequency and every obvious video attribute. A Pearson coefficient runs from -1 to 1, and anything hovering around zero means the two things move independently. Here is the table, and it is the most useful thing published about video discovery this year.
Views: -0.03. Likes: -0.02. Subscribers: -0.03. Duration: 0.02. Every metric that YouTube puts on the front of your dashboard sits at statistical zero. A negative 0.03 does not mean views hurt you, it means there is no relationship at all.
Then two things actually move. Description length: 0.31. Descriptions containing hashtags: 0.20. Recency: around 0.3. A 0.31 is a weak to moderate positive correlation, which in a dataset of 100 million observations is a real signal rather than noise.
Correlation is not causation, and a longer description does not magically summon a citation. The honest reading is mechanical: retrieval systems work on text. The description is the largest block of author written, machine readable text attached to a video that is not the transcript. A video with a two line description gives the retrieval layer almost nothing to match a question against. A video with 300 words of specific, factual description gives it a lot.

The description is 334 words long, and almost nobody writes it
The average cited video carries a description of 334 words and a title of 19 words. Go and look at your last upload. If your description is a link to Instagram and a line of hashtags, you have published a video with no text surface, and the retrieval layer has nothing to grab.
Nineteen words in a title is also longer than most creators write. That is not clickbait length, that is specific length: the full question the video answers, not a three word tease. The two studies agree on the underlying shape here. The videos being quoted are the ones that read like reference material.
This is where the manual cost becomes real. Writing 334 words of accurate, non repetitive description for every upload is roughly 20 to 30 minutes of work per video once you factor in getting the facts right. At 30 videos a month that is 10 to 15 hours of pure typing, on top of scripting, narration and editing, and it is the first thing that gets dropped when the week gets tight.
It is also the single cheapest thing to automate, because the description is derived from the script that already exists. FalconVid writes the title, the description and the tags from the finished script as part of the same run that produces the video, in the video's own language, and it does it for every upload without anyone remembering to. The thing most likely to be skipped by a human is the thing a pipeline never skips.
Timestamps are the repeat citation machine
Only 31% of cited videos carry timestamps. Inside Google AI Overviews, 73% of YouTube citations point at a timestamp rather than the video as a whole, and in AI Mode it is 27%. On ChatGPT, Perplexity, Copilot and Gemini the figure is zero: no timestamped YouTube citations appeared at all.
The number that matters commercially is this one: 78% of timestamped videos were cited more than once, across two to five different chapters. A single video with clean chapters is not one citation opportunity, it is two to five, because each chapter answers a different question and gets retrieved independently.
Read against the correlation table, this is the clearest instruction in either study. Views do nothing. Chapters multiply. And chapters are free.
We wrote about chapters before as a retention tool, and the honest finding then was that chapters and timestamps help navigation more than watch time. That conclusion still stands for retention. What changed in 2026 is that the same markup acquired a second job that pays better than the first one, and the work to add it did not change.
Shorts are almost invisible to AI answers
The format split in the Otterly data is not close. Long form video accounts for 94% of AI citations. Shorts account for 5.7%. Playlists, channels and livestreams together make up 0.3%.
The reason is the same reason the description matters: a 45 second Short produces perhaps 120 words of transcript, which is not enough for a retrieval system to build a confident answer on. A 12 minute video produces around 1,700 words of spoken text. One is a claim, the other is a source.
The median duration of cited videos is under 8 minutes, and the most common bucket is 10 to 20 minutes at 32.1% of citations. That is squarely the long form range, and it lines up with the duration bands that already worked for retention and for mid roll ads.
So the strategic picture for 2026 is not new, it is reinforced. Shorts remain the cheapest way to be discovered by a human scrolling. Long form remains the only way to be discovered by a machine answering a question, and it is also the only format that qualifies you for watch hours. If you were already prioritising long form, this is a third independent reason to keep going.
The three assets a FalconVid channel ships without being asked
Everything above reduces to three artefacts: a long form video with a real transcript, a long and specific description, and clean chapters. None of them are creative decisions. All three are the kind of work that gets skipped at 11pm on a Thursday.
That is exactly the shape of work the AI YouTube analyst side of the platform was built around. A FalconVid run produces the script first, then the narration, then the edit, and the SEO block comes out of the same script rather than being retrofitted afterwards: title, description, tags, and chapter markers that follow the actual structure of the script, because the pipeline knows where each act starts.
The reference video length in the platform is 12 minutes, which puts every default run inside the 10 to 20 minute band that collects 32.1% of citations. The narration is real spoken text in 63 languages, which means the transcript exists in the viewer's language and not as a machine translation bolted on later. And because the whole channel runs on a calendar you approve rather than on a night you find free, the 334 word description happens on video 40 exactly as it happened on video 1.
The economics are not exotic. On the ruler in production today a finished 12 minute video costs 1,731 credits in economy mode, 4,624 balanced and 16,158 premium. At US$ 0.003133 a credit that is US$ 5.42, US$ 14.49 and US$ 50.63 per finished video. The Starter plan at US$ 47 carries 15,000 credits a month, which is 8 videos in economy or 5 if two of them go balanced. Nobody runs a whole month in one mode, so quote the mode alongside the number or the number is dishonest.
What to actually do this week
There are only four moves, and three of them are retroactive, which is the good part. AI answers are re-crawled continuously, so a description you rewrite today can be cited on a video you uploaded last year.
First, pick your ten best performing evergreen videos and rewrite the descriptions to 300 words or more of specific, factual prose. Not hashtags, not links, prose that states what the video establishes. Second, add chapters to those same ten. Two to five chapters is the range that generated repeat citations. Third, lengthen thin titles toward the specific question the video answers rather than the tease. Fourth, from now on, publish long form with all three by default.
You will not see this in YouTube Analytics, and that is the frustrating part. There is no citation report. What you can watch is the traffic source labelled external and the referrers inside it, plus the slow shape of impressions on evergreen videos that stop decaying. Both are lagging and noisy. Treat this the way you would treat the long tail of views that keeps arriving years after upload: the payoff is real and the feedback loop is slow.
The manual ceiling on all of this is honest and low. One person writing 334 word descriptions and chaptering every upload can sustain maybe 8 to 12 long form videos a month before the quality of the descriptions collapses, and it is the descriptions that collapse first because nobody sees them. That is the ceiling of doing it by hand. It is not the ceiling of doing it at all: on Business at US$ 297 the same operation ships 54 long form videos a month with the description and the chapters written every single time, on Agency at US$ 597 it is 109 across up to 25 channels, and on Scale at US$ 997 it is 184 across 50, with 50 pipelines running at the same time. The limit was never the strategy. It was the typing.

