Blog

AI answers cite YouTube more than any website, and views have nothing to do with it

Two independent 2026 studies pulled apart more than 146 million citations between them. The thing that predicts whether a video gets quoted is not views, not subscribers, not likes. It is the part of the upload almost nobody bothers to fill in.

Ricardo AlmeidaFounder14 min read
Illustration of many dim video frames with only a few feeding a glowing answer panel

Four out of ten cited videos have fewer than 1,000 views

Otterly.ai published a study on 2 March 2026 that looked at more than 100 million AI citation instances across a 30 day window, spanning ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, Microsoft Copilot and Gemini. One number in it should stop anyone who has ever felt too small to compete: 40.83% of the YouTube citations went to videos with under 1,000 views.

Not 4%. Forty point eight three percent. Four in every ten times an AI system quoted a YouTube video as its source, that video had an audience you could fit in a school assembly. And 35% of the channels being cited had fewer than 10,000 subscribers. The median cited channel had 41 videos published in total.

This is not a rounding error or a quirk of one platform. It is a structural difference between how a recommendation engine works and how a retrieval engine works. The YouTube homepage is a popularity machine: it shows you what people like you watched to the end. An AI answer is a reference machine: it needs a source that states the fact cleanly, and it does not care who else has watched it.

The practical consequence is uncomfortable for anyone who spent two years chasing the algorithm, and very good news for anyone starting now. There is one distribution channel in 2026 where a brand new channel and a 500,000 subscriber channel enter on genuinely equal footing, and almost nobody is optimising for it.

Where the citations actually come from, platform by platform

The second study is Surfer's, published in May 2026, built on 46 million citations pulled from 36 million AI Overviews between March and August 2025. It answers a different question: of everything Google's AI Overview cites, what share is YouTube? The answer is 23.3%. Wikipedia is second at 18.4% and Google.com third at 16.4%. Every legacy news brand sits below those three.

So YouTube is the single most cited domain in Google's AI answers. That is worth reading twice, because it means video is not a side channel to the written web any more. It is the primary reference layer that the written answer is built on top of.

The Otterly numbers then split YouTube citations by which assistant is doing the citing, and the spread is brutal:

  • Perplexity: 38.7% of all YouTube citations
  • Google AI Overviews: 36.6%
  • Google AI Mode: 19.6%
  • ChatGPT: 4.4%
  • Microsoft Copilot: 0.5%
  • Gemini: 0.2%

Read that table before you build a strategy around ChatGPT

Three quarters of the opportunity sits inside Google's own surfaces once you add AI Overviews and AI Mode together, and Perplexity punches far above its user count because it is built citation first. ChatGPT, which is the assistant most people picture when they say AI search, accounts for 4.4% of YouTube citations. That does not make it worthless, it makes it a different game with different mechanics.

There is also a measurable payoff attached to being cited at all. Analyses of AI Overview behaviour in 2026 found that brands cited inside an AI Overview earn roughly 120% more organic clicks per impression than uncited brands appearing on the same queries. Meanwhile zero click searches on Google reached 68% in early 2026. Both things are true at once: fewer people click, and the ones who do click go disproportionately to whoever got quoted.

None of these are official YouTube or Google figures. Neither company publishes citation share. These are third party measurements from tools that watch AI answers at scale, and they should be read as strong directional evidence, not as a published rate card. That distinction matters more here than in most posts, because the whole field is eighteen months old.

If you want the grounding on how impossible it is to get a real number out of the platform itself, we went through that in detail when we looked at what YouTube search volume data actually exists. The short version: the platform never published volume, and every tool that shows you one is modelling.

What correlates with being cited, and what does not

Otterly ran Pearson correlations between citation frequency and every obvious video attribute. A Pearson coefficient runs from -1 to 1, and anything hovering around zero means the two things move independently. Here is the table, and it is the most useful thing published about video discovery this year.

Views: -0.03. Likes: -0.02. Subscribers: -0.03. Duration: 0.02. Every metric that YouTube puts on the front of your dashboard sits at statistical zero. A negative 0.03 does not mean views hurt you, it means there is no relationship at all.

Then two things actually move. Description length: 0.31. Descriptions containing hashtags: 0.20. Recency: around 0.3. A 0.31 is a weak to moderate positive correlation, which in a dataset of 100 million observations is a real signal rather than noise.

Correlation is not causation, and a longer description does not magically summon a citation. The honest reading is mechanical: retrieval systems work on text. The description is the largest block of author written, machine readable text attached to a video that is not the transcript. A video with a two line description gives the retrieval layer almost nothing to match a question against. A video with 300 words of specific, factual description gives it a lot.

Bar chart showing views, likes and subscribers flat against a taller description bar

The description is 334 words long, and almost nobody writes it

The average cited video carries a description of 334 words and a title of 19 words. Go and look at your last upload. If your description is a link to Instagram and a line of hashtags, you have published a video with no text surface, and the retrieval layer has nothing to grab.

Nineteen words in a title is also longer than most creators write. That is not clickbait length, that is specific length: the full question the video answers, not a three word tease. The two studies agree on the underlying shape here. The videos being quoted are the ones that read like reference material.

This is where the manual cost becomes real. Writing 334 words of accurate, non repetitive description for every upload is roughly 20 to 30 minutes of work per video once you factor in getting the facts right. At 30 videos a month that is 10 to 15 hours of pure typing, on top of scripting, narration and editing, and it is the first thing that gets dropped when the week gets tight.

It is also the single cheapest thing to automate, because the description is derived from the script that already exists. FalconVid writes the title, the description and the tags from the finished script as part of the same run that produces the video, in the video's own language, and it does it for every upload without anyone remembering to. The thing most likely to be skipped by a human is the thing a pipeline never skips.

Timestamps are the repeat citation machine

Only 31% of cited videos carry timestamps. Inside Google AI Overviews, 73% of YouTube citations point at a timestamp rather than the video as a whole, and in AI Mode it is 27%. On ChatGPT, Perplexity, Copilot and Gemini the figure is zero: no timestamped YouTube citations appeared at all.

The number that matters commercially is this one: 78% of timestamped videos were cited more than once, across two to five different chapters. A single video with clean chapters is not one citation opportunity, it is two to five, because each chapter answers a different question and gets retrieved independently.

Read against the correlation table, this is the clearest instruction in either study. Views do nothing. Chapters multiply. And chapters are free.

We wrote about chapters before as a retention tool, and the honest finding then was that chapters and timestamps help navigation more than watch time. That conclusion still stands for retention. What changed in 2026 is that the same markup acquired a second job that pays better than the first one, and the work to add it did not change.

Shorts are almost invisible to AI answers

The format split in the Otterly data is not close. Long form video accounts for 94% of AI citations. Shorts account for 5.7%. Playlists, channels and livestreams together make up 0.3%.

The reason is the same reason the description matters: a 45 second Short produces perhaps 120 words of transcript, which is not enough for a retrieval system to build a confident answer on. A 12 minute video produces around 1,700 words of spoken text. One is a claim, the other is a source.

The median duration of cited videos is under 8 minutes, and the most common bucket is 10 to 20 minutes at 32.1% of citations. That is squarely the long form range, and it lines up with the duration bands that already worked for retention and for mid roll ads.

So the strategic picture for 2026 is not new, it is reinforced. Shorts remain the cheapest way to be discovered by a human scrolling. Long form remains the only way to be discovered by a machine answering a question, and it is also the only format that qualifies you for watch hours. If you were already prioritising long form, this is a third independent reason to keep going.

The three assets a FalconVid channel ships without being asked

Everything above reduces to three artefacts: a long form video with a real transcript, a long and specific description, and clean chapters. None of them are creative decisions. All three are the kind of work that gets skipped at 11pm on a Thursday.

That is exactly the shape of work the AI YouTube analyst side of the platform was built around. A FalconVid run produces the script first, then the narration, then the edit, and the SEO block comes out of the same script rather than being retrofitted afterwards: title, description, tags, and chapter markers that follow the actual structure of the script, because the pipeline knows where each act starts.

The reference video length in the platform is 12 minutes, which puts every default run inside the 10 to 20 minute band that collects 32.1% of citations. The narration is real spoken text in 63 languages, which means the transcript exists in the viewer's language and not as a machine translation bolted on later. And because the whole channel runs on a calendar you approve rather than on a night you find free, the 334 word description happens on video 40 exactly as it happened on video 1.

The economics are not exotic. On the ruler in production today a finished 12 minute video costs 1,731 credits in economy mode, 4,624 balanced and 16,158 premium. At US$ 0.003133 a credit that is US$ 5.42, US$ 14.49 and US$ 50.63 per finished video. The Starter plan at US$ 47 carries 15,000 credits a month, which is 8 videos in economy or 5 if two of them go balanced. Nobody runs a whole month in one mode, so quote the mode alongside the number or the number is dishonest.

What to actually do this week

There are only four moves, and three of them are retroactive, which is the good part. AI answers are re-crawled continuously, so a description you rewrite today can be cited on a video you uploaded last year.

First, pick your ten best performing evergreen videos and rewrite the descriptions to 300 words or more of specific, factual prose. Not hashtags, not links, prose that states what the video establishes. Second, add chapters to those same ten. Two to five chapters is the range that generated repeat citations. Third, lengthen thin titles toward the specific question the video answers rather than the tease. Fourth, from now on, publish long form with all three by default.

You will not see this in YouTube Analytics, and that is the frustrating part. There is no citation report. What you can watch is the traffic source labelled external and the referrers inside it, plus the slow shape of impressions on evergreen videos that stop decaying. Both are lagging and noisy. Treat this the way you would treat the long tail of views that keeps arriving years after upload: the payoff is real and the feedback loop is slow.

The manual ceiling on all of this is honest and low. One person writing 334 word descriptions and chaptering every upload can sustain maybe 8 to 12 long form videos a month before the quality of the descriptions collapses, and it is the descriptions that collapse first because nobody sees them. That is the ceiling of doing it by hand. It is not the ceiling of doing it at all: on Business at US$ 297 the same operation ships 54 long form videos a month with the description and the chapters written every single time, on Agency at US$ 597 it is 109 across up to 25 channels, and on Scale at US$ 997 it is 184 across 50, with 50 pipelines running at the same time. The limit was never the strategy. It was the typing.

FAQ

Got questions? We've got answers.

Is YouTube really the most cited source in Google's AI Overviews?

According to Surfer's May 2026 analysis of 46 million citations, YouTube accounts for 23.3% of Google AI Overview citations, ahead of Wikipedia at 18.4% and Google.com at 16.4%. That is a third party measurement, not a figure Google publishes, and it reflects the March to August 2025 window the study covers.

Do views or subscribers help a video get cited by AI?

The measured correlation is effectively zero: views at -0.03, likes at -0.02 and subscribers at -0.03 in Otterly's 100 million citation dataset. 40.83% of cited videos had under 1,000 views and 35% of cited channels had under 10,000 subscribers. Popularity and citation are separate systems.

How long should a description be to have a chance of being cited?

The average cited video carries 334 words. Description length showed a 0.31 correlation with citation frequency, the strongest of any attribute measured. That is a weak to moderate positive, not a guarantee, but it is the only lever in the dataset that moved at all.

Do YouTube Shorts get cited?

Rarely. Long form video is 94% of AI citations against 5.7% for Shorts. A 45 second Short produces too little transcript for a retrieval system to answer from confidently. Shorts still work for human discovery, they just do not function as a source.

Do I have to appear on camera for any of this?

No, and nothing in either study touches presence on camera. What gets cited is text: the transcript, the description and the chapter titles. A faceless channel with a strong narration script has exactly the same surface area as a presenter led one. FalconVid produces the narration, the description and the chapters from the script without anyone filming anything.

Will an AI written description hurt me with YouTube?

There is no rule against it. YouTube's monetisation policies target reused and inauthentic content, not the tool used to write metadata. What gets punished is a description that lies about the video or is copied wholesale across uploads. A description generated from the actual script of that specific video is the opposite of that, and it is what the pipeline produces by default.

Can I get citations without publishing more videos?

Yes, and this is the part most people miss. Rewriting descriptions and adding chapters to videos already published is retroactive work: the AI systems re-crawl continuously, so an upload from last year can start being cited after you improve its text surface. Ten old videos fixed properly is usually a better week's work than one new upload.

Which assistant should I optimise for?

Google's surfaces, by a wide margin. AI Overviews at 36.6% and AI Mode at 19.6% are 56.2% of YouTube citations combined, and Perplexity adds 38.7%. ChatGPT is 4.4%. Google is also the only place where timestamped citations appear at all, which is why chapters pay there and nowhere else.

Ship the transcript, the description and the chapters on every video, without typing them

FalconVid writes the script, narrates it in 63 languages, edits the video and produces the title, the 300 word description and the chapter markers from that same script, on a calendar you approve. On today's ruler a finished 12 minute video is 1,731 credits in economy, 4,624 balanced and 16,158 premium. Starter at US$ 47 carries 15,000 credits a month, which is 8 videos in economy or 5 when two go balanced. Pro at US$ 97 is 17 in economy across 5 channels, Business at US$ 297 is 54 across 10, Agency at US$ 597 is 109 across 25 and Scale at US$ 997 is 184 across 50 with 50 pipelines at once. The free plan gives you 4 videos of up to 3 minutes every 30 days with no card, so you can see the metadata a run produces before you decide anything. Seven day guarantee, and the annual plan charges 10 months for 12.

Start free, no credit card

Charged today · 7-day guarantee · Cancel anytime

Keep reading