Clamor DataRequest access

Research

What we have measured, with the sample size attached.

These are the findings the pipeline is built on. Each one is a measurement with a stated n, not an estimate. Anything we have not measured is not here.

01

30%

n = 380, uniform sample

Most short-form video carries no usable speech

of sampled videos contain speech or on-screen text worth processing

Catalogue music, reused viral audio and silent photo carousels account for the rest. Classifying this before transcription is what makes full-population coverage affordable.

02

6%

n = 603, uniform US sample

Discourse worth measuring is a thin slice of the platform

of videos fall inside the monetisable topical scope

Politics, economics, work, technology, health and social issues. Entertainment, lip-sync and performance dominate the platform by volume and are excluded by design.

03

28%

n = 500, with live controls

Platform-native labels are precise but miss most political speech

of a political-by-construction corpus carried a political platform label

What the platform labels as political is genuinely political, so the field is useful for inclusion. It is useless for exclusion: commentary shaped like a rant or a late-night segment is filed under entertainment.

04

26%

n = 500, controls interleaved

The record disappears, which is why history cannot be backfilled

of a 2023–24 political corpus no longer existed by 2026

Live positive controls separate deletion from blocking. A labeled series of public discourse has to be collected as it happens; it cannot be reconstructed later at any price.

Request a sample dataset.

Tell us the market and the question. We return a real slice — schema, sample sizes and confidence intervals included — not a brochure.

Request accessRead the methodology