Early access. Launching soon.

Polly
PollyX Assistant · Online
Hi! I'm Polly 👋 I can answer any questions about PollyX - how reports work, what's included, pricing, or anything else. What would you like to know?

Suggested questions

Powered by AI · PollyX

Methodology

How the score is produced

A number with no method behind it is decoration. This page says exactly what goes into a PollyX report, how the 0 to 10 score is arrived at, and - just as importantly - what it cannot tell you. If something here contradicts what you see in a report, the report is wrong and we want to hear about it.

1. What gets collected

Every report draws on five independent layers, gathered in parallel:

  • Global news via the GDELT Project: up to 365 days of global news, spanning 100+ countries and 65+ languages. Queries run at several time horizons and sort orders (most recent, most relevant, most negative, most positive) so a report is not built only from whichever articles happen to rank first.
  • YouTube via the YouTube Data API: videos discussing the brand, plus their comment threads, which are frequently more candid than the videos themselves.
  • Live web search: at most 12 searches per report across Trustpilot, Reddit, Glassdoor, forums and similar public platforms. Amazon star ratings are fetched directly where the brand sells products there, and Google Reviews where it has physical locations - both are skipped entirely, rather than guessed at, when the brand has neither. This is a hard ceiling, not a target.
  • Community APIs: Bluesky, Mastodon, Hacker News, and Apple App Store ratings and reviews, where the brand has a presence on them.
  • Social platforms: public posts from X, TikTok and Instagram, fetched directly. TikTok contributes video captions and engagement counts: we do not watch the footage, and weight it accordingly. Instagram comes from the brand's hashtag feed, which skews to creators and sellers, so it counts as buzz rather than customer satisfaction. Reposts, declared ads and promotional spam are removed before analysis. Coverage on these platforms varies by brand and by day.

That totals typically 200 to 1,200 sources per report. The figure varies by how widely a brand is covered, which is why every report and every score on this site displays its own actual source count and generation date rather than a marketing number.

2. What gets thrown away

Brand-owned content is excluded from sentiment.Press releases, newswire syndication and the brand's own posts tell you what a company says about itself, not what the public thinks of it. Where that material is interesting it may be noted as context, but it never counts toward a score, a theme or a quote.

Sources that turn out to be about a different entity with the same name are dropped. If you supply "also known as" or "not about" terms when creating the report, those feed the search plan and this relevance filter directly.

Other markets, when you ask for one market only. By default a report covers the brand worldwide with its home market foregrounded. Choosing this market onlywhen you create a report is a hard restriction rather than a preference: news is filtered by the publisher's country, web and video search are run against that country, Amazon ratings come from that country's marketplace (and are skipped where it has none), and coverage of the same brand elsewhere is discarded rather than down-weighted. Reddit, X, TikTok, Instagram, Bluesky, Mastodon, Hacker News and App Store posts carry no country, so those are judged individually on whether the poster is writing about your market. A restricted report is usually a thinner report - that is the trade-off, and where the evidence is thin it will say so instead of quietly refilling from elsewhere. Restricting requires a location; without one the report stays worldwide.

3. How the score is assigned

This is the part most tools are vague about, so plainly: there is no fixed formula and no fixed weighting between the news, video, web and community layers. The score is not a weighted average, and GDELT's -10 to +10 tone values are not arithmetically rescaled onto 0 to 10.

What actually happens is that all five layers - including GDELT's tone values and article volumes - are supplied as evidence to an AI analyst, which judges the balance of positive against negative opinion and places the brand on this scale:

7 to 10 Positive5 to 7 Mixed or neutral0 to 5 Negative

It weighs how many people are talking, how strongly they feel, how recent the coverage is, and how credible the source is. One deliberate correction is applied: review platforms and complaint forums heavily over-represent dissatisfied customers, because unhappy people write reviews far more often than satisfied ones. A handful of complaint pages is not allowed to drag down a brand that is broadly well regarded.

The honest consequence of a judged rather than computed score is that it is reproducible in direction but not to the decimal. Two runs a day apart on the same brand should agree on whether sentiment is positive, mixed or negative, and should broadly agree on the themes. They may differ by a few tenths. Treat the first digit as the signal and the decimal as resolution, not precision.

4. Sample size, window, and confidence

Every score is published with the number of sources behind it and the date it was generated. A score with neither is not a measurement, so we do not show one.

The news window reaches back 365 days at its widest, with the most recent coverage weighted most heavily. Web and community results reflect what was publicly visible at generation time.

Each report also carries a confidence level of high, medium or low, reflecting how much real public evidence backs it. A brand with thin coverage gets a report marked low confidence rather than a confident-looking number. Where there is not enough public data to produce an honest report at all, we say so and refund you rather than generating something that reads authoritative and is not.

5. What this score cannot tell you

  • It is an AI-generated estimate, not a statistic with a confidence interval. There is no margin of error to quote because there is no sampling frame.
  • It measures publicly expressed sentiment. People who feel strongly enough to post are not a representative sample of your customers, and private opinion is invisible to it.
  • Coverage is uneven by language, country and platform. A brand strong in a market GDELT indexes thinly will look quieter than it is.
  • It is a snapshot. Comparing two reports on the same brand over time is more informative than any single number, which is why the trend chart exists.
  • Language models can misread sarcasm, misattribute a quote, or state something with more confidence than the evidence supports. Every claim in a report links to its source so you can check it, and we would rather you did.

6. Corrections

If a report contains a factual error, a misattributed quote, or a source that is about a different company, email contact@pollyx.org with the report link. We will correct it and, where the error was material, regenerate the report at no charge.

200 to 1,200 sources per report. Last reviewed 31 July 2026.