logoPandorLabs
review data

What Customers Actually Said

Reviews from marketplaces, app stores, and review platforms with rating, full text, verified-purchase status, and product match — normalised so a one-to-five star and a ten-point score sit in the same column.

at a glance

Product
Review Data
Formats
JSON · CSV · Parquet
Delivery
Webhooks, S3, Snowflake, BigQuery
Refresh
Continuous to daily
Scope
Public data only

the problem

Why this is harder than it looks

Reviews are the largest body of unprompted product feedback that exists, and most companies use almost none of it. The reason is rarely indifference — it is that reviews are scattered across marketplaces, app stores, and review platforms with a different rating scale, a different product identifier, and a different notion of what counts as verified on each one. Reading your own reviews on one site is easy and tells you little. Comparing your review corpus against three competitors across five platforms over eighteen months is where the actual insight is, and that requires the normalisation work that nobody wants to do.

who buys it

Who this is built for

Product teams looking for feature-level complaints before they reach churn, e-commerce and brand teams tracking rating trajectory against competitors, and AI teams who want labelled sentiment data where the star rating is a ground-truth label written by the customer.

sources covered

What we collect from

E-commerce marketplaces

Product reviews at SKU level, with verified-purchase status preserved where the platform exposes it.

App stores

Mobile reviews with app version attached, which is what lets a rating drop be tied to the release that caused it.

Software review platforms

B2B review sites where reviews are long, structured, and frequently disclose company size and role — unusually rich for competitive work.

Retailer product pages

Reviews left on a retailer’s own site rather than a marketplace, which brands routinely miss because they only monitor where they sell directly.

Travel and hospitality platforms

Property and experience reviews with stay dates and traveller type, where seasonality dominates the signal.

Local and service reviews

Location-level reviews for multi-site businesses, aggregated so a single underperforming branch is visible rather than averaged away.

what you get back

Fields in the delivered schema

Agreed with you before collection starts, and held stable afterwards — the sites change underneath, your columns do not.

  • Full review text and title
  • Rating, normalised to a common scale with the original scale preserved alongside it
  • Review date and, where exposed, the purchase or stay date
  • Verified-purchase flag where the platform provides one
  • Product or listing identifier, matched to your own catalogue
  • Reviewer display name, review count, and badge status where public
  • Helpful and unhelpful vote counts
  • Merchant or brand response, with its own timestamp
  • App version, variant, or size where the platform records it
  • Detected language and source platform

Public surfaces only

We collect what a visitor can see, honour a site's stated crawling preferences, and never bypass authentication. Provenance is recorded on every record.

One schema across sources

Records from any source arrive with the same field names, so adding a source does not mean rewriting anything downstream.

Compliance built in

GDPR and CCPA handling, a DPA signed before delivery, configurable retention, and deletion at source propagating through to your feed.

SKU-level
Product matching
Normalised
Rating scales
Verified
Purchase flags
99.9%
Uptime SLA

applications

What teams build with review data

Feature-level complaint detection

Aggregated review text surfaces the specific failure — a strap that breaks, a sync that drops, a size that runs small — long before it shows up as a return rate or a churn number.

Competitive rating benchmarking

Your rating trajectory against a named competitor set, per product and over time. The direction of travel is almost always more informative than the absolute number.

Release-quality monitoring

App store reviews carry the version they were left against, which ties a rating drop to a specific release and makes the regression obvious rather than mysterious.

Review-manipulation detection

Reviewer history, timing clusters, and verified-purchase ratios expose the shape of purchased reviews on competitor listings — and protect you from benchmarking against a number that was bought.

Labelled training data for sentiment work

Review text paired with a star rating is sentiment data with a ground-truth label the customer wrote themselves, which is considerably better than anything a human annotator will produce at scale.

Multi-location quality management

For franchises and multi-site businesses, per-location review streams show which sites are generating complaints while the aggregate average still looks healthy.

related products

Usually bought alongside

Everything below delivers on the same schema and the same infrastructure, so combining them is a configuration change rather than a project.

questions

Frequently Asked Questions

Everything you need to know before you send us your first request.

Every rating is delivered twice: once on a common scale so platforms can be compared, and once exactly as the source recorded it. Keeping the original matters, because a five-point scale and a ten-point scale do not carry the same distribution and a conversion that discards the source value cannot be audited later. Which normalisation you want is a scoping decision, since the right one depends on whether you are comparing across platforms or tracking within one.
Yes, and it is usually the most valuable part of the engagement. You supply your catalogue with whatever identifiers you hold, and reviews are matched to your SKUs across platforms — which is what turns a pile of review text into something a product team can act on per line item. Matching accuracy is validated against a sample before rollout rather than asserted.
Where the platform exposes it, yes, and it is one of the more important fields in the set. Unverified reviews are not worthless, but they behave differently enough that most analysis should either weight or segment on this flag. Where a platform provides no such signal, the field is explicitly null rather than defaulted to false.
We deliver the raw signals that detection depends on — reviewer history, posting-time clusters, verified-purchase ratios, rating distribution shape, and text duplication — rather than a single verdict field. We are deliberate about that: a proprietary fake-review score you cannot inspect is not something you should be making competitive decisions on. Customers apply their own thresholds to signals they can see.
For most platforms, the complete public review history on a listing, which for established products can run to years. Historical backfill is scoped and quoted separately from ongoing monitoring, and depth is confirmed per platform during scoping since some paginate their archives far more shallowly than others.
We treat them as though they do. Reviews are user-generated content that frequently contains identifying detail, so they are handled under a DPA with a documented lawful basis, deletion at source propagates to subsequent deliveries, and reviewer identifiers can be pseudonymised or dropped entirely at delivery if your use case does not require them.
Into the systems you already run. Records land as JSON, CSV, or Parquet in S3, GCS, or Azure Blob, or straight into Snowflake or BigQuery, on whatever cadence you set. Webhooks push new records to your endpoints as they are detected, which is how most customers wire alerting. A solutions engineer sets the schema, cadence, and destination with you during onboarding rather than handing you a docs page and wishing you luck.
Keeping the collectors working is our responsibility, not yours. Extractors are monitored continuously, and when a site changes we patch upstream while the schema we deliver to you stays fixed — so nothing downstream needs to be touched. If a change causes a genuine coverage gap we tell you which window was affected rather than quietly returning fewer records.
Yes, and we would rather you did. Every engagement starts with a free sample run against your own targets — your competitors, your catalogue, your keywords — and we hand back the actual dataset. Judging real output against your real requirements is the only useful evaluation, and it is a much better use of a first conversation than a slide deck.

Still have questions?

Talk to an engineer

Ready to Get Started?

Talk to us about your sources and volume. We'll return a sample dataset from your target sites before you commit to anything.

SOC 2 Type II
GDPR & CCPA compliant
99.9% uptime SLA

© 2026 PandorLabs, Inc. All rights reserved.