Choosing a B2B Data Provider in 2026
The five questions to ask any B2B data provider, a comparison of the provider types, and how scraper.io answers each one.
Written by
Scraper.io
Editorial
Choosing a B2B data provider is mostly a matter of working out which of four very different businesses you are actually talking to, because they are priced alike and fail differently. Five questions — provenance, refresh cadence, per-row source, licence and price model — separate them faster than any demo, and every one of them can be answered from a sample.
What a B2B data provider actually sells
The phrase covers at least four businesses. One sells a contact database it has accumulated and resells to everyone. One sells you the tools to collect data yourself and charges by the request. One is a marketplace or broker, listing sets it did not gather and cannot vouch for. One builds a set to your specification and maintains it.
They are frequently indistinguishable on a website, and they fail in completely different ways. A contact database goes quietly stale. A toolkit works perfectly and hands you an engineering liability. A broker's listing is accurate about its row count and silent about its age. A managed provider is the most expensive per row and the only one that can be asked to fix a wrong value.
The mistake that costs the most is not picking the wrong type. It is buying a large set from any of them without establishing what a single row can prove. Row counts are the easiest number to inflate and the least connected to whether a decision made on the data will be right.
Five questions to ask a B2B data provider
Ask all five before you look at coverage, and ask them of a sample rather than of a salesperson. If a provider cannot answer them from twenty rows, the answer for twenty million rows is not going to be better.
The questions are deliberately mechanical. None of them requires you to know the vendor's domain, and none of them can be satisfied with a description of a process.
- Provenance: pick one row and one field. Which exact URL or document was that value read from? If the answer is the dataset's name, there is no provenance.
- Refresh cadence: how often does this set re-run, and is the cadence per set or per field? One export date for a whole file is not a cadence.
- Per-row source and observation date: does every row state when it was last confirmed, or does the file state when it was generated? These are very different guarantees.
- Licence: what may you do with it — internal use, redistribution, training, resale in a product? Ask what happens to your rights when the subscription ends.
- Price model: per row, per seat, per set, per job, or annual platform fee? Ask what happens at renewal, and whether the price is a function of usage or of last year's price.
Every one of the five is answerable from a sample of twenty rows. A provider who needs a call to answer them is telling you something about the data, not about their process.
The four types of B2B data provider compared
The table compares the types rather than named vendors, and it deliberately carries no competitor pricing: published prices move, are frequently negotiated away, and quoting them here would be out of date before the quarter ends. What does not move is the structure of each offer.
Read across the row that matches whoever you are currently talking to. The failure mode column is the one worth taking seriously, because it is what you will be living with in month four.
| Type | What you are buying | Typical refresh | Per-row provenance | Usual failure mode |
|---|---|---|---|---|
| Contact and firmographic database | A licence to search an accumulated store everyone else can also license | Continuous but opaque; rarely stated per field | Rare — sources are aggregated and not exposed per value | Quiet decay; nobody can say when a given row was last confirmed |
| Scraping platform or proxy toolkit | Infrastructure and requests; you build and run the collection | Whatever you schedule and maintain | Whatever you build in, which is usually nothing at first | It works, then it becomes your maintenance backlog |
| Marketplace or data broker | A listed file assembled by a third party | Often a one-off export with a generation date | Usually none; the listing describes the file, not the values | Accurate row count, unknowable age, no route to a correction |
| Managed data provider | A set built to your specification and kept running | A stated cadence per set, agreed up front | Possible, and the main reason to pay more | Highest per-row cost; scope has to be written down properly |
Types, not vendors. No competitor prices are quoted anywhere on this page, because published data pricing changes frequently and is often negotiated; ask each provider directly and compare the answers against the five questions above.
How scraper.io answers the five questions
We are the fourth type, so it is only fair to answer our own questionnaire in the same order, in plain terms, with the parts we do not do stated as clearly as the parts we do.
On provenance: every value carries a receipt — the source URL it was read from, the time it was fetched, and the hash of the page it came from. The hash is the part that does the work people expect a timestamp to do: a URL tells you where to look, and a hash tells you whether what is there now is the same document we read. If it no longer matches, the value is not necessarily wrong, but it is no longer supported by the evidence on file and gets re-checked rather than re-asserted.
On per-row source and observation date: both are on the row, not on the file. On uncertainty: conflicts between sources are published as conflicts, and a missing value is published as missing rather than as a zero, an empty string or a plausible default. Confidence is an explicit field — CONFIRMED where two or more independent origins agree, SIGNAL where a single origin is all there is — so software can branch on it.
On what we do not claim: coverage is bounded by the sets that exist. A managed provider is the wrong choice if you want to search everything about everyone; it is the right one if you have a narrow question you will ask every week and act on.
Refresh cadence, feed by feed
Cadence is the question most often answered with an adjective, so here it is as a table. Each row is a live feed with the number of rows it publishes in its current snapshot after its sanity filter, and the schedule it is configured to re-run on.
The row counts are published rows, not raw fetches. Rows that fail the sanity filter — an implausible number, a duplicate primary key, a version stamp that leaked into a price column — are withheld and counted rather than silently corrected. API Deprecation Watch is the clearest example: its guard drops changelog cross-references and “now GA” notices, taking a 40-row snapshot down to the 37 it publishes.
| Feed | Rows published | Cadence |
|---|---|---|
| Frontier Lab Founders | 93 | Weekly |
| API Deprecation Watch | 37 | Daily |
| Tender & Award Watch | 40 | Daily |
| Cornwall & Devon — Electrical Works Tenders | 23 | Daily |
| Competitor Pricing Diffs | 40 | Daily |
| Champion Moves | 40 | Daily |
| UK Distressed Property | 39 | Weekly |
| Stealth Startup Radar | 40 | Daily |
| Sanctions Diffs | 26 | Daily |
Counts are the rows each feed publishes in its current snapshot after its sanity filter, read from the feed catalogue; cadence is each feed's configured re-run schedule. Several feeds hold more records in a full run than a published snapshot shows, and each feed page states both.
How scraper.io is priced
Two things are for sale, and they are priced differently because they are different products. A job is a one-off: you describe the set, we build it, and you get the rows with their receipts. Jobs start at £150 and are quoted per job, because the real cost driver is not row count but how hard the sources fight back — a clean public register and a site that changes its markup weekly are not the same work.
A feed is the same set kept alive on the cadence in the table above, billed monthly per feed. We are not quoting a per-feed figure on this page, because feed pricing is being reset; the current numbers live on the pricing page, which is the page to trust if this one ever disagrees with it.
The reason both exist is that most buyers cannot tell at the outset which they need. Buy the job, use the rows, and find out whether the answer decays. If it does, the feed behind it is already running. If it does not, you paid once for something you only needed once, which is the correct outcome and not one most data contracts allow.
If you would rather start from something already built than describe a set, the shipped sets are in the dataset catalogue, each with its schema and a sample on the page.
A short due-diligence checklist before you sign
Run this on a sample, in this order, before the row count enters the conversation at all. A small set that passes all six is worth more than a large one that fails the first two, because software can act on the small one and a person has to check the large one.
The last item catches more problems than the rest combined, and almost nobody runs it.
- Take one row and one field, and ask for the exact URL it was read from.
- Ask when that specific value was last confirmed, not when the file was produced.
- Ask how a blank is represented, then check it is not being served as a zero or an empty string.
- Ask what happens when two sources disagree, and whether you get to see the disagreement.
- Ask what the row is keyed on, then rename something and check it does not appear as a second entity.
- Ask for the same sample twice, a fortnight apart, and diff them. Whatever the diff shows is your real refresh cadence.
Frequently asked questions
- What should I ask a B2B data provider before buying?
- Five things, all answerable from a sample: which exact URL each value was read from, how often the set re-runs, whether the observation date is per row or per file, what the licence permits after the subscription ends, and how the price is calculated at renewal.
- How often should B2B data be refreshed?
- It depends on the field, not the file. A company name is stable for months; a tender deadline, a price or a sanctions listing is not. Cadence should be stated per set at minimum, and a provider who gives one export date for everything is not answering the question.
- How much does scraper.io cost?
- A one-off job starts at £150 and is quoted per job, because difficulty rather than row count drives the cost. Feeds are the same set kept running and are billed monthly per feed; the current feed pricing is on the pricing page.
The provider worth buying from is the one whose sample survives the five questions, and the row count is the last thing to look at rather than the first. If a single value cannot be traced back to a page and a date, everything built on top of it is a claim rather than a fact.