"Verified" is the loosest word in the data industry, because every vendor defines it privately and none of them publish the definition. Ours means three things: deduplicated, scored on how many fields carry a value, and filtered to the project types and thresholds you asked for. It does not mean every value has been independently confirmed. The harder question is what happens when you want one file across many jurisdictions, none of which publish the same way. That reconciliation is the actual product, and it is worth knowing how a vendor does it before you buy from them.
The loosest word in the data industry
Every construction data vendor says its records are verified. None of them say what the word means, which is what makes it useful to them and useless to you. It can mean a human looked at a sample. It can mean a script checked that a field was not empty. It can mean nothing at all was checked and the word tested well on the pricing page.
The tell is a headline accuracy number with no method attached. Accurate against what, measured how, on what sample, and when? A figure that cannot answer those four questions is a marketing number wearing a lab coat, and it is almost impossible to disprove, which is precisely why it gets printed. A figure that can answer them is worth having — our own accuracy target, and how we sample-audit against it, is set out on the FAQ.
What this article adds to that number is the method — what actually happens to a record between the building department and your file, and which part of that is genuinely difficult. If you are evaluating what sits on each record, the method is the part you can check against a sample yourself, which is why it is worth publishing alongside the figure rather than instead of it. The same discipline runs through how we describe what the records support on a single project and what they support across a whole market.
Completeness and correctness are different measurements
Here is what happens to a record between collection and delivery, in full.
The same project frequently appears more than once — resubmissions, amendments, records that exist in more than one feed from the same authority. Duplicates get removed. This matters more than it sounds: a duplicate is a nuisance in a prospecting list and a distortion in any share calculation, because it inflates whichever contractor or category it lands in.
Each record is scored on how many of its up to twenty fields carry a value. That score is what lets a thin record be identified as thin before it reaches you, rather than after your team has spent a morning on it. It is a measure of how much of the record is present.
The set is then cut to the fields, record types and thresholds you specified. If contractor phone is the field your process depends on, you can require it and receive only records that have one. The filtering runs against your thresholds, not a generic house standard.
That is the whole process. What it does not include is independent confirmation of each individual value — no dialing every number to see whether it rings, no checking every license against a state board, no confirming that a contact still works where the filing said they did.
We could describe scoring as verification and most readers would never know the difference. We would rather say it plainly: completeness measures how much of a record is present. Correctness measures whether each value is true. They are different measurements and we report the one we actually make.
Scoring how many fields are populated tells you nothing about whether the phone number rings. Any vendor conflating those two is selling you the easier measurement at the price of the harder one.
The reason this matters commercially rather than philosophically: a completeness figure is reproducible. You can take a sample, count the populated fields yourself, and get our number. An accuracy figure is not reproducible without doing the vendor's work over from scratch, which nobody does, which is why the figure survives. The rest of what we do and do not claim is written the same way, and real rows from your own markets are the fastest way to check any of it.
No two authorities publish the same way
Here is the part of this work nobody markets, because it does not sound like much until you try it.
Take one job — a house getting a new roof — and look at how four different building departments record it. One files it as a property type of residential with a sub-type of roofing. One calls it a work class of "Res-Reroof." One assigns a numeric category code and publishes nothing else. One puts it nowhere structured at all, and the only evidence of what the job is sits in a free-text description a clerk typed. Four authorities, one roof, four descriptions that share not a single common value.
None of them is wrong. Each is internally consistent and perfectly usable if that county is your whole world. They only become a problem the moment you want to ask one question across all four — which is exactly the moment a supplier or builder operating in more than one market starts caring.
Any single jurisdiction's data is easy to work with. The difficulty is not in any one of them. It is in the fact that no two of them agree, and there is no authority above them that makes them.
That is the whole shape of the problem, and it compounds in ways that are not obvious from one file. Field names differ. Date formats differ. Some publish a status vocabulary of six values and some of sixty. Some export clean columns and some export a spreadsheet where the headers and the data have quietly drifted out of alignment — which we have found in files from real authorities, and which will silently corrupt any analysis run on top of it.
Reconciling all of that into one specification is the product. Not the collection — the records are public and anyone determined enough can scrape a county. What is impractical is doing it across three hundred authorities and ending up with a file where "commercial" means the same thing in Florida as it does in Arizona, where a date sorts correctly across every row, and where a status filter returns what you expected.
This is also why we publish the specification rather than describing it on a call. A normalized spec is a commitment: it says every record you receive will speak this vocabulary regardless of which of three hundred authorities it came from. A vendor who will not put that in writing is telling you their file still speaks three hundred languages, and that looking at real rows will show it.
What an incomplete field actually costs
Where a field you depend on is missing, the cost is real but rarely the one people quote. It is not the phone call itself — a dial that goes nowhere costs a couple of minutes, not an hour. It sits in three quieter places.
Rework. Every field your process depends on that arrives empty becomes a research task for somebody. Across a file of any size that work compounds quickly, and at sales compensation it is expensive time that looks like selling on a calendar and is not. The fix is not more effort from the rep; it is deciding up front which fields your process genuinely requires and filtering to them before the file ships, rather than discovering the gap afterwards.
System pollution. Thin records entering a CRM corrupt everything measured downstream. Conversion rates fall because the denominator includes records nobody could ever have actioned. Territory comparisons skew toward whichever market happens to publish more contact detail, which reads as better sales execution when it is actually just a more forthcoming authority.
Confidence. This one is soft and it is the most expensive. A rep who works two lists that go nowhere stops trusting the third, and quietly returns to the accounts they already know. Once a sales team decides that data-driven prospecting does not work, the argument is very hard to reopen — and it is usually reopened by showing them a file cut to the fields their process actually needs rather than by asking them to try harder. That is a filtering problem, and it is solvable at the jurisdiction level once you know which authorities carry what, which is the whole argument for looking at real rows first.
How to test verified construction data before you buy
Four questions, in order. They work on us and on anyone else.
What does "verified" mean in your process, step by step? Not the adjective — the steps. If the answer is a paragraph of reassurance rather than a list of operations, there is no process to describe.
What share of my segment carries the fields I need, in the jurisdictions I buy? A single headline number across a national set is an average that hides exactly the variation you are purchasing. Ask for that breakdown by authority, not as a national average.
When a field is empty, is that you or the authority? The distinction decides whether a different vendor would help. If the jurisdiction never published it, no vendor has it, and that is worth knowing before you switch rather than after. Which authorities you buy moves field completeness more than which vendor you buy from.
Give me a sample and let me check it myself. Count the populated fields yourself, authority by authority, and compare that against whatever rate the vendor stated. Attempt contact on twenty rows. The published field list tells you which rows are worth testing. This is the only step that produces evidence rather than assurances, and a vendor reluctant to allow it has told you something. Commercial terms should follow that test, not precede it.
You should not have to verify a file before your team can use it. But you should absolutely be able to, and you should insist on doing it once before you sign anything. Ask for the numbers, then take a sample and check them yourself.
Why Alliance Data Solutions
Everything above applies to any complete records set. Here is the case for this one, in the same terms the article has used throughout.
Coverage defined at the level that actually matters. 39+ states and 300+ jurisdictions, counted by issuing authority rather than by state — because a state is not a market and a county label is not always a whole county. That distinction is why the list is published rather than summarized as a claim.
The specification is public. Twenty data points, named. Which of them a given record carries depends on what its authority published — that is what the scoring described earlier measures, and it is why the field list is worth reading before you buy rather than after. Most vendors describe their fields on a call, which is a way of not committing to them. Ours is on a page you can hold us to.
Timing you can plan against. A record filed at its source is generally available in a delivered file the next day, and within a few days for authorities that publish on slower cycles. The constraint is the authority's own schedule, not our collection.
Processing happens before delivery, not after. Deduplicated, scored on how many fields carry a value, then filtered to the record types, categories and thresholds you set. The alternative — which is what most files amount to — is a rep opening four thousand rows and spending Monday deciding which two hundred are worth touching. You are buying the triage as much as the data.
Delivered how your team already works. CSV, Excel or SFTP, cut to your specification, on the cadence written into your agreement. API access is on the 2026 roadmap and is not live today, and we would rather say so than imply otherwise. Published plans run from a single metro up to multi-state; nationwide programs and data licensing are custom-quoted, with territory, volume and frequency set in the agreement.
And the limit, stated the same way it is stated everywhere else on this site: this is public building activity, assembled, deduplicated, scored and filtered at national scale. It is not a proprietary intelligence network and we do not claim one. What makes it worth buying is that doing it yourself across three hundred authorities is impractical, not that the underlying filings are secret.
Common questions
What does verified mean for construction data?
At Alliance Data Solutions it means three specific things: records are deduplicated, each record is scored on how many of its up to twenty fields carry a value, and the set is filtered to the fields and thresholds the customer specified. It does not mean each individual value has been independently confirmed against an outside source. Completeness and correctness are different measurements and we report the one we make.
Why is construction data hard to use across multiple jurisdictions?
Because every issuing authority publishes on its own terms and no authority above them enforces a standard. The same job can appear as a property type plus sub-type in one county, a work class string in another, a numeric category code in a third, and only as free text in a fourth. Field names, date formats and status vocabularies all differ too. Any one jurisdiction is straightforward; reconciling many into a single specification is the work.
What do construction records not tell you?
Records describe filings, not commercial outcomes. They do not carry bid stage, award, material specification, budget split, supplier relationships or financing. A record confirms that a project was filed at a stated value in a stated category in a stated place. Everything past that is inference, and worth treating as inference.
How should a buyer test construction data quality?
Ask what verified means as a list of operations rather than an adjective, ask how records from different authorities are reconciled into one specification, and ask whether that specification is published. Then take a sample spanning several jurisdictions at once and check that a category filter, a status filter and a date sort behave the same way across all of them. A vendor whose file still speaks each authority's own vocabulary will fail that test in about ten minutes.