Where European B2B contact data actually comes from

Every database sells the same four sources with a different logo. Knowing which source produced a field tells you whether to trust it, and what you may lawfully do with it.

Insights cover: four streams of unequal weight converging into one contact row.

Every B2B database sells you the same thing with a different logo on it. Understanding where the rows actually come from tells you which fields to trust, which to verify, and which to ignore. It also tells you what you can lawfully do with them in Europe, which is a narrower set than the sales page implies.

There are four sources of European B2B contact data. Every vendor is a blend of them.

Source one: official registers

Every EU member state runs a business register, and most of them publish at least the company name, registration number, legal address, incorporation date and directors. Some publish annual accounts. This is the most reliable data available and almost nobody in sales uses it, because it carries no email addresses.

Company registers worth pulling from directly
CountryRegisterBulk access
United KingdomCompanies HouseFree API and monthly full dumps
LithuaniaRegistrų centrasOpen data sets, some paid extracts
NetherlandsKVKPaid API, free search
FranceINSEE SIRENEFree API and full open data download
PolandKRS and CEIDGFree API
EU-wideBRISSearch only, no bulk export

Registers give you the spine of a list: which companies exist, how old they are, how big, and where. SIRENE alone covers roughly 30 million French establishments with an activity code and a headcount band. Start here when you need coverage of a market rather than a handful of names.

The weakness is currency. A register tells you a company was incorporated and who signed the papers. It rarely tells you who runs procurement today.

Source two: web crawling

Vendors crawl company sites, job boards, press releases and public profiles, then infer structure from what they find. This produces technographics, headcount estimates, hiring signals and the "about us" page contacts.

Crawled data is fresh and wrong in interesting ways. A company that still lists a 2023 tech stack on a case study will be tagged with that stack. A regional office with its own site becomes a separate company. Treat every crawled field as a hypothesis with a date attached.

It is also the only source that produces dated events, which is why it matters more than the accuracy suggests. A job advert posted eleven days ago is a fact about this week. A headcount estimate is a guess about last year.

Source three: contribution networks

This is where most work email addresses come from, and vendors describe it in careful language. A user installs a browser extension or connects a mailbox in exchange for free credits. The vendor reads the contacts, the signatures, and sometimes the calendar. Those contacts enter the database and get sold back to the market.

The data is good. Signature blocks carry real titles, real direct dials and real addresses, verified by the fact that somebody actually corresponded with that person. The problem is that neither you nor the vendor has any relationship with the person in that row.

Under Article 14 of the GDPR, when you obtain personal data from somewhere other than the person, you owe them a notice within a month or at first contact, whichever comes first. Naming the source in your first email satisfies more of that obligation than most senders realise, and it costs one sentence.

Source four: pattern guessing

Given a name and a domain, a tool generates the twelve plausible formats, then tests which one the mail server accepts. This is how a vendor produces an address for a person who has never appeared in any of the three sources above.

Guessed addresses look identical to real ones in a CSV. The only tell is the confidence score, which most tools expose and most buyers ignore. Anything below about 90 percent belongs in a separate low-volume segment, because a bounce costs you more than the contact was worth.

Guessing fails completely against catch-all domains, which accept everything and verify nothing. Roughly a fifth of European mid-market domains are catch-all. Those rows will pass verification and bounce anyway.

What each source is actually good for

Trust each source for what it knows
You needUseDo not use
Does this company exist and how big is itNational registerCrawled headcount estimates
Is something happening there right nowCrawled job ads, filings, pressAny static firmographic field
Who holds this job todayContribution network, then verifyRegister director lists
What is their emailContribution network, then pattern as fallbackPattern guessing on catch-all domains
What do they buy and from whomPublic tender portals, filings, case studiesVendor technographic tags without a date

The legal position, briefly

Business contact data is personal data when it identifies a person. info@company.lt is not personal data. d.liaudanskas@company.lt is, and so is a named row in a database with a job title next to it.

You can process it under legitimate interest for B2B prospecting, which Recital 47 of the GDPR supports directly. What legitimate interest requires from you is a balancing test you have written down, a notice at first contact naming where you got the data, and an opt-out you honour on the first request. What it does not do is override the ePrivacy rules on unsolicited electronic mail, which vary by member state and are stricter than the GDPR in Germany, Italy, Spain, Poland and the Baltics.

Keep a record of the source per row. When a recipient asks where you got their address, the answer "a database we bought" is worse than no answer. The answer "your company register entry, plus your published team page" ends the conversation.

How to assemble a list that holds up

Start with a register pull to define the universe. Filter it on an activity code and a size band. Then layer a dated crawled signal on top to decide who gets contacted this month rather than someday. Only after that do you spend money resolving contacts, because contact enrichment is the expensive step and you have just cut the volume by 90 percent.

Verify every address. Segment anything under 90 percent confidence. Record the source. Then send.

Most teams run this in the opposite order: buy 20,000 contacts, filter by industry dropdown, send to all of them. That produces the bounce rate, the complaint rate and the reply rate you would expect from a list nobody thought about.

Related work

Ripe Leads builds lists from registers and dated signals

Ripe Leads is KoFi Tech's outbound arm. It starts from official registers and public filings, layers a dated buying signal, and resolves contacts last. Every row carries its source, which is what makes the Article 14 notice honest.

Visit Ripe Leads