Where European B2B contact data actually comes from
Every database sells the same four sources with a different logo. Knowing which source produced a field tells you whether to trust it, and what you may lawfully do with it.
Every B2B database sells you the same thing with a different logo on it. Understanding where the rows actually come from tells you which fields to trust, which to verify, and which to ignore. It also tells you what you can lawfully do with them in Europe, which is a narrower set than the sales page implies.
There are four sources of European B2B contact data. Every vendor is a blend of them.
Source one: official registers
Every EU member state runs a business register, and most of them publish at least the company name, registration number, legal address, incorporation date and directors. Some publish annual accounts. This is the most reliable data available and almost nobody in sales uses it, because it carries no email addresses.
| Country | Register | Bulk access |
|---|---|---|
| United Kingdom | Companies House | Free API and monthly full dumps |
| Lithuania | Registrų centras | Open data sets, some paid extracts |
| Netherlands | KVK | Paid API, free search |
| France | INSEE SIRENE | Free API and full open data download |
| Poland | KRS and CEIDG | Free API |
| EU-wide | BRIS | Search only, no bulk export |
Registers give you the spine of a list: which companies exist, how old they are, how big, and where. SIRENE alone covers roughly 30 million French establishments with an activity code and a headcount band. Start here when you need coverage of a market rather than a handful of names.
The weakness is currency. A register tells you a company was incorporated and who signed the papers. It rarely tells you who runs procurement today.
Source two: web crawling
Vendors crawl company sites, job boards, press releases and public profiles, then infer structure from what they find. This produces technographics, headcount estimates, hiring signals and the "about us" page contacts.
Crawled data is fresh and wrong in interesting ways. A company that still lists a 2023 tech stack on a case study will be tagged with that stack. A regional office with its own site becomes a separate company. Treat every crawled field as a hypothesis with a date attached.
It is also the only source that produces dated events, which is why it matters more than the accuracy suggests. A job advert posted eleven days ago is a fact about this week. A headcount estimate is a guess about last year.
Source three: contribution networks
This is where most work email addresses come from, and vendors describe it in careful language. A user installs a browser extension or connects a mailbox in exchange for free credits. The vendor reads the contacts, the signatures, and sometimes the calendar. Those contacts enter the database and get sold back to the market.
The data is good. Signature blocks carry real titles, real direct dials and real addresses, verified by the fact that somebody actually corresponded with that person. The problem is that neither you nor the vendor has any relationship with the person in that row.
Under Article 14 of the GDPR, when you obtain personal data from somewhere other than the person, you owe them a notice within a month or at first contact, whichever comes first. Naming the source in your first email satisfies more of that obligation than most senders realise, and it costs one sentence.
Source four: pattern guessing
Given a name and a domain, a tool generates the twelve plausible formats, then tests which one the mail server accepts. This is how a vendor produces an address for a person who has never appeared in any of the three sources above.
Guessed addresses look identical to real ones in a CSV. The only tell is the confidence score, which most tools expose and most buyers ignore. Anything below about 90 percent belongs in a separate low-volume segment, because a bounce costs you more than the contact was worth.
Guessing fails completely against catch-all domains, which accept everything and verify nothing. Roughly a fifth of European mid-market domains are catch-all. Those rows will pass verification and bounce anyway.
What each source is actually good for
| You need | Use | Do not use |
|---|---|---|
| Does this company exist and how big is it | National register | Crawled headcount estimates |
| Is something happening there right now | Crawled job ads, filings, press | Any static firmographic field |
| Who holds this job today | Contribution network, then verify | Register director lists |
| What is their email | Contribution network, then pattern as fallback | Pattern guessing on catch-all domains |
| What do they buy and from whom | Public tender portals, filings, case studies | Vendor technographic tags without a date |
The legal position, briefly
Business contact data is personal data when it identifies a person. info@company.lt
is not personal data. d.liaudanskas@company.lt is, and so is a named row in a database
with a job title next to it.
You can process it under legitimate interest for B2B prospecting, which Recital 47 of the GDPR supports directly. What legitimate interest requires from you is a balancing test you have written down, a notice at first contact naming where you got the data, and an opt-out you honour on the first request. What it does not do is override the ePrivacy rules on unsolicited electronic mail, which vary by member state and are stricter than the GDPR in Germany, Italy, Spain, Poland and the Baltics.
Keep a record of the source per row. When a recipient asks where you got their address, the answer "a database we bought" is worse than no answer. The answer "your company register entry, plus your published team page" ends the conversation.
How to assemble a list that holds up
Start with a register pull to define the universe. Filter it on an activity code and a size band. Then layer a dated crawled signal on top to decide who gets contacted this month rather than someday. Only after that do you spend money resolving contacts, because contact enrichment is the expensive step and you have just cut the volume by 90 percent.
Verify every address. Segment anything under 90 percent confidence. Record the source. Then send.
Most teams run this in the opposite order: buy 20,000 contacts, filter by industry dropdown, send to all of them. That produces the bounce rate, the complaint rate and the reply rate you would expect from a list nobody thought about.
Ripe Leads builds lists from registers and dated signals
Ripe Leads is KoFi Tech's outbound arm. It starts from official registers and public filings, layers a dated buying signal, and resolves contacts last. Every row carries its source, which is what makes the Article 14 notice honest.
Visit Ripe Leads