Why the Largest Public Cookie Database is Missing 60% of Cookies
Your latest privacy scan finds 312 cookies. You recognize maybe 40 of them. The rest: unknown names, unknown owners, unknown purpose.
So the cookie identification detective work begins. Google the name. Check a public cookie database. Dig through internal documentation. Ping Marketing on Slack. Ask Engineering. Somewhere, someone who left the company two years ago probably knows the answer.
Sound familiar? This isn’t unusual. Most privacy teams know the drill.
Finding where cookies are on your site is hard enough. Understanding what they actually do is harder, and many of the tools built to help don’t fully solve either problem.
Privacy regulations like GDPR, the ePrivacy Directive, and CCPA/CPRA converge on the same requirements: you have to know what each cookie does, its purpose, and where that data goes.
An unknown or miscategorized cookie means you can’t inform a regulator, a customer, or your own legal team. Cookie identification is the starting point for consent accuracy, audit readiness, and vendor accountability.
Most Cookie Databases Aren’t Keeping Up with Your Website
For years, privacy teams have leaned on public cookie databases to answer one simple question: what is this cookie doing on my website? The problem is that most of those databases were never built to keep up with how enterprise websites actually change.
New vendors get added. Campaigns launch and die. Tags get deployed and forgotten. A cookie that exists today may vanish tomorrow.
Researchers running a 2024 study on cookie banner tools crawled 30,000 websites and needed to classify the cookies they found. Using Cookiepedia, one of the largest public cookie databases available, they could only match 38% of the distinct cookies to an entry. So, that’s less than 40% of known cookies catalogued by a source privacy teams have relied on for years as if it were comprehensive. That leaves the other 60% unaccounted for.
Will LLMs Replace Cookie Databases?
It’s tempting to think AI could solve the cookie identification problem. Just ask an LLM what a cookie does, right? The very nature of LLMs handicaps their results. Why aren’t LLMs a good solution to identify cookies?
Capability
LLMs don’t natively recognize cookies, tags, or tracking pixels. They weren’t trained to detect those patterns, and they often rely on those very same outdated cookie databases to know what each one does. If it doesn’t exist in one of those databases, the database just returns no answer.
Determinism
LLMs are probabilistic, not factual. They generate the most statistically likely answer, not the correct one, which is why they hallucinate. In our own testing, LLMs fabricated cookies and tags that didn’t exist more than 50% of the time and correctly categorized real ones less than 50% of the time. Preventing that kind of error requires a deterministic tool built on verified data.
Consistency
Ask an LLM the same question twice, and you may get two different answers. Inconsistent output means false positives, and false positives mean analysts waste time chasing down errors.
Cookie Governance Requires More Than Categorization
Cookie categories are an important first step in collecting user consent, allowing them to agree to the type of cookies they will accept. Cookie consent categories sort cookies by function, typically: necessary, performance, functionality, and targeting or advertising. They tell the visitor broadly what a cookie is for, so websites must categorize a cookie correctly.
Even a correctly categorized cookie doesn’t answer the questions privacy teams actually need to manage. Who approved this cookie? Which team owns it? Why was it added? When was it last reviewed? Is it still needed? Can it be safely removed?
Those answers don’t live in any external database. They live inside your organization, usually in someone’s memory. When that person changes roles or leaves the company, the knowledge can often leave with them. The next privacy review would start from the beginning. The same questions get asked again. The same research gets repeated, and resources are wasted.
That’s a clear governance problem, and no existing database, however fast or accurate, was built to solve it. Until now.
The Cookie Database Built from Billions of Real Scans
Getting a real answer isn’t about a faster way to look something up. It’s about starting from more complete data in the first place, and being honest about how confident an answer actually is.
That’s the bet behind ObservePoint’s Cookie and Tag Database: built from more than 10 years of real ObservePoint scans across enterprise websites, not a public crawl or a one-time research pass, and not a model guessing. Every entry comes from an actual cookie observed in the wild and includes what tags set it, what it’s used for, what category it’s most commonly found in, its risk level and why, and how prevalent it is. Every entry is also reviewed by a person before it reaches a customer.
The result is deterministic: the same cookie returns the same answer every time.
ObservePoint’s Cookie Database also allows users to edit the institutional record next to the definition itself: who approved a cookie, who owns it, when it was last reviewed, so that answer survives longer than the person who gave it.
The resulting cookie database is the most comprehensive and most accurate database you can find.
- 4.2 billion cookies from real websites were scanned
- 6.35 million cookies defined, categorized, and risk-assessed
- 125K tags identified
- 135K vendors mapped
And, you can leverage ObservePoint’s other features like patented cookie origin reporting and tag and cookie initiators.
The real fix is never guessing at all. And if you ever find a cookie we haven’t seen yet, tell us. We’ll add it.