{"id":6120,"date":"2026-07-10T23:21:24","date_gmt":"2026-07-10T23:21:24","guid":{"rendered":"https:\/\/tokenmetrics.com\/blog\/how-token-metrics-evaluates-ai-crypto-research-tools\/"},"modified":"2026-07-14T09:04:48","modified_gmt":"2026-07-14T09:04:48","slug":"how-token-metrics-evaluates-ai-crypto-research-tools","status":"publish","type":"post","link":"https:\/\/tokenmetrics.com\/blog\/how-token-metrics-evaluates-ai-crypto-research-tools\/","title":{"rendered":"How Token Metrics Evaluates AI Crypto Research Tools"},"content":{"rendered":"<style id=\"tm-mobile-overflow-containment-v1\">.single .entry-content,.single .entry-content *{box-sizing:border-box}.single .entry-content{max-width:100%;overflow-x:hidden}.single .entry-content img,.single .entry-content iframe,.single .entry-content pre{max-width:100%}@media(max-width:760px){.single .entry-content table{display:block!important;width:100%!important;max-width:100%!important;overflow-x:auto!important;-webkit-overflow-scrolling:touch}.single .entry-content th,.single .entry-content td{overflow-wrap:anywhere}.single .entry-content a{overflow-wrap:anywhere;word-break:break-word}}<\/style>\n<style>.single .entry-content .tm-b3{box-sizing:border-box;margin:28px 0;padding:clamp(18px,3vw,28px);border:1px solid rgba(17,17,17,.12);border-radius:24px;background:linear-gradient(145deg,#fffdf3,#fff 58%,#f4f4f5);box-shadow:0 14px 34px rgba(17,17,17,.07)}.single .entry-content .tm-b3 *{box-sizing:border-box}.single .entry-content .tm-b3-grid{display:grid;grid-template-columns:repeat(3,minmax(0,1fr));gap:12px}.single .entry-content .tm-b3 .tm-nd-card{min-width:0;padding:16px;border:1px solid rgba(17,17,17,.1);border-radius:16px;background:#fff}.single .entry-content .tm-b3 a{overflow-wrap:anywhere}.single .entry-content .tm-b3 select{width:100%;max-width:100%;padding:10px;margin-top:6px;border:1px solid #a3a3a3;border-radius:8px;background:#fff}.single .entry-content .tm-b3 button{border:0;cursor:pointer}@media(max-width:700px){.single .entry-content .tm-b3-grid{grid-template-columns:1fr}.single .entry-content .tm-b3{padding:18px;border-radius:18px}}<\/style>\n<p><strong>Quick answer:<\/strong> We judge crypto research tools by the question they help answer and the proof a reader can inspect. We do not treat an AI label as proof of accuracy. We do not publish a numeric score without dated inputs, fixed weights, and a reproducible calculation.<\/p>\n<p class=\"tm-editorial-attribution\"><strong>Prepared by Token Metrics Research Team.<\/strong> Source review date: July 14, 2026. Product terms can change; check the official sources below.<\/p>\n<h2 class=\"wp-block-heading\">Scope<\/h2>\n<p>This method covers on-chain search, entity and address labels, alerts, portfolio views, data queries, and API access. It also covers systems that summarize data with a model. It does not test future returns, trade signals, legal status, or the identity of a wallet owner.<\/p>\n<p>The goal is a sound research trail. A useful tool should help a reader move from a question to evidence, note uncertainty, and check the result.<\/p>\n<h2 class=\"wp-block-heading\">Evidence hierarchy<\/h2>\n<p>Official API docs, method notes, privacy terms, changelogs, and product help come first. Raw chain data comes next for a known test. A product output is an interpretation of that data, not a replacement for it.<\/p>\n<p>A vendor blog can explain a feature. It cannot prove a label is right in every case. We separate \u201cdocumented,\u201d \u201cobserved in a test,\u201d and \u201cverified against a primary record.\u201d<\/p>\n<ul>\n<li>First: official docs, methods, terms, and change logs.<\/li>\n<li>Second: saved tests with known addresses, times, and filters.<\/li>\n<li>Third: raw chain records and other direct evidence.<\/li>\n<li>Context only: third-party reviews and social posts.<\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\">Criteria and tests<\/h2>\n<p>Question fit asks whether the tool can answer the named task. Traceability asks whether the reader can see the address, transaction, time, filter, and source behind an output. Coverage checks the relevant chain and field. Control checks permissions, rate limits, exports, retention, and deletion.<\/p>\n<p>For labels, we look for method notes and a way to challenge or cross-check the label. For alerts, we save the rule and compare the event time with the chain. For APIs, we test auth, errors, pagination, limits, and a small known response.<\/p>\n<h2 class=\"wp-block-heading\">Scoring policy<\/h2>\n<p>This cluster uses no numeric product score. The evidence is not a single benchmark with common fixtures and public weights. A precise number would hide the difference between address research, alerts, dashboards, and API work.<\/p>\n<p>We instead use pass, partial, fail, and not tested for each named job. If a later score is added, the page must show the formula, weights, fixture version, date, missing-data rule, and per-product inputs.<\/p>\n<h2 class=\"wp-block-heading\">Research process<\/h2>\n<p>We write the research question first. We then list what evidence would change the answer. We collect primary sources, run a narrow test, save the query or address, and note the date. A second person should be able to follow the trail.<\/p>\n<p>We do not enter private keys, seed phrases, client data, or material nonpublic data. API keys stay in a secret store. A model summary is checked against cited records before use.<\/p>\n<h2 class=\"wp-block-heading\">Freshness and updates<\/h2>\n<p>We review after a change to supported chains, endpoints, labels, alert rules, prices, privacy terms, or access. Official changelogs help, but a new page alone does not prove old behavior still works.<\/p>\n<p>Time-sensitive claims use a date. We avoid \u201creal time,\u201d \u201calways,\u201d and fixed coverage counts unless the source defines and updates them.<\/p>\n<h2 class=\"wp-block-heading\">Limits and conflicts<\/h2>\n<p>On-chain data is public, but entity labels are inferences. One entity may use many addresses. One address may serve many users. Bridges and contracts can make flows hard to read. A tool can be useful and still be wrong.<\/p>\n<p>Token Metrics may have a commercial interest in its own product or partner links. We disclose that fact. It does not count as evidence or add weight.<\/p>\n<h2 class=\"wp-block-heading\">Worked example<\/h2>\n<p>A reader wants to know whether a public address sent funds to an exchange. The reviewer saves the address, chain, time range, and transaction hash. The tool label is one clue. The chain record confirms the transfer. The exchange label still needs a documented basis.<\/p>\n<p>The result should say what is known, what is inferred, and what is not known. It should not turn one transfer into a claim about the owner, motive, or next price move.<\/p>\n<h2 class=\"wp-block-heading\">Frequently asked questions<\/h2>\n<h3 class=\"wp-block-heading\">Do AI features make a tool more accurate?<\/h3>\n<p>Not by themselves. Accuracy needs a defined task, known examples, error review, and a way to trace the output.<\/p>\n<h3 class=\"wp-block-heading\">Can a wallet label prove identity?<\/h3>\n<p>No. Treat it as a research lead unless strong primary evidence supports it.<\/p>\n<h3 class=\"wp-block-heading\">Why is there no ranking?<\/h3>\n<p>The tools serve different jobs and the current evidence does not support one public, reproducible score.<\/p>\n<h3 class=\"wp-block-heading\">What should I save from a test?<\/h3>\n<p>Save the question, address or query, chain, filters, time, output, source links, and any manual check.<\/p>\n<h2 class=\"wp-block-heading\">Editorial release gate<\/h2>\n<p>Before a page ships, an editor checks every product claim against the linked source. A second pass removes stale counts, vague superlatives, and words such as \u201calways,\u201d \u201cinstant,\u201d or \u201cguaranteed.\u201d The editor also checks that an example is clearly an example and not a hidden claim about a live account.<\/p>\n<p>The final copy must state what was not tested. It must show why a product may fit one reader and fail another. Links are opened again at release. If a source has moved, the claim is held until the new primary page is found.<\/p>\n<p>A score cannot pass this gate by sounding precise. It needs a public formula and saved inputs. When that proof is absent, \u201cnot scored\u201d is the correct result. This rule keeps a clean card design without turning an opinion into a fact.<\/p>\n<p>The release record names the editor, date, source set, and next review trigger. That record is part of the method, not a claim that a named outside advisor approved the work.<\/p>\n<h2 class=\"wp-block-heading\">What would change this decision<\/h2>\n<p>The method should change when a test no longer matches the reader job, a better primary source exists, or the same failure appears across products.<\/p>\n<p>A change is not accepted from a headline alone. Open the primary page, note the date, and repeat the hard case. Keep the old finding until the new result can be checked. If the change affects only one region, plan, chain, or platform, say so rather than rewriting the whole verdict.<\/p>\n<p>Readers should also revisit the choice when their own job changes. A tool selected for a quick view may be wrong for an audit trail. A tool selected for one public address may be wrong for a team API. Fit belongs to the task, so the update record must name the task.<\/p>\n<h2>Related guides in this research cluster<\/h2>\n<ul class=\"tm-cluster-links\">\n<li><a href=\"https:\/\/tokenmetrics.com\/blog\/best-ai-crypto-research-tools-2026\/\">AI crypto research tools hub<\/a><\/li>\n<li><a href=\"https:\/\/tokenmetrics.com\/blog\/nansen-profile-2026\/\">Nansen profile<\/a><\/li>\n<li><a href=\"https:\/\/tokenmetrics.com\/blog\/arkham-profile-2026\/\">Arkham profile<\/a><\/li>\n<\/ul>\n<h2 class=\"wp-block-heading\">Sources checked<\/h2>\n<ul>\n<li><a href=\"https:\/\/docs.nansen.ai\/\" target=\"_blank\" rel=\"nofollow noopener\">Nansen API introduction<\/a><\/li>\n<li><a href=\"https:\/\/docs.nansen.ai\/api\/profiler\/address-labels\" target=\"_blank\" rel=\"nofollow noopener\">Nansen address-label endpoint<\/a><\/li>\n<li><a href=\"https:\/\/info.arkm.com\/arkham-intel-api\" target=\"_blank\" rel=\"nofollow noopener\">Arkham Intel API<\/a><\/li>\n<\/ul>\n<section class=\"tm-nd-disclosure\">\n<p><strong>Disclosure:<\/strong> Token Metrics may earn compensation when readers use some partner links. Compensation does not decide what we cover. This page gives general educational information, not investment, legal, tax, custody, compliance, or engineering advice.<\/p>\n<\/section>\n<p><a class=\"tm-nd-cta\" data-tm-track=\"tm-cta-click\" data-tm-cta=\"daily-pulse\" href=\"https:\/\/tokenmetrics.com\/?utm_source=tokenmetrics_blog&amp;utm_medium=ai_research_tools&amp;utm_campaign=nerdwallet_guide_2026\">Start with the free Daily Pulse<\/a><\/p>\n<p><script type=\"application\/ld+json\">{\"@context\":\"https:\/\/schema.org\",\"@type\":\"Article\",\"headline\":\"How Token Metrics Evaluates AI Crypto Research Tools\",\"publisher\":{\"@type\":\"Organization\",\"name\":\"Token Metrics\"}}<\/script><\/p>\n","protected":false},"excerpt":{"rendered":"Quick answer: We judge crypto research tools by the question they help answer and the proof a reader&hellip;","protected":false},"author":1,"featured_media":6211,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_jetpack_newsletter_access":"","_jetpack_dont_email_post_to_subs":false,"_jetpack_newsletter_tier_id":0,"_jetpack_memberships_contains_paywalled_content":false,"_tm_paid_cta_tier":"","_tm_paid_cta_heading":"","_tm_paid_cta_body":"","csco_display_header_overlay":false,"csco_singular_sidebar":"","csco_page_header_type":"","csco_page_load_nextpost":"","csco_page_reading_time":"","csco_page_toc_navigation":"","csco_post_video_location":[],"csco_post_video_location_hash":"","csco_post_video_url":"","csco_post_video_bg_start_time":0,"csco_post_video_bg_end_time":0,"csco_post_video_bg_volume":false,"_jetpack_memberships_contains_paid_content":false,"footnotes":"","jetpack_post_was_ever_published":false},"categories":[6],"tags":[],"sections":[],"entities":[],"class_list":["post-6120","post","type-post","status-publish","format-standard","has-post-thumbnail","category-guides","cs-entry","cs-video-wrap"],"jetpack_sharing_enabled":true,"jetpack_featured_media_url":"https:\/\/tokenmetrics.com\/blog\/wp-content\/uploads\/2026\/07\/nerdwallet-6120-editorial-20260713.webp","_links":{"self":[{"href":"https:\/\/tokenmetrics.com\/blog\/wp-json\/wp\/v2\/posts\/6120","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/tokenmetrics.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/tokenmetrics.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/tokenmetrics.com\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/tokenmetrics.com\/blog\/wp-json\/wp\/v2\/comments?post=6120"}],"version-history":[{"count":4,"href":"https:\/\/tokenmetrics.com\/blog\/wp-json\/wp\/v2\/posts\/6120\/revisions"}],"predecessor-version":[{"id":6322,"href":"https:\/\/tokenmetrics.com\/blog\/wp-json\/wp\/v2\/posts\/6120\/revisions\/6322"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/tokenmetrics.com\/blog\/wp-json\/wp\/v2\/media\/6211"}],"wp:attachment":[{"href":"https:\/\/tokenmetrics.com\/blog\/wp-json\/wp\/v2\/media?parent=6120"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/tokenmetrics.com\/blog\/wp-json\/wp\/v2\/categories?post=6120"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/tokenmetrics.com\/blog\/wp-json\/wp\/v2\/tags?post=6120"},{"taxonomy":"section","embeddable":true,"href":"https:\/\/tokenmetrics.com\/blog\/wp-json\/wp\/v2\/sections?post=6120"},{"taxonomy":"entity","embeddable":true,"href":"https:\/\/tokenmetrics.com\/blog\/wp-json\/wp\/v2\/entities?post=6120"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}