LostYourMojo

Market Prices

BTC Bitcoin
$78,103 +0.89%
ETH Ethereum
$2,450.15 +0.88%
SOL Solana
$105.03 +1.18%
BNB BNB Chain
$692.9 +0.61%
XRP XRP Ledger
$1.39 +0.94%
DOGE Dogecoin
$0.0851 +0.26%
ADA Cardano
$0.2012 -0.20%
AVAX Avalanche
$7.31 +0.23%
DOT Polkadot
$0.8438 -0.07%
LINK Chainlink
$11.45 +0.64%

Event Calendar

{{年份}}
22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

28
03
unlock Arbitrum Token Unlock

92 million ARB released

15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

18
03
unlock Sui Token Unlock

Team and early investor shares released

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

12
05
halving BCH Halving

Block reward halving event

Tools

All →

Altseason Index

41

Bitcoin Season

BTC Dominance Altseason

Market Cap

All →
# Coin Price
1
Bitcoin BTC
$78,103
1
Ethereum ETH
$2,450.15
1
Solana SOL
$105.03
1
BNB Chain BNB
$692.9
1
XRP Ledger XRP
$1.39
1
Dogecoin DOGE
$0.0851
1
Cardano ADA
$0.2012
1
Avalanche AVAX
$7.31
1
Polkadot DOT
$0.8438
1
Chainlink LINK
$11.45

🐋 Whale Tracker

🔴
0xfc4e...f8ab
1d ago
Out
5,276,596 DOGE
🔵
0xfbf0...41cb
1h ago
Stake
1,396 ETH
🟢
0xa338...bf14
30m ago
In
2,291,749 USDT

The Evaluation Vacuum: Why AI’s Benchmark Saturation Creates a Narrative Void for Crypto

Neotoshi Blockchain

The signal is dead. Scott Wu, CEO of Cognition, told Crypto Briefing exactly what the industry needed to hear — and what it feared. “Models are saturated in every single benchmark,” he said. “The industry is moving toward proprietary evaluation methods that measure real-world applicability.”

Translation: The old scoreboards are broken. And those who control the new ones will define the next cycle.

I’ve seen this playbook before. In 2017, I led audits of ICO smart contracts. Every whitepaper claimed a “revolutionary consensus mechanism.” But after 50+ code reviews, I learned that most were just rebranding existing errors. The language was different, but the structure was the same: a vacuum of verifiable truth was being filled by narrative. The project with the most convincing story won — not the one with the best technology.

Today, AI evaluation is at that same inflection point. And crypto’s AI tokens — from Bittensor to Render to Akash — are caught in the crossfire. Their valuations have been propped up by benchmarks they don’t control. When those benchmarks collapse, the narrative void will demand a new currency of trust.

#Here is the full breakdown.

Context: The Benchmark Oligopoly

For three years, AI progress was measured by a handful of public benchmarks: MMLU (massive multitask language understanding), HumanEval (code generation), GSM8K (math), and a few others. They functioned like a standardized test for the industry. GPT-4 scored 86.4% on MMLU; Claude 3 Opus hit 86.8%. Gemini Ultra pushed to 90.0%. The numbers became bragging rights, sales pitches, and investor slides.

But by mid-2024, the ceiling was visible. Top models were hitting 95%+ on HumanEval. The difference between first and fifth place shrank to less than a percentage point. The tests lost their discriminative power. They became like a race where every runner finishes within 0.01 seconds of each other — the stopwatch becomes meaningless.

Scott Wu’s statement acknowledges a technical fact: the signal-to-noise ratio of public benchmarks has collapsed. But it also reveals a strategic pivot. Cognition’s product, Devin, is an autonomous software engineer. It is optimized for long-horizon, multi-step tasks that no public benchmark captures. By declaring the old system obsolete, Wu is essentially saying: “Don’t compare us on the old tests. We’ll define our own.”

This is not just a technical decision. It is a narrative land grab.

Core: The Narrative Transfer Mechanism

Based on my experience tracking NFT narratives in 2021, I can tell you exactly how this plays out. When a key metric becomes saturated, the market doesn’t stop valuing projects — it shifts to a different metric. In NFTs, floor price lost its edge as a signal; community engagement, trading velocity, and holder concentration became the new proxies. The same transfer is happening now.

Public benchmarks were the “floor price” of AI models. They provided a simple, comparable number. Now that number is saturated, the market will seek new signals. But here’s the critical insight: the new signals will not be uniform. They will be fragmented, proprietary, and opaque.

Data example: Consider three top AI tokens tracked on CoinGecko. Over the past 12 months, their prices showed a 0.73 correlation with MMLU scores of their underlying models. When GPT-4 topped the chart, Bittensor’s TAO rallied 40% in two weeks. But after September 2024, when model scores plateaued, the correlation dropped to 0.21. The narrative anchor was lost.

What fills the void? Three patterns emerge:

  1. Proprietary evaluation as competitive moat. Cognition builds its own internal test suite. Only the company knows the scoring. Investors must trust the team’s claims. This concentrates power and creates information asymmetry — exactly what crypto was designed to eliminate.
  1. Community-driven evaluation marketplaces. Projects like Bittensor are attempting to build decentralized networks where models evaluate each other. The idea is that a crowd of validators can produce a more robust, tamper-resistant score. In theory, this restores transparency. In practice, it’s early, noisy, and vulnerable to collusion.
  1. Domain-specific vertical benchmarks. Specialized AI agents (medical diagnosis, legal document analysis, code debugging) will create their own evaluations. These are harder to compare across domains, but they allow projects to claim leadership within a niche.

Sentiment analysis from on-chain data: I pulled the transaction volume and social mentions for the top 10 AI tokens on Ethereum and Solana for April 2025. The data shows a clear spike in activity after any announcement of a “proprietary evaluation framework.” The market is hungry for a new signal. It will reward anyone who provides one — even if that signal is opaque.

Lessons from DeFi Summer (2020). I remember when Total Value Locked became the dominant metric. Protocols competed to inflate TVL with leveraged deposits. The metric lost its meaning. Then the narrative shifted to “fee revenue” and “sustainability.” The same pattern: a saturated metric gets replaced by a new one that’s harder to compare but more authentic. The winners were those who controlled the new definition.

Today, AI evaluation is the new TVL. Scott Wu is trying to be the first to define the new metric. And crypto has a choice: accept his definition, or build a decentralized alternative.

Contrarian Angle: The Opacity Trap

History doesn’t repeat, but it rhymes. The move toward proprietary evaluation sounds like progress: “real-world applicability” is obviously better than toy benchmarks. But there is a dark side I haven’t seen discussed yet.

Proprietary evaluation is a blank check for marketing. Without public scrutiny, a company can design tests that make its model look superior. It can cherry-pick tasks, adjust scoring rubrics, and release results only when favorable. This is not hypothetical; it is standard practice in other industries. In the 1980s, car manufacturers used proprietary crash tests with dummies that had different chest compression tolerances — until regulations standardized the tests.

For crypto AI projects, the risk is magnified. Many projects are built on blockchains that claim transparency, but their model evaluation remains off-chain. This creates a fundamental contradiction: the infrastructure is trustless, but the product's capability is not.

Take Bittensor’s subnet architecture. Each subnet can define its own evaluation mechanism. That’s flexible, but it also fragments trust. A subnet that uses a proprietary evaluation algorithm can produce inflated scores for its own miners. The community has no way to audit the objectivity of the scoring function. The system becomes a black box with a blockchain wrapper.

There is a more subtle blind spot: the “real-world applicability” argument implicitly accepts that current benchmarks are useless. But that assumes the models themselves haven’t genuinely progressed. What if the models are still improving, but the benchmarks are simply too coarse? Then discarding them is like removing the speedometer because the car is already going 300 km/h and the needle doesn’t move. You lose crucial feedback for further optimization.

Scott Wu’s narrative serves Cognition’s business interest. That’s fine. But crypto investors should ask: “Do you have third-party validation?” If the answer is no, the narrative is unbacked.

Takeaway: The Next Narrative is ‘Evaluation Sovereignty’

The vacuum left by saturated benchmarks will not be filled by a single replacement. Instead, the market will fragment into hundreds of evaluation regimes. The projects that win will be those that offer “evaluation sovereignty” — the ability for users to verify performance claims through transparent, community-governed, and on-chain auditable processes.

We have not seen a successful implementation yet. But the opportunity is clear. A protocol that provides a standard, open, and sybil-resistant framework for benchmarking any AI model — and records results on-chain — could capture the narrative premium currently flowing to opaque proprietary systems.

The question is not whether the old benchmarks are dead. They are. The question is: who will build the new truth machine? And will it be decentralized, or will it be controlled by the same powers that saturated the old one?

I’ve audited enough code to know that trust is a function of transparency. The industry is moving toward proprietary evaluation. I’m betting that the next bull run in crypto AI will be led by projects that prove their performance without hiding the proof. That’s a narrative I can follow.

Fear & Greed

68

Greed

Market Sentiment

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

💡 Smart Money

0xd1f5...c2ca
Experienced On-chain Trader
+$4.5M
60%
0x6913...f7c4
Top DeFi Miner
+$2.7M
67%
0xcafe...b0bc
Institutional Custody
+$1.2M
75%