Bank of America just launched an AI tracker. But who audits the auditor?
Last week, a press release crossed my terminal. Bank of America, the second-largest bank in the US, rolled out a tool to track “model intelligence and costs” across the AI landscape. The crypto community barely blinked. But I did. Not because this is a breakthrough—it’s not. Because it’s a perfect case study in how centralized information asymmetry sneaks into the very fabric of the emerging AI economy.
Let me be clear: I’m not impressed by the launch. I’m alarmed by the lack of scrutiny surrounding it. As someone who spent 14 nights manually auditing TheDAO’s successor contracts in 2017, I learned one thing: code does not lie, but it does hide. The same principle applies to financial research tools. The question is not whether Bank of America can aggregate benchmarks. It’s whether we can verify the data, logic, and incentives behind that aggregation.
Context: What We Know (and What We Don’t)
From the sparse announcement, we know the tool covers two dimensions: “model intelligence” (likely a composite of public benchmarks like MMLU, HumanEval, MATH) and “costs” (probably API pricing per million tokens). The intended audience is institutional investors and corporate decision-makers. The implicit goal is to help them compare AI models on a standardized scale.

But here’s the gaping hole: no methodology, no data sources, no update frequency, no open API. The tool is a black box wrapped in a bank’s brand. In crypto, we call that a “trust me” model. And we know where that leads.
Core: The Architecture of Opaque Scoring
Let’s reverse-engineer what this tool likely looks like under the hood. Based on my experience stress-testing Curve Finance’s invariant calculations in 2020, I’d bet the tech stack is straightforward:
- Data ingestion layer: Scrapes public benchmarks from Hugging Face, LMArena, and API pricing from model providers. Some data may come from proprietary Bloomberg terminals.
- Normalization engine: Transforms diverse scores (0-100, pass@1, accuracy) into a single “intelligence” metric. This is where the real magic—and manipulation—happens.
- Weighting function: Assigns importance to each benchmark. Does MMLU count more than HumanEval? If so, a model strong in reasoning but weak in coding gets unfairly penalized.
- Cost integration: Creates a simple ratio (intelligence / cost) to highlight “high value” models.
This is a classic “combinatorial innovation.” It’s not new. Artificial Analysis and Vellum already do this. The difference is distribution: Bank of America can shove this into quarterly earnings calls.
But here’s the contrarian truth: this tool doesn’t reduce information asymmetry—it centralizes it. The bank now controls the lens through which institutional investors see AI models. If they decide to lower the weight of a benchmark that favors a competitor’s model, they can. If they want to promote a model from a company they’re advising on an IPO, they can. No code is published. No audit trail exists.
Contrarian: The Blind Spots That Matter More Than Benchmarks
Every AI evaluation framework suffers from benchmark overfitting. Models train to the test. But the real risk isn’t technical—it’s structural. Bank of America is both a financier of AI companies (via investment banking) and a judge of their models. This is a textbook conflict of interest, yet no one is asking for a decentralized audit.
Tracing the noise floor to find the alpha signal.
In my 2021 work analyzing NFT metadata storage, I found that 40% of “decentralized” NFTs had centralized links. The same pattern is repeating here: a “standardized” tracker that is anything but standard. The bank’s weightings are trade secrets. The data sources are proprietary. The model selection is opaque.
What about the models that don’t appear on the tracker? Open-source models like Llama 3.1, fine-tuned variants, or Chinese models like DeepSeek? If they’re excluded, the tracker misleads investors into thinking the AI landscape is narrower than it is.

And what about decentralized AI? Projects like Bittensor, Render, or Akash offer compute and model inference outside the traditional API model. The tracker likely ignores them because they don’t fit the “cost per token” framework. This is the same mistake Wall Street made with crypto in 2013: evaluating a new paradigm using old metrics.
Takeaway: The Real Value Lies in Verifiable, On-Chain Evaluation
Redundancy is the enemy of scalability. But so is opacity. The Bank of America AI tracker is a symptom of a larger problem: the lack of transparent, verifiable, and decentralized model evaluation infrastructure. The next wave of crypto AI projects will need to build what this bank’s tool pretends to be—a trustless, community-audited index of model intelligence and cost.
Until then, treat this tracker like a single point of failure. Diversify your sources, question the weights, and demand the code. Because in the end, logic gates are the new legal contracts. And Bank of America just signed a contract without publishing the terms.