In the quiet of a Singapore evening, I found myself staring at a Bloomberg terminal. The numbers blinked—APY, TVL, gas fees—but the newest addition was a tracker: model intelligence and cost. Why does a bank need to rate AI? Because the same institutions that defined the value of debt now want to define the value of thought. My code was the covenant, not just the contract. But here, the covenant was written by a committee. The screen glowed with a promise of clarity, but my mind drifted to a different kind of light—the soft glow of a bear market, where the silence taught us the truth. In the silence of the bear, we heard the truth. And now, with this tracker, I wonder if the noise is returning.
Bank of America, the giant of Charlotte, has launched a tool that tracks AI model intelligence and costs. The facts are sparse: it is a research product for institutional clients, likely aggregating public benchmarks like MMLU, HumanEval, and API pricing data. It is not a new model, but a lens—a way to compare the brains and the budgets of the AI economy. The crypto industry knows this pattern well. We have our own trackers: DeFiLlama for TVL, CoinGecko for prices, L2Beat for rollups. But this is different. This is a bank inserting itself into the evaluation of intelligence, a move that echoes the credit rating agencies of 2008. The tool will probably be free for clients, monetized indirectly through trading commissions and investment banking fees. The target is the C-suite: CTOs, CFOs, and investors who need a single number to decide which AI model to trust. The problem is that a single number can never capture the soul of a model.
Let me tell you a story from my own code. In 2020, I spent 300 hours auditing Uniswap V2. I did not just look for bugs; I looked for the philosophy. The fair-launch, the immutable code, the lack of a central authority—that was the covenant. Every broken token taught me how to hold value. And I learned that value is not a static number. It is a relationship between trust, performance, and context. Bank of America's tracker tries to compress that relationship into a score. But a score is a lie. A model's intelligence is not a universal truth; it is a function of the task, the data, the deployment. The same way a DeFi protocol's TVL does not tell you if it is secure, a model's benchmark score does not tell you if it is safe.
Yet, the tool will have impact. It will reduce information asymmetry between AI model providers and enterprise buyers. That is a good thing. But it will also create a new form of centralization: a single point of evaluation. In crypto, we call this the oracle problem. A centralized oracle can be manipulated, bought, or simply wrong. The same risk applies here. If Bank of America becomes the de facto rating agency for AI models, then its methodology, its data sources, and its biases will shape the entire industry. I have seen this before. In the DeFi summer of 2020, projects with high APY attracted TVL, but when the incentives stopped, the users vanished. The metric was hollow. The same will happen with AI scores. A model that scores high on MMLU may fail in a real-world financial application because of latency or security. The tracker will not capture that.
Let me dig deeper into the technical architecture. The tool likely uses a weighted combination of benchmark scores (MMLU, HumanEval, MATH, etc.) and cost data (API price per million tokens). It may also include deployment costs, but that is less certain. The innovation is not in the metrics—they are public—but in the aggregation and the branding. Bank of America's research team, with its global reach, can push this into boardrooms and investor calls. That is a powerful distribution channel. The tool could become a "Gartner Magic Quadrant" for AI, but with financial stakes. Companies that score high will see their stock rise; those that score low may struggle to raise capital. This is a double-edged sword. On one hand, it encourages competition on performance and cost. On the other hand, it creates a winner-take-all dynamic where the rating defines the reality.
Now, the contrarian angle. Perhaps this tool is actually good for decentralization. How? By commoditizing AI model evaluation, it could accelerate the adoption of open-source models. Open-source models like Llama, Mistral, and DeepSeek often offer lower costs and competitive performance. If the tracker highlights the cost-efficiency of these models, enterprises may shift away from proprietary APIs. This could reduce the power of the large AI labs. But the tracker is controlled by a single bank, which is still a centralized authority. The real decentralization would come from an on-chain, community-driven evaluation platform that uses zero-knowledge proofs to verify model performance without revealing the data. That is the future I dream of. But for now, we have a bank's tracker.
Let me talk about the risks. First, the benchmark overfitting problem. Models are trained to perform well on benchmarks, not necessarily on real-world tasks. The tracker will only measure what it measures, and what it does not measure—like safety, bias, or robustness—will be ignored. This is the same problem that plagued the credit rating agencies before 2008. They rated mortgage-backed securities based on historical data, ignoring the tail risks. The result was a collapse. The same could happen here if a model scores high but fails catastrophically in a high-stakes application. Second, the conflict of interest. Bank of America is also a banker to AI companies. If it gives a low rating to a client's model, that client may take its business elsewhere. Conversely, a high rating could be seen as a favor. The tool's credibility will depend on its independence, but independence is hard when the same institution also underwrites the companies being rated. In crypto, we have a term for this: "rehypothecation." Trust is compiled, not claimed.
Third, the data source problem. The tracker likely relies on publicly available benchmark results, which are self-reported by model providers. There is no verification. A provider could cherry-pick the best results or omit failures. In crypto, we solved this with on-chain data and verifiable computation. For AI, we need something similar: a decentralized evaluation network where models are tested on hidden tasks, and the results are posted on-chain. Projects like the Open LM benchmark and the Continuum project are moving in this direction, but they are not yet mainstream. Bank of America's tracker, by contrast, is a black box. Its methodology is not public, and its updates are slow. In a market where models improve every week, a lagging tracker is worse than none.
So what is the takeaway? The tracker is a symptom, not a solution. It reflects the growing need for standardized AI evaluation, but it also reflects the gravitational pull of centralized institutions. The same forces that brought us the 2008 crisis are now shaping the AI economy. The antidote is not to reject evaluation, but to build it differently. We need an open, transparent, and decentralized evaluation layer that is owned by the community, not by a bank. Think of it as a "data availability oracle" for AI performance—a layer that any model can submit to, and any user can verify. The technology exists: zero-knowledge proofs, decentralized storage, and token incentives. The question is whether we will build it before the banks own the truth.
In the silence of the bear, we heard the truth. The truth was that the market was a mirror of our collective fear and greed. Now, the truth about AI will be reflected in a bank's tracker. But I do not want to see my reflection in a bank's glass. I want to see it in a code I can read, a contract I can trust, and a community I can join. The tracker is the start of a conversation. The real work is just beginning. Every broken token taught me how to hold value. And value is not a number on a screen. It is the agreement between people to trust a system. Bank of America's tracker has a system. But is it trustworthy? The answer lies not in the metrics, but in the hands that build them.
Let us build those hands ourselves. Let us create decentralized oracles for AI intelligence, where the score is not a black box but a transparent proof. Let us ensure that the covenant of code remains the covenant of the people, not the committee. The bear market taught us patience. The bull market taught us greed. And now, the AI market will teach us something new: the value of decentralized truth. My code was the covenant, not just the contract. Will we write the next covenant together?

