Scott Wu, CEO of Cognition, told a reporter that every public AI benchmark is saturated. He said the industry is shifting toward proprietary evaluation methods that emphasize real-world applicability. This statement, buried in a Crypto Briefing interview, is not a neutral observation. It is a strategic move in a larger war over narrative control and valuation within the AI-crypto convergence space.
For the past 18 months, I have audited five AI-crypto projects claiming decentralized compute. Four of them ran their model evaluations on centralized AWS clusters. They published MMLU and HumanEval scores to pretend decentralization. The fifth—a small team in Singapore—couldn't even reproduce their claimed accuracy. When I asked for the evaluation pipeline, they sent a spreadsheet with handwritten labels. This is the reality behind the benchmark hype.
Wu is correct that public benchmarks have lost discriminating power. GPT-4, Claude 3, and Gemini Ultra all score above 90% on MMLU. HumanEval pass rates for code generation top 85%. The variance is within noise for most buyers. Yet the industry continues to fundraise on these numbers. Cognition's Devin—an autonomous software engineer agent—cannot be fairly evaluated by a static multiple-choice test. But the problem runs deeper than technical adequacy.
The shift to proprietary evaluation creates an information asymmetry that benefits well-funded incumbents. It allows companies like Cognition to define their own success metrics. It lets them hide failure rates behind curated test suites. In crypto terms, this is equivalent to a token team controlling the oracle for their own price feed. No independent verifier can challenge the claim.
During my 2022 DeFi collapse audit, I saw the same pattern with TVL. Projects touted total value locked as a proxy for usage. When I traced the addresses, 40% of the TVL came from the team's own wallet chain. The narrative controlled the metric. The metric controlled the narrative. Now AI companies are learning the same playbook: change the benchmark, change the story.
Wu's interview contains a second layer that the reporter missed. He said the industry is 'shifting toward proprietary evaluation.' He did not say 'shifting toward transparent evaluation.' The absence of that word is the real headline. In blockchain, we have learned that without transparency, any claim becomes a sales pitch. The recent collapse of a large AI data labeling token project—where the founding team admitted to fabricating 70% of their throughput metrics—proves this. The benchmark was internal. The audit was internal. The explosion was public.
Your alpha is someone else. The irony is that Wu himself is likely running internal benchmarks that show Devin outperforming Claude and GPT-4 on software engineering tasks. But he cannot release those benchmarks for competitive reasons. Each company is building a secret evaluation lab. The winner will be the one who convinces the market that their secret lab is the only valid one.
The contrarian view: The bulls might say this is a necessary evolution. Custom evaluation is how every industry works. A pharmaceutical company runs its own clinical trials. A hedge fund builds its own backtest engine. Why should AI be different? This argument has surface logic. But it ignores the fact that AI models are being sold as general-purpose products. A drug trial measures one effect in one population. An AI model is expected to perform across millions of tasks. Proprietary evaluation for a model the size of GPT-4 is like a drug company doing a clinical trial with only 10 patients and claiming statistical significance. The math does not hold.
In the crypto context, this evaluation shift creates both risk and opportunity. The risk: investors will back projects based on cherry-picked scores from black-box test suites. The opportunity: neutral third-party evaluation services will become critical infrastructure, analogous to blockchain explorers. If a project cannot pass an auditable, public benchmark in its claimed domain, treat its valuation with extreme skepticism. I have seen three AI-crypto deals in 2025 that cited proprietary benchmarks to justify a 10x premium. I declined all three. Two have since revised their metrics downward.
The takeaway is not that benchmarks are dead. It is that the death of public benchmarks is being weaponized. When a new AI-crypto project tells you they have a proprietary evaluation that proves their superiority, ask for the source code. Ask for the test set. Ask for the reproducibility of one sample result. If they say it's confidential, walk. The market is entering a phase where the cost of acquiring reliable information is rising. That is exactly when due diligence pays its highest premiums.
I have spent 13 years watching narratives distort on-chain reality. The benchmark saturation announcement is the clearest signal yet that we are at the inflection point for AI-crypto evaluation. The next 12 months will separate projects that build verifiable evaluation infrastructure from those that just talk about it. The alpha belongs to the ones who can decode the hidden assumptions inside a test score.
Your alpha is someone else. Choose your benchmark carefully.


