
The $40M Bet on AI Evaluation: Verifiability Theater or the Next Security Layer?
CryptoVault
I’ve been staring at the Crypto Briefing report on Vals AI’s $40 million Series A, led by a16z. The headline is loud: “AI evaluation tool raises $40M to make AI reliable.” But the article is a ghost. Five sparse information points—funding, a vague product launch, and a generic thesis about trustworthy AI. No technical depth. No code. No data. Just a story wrapped in market hype.
Yet, as a researcher who has spent years excavating truth from the buried layers of smart contracts and DeFi protocols, I’ve learned that the most dangerous narratives are the ones that feel comfortable. The $40M investment in Vals AI is a signal. But what exactly is it signaling? Is this a genuine infrastructure need, or is it the beginning of a new kind of theater—where evaluation tools become the “audit reports” of the AI age, providing comfort without cryptographic proof?
Every bug is a story waiting to be decoded. The story here is about the AI industry’s transition from model training to model deployment. In 2024 and 2025, enterprises moved from proof-of-concept to production, only to realize that “demo AI” is not “production AI.” The gap is trust. How do you know the model will behave correctly when it’s processing a customer refund or diagnosing a medical image? The answer, according to the market, is evaluation tools. They act as the quality assurance layer—testing, benchmarking, and flagging errors before the model touches real data.
But here’s the rub: the evaluation tools themselves are black boxes. The analysis of Vals AI’s limited public information suggests their technology stack likely relies on calling external frontier models—GPT-4o, Claude, Gemini—as the “judge” to evaluate other models. This is the industry standard. It’s also a logical loop. You’re using one AI to judge another AI. Who evaluates the evaluator? The Latin phrase “Quis custodiet ipsos custodes” comes to mind. Who watches the watchmen?
In my 2026 work on the AI-ZK convergence framework, I mapped out a solution: zero-knowledge proofs for verifiable computation. If an evaluation tool produces a score, that score should be a cryptographic attestation—provable without revealing the underlying model or data. Today, Vals AI’s evaluation process is a closed system. The methodology is proprietary, the datasets are hidden, and the results are claims. In a bear market where survival matters more than gains, this is a risk. Enterprises that rely on these tools for compliance are building on sand.
Let me take you deeper into the technical trade-offs. The analysis of the report reveals a critical missing piece: the evaluation dimensions. Does Vals AI test for accuracy, robustness, alignment, or jailbreak resistance? The report doesn’t say. But based on my experience reverse-engineering Solidity code in 2017, I know that the devil is in the granularity. A tool that claims “reliable AI evaluation” but only tests for surface-level accuracy is like a smart contract audit that only checks for integer overflow—it misses the reentrancy attack.
From my DeFi composability cartography work in 2020, I learned that systemic risks emerge from interactions. The same applies here. An evaluation tool that only looks at individual model outputs cannot capture the cascading failures in multi-agent systems. The future of AI is agentic—multiple models calling each other, sharing state, and executing complex workflows. Vals AI’s new product, if it’s to be relevant, must support this. But the report gave no evidence.
Now, the contrarian angle. The $40 million from a16z is a strong signal, but it’s also a trap. The venture capital machine often conflates “funding” with “validation.” In the blockchain world, we saw this with ICOs—projects raising millions on whitepapers, only to deliver nothing. The same pattern is emerging in AI infrastructure. a16z is betting on the narrative of “AI governance” as a category. Vals AI is their ticket. But the security blind spots are real. The report itself notes that the company operates in the “evaluation tool” layer, not the base model layer. That means their moat is not technology—it’s data accumulation and customer lock-in. If a bigger player like OpenAI or Anthropic decides to bundle evaluation into their API, Vals AI’s independent value proposition evaporates.
Navigating the labyrinth where value flows unseen is my specialty. In this case, the value is in the trust that evaluation tools produce. But trust without verification is just faith. The report highlights that a16z’s investment is a signal of “AI safety infrastructure” demand. However, I see a more dystopian possibility: the creation of a compliance theater where companies buy evaluation reports to satisfy regulators, but the reports themselves are not cryptographically auditable. This is the same problem I identified in my 2022 modular research on Celestia—security is secondary to availability. Here, availability of a “trust score” becomes a substitute for actual security.
What does this mean for the broader crypto-AI convergence? The industry is moving toward on-chain AI agents and decentralized inference. Projects like Bittensor, Akash, and Ritual are building markets for AI compute and model exchange. But for these markets to be trustless, they need provable evaluation. A model that wins a benchmark on a centralized evaluation tool cannot be trusted in a decentralized context. The evaluator must be a smart contract—transparent, immutable, and auditable.
My takeaway? The $40 million is a bet on the need for evaluation, but the winner will be the one who makes evaluation cryptographically verifiable. Vals AI has the capital, but without addressing the verifiability gap, they are building a beautiful facade. In the next 18 months, I predict a shift: the emergence of “ZK-evaluators”—zero-knowledge proofs that certify model outputs without revealing the model or the evaluation methodology. This is the convergence of my two worlds: blockchain and AI. The labs that master this will not just raise $40 million; they will define the infrastructure of the AI age.
Until then, I’ll keep excavating truth from the code’s buried layers—and from the silence in the press releases.