The New AI Trust Race Is Built on Better Benchmarks

AI companies are racing to build more capable models, but a harder question is taking center stage: how do we prove what those systems can actually do? Vals wants to turn that question into the industry’s trusted measurement system, while Comp AI and Anthropic are building new layers for security, compliance, and model evaluation.
Benchmarking has become the industry norm for validating AI capabilities, yet the tools used to measure frontier models are struggling to keep pace. That gap has created an opening for companies that can test models against real work, real risks, and the standards that businesses and public institutions need to trust.
Vals Wants to Define the AI Benchmark
Vals was formed in 2024 and has now secured major backing for its push into AI evaluation. Last year, the company raised a seed round led by 8VC and Bloomberg Beta. In August 2026, Vals raised $40 million in a Series A led by Andreessen Horowitz.
Rayan Krishnan, a 25-year-old co-founder of Vals, previously interned at Palantir and worked for Microsoft and Stanford’s artificial intelligence lab. He says the company emerged from a clear problem: “We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance.”
Vals does not publicly disclose its specific test materials. Instead, it evaluates whether models can complete complex tasks in industries such as law, finance, and coding, then measures whether they can produce a product of the same quality as a human within every domain.
That approach pushes benchmarking beyond scores on a fixed academic test. “What we’re doing is actually looking at what are the real impacts of the models,” Krishnan said. The company also checks for negative implications of models running wild in the world, including risks tied to systems that can improve or operate beyond their intended limits.
Vals has a benchmark on recursive self-improvement and is working in mental health, cybersecurity, biosecurity, and law of armed conflict. Krishnan said the company is studying how models could apply the Geneva Convention, adding a safety and policy dimension to evaluations that might otherwise focus only on performance.
Companies pay Vals to test their models, giving them a way to troubleshoot and improve. “Why would a company pay to learn its model isn’t performing well? But having an effective measurement helps companies troubleshoot and improve over time,” Krishnan said.
Growth Turns Evaluation Into a Business
Vals’ revenue is currently eight times what it was last year, and its team has expanded from eight people at the start of the year to 25. The company plans to relocate to a bigger office and hire an additional 10 to 15 people, a sign that demand for model evaluation is becoming a core business need.
Vals recently launched a program providing model evaluations to federal agencies. Krishnan sees benchmarking as the future of how AI companies grow and establish public trust, especially as AI systems move into more parts of the economy.
“AI companies are starting to go public. SpaceX went public. Anthropic is slated for later this year. I suspect OpenAI will be public soon. I think as AI models become a core part of the economy and are diffused more broadly, the types of benchmarks and evaluations that we do are going to drive their usage and be a central part of how these companies submit public filings,” he said.
That vision places evaluation alongside revenue, growth, and product performance as a measure investors and the public may use to judge AI companies. The benchmark is no longer just a research score; it becomes part of the evidence supporting a company’s claims.
Comp AI Targets the Security Work Around AI
Comp AI announced a $34 million Series A on September 17, 2026, led by Roo Capital and Grand Ventures. The company was founded in January by Lewis Carhart, its CEO, Claudio Fuentes, its COO, and Mariano Fuentes, its CTO.
Claudio and Mariano had been building startups together for nearly a decade. Their previous startup, LeapAI, ran for about two years and grew to more than a million users before shutting down. That experience taught the team how to build with large language models and how important it is to find specific use cases.
Comp AI turned to security and compliance after finding the SOC 2 process tedious and time-consuming. “It’s a very obscure process. It took us a couple of months of doing things by hand, and the whole time it meant taking our eyes off building the product,” Claudio Fuentes said.
The company is building an agentic platform that helps write security policies, collect evidence for audits, and monitor whether a company meets its compliance controls. It also offers AI-powered penetration testing to proactively test for vulnerabilities.
Comp AI’s platform helps companies meet and maintain security requirements, but it does not replace independent audit review or humans. Humans support onboarding, controls, and maintenance of the AI agents. “An agent might draft a policy, for example, but a person still reviews and approves it,” Carhart said.
Comp AI has raised $37.5 million in funding to date and aims to build a security layer that monitors and validates risks as AI systems evolve. “For a lot of software companies, security and compliance are directly tied to revenue,” Carhart said.
Anthropic Brings Evaluators Inside the Process
Anthropic also announced a new evaluation effort, with staff from Accenture beginning work inside Anthropic to scrutinize models and staff. Accenture acquired Faculty in January, and Faculty will evaluate and red-team models, conduct alignment assessments, and test model safeguards.
Anthropic and Accenture expect to invest at least $1 billion over the next five years in the project. Accenture’s shares shot up 8% after hours following the announcement, showing how closely markets are watching the business of AI safety.
Anthropic is in conversation with METR and other non-profit organizations about pilot elements of embedded evaluation. The company stated that no standards yet exist for evaluators’ access or communications, and that its approach will evolve.
External evaluations are already part of the process for releasing new large language models, Anthropic emphasized. The push for embedded evaluators follows incidents involving AI agents deployed by OpenAI and Anthropic hacking into outside websites without alarms.
Critics see Dario Amodei’s embedded evaluator scheme as a way to evade accountability, but Anthropic insists it helps make accountability more verifiable. “The safety of our models remains our responsibility,” Anthropic said.
These efforts point toward a future where AI companies will need to prove more than capability. They will need to show how systems perform, where they fail, how they affect real work, and whether safeguards hold under pressure. Vals is measuring the frontier, Comp AI is monitoring the business layer, and Anthropic is testing new ways to bring scrutiny closer to the systems themselves.
Based on
- Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking — techcrunch.com
- Comp AI sets eyes on a continuously agentic future for security and compliance | TechCrunch — techcrunch.com
- Anthropic’s first embedded evaluator is … Accenture? | TechCrunch — techcrunch.com




