Vals raises $40 million to expand AI model evaluations

Vals secures Series A funding for AI evaluation
AI benchmarking startup Vals has raised a $40 million Series A led by Andreessen Horowitz. The company, founded in 2024, said the financing follows a seed round led by 8VC and Bloomberg Beta and comes during a period of rapid growth.
Vals evaluates AI models on work associated with particular industries rather than treating intelligence as a purely abstract question. Its tests cover tasks in fields including law, finance and coding, with the stated aim of determining whether a model can produce work of the same quality as a human in a given domain.
Co-founder Rayan Krishnan said the company was created in response to the speed at which capable new models have reached the market, while academic benchmarks have struggled to keep pace. He argues that evaluations should verify the capabilities vendors advertise as AI is integrated more widely into business and society.
Private test materials and broader risk checks
A central distinction in Vals' approach is that it does not publicly disclose its specific test materials. Publicly available benchmarks can enable companies to train models against the questions, potentially improving benchmark results without demonstrating the same performance on unseen work.
The startup also says it looks beyond positive outputs. Krishnan described evaluating the possible negative implications if models operated freely in real-world settings. Vals has added benchmarks related to recursive self-improvement and is working in mental health, cybersecurity, biosecurity and the law of armed conflict, including how models apply the Geneva Convention.
Companies pay Vals to test their models, a process Krishnan compared to a student paying the College Board to take the SAT. The assessment is intended to help model providers identify shortcomings, troubleshoot systems and improve them over time. The resulting evaluations are also becoming a factor for organizations considering which AI models to acquire.
Growth plans include federal agency evaluations
Vals said its revenue is currently eight times its level a year earlier. Its headcount has risen from eight at the start of the year to 25, and the company plans to move to a substantially larger office while adding another 10 to 15 employees.
The company has also launched a programme to provide model evaluations to federal agencies. Krishnan expects benchmarking and evaluation to become more important as AI models become a larger part of the economy and as AI companies use evidence of capabilities in public filings and investment discussions.
For businesses adopting AI, the practical implication is to evaluate a model against the specific tasks it will perform, while considering both its useful outputs and the risks identified by domain-relevant testing.

