Vals Raises $40 Million to Build AI Benchmarking Platform for Enterprise Models
AI benchmarking startup Vals raises $40 million led by Andreessen Horowitz to develop independent evaluations for testing AI models across industries.
AI benchmarking startup Vals has raised $40 million in a Series A round led by Andreessen Horowitz as it develops new evaluation methods to measure how artificial intelligence models perform in real-world situations.
The company, founded in 2024, is focused on improving AI testing systems that have struggled to keep pace with rapidly advancing models. Vals previously raised a seed round led by 8VC and Bloomberg Beta before securing its latest funding, according to reports on the Series A investment.
Building a New Approach to AI Benchmarking
AI benchmarks have become a common way for companies to compare model performance, but researchers have raised concerns that some existing tests are outdated or can be optimised against. Recent analysis of AI benchmarks has highlighted challenges in measuring modern AI capabilities.
Vals was created to address these limitations by evaluating AI systems on practical tasks instead of only measuring general knowledge. The startup keeps its specific tests private to reduce the possibility of models being trained specifically for benchmark results, a concern also discussed in reporting on AI model evaluations.
The company tests whether AI models can complete complex work in areas such as law, finance and coding, while also examining potential failures and risks. Founder Rayan Krishnan said the goal is to measure the real-world impact of AI systems rather than only their performance on academic-style tests.
Vals Expands Enterprise AI Evaluations
Rayan Krishnan founded Vals after interning at Palantir and working with Microsoft and Stanford’s artificial intelligence lab. The company grew from Krishnan’s view that AI evaluation methods were not advancing as quickly as the models they were designed to measure.
The startup has expanded its benchmarks into areas including cybersecurity, mental health, biosecurity and specialised applications such as evaluating AI systems against legal and operational standards.
Companies pay Vals to evaluate their AI models and identify areas for improvement. The company compares its role to standardised testing organisations, where independent measurement helps organisations understand capabilities before making decisions.
Vals recently said its revenue grew significantly compared with the previous year and that it expanded its team from eight employees to 25. The company also launched a program to provide AI model evaluations for federal agencies through its evaluation program announcement.
The startup has shared growth updates through its official X account as it continues developing its benchmarking platform.
Vals operates from a San Francisco office located in a historic building that previously served as a brewery, according to local landmark information.
As AI systems become more widely used across industries, Vals aims to provide independent testing methods that help companies evaluate reliability, performance and practical usefulness before deployment.
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Angry
0
Sad
0
Wow
0