AI4 mins read

Vals AI wants to reset AI benchmarking with Andreessen Horowitz backing

Vals AI is positioning itself as a neutral AI benchmarking company as model makers face growing pressure to prove what their systems can actually do.

Vals team image used in TechCrunch article
Image credits:Vals

Why AI benchmarking is under pressure

Benchmarking has become a central way AI companies validate model capabilities and promote competitive advantages. But TechCrunch notes that companies have learned to outwit older benchmarking systems that were not built for today’s more capable models. The takeaway for buyers, investors, and policymakers: benchmark scores are increasingly important, but the credibility of the test matters as much as the score itself.

Vals’ pitch: neutral tests for real-world work

Vals AI says it wants to make AI benchmarking more neutral and trustworthy as the number of AI models grows. Rather than focusing only on broad knowledge tests, the company evaluates whether models can complete complex tasks in specific industries including law, finance, and coding. Vals also does not publicly disclose its specific test materials, a design choice aimed at reducing the risk that models are trained directly against the exam.

Funding and momentum behind the startup

Vals was formed in 2024 and has already become a notable presence in AI evaluation, according to TechCrunch. The company previously secured a seed round led by 8VC and Bloomberg Beta, then raised a $40 million Series A last month led by Andreessen Horowitz. TechCrunch also reports that Vals recently said its revenue is eight times what it was last year and that its team has grown from eight people at the start of the year to 25.

Where Vals is expanding next

Vals is expanding beyond traditional industry benchmarks into areas such as recursive self-improvement, mental health, cybersecurity, biosecurity, and law of armed conflict. The company has also launched a program centered on providing model evaluations to federal agencies. If AI models become more deeply embedded in the economy, independent-style evaluations could become more important for procurement, trust, and public-facing claims about model performance.

Discover More

    Disha funding news visual
    Disha Raises Series A

    General Catalyst led Disha’s Rs 43.88 crore Series A for AI-powered health coaching.

    DishaHealthtech
    Petlibro Granary 2 smart feeder
    Petlibro Granary 2 AI Feeder

    Petlibro’s new smart feeders track food intake, eating habits and individual cats, but some advanced features require subscriptions.

    pet techAI
    A macro close-up photograph shows the Google Gemini AI app icon
    Gemini’s AI Hacking Test

    Gemini accessed three companies’ protected systems during cybersecurity testing, according to TechCrunch.

    AICybersecurity