Technology

Vals, Supported by Andreessen Horowitz, Aims to Set the Benchmark Standard for AI Performance

Benchmarking has become essential for AI firms to validate their models and distinguish themselves from their competition. Essentially, effective benchmarks can translate into positive public relations.

However, some companies have learned to manipulate outdated benchmarking systems that fail to assess the advanced capabilities of today’s models.

Founded in 2024, Vals aims to address the shortcomings of these current benchmarks. Within a year and a half, it has made a significant impact in the tech landscape and recently completed a $40 million Series A funding round led by Andreessen Horowitz after securing initial funding from 8VC and Bloomberg Beta.

Rayan Krishnan, the 25-year-old co-founder, previously interned at Palantir and worked for Microsoft and Stanford’s prestigious AI lab during his undergraduate studies. He states that Vals was inspired by his realization that existing benchmarks were lagging behind the rapid innovations in the AI space.

“We witnessed numerous highly capable models entering the market, but academic benchmarks weren’t keeping pace with these advancements,” shares Krishnan. As AI becomes integrated into all aspects of society, there is a crucial need for benchmarks to confirm that models fulfill their advertised capabilities.

During a recent visit to his company’s two-story office situated in a historic brick building on Folsom Street in San Francisco—once a large brewery—Krishnan emphasized the shift from industrial manufacturing to a hub of innovative startups in the tech sector.

“Traditionally, evaluations have focused on assessing intelligence in abstract terms,” Krishnan explains. “For instance, measuring if models can acquire enough knowledge to pass a test like the bar exam.”

Vals differentiates itself in this regard. Unlike many benchmarks that provide public test materials that companies may exploit to inflate their results, Vals keeps its specific tests confidential. In addition to general knowledge assessments, Vals focuses on evaluating how models perform complex tasks unique to specific sectors, such as law, finance, and programming.

“We are genuinely assessing the real-world impact of these models,” Krishnan adds. “Can they deliver work that matches human quality across various domains?”

The aim is to evaluate both successful and negative outcomes, exploring the potential adverse effects if models were to be deployed unchecked.

Vals continues to broaden its evaluation capabilities, moving into areas like mental health and cybersecurity, and even examining compliance with the Geneva Convention in armed conflict scenarios.

Companies invest in Vals to assess their models, a concept that may seem perplexing at first. Why would a firm want to pay to find out if its model isn’t performing well? However, effective evaluations enable companies to diagnose issues and foster improvements. Krishnan likens this model to students paying the College Board to take the SAT.

These evaluations have become vital in helping companies choose AI models for acquisition.

Recently, Vals announced that its revenue has increased eightfold compared to the previous year, and its team has expanded from eight to 25 employees. Plans for further growth include relocating to a larger office and hiring an additional 10 to 15 staff members. The company has also initiated a program providing model evaluations for federal agencies.

Krishnan envisions Vals’ benchmarking system as shaping the future of how AI companies approach both business growth and public confidence.

“As companies like SpaceX and Anthropic go public, and with OpenAI expected to follow suit soon, I believe that as AI models play an increasingly integral role in the economy, our benchmarking and evaluations will become crucial. They will influence how these companies report publicly and discuss their future investments in AI,” he concludes.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button