{"id":33119,"date":"2026-09-19T16:35:13","date_gmt":"2026-09-19T14:35:13","guid":{"rendered":"https:\/\/www.mixtv1.com\/index.php\/2026\/09\/19\/vals-backed-by-andreessen-horowitz-is-looking-to-become-the-gold-standard-for-ai-benchmarking\/"},"modified":"2026-09-19T16:36:44","modified_gmt":"2026-09-19T14:36:44","slug":"can-andreessen-horowitz-backed-vals-set-the-gold-standard-for-ai-benchmarking","status":"publish","type":"post","link":"https:\/\/www.mixtv1.com\/index.php\/2026\/09\/19\/can-andreessen-horowitz-backed-vals-set-the-gold-standard-for-ai-benchmarking\/","title":{"rendered":"Can Andreessen Horowitz-Backed Vals Set the Gold Standard for AI Benchmarking?"},"content":{"rendered":"<h2>The Crisis of Credibility in AI Model Evaluation<\/h2>\n<p>In the current artificial intelligence landscape, benchmarking has evolved into the primary mechanism for validating model performance. For developers, these metrics serve a dual purpose: they act as a technical yardstick and, perhaps more importantly, as a powerful marketing tool. When a company\u2019s model tops the charts, it translates directly into favorable press and a perceived competitive edge. Essentially, high benchmark scores have become the gold standard for AI public relations.<\/p>\n<p>However, this reliance on traditional evaluation frameworks is increasingly problematic. Many of the legacy systems currently in use were designed for a different era of computing and are fundamentally ill-equipped to assess the nuanced, multi-modal capabilities of modern generative AI. This mismatch has created a &#8220;gaming the system&#8221; culture, where developers optimize models specifically to excel on outdated tests rather than improving real-world utility.<\/p>\n<h2>Vals: A New Paradigm for Model Validation<\/h2>\n<p>Recognizing the widening gap between AI progress and evaluation accuracy, a startup named Vals emerged in 2024 with a clear objective: to overhaul the broken benchmarking ecosystem. The company has rapidly ascended within the tech sector, moving from a nascent project to a major industry player in under 24 months.<\/p>\n<p>The startup\u2019s momentum is evidenced by its aggressive fundraising trajectory. Following an initial seed round supported by 8VC and Bloomberg Beta, the company recently closed a $40 million Series A funding round led by Andreessen Horowitz. This significant capital injection underscores the industry&#8217;s urgent need for more reliable, transparent, and rigorous evaluation standards.<\/p>\n<h2>The Vision Behind the Benchmarking Shift<\/h2>\n<p>The impetus for Vals stems from the firsthand experiences of its co-founder, Rayan Krishnan. At just 25 years old, Krishnan brings a wealth of technical pedigree to the venture, including internships at Palantir and significant research contributions at Microsoft and Stanford University\u2019s renowned AI laboratory. During his time in these high-level environments, Krishnan observed a recurring frustration: the tools used to measure AI intelligence were failing to keep pace with the rapid evolution of the models themselves.<\/p>\n<p>As the industry moves toward more complex agentic workflows, the limitations of static benchmarks become even more apparent. Current data suggests that as models become more capable, the &#8220;saturation effect&#8221;-where models score near-perfectly on legacy tests-makes it nearly impossible for enterprises to distinguish between top-tier providers. By developing dynamic, context-aware evaluation tools, Vals aims to provide the clarity that stakeholders, from investors to enterprise CTOs, desperately need to make informed decisions.<\/p>\n<p><a class=\"echo_read_more\" href=\"https:\/\/techcrunch.com\/2026\/09\/19\/vals-backed-by-andreessen-horowitz-is-looking-to-become-the-gold-standard-for-ai-benchmarking\/\" target=\"_blank\"> \u00bb More Info >>><\/a><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Benchmarking has become the gold standard for AI companies looking to prove their mettle-and, more importantly, to craft a winning narrative. When the metrics tilt in their favor, these numbers transform into powerful marketing ammunition, signaling dominance in a crowded field. In short, a high score is the ultimate PR flex. The problem? Companies have cracked the code, learning exactly how to game these legacy systems to their advantage<\/p>\n","protected":false},"author":55,"featured_media":33120,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"wpai_generated_summary":"","wpai_meta_description":"","footnotes":""},"categories":[1399],"tags":[348,4483,36,353],"class_list":["post-33119","post","type-post","status-publish","format-standard","has-post-thumbnail","category-techplus","tag-ai","tag-andreessen-horowitz","tag-mixtv","tag-startups"],"_links":{"self":[{"href":"https:\/\/www.mixtv1.com\/index.php\/wp-json\/wp\/v2\/posts\/33119","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.mixtv1.com\/index.php\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.mixtv1.com\/index.php\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.mixtv1.com\/index.php\/wp-json\/wp\/v2\/users\/55"}],"replies":[{"embeddable":true,"href":"https:\/\/www.mixtv1.com\/index.php\/wp-json\/wp\/v2\/comments?post=33119"}],"version-history":[{"count":0,"href":"https:\/\/www.mixtv1.com\/index.php\/wp-json\/wp\/v2\/posts\/33119\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.mixtv1.com\/index.php\/wp-json\/wp\/v2\/media\/33120"}],"wp:attachment":[{"href":"https:\/\/www.mixtv1.com\/index.php\/wp-json\/wp\/v2\/media?parent=33119"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.mixtv1.com\/index.php\/wp-json\/wp\/v2\/categories?post=33119"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.mixtv1.com\/index.php\/wp-json\/wp\/v2\/tags?post=33119"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}