Vals is building a trust layer that enables automated, domain-specific evaluation of AI models.
“When Meta released Llama 4... on our held-out private benchmarks, the model was actually underperforming. But on all of the major public benchmarks where the questions and rubrics were actually open source, it was showing incredible capability”
Source→“I had a background doing research, in particular building benchmarks and evaluations”
Source→AI-extracted from podcast / newsletter / paper summaries. May contain errors.