In AI benchmarking, a test set is a collection of data used to evaluate the performance of a model after training. Training on the test set can misleadingly inflate a model's benchmark scores, making the model appear more powerful than it actually is.
Over the weekend, an unconfirmed rumor began circulating on X and Reddit that Meta was artificially boosting the benchmark results of its new model. The rumor appears to have originated from a post on a Chinese social media site written by a user who claimed to have resigned from Meta in protest of the company's benchmarking practices.
Reports that Maverick and Scout performed poorly on certain tasks fueled the rumors, as did Meta's decision to use an unreleased experimental version of Maverick to achieve better scores on the benchmark LMArena. Researchers on X observed clear differences between the behavior of publicly downloadable Maverick and models hosted on LMArena.
Al-Dahle acknowledged that some users have found the quality of Maverick and Scout to be "varied" between different cloud providers of hosting models.
"Because we remove models as soon as they are ready, we expect all public implementations to take several days to complete," Al-Dahle said. "We will continue to work hard to fix bugs and engage partners."
Related articles: