Why Does a Model Look Great on Vectara but Bad on AA-Omniscience?
https://victor-wiki.win/index.php/Fusion_mode_for_quick_multi-perspective_consensus
In the past decade of evaluating Large Language Models (LLMs), I have seen the same cycle repeat itself