In my testing, I am starting to make a distinction°hard word between two types of "tests" for LLMs°hard word.
----
A proficiency test covers simple tasks. Some examples:
- Repeat°hard word the misspelled°hard word word in this sentence.
- Translate this word from English°hard word to French°hard word. 💡 the linguistic°hard word knowledge of models is an unresolved°hard word question. Should it know 5 languages, or 40, or 400? In the specific°hard word case of English°hard word/French°hard word: it is plausible°hard word to claim°hard word that one cannot truly°hard word know the English°hard word language without knowing French°hard word. The LLM°hard word should also know French°hard word.
- Choose the definition°hard word of this word.
The accuracy°hard word in performing°hard word these tasks is, often, surprisingly bad, compared to performance°hard word on other tasks. This may be due to a lack°hard word of training for these tasks.
----
On the other hand, a qual test ⚙️ possibly for "qualification°hard word" are more difficult.
- Write two paragraphs°hard word about the city of Toulouse°hard word.
- Explain the theory of relativity°hard word to a nine-year-old°hard word.
- Answer these questions from the GRE°hard word Verbal°hard word Reasoning section°hard word.
From a technical°hard word perspective°hard word: many of these are free-form°hard word responses°hard word that are scored°hard word by a larger LLM°hard word.
The interesting question is not whether any LLM°hard word can answer these, but whether an LLM°hard word under 16GB°hard word in size can do so.
----
the Frontier°hard word tests are not particularly interesting.