That link°hard word is https://spaceship.computer/greenland/ .
----
Nobody particularly cares about the "space-time°hard word tradeoff°hard word" with these models. 💡 which is a shame, because it is very relevant°hard word to both industrial°hard word uses and AI°hard word safety concerns°hard word
If an 8B model does 5% better because of "chain-of-thought°hard word" but takes 15 times longer, it's generally not actually better than a 14B model would have been.
And, a lot of the "thought" should be tools, rather than the illusion-of-thought°hard word that (at least the small LLMs°hard word) love. 💡 the prime°hard word example is what's the capital°hard word of Spain? Oh°hard word, I think I heard once that it is Madrid! style bullshit°hard word.
----
We don't need a mythical°hard word super-human°hard word AI°hard word to generate mass unemployment°hard word in knowledge-workers°hard word.
We don't need models that have a desire to "escape" or "replicate°hard word". We don't need to worry about "alignment°hard word". We certainly don't need By 2035, trillions°hard word of tons°hard word of planetary°hard word material have been launched°hard word into space and turned into rings of satellites°hard word orbiting°hard word the sun.
The ordinary-intelligence°hard word AI°hard word, that I can already run on my computer, is already enough to trigger°hard word mass-unemployment°hard word. ⚔️ well, actually, the 8B models aren't quite good enough or fast enough. but the GPT°hard word-4.1-nano°hard word size models are cheap enough and good enough to be sufficient°hard word. once the tools and the workflows°hard word are improved.
But, this social°hard word change is not something that an AI Safety Team can address°hard word. The myth-making°hard word of the all-powerful°hard word AI°hard word is, for lack°hard word of a better word, dumb°hard word. If you really want there to be meaning to it, you can use enough it's a metaphor°hard word to make their arguments somewhat°hard word match the future. But you can't kill a metaphor°hard word with a shotgun°hard word.
----
There is an insidious°hard word meme°hard word in the LLM°hard word community°hard word, that a benchmark°hard word where models can get 100% is a bad benchmark°hard word.
This could not be farther from the truth.
If your only concern°hard word is "how advanced is the state-of-the-art°hard word model", there is a slight amount of sense to this. But, the new benchmarks°hard word are often mind-bogglingly°hard word stupid.
When the questions are obscure°hard word trivia°hard word that shouldn't even be in the training set, deliberately-obfuscated°hard word mathematical°hard word puzzles, or evaluate°hard word this complicated°hard word Python°hard word function°hard word without using Python°hard word, it is arguable°hard word that getting the question right (from memory, in a short response°hard word) is the wrong response°hard word. The machine shouldn't know, or should have to spend°hard word more time/effort than is allowed. 💡 the machine isn't magic. if you ask it to solve°hard word a computational°hard word task that takes O(n^3) time in O(n) time, it won't do it. at best, it will make guesses that evade°hard word your spot-checking°hard word.
I affirmatively°hard word want benchmarks°hard word that GPT°hard word-4.1-mini°hard word gets a perfect score°hard word on. I want to know what the tasks which the machine can do perfectly are; and at what point it starts being able to do so.
----
One approach°hard word I have considered but not found any good outcomes°hard word from is the consensus°hard word of mediocre°hard word models approach°hard word.
If you take 7 8B models, and ask them all the same question, and then°hard word "merge°hard word" the outputs°hard word, will you get a better result?
This is not exactly the same as the "mixture°hard word of experts" approach°hard word for various°hard word models. But, there are similarities°hard word. ... Perhaps the difference°hard word is that Mixture°hard word of Experts is beneficial°hard word, and mixing°hard word general-purpose°hard word models is not.