greenland, a post-mortem, part 3

Hide hard words

JSON°hard word output°hard word is almost a necessity°hard word for an LLM°hard word to be usable°hard word today. All of the major LLM°hard word platforms°hard word have it in some form. But, if you are using a model from 2023, it might not support it, or it might not work very well.

While many of the improvements°hard word from 2 years ago are in the tools running the LLM°hard word 💡 such as the token-selection°hard word algorithm°hard word, there is some amount of understanding of the output-format°hard word that needs to be trained into the model.

----

Trying to test Phi°hard word-2 (December 2023, 2.7B params°hard word) or Mistral°hard word-0.3 (September 2023, 7B params°hard word) seems unlikely°hard word to be worth°hard word any time/effort. I know there are newer models that are better; and I'm not sure there will be usable°hard word results at all.

Does that mean the models we have today will be useless°hard word in 18 months? Probably not. Maybe there will be a GPT°hard word-4.1-nano°hard word quality model that is 2c IN / 5c OUT per°hard word million°hard word tokens°hard word 💡 currently GPT°hard word-4.1-nano°hard word is 10c IN / 40c OUT per°hard word million°hard word tokens°hard word. For almost all personal°hard word uses, this is not a substantial improvement°hard word.

----

Whether Falcon°hard word 3 ⚙️ https://huggingface.co/blog/falcon3 is worth°hard word considering is a different question.

Their press-release°hard word has benchmarks°hard word showing them as slightly better than earlier systems of similar size. But nothing ground-breaking°hard word; and in fact we know there can't be anything°hard word too unique°hard word. If there were, it would have already been copied.

It is "just another model". 💬 if you want to build a forest, it helps to have many different trees

----

What about Granite (the IBM°hard word offering)? ⚙️ https://www.ibm.com/granite/docs/

This one I happened to already test. The results were very unremarkable°hard word. Like most 8B models, this 8B model gave acceptable°hard word results for tasks that did not require°hard word deep insight°hard word or precision°hard word.

----

The highest-profile°hard word "local°hard word models" are Gemma°hard word ⚙️ Google's°hard word latest model, Llama°hard word ⚙️ Facebook's°hard word latest model, QWEN°hard word ⚙️ Alibaba's°hard word latest model, and Phi°hard word ⚙️ Microsoft's°hard word latest model. 🔥 Amazon°hard word and Apple do not seem to be releasing°hard word their own models. Netflix°hard word is not, either. 💡 there are others; Mistral°hard word is probably the leading European°hard word provider°hard word. 🔥 I still don't care about Deepseek°hard word; the "thought" is largely a party-trick°hard word that people will see through soon enough ... also most other models also do that in some way now.

And, all of these seem to be hitting limits at the 8B param°hard word size. The latest releases°hard word are more interesting at the 24-40B param°hard word size. Which can be run on a local°hard word machine ... just not the ones I own.

----

The 1.5B parameter°hard word models are useful°hard word for speculative°hard word decoding°hard word ⚙️ https://research.google/blog/looking-back-at-speculative-decoding/ , which is where you use one model to make a cheap "guess" for the larger model, allowing more tokens°hard word to be calculated°hard word at once.

Beyond that, they are largely toys. With fine-tuning°hard word and testing, you can probably use a model for a single useful°hard word task. But the 1.5B models are not general-purpose°hard word AI°hard word, and they probably never will be.

----

For "cloud" models, there is Gemini°hard word ⚙️ Google°hard word, GPT°hard word ⚙️ OpenAI°hard word, and Claude°hard word ⚙️ Anthropic°hard word. And, several others that I haven't bothered°hard word with. 💡 Perplexity°hard word has an API°hard word called Sonar°hard word. Amazon°hard word has something called Nova°hard word. And there is still TSFKAT's°hard word offering. ⚙️ TSFKAT°hard word = "the site°hard word formerly°hard word known as Twitter°hard word"

And ... without a specific°hard word work-task°hard word, it is unlikely°hard word that benchmarking°hard word / testing these models will come up with any useful°hard word data.