another generation

Hide hard words

Gemma°hard word3 is out: https://blog.google/technology/developers/gemma-3/

Also out, since my last evaluations°hard word: Claude°hard word 3.7, ChatGPT°hard word 4.5, QwQ°hard word-32B.

----

There are a few "smoke tests" I want to run. But, beyond that, I'm not certain I will have the time or interest to do any deep evaluations°hard word.

I already know that "8B" models can do some tasks at a reasonable°hard word speed, and can't do other tasks. It is very unlikely°hard word that the new models will move the needle.

As far as the new very-large°hard word models are concerned°hard word: my initial°hard word impressions°hard word have not shown them to be a substantial improvement°hard word. There is more "DeepSeek°hard word" style internal°hard word narrative°hard word, but the results are often worse as a result. 💡 was it a bad test? do I need to change the prompts°hard word? or are they privileging°hard word "results that make stupid people think the machine is smart°hard word" over accurate°hard word results?