phi4 preliminary thoughts

Hide hard words

Initial°hard word indications°hard word are that phi°hard word4:14b is substantially slower than gemma°hard word2:9b on a 24GB°hard word MacBook°hard word Pro°hard word, with no noticeable°hard word performance°hard word difference°hard word. 💡 part of the slowness°hard word is related°hard word to the "two phase°hard word response°hard word" code°hard word I added to work around Ollama's°hard word inability°hard word to do structured°hard word JSON°hard word responses°hard word correctly. Even adjusting for that, it is twice as slow as other models

It also still has some of the "excessively°hard word long responses°hard word" problems that made phi°hard word3.5 nearly-unusable°hard word on benchmark°hard word tasks, and completely unreliable°hard word for production°hard word tasks. At least these seem to be sensible°hard word responses°hard word rather than end-of-message°hard word token°hard word bugs°hard word.

Some of the issues°hard word may be related°hard word to memory pressure, but without Safari°hard word running there should be plenty°hard word of RAM°hard word.

I see no reason to use this model over gemma°hard word2:9b or qwen°hard word2.5:7b locally°hard word.