CPU inference for Deepseek V4 flash?

Has anyone tested performance of deepseek v4 flash on AmpereOne, on CPU, on llama.cpp?

I am wondering how some of the larger MoE models perform.

I use Qwen 3.6 35b a3b on Altra and it is very decent, but I don’t have enough ram for Deepseek flash.

Guessing something like GLM would be a non starter? Has anyone tried it for fun?

1 Like

@binh, @quocbao, @joneill @TheComputerGuy would any of you have an answer for this?

1 Like

Hmm, I’m not entirely sure. I use llama.cpp but mostly load things onto my RTX 5070. I have been using Qwen3.6 35B and it works pretty well with CPU MoE. I get about 5 - 6 t/s, not the fastest but good enough for background tasks.

I just noticed that Unsloth published a GGUF version of the model three days ago. I’m trying it out now.

2 Likes

Definitely interested in this, and your llama.cpp command line params

1 Like

+1 (apparently replies must be at least 20 characters)