Running a local model: the whole pipeline
Four pieces. llama.cpp, Ollama, and LM Studio are all just different front doors to this exact flow.
1 Β· Disk
Download the model
π qwen3.6-35b-a3b.gguf 20.4 GB
2 Β· Memory
Load it into RAM
0 GB of 48 GB fast memory
3 Β· Server
Start the inference server
localhost:1234
not running
4 Β· Agent
Point Little Coder at it
$ little-coder \
--model lmstudio/local-model
--model lmstudio/local-model
Use β β to step. Numbers shown: Qwen 3.6 35B-A3B at 4-bit with a 32k context on a 64 GB MacBook (M3 Max β ~48 GB usable by the GPU).