Running Local LLMs on RK3588: Real Benchmarks and What to Expect
Share
Short answer: an RK3588 board runs small local language models at a usable but modest pace — roughly 8–11 tokens/s at 1–1.5B, 5–7 tokens/s at 3B, and 3–4 tokens/s at 8B when using the Rockchip NPU via RKLLM. That is fast enough for a responsive edge assistant, and roughly 1.5–3× faster than running the same model on the board's own CPU cores. It is not fast enough to feel like a desktop GPU.
The binding constraint is memory bandwidth, not TOPS. Text generation reads the entire model from RAM for every single token, so a 6 TOPS NPU does not translate into 6 TOPS of generation speed.
Measured results (third-party, on real RK3588 boards)
| Model | Quant. | Speed | Board / conditions | Source |
| DeepSeek-R1 1.5B | RKLLM | 11.5 tok/s | Indiedroid Nova (RK3588S, 16 GB), RKLLM 1.2.1, NPU 60% avg | Indiedroid Nova LLM benchmarks |
| Qwen2.5 3B Instruct | RKLLM | 7.0 tok/s | Same board, NPU 68% avg | same |
| Llama 3.1 8B Instruct | RKLLM | 3.72 tok/s | Same board, NPU 79% avg, 8.5 GiB RAM | same |
| TinyLlama 1.1B Chat | W8A8 | 10–15 tok/s (Rockchip figure), TTFT 200–500 ms | Orange Pi 5 Max, RKLLM 1.2.2, RKNPU 0.9.8 | tinycomputers.io |
| TinyLlama 1.1B | 4-bit | 8–11 tok/s (NPU) vs 6–8 (CPU llama.cpp) | Community reports, Orange Pi 5 Plus / ROCK 5B | bigiron.cc |
| Qwen2.5 1.5B | 4-bit | 6–8 tok/s (NPU) vs 4–6 (CPU) | same | same |
| Phi-3-mini / Qwen2.5 3–4B | 4-bit | 3–5 tok/s (NPU) vs 2–4 (CPU) | same | same |
Figures are from independent third-party testing, not our own lab. Speeds vary with RAM size, thermals, driver version and model conversion.
NPU vs CPU on the same board: about 2×, at identical quality
The most useful head-to-head comes from a 14-model test on a Radxa ROCK 5B+ (RK3588, 24 GB). Running Qwen2.5-3B on the NPU versus the CPU produced identical answer quality (11/14 benchmark tasks) but half the latency — 44.8s versus 87.0s end-to-end (source).
In other words: RKLLM quantisation did not degrade output quality. It just ran faster. If a model exists in both GGUF and RKLLM form, prefer the NPU version.
Two hard technical constraints
- RK3588 uses W8A8 for LLM inference (8-bit weights, 8-bit activations) via RKLLM — not W4A16. You cannot simply pick your favourite 4-bit GGUF.
-
You cannot offload GGUF layers to the NPU. The NPU only runs dedicated
.rkllmmodels through the RKLLM SDK. GGUF and RKLLM are two separate pipelines — so budget time for conversion.
When this makes sense — and when it doesn't
| Good fit | Poor fit |
| Offline / air-gapped assistants | Large-model chat that must feel instant |
| Fixed-task text generation at 1–3B | Frequently changing models (re-conversion cost) |
| Privacy-sensitive deployments | Workloads needing 7B+ with long context |
| Low power budget (typically 5–18 W) | Anything needing CUDA-only tooling |
The honest caveat: this NPU is built for vision, not language
Independent reviewers converge on the same point: LLM decode is the wrong workload to point an RK3588 NPU at. The NPU's real strength is object detection — fixed input tensors, compute-dense convolutions, small weights that fit in the working set. A single RK3588 NPU can run real-time detection across a double-digit camera count at single-digit watts, which no discrete GPU matches on cost.
So if you are also running camera inference on the same board, expect contention: both workloads compete for the same memory bandwidth and NPU cores.
Practical guidance
- Start at 1–3B. That is the sweet spot for usable token rates.
- Choose RAM over marginal clock speed. 16 GB gives real headroom; 8 GB limits you to smaller models.
- Prefer boards with NVMe if you swap models often — model loading is I/O bound.
- Watch sustained thermals. Long generations load the CPU and NPU continuously; throttle and your token rate drops with it.
- Measure before you promise. Your own model, quantisation and context length will move these numbers.
Frequently asked
Can RK3588 run Llama 3.1 8B? Yes — measured at about 3.7 tokens/s with ~8.5 GB RAM in use. Usable for batch or non-interactive work, sluggish for chat.
Is the NPU worth the conversion hassle? For text generation, expect roughly 1.3–2× over the board's own CPU cores. Worth it if your model is stable and you have converted once; less so if you change models weekly.
How much faster is it than a Raspberry Pi 5? Community testing shows roughly 1.5–2× at 1.5B and 2–3× at 3B, since the Pi 5 has no NPU and runs everything on CPU.
Do you supply pre-flashed images and SDKs? Yes — Ai-Paipai provides local engineering support, datasheets, SDKs and pre-flashed images where applicable, plus OEM / ODM options for volume projects.
Sources
- Indiedroid Nova LLM benchmarks — RK3588S, 16 GB, RKLLM 1.2.1
- Rockchip RK3588 NPU deep dive — Orange Pi 5 Max, RKLLM 1.2.2
- 14 models benchmarked on RK3588 (CPU vs NPU) — ROCK 5B+, 24 GB
- RK3588 NPU LLM inference: the honest check — community-reported figures
Browse RK3588 boards and edge AI hardware, or read RK3588 vs NVIDIA Jetson: which to choose.



