Running Local LLMs on RK3588: Real Benchmarks and What to Expect

Short answer: an RK3588 board runs small local language models at a usable but modest pace — roughly 8–11 tokens/s at 1–1.5B, 5–7 tokens/s at 3B, and 3–4 tokens/s at 8B when using the Rockchip NPU via RKLLM. That is fast enough for a responsive edge assistant, and roughly 1.5–3× faster than running the same model on the board's own CPU cores. It is not fast enough to feel like a desktop GPU.

The binding constraint is memory bandwidth, not TOPS. Text generation reads the entire model from RAM for every single token, so a 6 TOPS NPU does not translate into 6 TOPS of generation speed.

Measured results (third-party, on real RK3588 boards)

Model Quant. Speed Board / conditions Source
DeepSeek-R1 1.5B RKLLM 11.5 tok/s Indiedroid Nova (RK3588S, 16 GB), RKLLM 1.2.1, NPU 60% avg Indiedroid Nova LLM benchmarks
Qwen2.5 3B Instruct RKLLM 7.0 tok/s Same board, NPU 68% avg same
Llama 3.1 8B Instruct RKLLM 3.72 tok/s Same board, NPU 79% avg, 8.5 GiB RAM same
TinyLlama 1.1B Chat W8A8 10–15 tok/s (Rockchip figure), TTFT 200–500 ms Orange Pi 5 Max, RKLLM 1.2.2, RKNPU 0.9.8 tinycomputers.io
TinyLlama 1.1B 4-bit 8–11 tok/s (NPU) vs 6–8 (CPU llama.cpp) Community reports, Orange Pi 5 Plus / ROCK 5B bigiron.cc
Qwen2.5 1.5B 4-bit 6–8 tok/s (NPU) vs 4–6 (CPU) same same
Phi-3-mini / Qwen2.5 3–4B 4-bit 3–5 tok/s (NPU) vs 2–4 (CPU) same same

Figures are from independent third-party testing, not our own lab. Speeds vary with RAM size, thermals, driver version and model conversion.

NPU vs CPU on the same board: about 2×, at identical quality

The most useful head-to-head comes from a 14-model test on a Radxa ROCK 5B+ (RK3588, 24 GB). Running Qwen2.5-3B on the NPU versus the CPU produced identical answer quality (11/14 benchmark tasks) but half the latency — 44.8s versus 87.0s end-to-end (source).

In other words: RKLLM quantisation did not degrade output quality. It just ran faster. If a model exists in both GGUF and RKLLM form, prefer the NPU version.

Two hard technical constraints

  • RK3588 uses W8A8 for LLM inference (8-bit weights, 8-bit activations) via RKLLM — not W4A16. You cannot simply pick your favourite 4-bit GGUF.
  • You cannot offload GGUF layers to the NPU. The NPU only runs dedicated .rkllm models through the RKLLM SDK. GGUF and RKLLM are two separate pipelines — so budget time for conversion.

When this makes sense — and when it doesn't

Good fit Poor fit
Offline / air-gapped assistants Large-model chat that must feel instant
Fixed-task text generation at 1–3B Frequently changing models (re-conversion cost)
Privacy-sensitive deployments Workloads needing 7B+ with long context
Low power budget (typically 5–18 W) Anything needing CUDA-only tooling

The honest caveat: this NPU is built for vision, not language

Independent reviewers converge on the same point: LLM decode is the wrong workload to point an RK3588 NPU at. The NPU's real strength is object detection — fixed input tensors, compute-dense convolutions, small weights that fit in the working set. A single RK3588 NPU can run real-time detection across a double-digit camera count at single-digit watts, which no discrete GPU matches on cost.

So if you are also running camera inference on the same board, expect contention: both workloads compete for the same memory bandwidth and NPU cores.

Practical guidance

  • Start at 1–3B. That is the sweet spot for usable token rates.
  • Choose RAM over marginal clock speed. 16 GB gives real headroom; 8 GB limits you to smaller models.
  • Prefer boards with NVMe if you swap models often — model loading is I/O bound.
  • Watch sustained thermals. Long generations load the CPU and NPU continuously; throttle and your token rate drops with it.
  • Measure before you promise. Your own model, quantisation and context length will move these numbers.

Frequently asked

Can RK3588 run Llama 3.1 8B? Yes — measured at about 3.7 tokens/s with ~8.5 GB RAM in use. Usable for batch or non-interactive work, sluggish for chat.

Is the NPU worth the conversion hassle? For text generation, expect roughly 1.3–2× over the board's own CPU cores. Worth it if your model is stable and you have converted once; less so if you change models weekly.

How much faster is it than a Raspberry Pi 5? Community testing shows roughly 1.5–2× at 1.5B and 2–3× at 3B, since the Pi 5 has no NPU and runs everything on CPU.

Do you supply pre-flashed images and SDKs? Yes — Ai-Paipai provides local engineering support, datasheets, SDKs and pre-flashed images where applicable, plus OEM / ODM options for volume projects.

Sources

Browse RK3588 boards and edge AI hardware, or read RK3588 vs NVIDIA Jetson: which to choose.

Boards that run the benchmarks above
Shipped from Shenzhen. Volume, OEM and custom-configuration pricing available on request.
Not sure which configuration fits your project? Email lixu@ai-paipai.com with your workload and budget — we reply within one business day. Browse the full catalogue at ai-paipai.store.
Back to blog