跳到正文
r/LocalLLaMA· /u/No-Doughnut6532·· 3 小时前AI 评分63

Orange Pi 5 Plus(RK3588,16GB)本地 LLM 实测:Ollama 吞吐、NPU 卸载、核心绑定与散热

[Benchmark] Running Local LLMs on Orange Pi 5 Plus (RK3588, 16GB): Ollama Tok/s, NPU Offloading, Core Pinning & Thermals

AI 导读

作者在 Orange Pi 5 Plus(RK3588,16GB LPDDR4x)上实测本地 LLM 推理,发现把 Ollama 限制在 4 个 A76 大核相比默认 8 线程最高提速 318%,Qwen 2.5 1.5B 从 3.46 tok/s 升至 14.48 tok/s。

正文

Disclosure: This unit was provided free of charge by Orange Pi for testing. No editorial review, no preconditions, no script. All data, bottlenecks, and thermal behavior are reported directly from hardware testing.


TL;DR

  • Core Pinning is critical on RK3588: Setting Ollama to 4 threads (A76 Big cores only) gives up to a +318% speedup over the default 8 threads, which stall waiting for the slower A55 Little cores.
  • Inference speeds (4T CPU): DeepSeek-Coder 1.3B hits 16.9 tok/s, Qwen 2.5 1.5B hits 14.5 tok/s, Llama 3.2 1B hits 14.6 tok/s, Phi-3 Mini 3.8B hits 6.6 tok/s, Llama 3.2 3B hits 7.3 tok/s.
  • The 8B memory wall: Llama 3.1 8B drops to 2.3 tok/s and pushes temperatures to 85°C. LPDDR4x bandwidth (~25-30 GB/s measured) is the hard physical ceiling.
  • NPU vs CPU: Ollama runs 100% on CPU. Using the native RKLLM runtime on the 6 TOPS NPU yields 21.55 tok/s on Qwen 1.5 0.5B with sub-100ms TTFT while keeping CPU load at ~0%.
  • Thermals: The board is sold bare-die without a cooler in standard retail packaging. Idle is 52.7°C, 1B-3B inference sits at 68-74°C, but 8B or sustained workloads hit the 85°C throttle ceiling without an active heatsink.

Hey r/LocalLLaMA,

I have been benchmarking an Orange Pi 5 Plus (RK3588, 16GB LPDDR4x, Samsung PM981a 256GB NVMe SSD with DRAM cache) running Ubuntu 22.04 LTS (Kernel 6.1.99-rockchip-rk3588).

The goal was to test whether an 8-core ARM SBC can realistically handle small 1B-3B models for 24/7 background agents or home automation without cooking itself or locking up the host system.

Here is the breakdown of CPU vs NPU performance, the big.LITTLE scheduling trap, and thermal limits.


1. Memory and Storage Architecture

When running local models on an SBC, two bottlenecks matter most:

  • Unified Memory Capacity vs Bandwidth: With 16GB of unified memory, context windows are not squeezed. You can load a quantized 3B or 7B model with an 8k-16k context window and still have ample RAM for Docker and OS services. However, the RK3588 uses a quad-channel 32-bit LPDDR4x bus (~34 GB/s theoretical, ~25-30 GB/s measured). In autoregressive CPU token generation, memory bandwidth is the primary ceiling.
  • Storage Ingestion (Samsung PM981a NVMe): Under direct I/O testing via fio, the M.2 PCIe 3.0 x4 slot delivered 2,862 MB/s sequential read and 197k 4K random read IOPS. Model weights load into system RAM in under a second (a 1.3GB model loads in ~0.6s).

2. Ollama & llama.cpp Inference Benchmarks (ARM64 CPU)

We tested Ollama (native ARM64 build) targeting the heterogeneous big.LITTLE topology (4x Cortex-A76 performance cores @ 2.26–2.4GHz + 4x Cortex-A55 efficiency cores @ 1.8GHz).

Prompt: Technical explanation of gradient descent and backpropagation (~200+ generated tokens).

Model Parameters Threading Configuration Eval (Generation) Rate Prompt Processing Rate TTFT (Time to First Token) Memory (RSS)
Llama 3.2: 1B 1.23B (Q4_K_M) Big Cores Only (4T) 14.62 tok/s 108.11 tok/s 425.5 ms ~1.3 GB
Llama 3.2: 1B 1.23B (Q4_K_M) All Cores Default (8T) 10.67 tok/s 79.91 tok/s 575.6 ms ~1.3 GB
DeepSeek-Coder: 1.3B 1.35B (Q4_0) Big Cores Only (4T) 16.90 tok/s 89.72 tok/s 1,025.5 ms ~1.4 GB
DeepSeek-Coder: 1.3B 1.35B (Q4_0) All Cores Default (8T) 4.52 tok/s 27.91 tok/s 3,295.7 ms ~1.4 GB
Qwen 2.5: 1.5B 1.54B (Q4_K_M) Big Cores Only (4T) 14.48 tok/s 70.24 tok/s 711.8 ms ~1.6 GB
Qwen 2.5: 1.5B 1.54B (Q4_K_M) All Cores Default (8T) 3.46 tok/s 42.33 tok/s 1,181.2 ms ~1.6 GB
Llama 3.2: 3B 3.21B (Q4_K_M) Big Cores Only (4T) 7.28 tok/s 28.13 tok/s 1,635.2 ms ~2.8 GB
Llama 3.2: 3B 3.21B (Q4_K_M) All Cores Default (8T) 1.99 tok/s 10.84 tok/s 4,245.2 ms ~2.8 GB
Phi-3 Mini: 3.8B 3.82B (Q4_K_M) Big Cores Only (4T) 6.56 tok/s 35.17 tok/s 909.8 ms ~3.1 GB
Phi-3 Mini: 3.8B 3.82B (Q4_K_M) All Cores Default (8T) 5.27 tok/s 32.97 tok/s 970.5 ms ~3.1 GB
Llama 3.1: 8B 8.03B (Q4_K_M) Big Cores Only (4T) 2.32 tok/s 10.24 tok/s 3,028.3 ms ~5.4 GB
Llama 3.1: 8B 8.03B (Q4_K_M) All Cores Default (8T) 2.10 tok/s 7.66 tok/s 4,044.8 ms ~5.4 GB

The big.LITTLE Scheduling Trap (+318% speedup with 4 threads)

  • Why num_thread: 4 is mandatory on RK3588: By default, Ollama spawns 8 threads across all cores. Because the 4 Little Cortex-A55 cores run at 1.8 GHz with smaller caches, thread barriers in llama.cpp cause severe synchronization stalls.
  • Restricting inference to the 4 Big Cortex-A76 cores yielded:
    • Llama 3.2: 1B: 10.67 -> 14.62 tok/s (+37%)
    • DeepSeek-Coder: 1.3B: 4.52 -> 16.90 tok/s (+274%, prompt rate +221%)
    • Qwen 2.5: 1.5B: 3.46 -> 14.48 tok/s (+318%)
    • Llama 3.2: 3B: 1.99 -> 7.28 tok/s (+265%, TTFT down from 4.2s to 1.6s)
    • Phi-3 Mini: 3.8B: 5.27 -> 6.56 tok/s (+24%)
    • Llama 3.1: 8B: 2.10 -> 2.32 tok/s (+10%, TTFT down by 1s)
  • The 8B limit: Running an 8B model on CPU is fundamentally memory-bandwidth bound. At ~5GB per token generation step, theoretical max is ~5 tok/s, making 2.32 tok/s the practical limit. It also pushed temperatures to 85.0°C uncooled.

3. CPU vs Hardware NPU (6 TOPS, 3 Cores)

Ollama compiles llama.cpp with ARM NEON SIMD instructions and runs 100% on the CPU. It does not touch the Rockchip NPU.

To test the 3-core 6 TOPS NPU, we compiled a native C++ runner (tools/rkllm_bench_v1) linked directly to Rockchip's librkllmrt.so runtime and kernel driver (/dev/rknpu_mem).

Metric / Dimension Ollama CPU Inference (ARM NEON) Rockchip NPU Hardware (RKLLM Runtime)
Compute Engine 4x Cortex-A76 @ 2.4GHz + 4x A55 @ 1.8GHz 3-Core Dedicated Neural NPU (6 TOPS INT8/INT4)
0.5B Model Eval ~20 - 24 tok/s 21.55 tok/s (Qwen 1.5 0.5B - Measured on-device)
1.3B - 1.5B Eval 16.90 tok/s (DeepSeek) / 14.48 (Qwen) ~16.69 tok/s (Qwen 2.5 1.5B - Reference Data)
3B - 4B Model Eval 6.56 tok/s (Phi-3) / 7.28 (Llama 3.2) ~7.45 tok/s (Phi-3 Mini 3.8B - Reference Data)
7B / 8B Model Eval 2.32 tok/s (Llama 3.1 8B) ~4.5 - 4.98 tok/s (Qwen 7B / ChatGLM - Reference Data)
CPU Utilization 100% Core Saturation (System frozen for other tasks) ~0% CPU Load (CPU 100% free for Docker/OS)
SoC Thermals Reaches 84.1°C – 85.0°C Runs drastically cooler (~60–68°C)
Model Ecosystem Any GGUF via Ollama / llama.cpp Requires .rkllm quantization via rkllm-toolkit

Key NPU trade-offs for homelab use:

  1. Zero CPU load: During NPU generation, CPU cores stay at ~0%. Home Assistant, Nextcloud, and other Docker containers remain fully responsive.
  2. Speedup on larger models: On 7B models, the NPU delivers ~4.8 tok/s vs 2.3 tok/s on CPU because dedicated matrix engines handle the tensor math without thrashing CPU caches.
  3. Sub-100ms latency: On compact models, Time to First Token (TTFT) drops to 96.4 ms on NPU.
  4. Format restriction: You cannot load arbitrary GGUFs; weights must be converted ahead of time to .rkllm using Rockchip's conversion toolkit.

4. Thermal Behavior & Power (Bare-Die / Uncooled Testing)

The standard retail package from Orange Pi is sold board-only (cooling accessories are sold separately as is standard for SBCs), so all tests evaluate out-of-the-box bare-die thermals on an open desk: - Idle (Ollama background daemon waiting): 52.7°C (~4–5W estimated SoC envelope) - Continuous 1B/3B Generation (4T Big Cores): 68–74°C (dissipating through PCB copper planes) - Sustained 8B Generation (8.03B params): Pushes the bare SoC directly to 84.1°C – 85.0°C (hitting the kernel DVFS limit). An aftermarket cooler or fan is required for sustained heavy loads. - Estimated wall power: ~12–16W under sustained multi-core inference.


5. Verdict: Is RK3588 Viable for Local AI?

Where it works well: - Background autonomous agents (summarizing feeds, home automation reasoning in Home Assistant, bot handlers) using Llama 3.2 1B, DeepSeek-Coder 1.3B, or Qwen 2.5 1.5B. - Low-latency function calling: at 14-17 tok/s, 1B models generate faster than reading speed. - Local embedding and vector search.

Where it falls short: - Running 8B+ models interactively (2.3 tok/s is too slow for back-and-forth chat). - Running without a heatsink under sustained compute.


6. Reproducibility & Test Scripts

All test scripts (tools/benchmark_ollama.py), raw JSON benchmark logs, and hardware configs are available in the repository: GitHub: Orange Pi 5 Plus Benchmarks

What models are you running on edge ARM boards? Anyone here running RKLLM in production vs pure llama.cpp?

submitted by /u/No-Doughnut6532
[link] [留言]

来源:r/LocalLLaMA · reddit.com