跳到正文
r/MachineLearning· /u/Ok_Cartographer5609·· 2 小时前AI 评分30

用 Rust 写的分块库 Chunkr,速度约为同类方案 20 倍

A chunking lib in Rust that is ~20x faster [P]

AI 导读

开发者用 Rust 构建了分块库 Chunkr,支持 Character、Recursive、Markdown header、Late chunking、Hierarchical chunking 等策略,并内置原生 PDF 加载器。

正文

Hey,

I wanted a faster chunking library for my system without affecting the overall accuracy. Did not find many options. So I've build https://github.com/d1pankarmedhi/chunkr

It has most of the chunking strategies like Character, Recursive, Markdown header, Late chunking, Hierarchical chunking, etc. It also supports native PDF loader, and other additional file types.

Some stats (MBA M4 16GB):

Test Case (matched parameters) Chunkr LangChain LlamaIndex Chonkie semchunk text-splitter
Recursive (1 MB, 1000/200) 2,264 MB/s 769 MB/s 10 MB/s 225 MB/s 42 MB/s 175 MB/s
Recursive (5 MB, 1000/200) 2,039 MB/s 696 MB/s — 201 MB/s 40 MB/s 46 MB/s
Fixed Char (1 MB, 1000/200) 750 MB/s 1.7 MB/s — 22 MB/s — —
Markdown (500 KB, 1000/150) 819 MB/s 67 MB/s 19 MB/s — — 40 MB/s
Python Code (200 KB, 1500/200) 3,232 MB/s 622 MB/s — — — 5.7 MB/s
Sentence (500 KB) 622 MB/s — 10 MB/s 20 MB/s — —
BPE Tokens (200 KB, cl100k_base, 512/50) 38 MB/s 43 MB/s 2.0 MB/s 151 MB/s — 7.2 MB/s
100 docs x 50 KB (parallel batch) 3,224 MB/s 679 MB/s — 213 MB/s — —

Extractor / Pipeline Latency Throughput Speedup vs PyPDF
Chunkr PDFLoader (Full Text) 747.9 ms 2,762 pgs/s 15.9x Faster
Chunkr PDFLoader (Page Documents) 721.0 ms 2,865 pgs/s 16.5x Faster
PyMuPDF (fitz) 2,616.8 ms 789.5 pgs/s 4.5x Faster
pypdf (pure Python) 11,900.5 ms 173.6 pgs/s 1.0x (baseline)
Chunkr End-to-End (PDF + Recursive) 798.1 ms 2,589 pgs/s 14.9x Faster
PyMuPDF + LangChain RecursiveTextSplitter 2,659.3 ms 776.9 pgs/s 4.5x Faster
pypdf + LangChain RecursiveTextSplitter 12,054.5 ms 171.4 pgs/s 1.0x (baseline)

Do check it out and share your feedback. Thanks!

submitted by /u/Ok_Cartographer5609
[link] [留言]

来源:r/MachineLearning · reddit.com