#nlp #japanese #tokenizer #ngrams

tiniestsegmenter

Compact Japanese segmenter

3 releases (breaking)

0.3.0 Sep 24, 2024
0.2.0 Jun 24, 2024
0.1.1 May 11, 2024
0.1.0 May 11, 2024

#2020 in Text processing

Download history 50/week @ 2024-10-08 3/week @ 2024-10-15 3/week @ 2024-10-22 37/week @ 2024-10-29 58/week @ 2024-11-05 34/week @ 2024-11-12 36/week @ 2024-11-19 12/week @ 2024-11-26 40/week @ 2024-12-03 39/week @ 2024-12-10 6/week @ 2024-12-17

185 downloads per month

Custom license

58KB
2K SLoC

TiniestSegmenter

A port of TinySegmenter written in pure, safe rust with no dependencies. You can find bindings for both Rust and Python.

TinySegmenter is an n-gram word tokenizer for Japanese text originally built by Taku Kudo (2008).

Usage

Add the crate to your project: cargo add tiniestsegmenter.

use tiniestsegmenter as ts;

fn main() {
    let tokens: Vec<&str> = ts::tokenize("ジャガイモが好きです。");
}

No runtime deps