#unicode-normalization #alignment #recomposition #unicode-text #text #decomposition #unicode #normalization

unicode-normalization-alignments

This crate provides functions for normalization of Unicode strings, including Canonical and Compatible Decomposition and Recomposition, as described in Unicode Standard Annex #15

1 unstable release

Uses old Rust 2015

0.1.12 Dec 30, 2019

#1088 in Text processing

Download history 52147/week @ 2024-12-14 23858/week @ 2024-12-21 29266/week @ 2024-12-28 51999/week @ 2025-01-04 57988/week @ 2025-01-11 50539/week @ 2025-01-18 59169/week @ 2025-01-25 61133/week @ 2025-02-01 65435/week @ 2025-02-08 65727/week @ 2025-02-15 79739/week @ 2025-02-22 76122/week @ 2025-03-01 87920/week @ 2025-03-08 78163/week @ 2025-03-15 71794/week @ 2025-03-22 58589/week @ 2025-03-29

309,557 downloads per month
Used in 172 crates (4 directly)

MIT/Apache

495KB
24K SLoC

unicode-normalization-alignments

Build Status Docs

This is a forked version of unicode-normalization wich provides alignment information during normalization.

Unicode character composition and decomposition utilities as described in Unicode Standard Annex #15.

This crate requires Rust 1.36+.

extern crate unicode_normalization_alignments;

use unicode_normalization_alignments::char::compose;
use unicode_normalization_alignments::UnicodeNormalization;

fn main() {
	assert_eq!(compose('A','\u{30a}'), Some('Å'));

	let s = "ÅΩ";
	let c = s.nfc().map(|(c, diff)| {
		match diff {
			0 => println!("Nothing changed here"),
			1 => println!("New character"),
			_ => println!("{} characters were removed", diff),
		}

		c
	}).collect::<String>();
	assert_eq!(c, "ÅΩ");
}

crates.io

You can use this package in your project by adding the following to your Cargo.toml:

[dependencies]
unicode-normalization-alignments = "0.1.12"

Dependencies