`stringi`

Fast and Portable Character String Processing in R (with the Unicode ICU)

A comprehensive tutorial and reference manual is available at https://stringi.gagolewski.com/.

Check out stringx for a set of wrappers around stringi with a base R-compatible API.

To learn more about R, check out Marek's open-access (free!) textbook Deep R Programming.

stringi (pronounced “stringy”, IPA [strinɡi]) is THE R package for string/text/natural language processing. It is very fast, consistent, convenient, and — thanks to the ICU – International Components for Unicode library — portable across all locales and platforms.

Available features include:

string concatenation, padding, wrapping,
substring extraction,
pattern searching (e.g., with Java-like regular expressions),
collation and sorting,
random string generation,
case mapping and folding,
string transliteration,
Unicode normalisation,
date-time formatting and parsing,

and many more.

Package Maintainer: Marek Gagolewski

Authors and Contributors: Marek Gagolewski, with contributions from Bartłomiej Tartanus and many others.

The package's API was inspired by that of the early (pre-tidyverse; v0.6.2) version of Hadley Wickham's stringr package (and since the 2015 v1.0.0 stringr is powered by stringi).

Homepage: https://stringi.gagolewski.com/

Citation: Gagolewski M., stringi: Fast and portable character string processing in R, Journal of Statistical Software 103(2), 2022, 1–59, https://dx.doi.org/10.18637/jss.v103.i02.

CRAN Entry: https://CRAN.R-project.org/package=stringi

System Requirements: R >= 3.4, ICU4C >= 61 (refer to the INSTALL file for more details)

License: stringi's source code is distributed under the open source BSD-3-clause license. For more details, see LICENSE.

This git repository also contains a custom subset of ICU4C source code which is copyrighted by Unicode, Inc. and others. A binary version of the Unicode Character Database is included. For more details on copyright holders, see LICENSE. The ICU project is covered by the Unicode license — a simple, permissive non-copyleft free software license, compatible with the GNU GPL. The ICU license is intended to allow ICU to be included in free software projects as well as in proprietary or commercial products.

Changes: see the NEWS file.

How to access the stringi C++ API from within an Rcpp-based R package

Name		Name	Last commit message	Last commit date
Latest commit History 1,691 Commits
.devel		.devel
.github/workflows		.github/workflows
R		R
datasets		datasets
docs		docs
inst		inst
man		man
src		src
tools		tools
.Rbuildignore		.Rbuildignore
.gitignore		.gitignore
.grepexclude		.grepexclude
.nojekyll		.nojekyll
CITATION.cff		CITATION.cff
CODE_OF_CONDUCT.md		CODE_OF_CONDUCT.md
DESCRIPTION		DESCRIPTION
INSTALL		INSTALL
LICENSE		LICENSE
Makefile		Makefile
NAMESPACE		NAMESPACE
NEWS		NEWS
README.md		README.md
TODO		TODO
cleanup		cleanup
configure		configure
configure.ac		configure.ac
configure.win		configure.win

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

Repository files navigation

`stringi`

Fast and Portable Character String Processing in R (with the Unicode ICU)

About

Uh oh!

Releases 39

Uh oh!

Contributors 14

Languages

License

gagolews/stringi

Folders and files

Latest commit

History

Repository files navigation

stringi

Fast and Portable Character String Processing in R (with the Unicode ICU)

About

Topics

Resources

License

Code of conduct

Uh oh!

Stars

Watchers

Forks

Releases 39

Uh oh!

Contributors 14

Languages

`stringi`