A self-learning, privacy-preserving toolkit for recognizing, classifying and extracting structure from messy text and records on the JVM. No GPU, no external AI, no neural networks — pure classical statistical machine learning, fully on the CPU.
The name (a coiled strand of yarn) reflects the purpose: untangling a skein of unstructured text into recognizable patterns. Each module is one thread that builds on the others.
- Shared text foundation — normalization, broken-word repair (
"apart ment" → "apartment"), and a typed tokenizer producing pattern signatures (<word> <date> <numeric>). - Classification — assign a label to a whole record, typo-tolerant, self-training, with privacy-preserving feature hashing so personal data is never stored in clear text. Predictions come with calibrated confidences, an abstain option, and an exact per-feature explanation of the score. Model quality is measurable: holdout and k-fold evaluation with per-class metrics.
- Extraction — pull structured values out of text via typed-token patterns and slot filling. A trained CRF token tagger can be saved and resumed.
A note on the privacy claim. Feature hashing keeps personal data out of classification models: the features are irreversible keyed hashes. The CRF tagger in
skein-extractis different — it learns features keyed by the token text itself, so a saved CRF model contains fragments of the training text unless written withFeatureRetentionEnum.STRUCTURAL_ONLY. Seeskein-extract.
skein-text (pure foundation, zero heavy deps)
╱ ╲
skein-classify skein-extract
│
skein-store-postgres
| Module | Published | Responsibility |
|---|---|---|
skein-bom |
yes | Version alignment for consumers (Bill of Materials) |
skein-text |
yes | Shared text foundation: normalization, broken-word repair, typed tokenizer, pattern signatures |
skein-classify |
yes | Record → label classification (Naive Bayes + logistic regression, active learning) |
skein-extract |
yes | Text → structured field extraction (pattern DSL, slot filling, CRF token tagging) |
skein-store-postgres |
yes | Optional PostgreSQL persistence adapter (AES-256-GCM at rest) |
skein-cli |
yes | Command-line tools: active-learning data labeling and batch classification |
examples |
no | Runnable samples, including transaction categorization |
./gradlew :examples:run demonstrates classify → route → extract over bank transactions —
see examples.
Published modules ship a main jar, a sources jar and a Dokka-generated javadoc jar with full POM
metadata (MIT license, SCM), to Maven Central (Sonatype Central Portal, GPG-signed via JReleaser)
and GitHub Packages. skein-bom aligns all module versions for consumers.
See CHANGELOG.md.
- Kotlin 2.3 (K2), JDK 25, Gradle 9.6 (wrapper committed).
- Versions live in
gradle/libs.versions.toml. - Shared build config under
config/gradle.
MIT.