Skip to content

Latest commit

 

History

53 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Skein

skein-banner.svg

A self-learning, privacy-preserving toolkit for recognizing, classifying and extracting structure from messy text and records on the JVM. No GPU, no external AI, no neural networks — pure classical statistical machine learning, fully on the CPU.

The name (a coiled strand of yarn) reflects the purpose: untangling a skein of unstructured text into recognizable patterns. Each module is one thread that builds on the others.

Building blocks

  1. Shared text foundation — normalization, broken-word repair ("apart ment" → "apartment"), and a typed tokenizer producing pattern signatures (<word> <date> <numeric>).
  2. Classification — assign a label to a whole record, typo-tolerant, self-training, with privacy-preserving feature hashing so personal data is never stored in clear text. Predictions come with calibrated confidences, an abstain option, and an exact per-feature explanation of the score. Model quality is measurable: holdout and k-fold evaluation with per-class metrics.
  3. Extraction — pull structured values out of text via typed-token patterns and slot filling. A trained CRF token tagger can be saved and resumed.

A note on the privacy claim. Feature hashing keeps personal data out of classification models: the features are irreversible keyed hashes. The CRF tagger in skein-extract is different — it learns features keyed by the token text itself, so a saved CRF model contains fragments of the training text unless written with FeatureRetentionEnum.STRUCTURAL_ONLY. See skein-extract.

Module layout

                 skein-text  (pure foundation, zero heavy deps)
                  ╱        ╲
        skein-classify    skein-extract
              │
     skein-store-postgres
Module Published Responsibility
skein-bom yes Version alignment for consumers (Bill of Materials)
skein-text yes Shared text foundation: normalization, broken-word repair, typed tokenizer, pattern signatures
skein-classify yes Record → label classification (Naive Bayes + logistic regression, active learning)
skein-extract yes Text → structured field extraction (pattern DSL, slot filling, CRF token tagging)
skein-store-postgres yes Optional PostgreSQL persistence adapter (AES-256-GCM at rest)
skein-cli yes Command-line tools: active-learning data labeling and batch classification
examples no Runnable samples, including transaction categorization

End-to-end example

./gradlew :examples:run demonstrates classify → route → extract over bank transactions — see examples.

Publishing

Published modules ship a main jar, a sources jar and a Dokka-generated javadoc jar with full POM metadata (MIT license, SCM), to Maven Central (Sonatype Central Portal, GPG-signed via JReleaser) and GitHub Packages. skein-bom aligns all module versions for consumers.

Changelog

See CHANGELOG.md.

Toolchain

License

MIT.


About

A self-learning, privacy-preserving toolkit for recognizing, classifying and extracting structure from messy text and records on the JVM. No GPU, no external AI, no neural networks — pure classical statistical machine learning, fully on the CPU.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages