Prism English (prism-en)
English UPOS tagging, morphological features, and lemmatization with
calibrated confidences, built for on-device, fully offline use.
A compact 18.0 M-parameter model on a ModernBERT-lineage (Ettin)
backbone β the same artifact contract and native runtimes as
prism-no, now for English.
This repository mirrors the versioned release artifacts of the Prism project. Prism ships native runtimes for Swift, C++, C, and Java/Kotlin that read these artifacts directly. Requires the language-independent Prism runtimes, library 0.6.0 or later.
Which folder to use
| Folder | Size | What it is |
|---|---|---|
prism-en-0.1.0/ |
β 94 MB | fp32 reference model |
This first English release ships a single precision (fp32). There is
no int8 -fast variant yet: ModernBERT's rotary-embedding indexing
does not lower cleanly through the int8 path, and a fast variant is
planned.
An application bundles the folder as-is. It is everything a Prism runtime needs; point the tagger API at its local path:
let tagger = try PrismTagger(artifactURL: artifactFolder) // Swift
prism::tagger::Tagger tagger("prism-en-0.1.0"); // C++
try (var tagger = PrismTagger.load(Path.of("prism-en-0.1.0"))) {} // Java
Quick starts for every binding: the
project README. The
artifact contract (programs, model.ptd weights, tokenizer and label
schemas, checksums) is documented in
docs/INTEGRATION.md.
Note: these are ExecuTorch
programs with the decoding policy and calibration baked in β not
transformers-loadable checkpoints. vocabulary.json is a standard
Hugging Face tokenizer.json (the ModernBERT tokenizer) and loads with
the tokenizers library.
Quality
Measured on the UD English-EWT test split with the official UD metrics, against UDPipe 2.17:
| Test F1 | prism-en 0.1.0 | UDPipe 2.17 |
|---|---|---|
| UPOS | 97.29% | 97.65% |
| UFeats | 97.71% | 98.16% |
| Lemmas | 97.36% | 97.99% |
Competitive on EWT web text β within ~0.4β0.6 pp of UDPipe β from an 18 M-parameter model that runs fully offline on a laptop CPU. The Ettin-400m teacher beats UDPipe on all three metrics; distilling down to the 18 M student gives up ~0.5 pp, a deliberate later-optimization target (0.1.0 ships the current student so English is available now).
Model description
An 18.0 M-parameter encoder student on
jhu-clsp/ettin-encoder-17m
(a 7-layer ModernBERT-lineage encoder, hidden 256), distilled from the
Ettin-encoder-400m teacher, with a character CNN feeding the morphology
and lemma heads, a structured morphology decoder, and per-head
temperature calibration. Label spaces: 18 UPOS Β· 21 morphology
features Β· 1,632 lemma edit rules; 50,368 byte-level BPE subwords.
Full technical reference:
docs/ARCHITECTURE.md.
Training data and attribution
- Gold treebank: UD English-EWT
(revision
c5baffde, UD 2.17) by the Universal Dependencies contributors β CC BY-SA 4.0. (The GUM treebank is deliberately excluded: its CC BY-NC-SA license is incompatible with this release.) - Silver corpora (unlabeled running text, labeled offline by the Prism teacher; the text itself is not redistributed): Project Gutenberg (public domain) and English Wikipedia (Wikimedia contributors β CC BY-SA 4.0).
- Backbone:
jhu-clsp/ettin-encoder-17mand its Ettin-encoder-400m teacher by the JHU CLSP group β MIT.
Pinned revisions and checksums travel inside the artifact
(manifest.json, LICENSES/).
License
Model weights: CC BY-SA 4.0 β the same share-alike license as the
UD English-EWT gold annotations. Using or bundling the unmodified
artifact β including commercially, in closed-source applications β is
fine (keep the LICENSES/ attribution); redistributed modified weights
must stay open. Prism source code is Apache 2.0.
- Downloads last month
- -