AI & ML interests
Bringing Hebrew Tiberian tradition to the AI age.
Recent Activity
Tiberian AI
We bring Geoffrey Khan’s reconstruction of the Tiberian pronunciation tradition of Biblical Hebrew into modern NLP — rule engines, distilled neural models, and interactive demos that map fully pointed Masoretic text (niqqud and teʿamim) to Tiberian IPA.
The Tiberian reading tradition is the oral system behind the vocalization and accent signs of the medieval Masoretes of Tiberias. Khan’s open-access volumes remain our linguistic source of truth; we encode that phonology in software and train models that imitate a carefully curated teacher, not a free-form “guess” at Biblical Hebrew sound.
What we ship
| Artifact | Role |
|---|---|
tiberianai/tiberian-hebrew-ipa-byt5 |
ByT5 distillation: Masoretic Hebrew → Tiberian IPA |
tiberianai/tiberian-hebrew-ipa-bhs |
Gated parallel corpus (manual approval; Dataset Viewer for approved users) |
Demo Space |
ZeroGPU Gradio app for interactive transcription |
| Rule-based teacher | Khan/Loder-aligned pipeline that labels training data |
Authors: John Locke, Teodor Bors
Intended use
- Input: Biblical Hebrew with niqqud and cantillation (teĘżamim)
- Output: Tiberian IPA as defined by our teacher (forte–lene stream; perpetual Qere for the Tetragrammaton, etc.)
- Not for: Modern Hebrew, unpointed text, or non-Tiberian liturgical traditions
Linguistic foundation
- Geoffrey Khan, The Tiberian Pronunciation Tradition of Biblical Hebrew (Open Book Publishers; CC BY)
- Charles Loder — hebrew-transliteration & havarotjs
- Text: Biblia Hebraica Stuttgartensia / WLC-style pointed sources (gated; not freely redistributed)
Links
High automatic metrics mean the model matches the teacher IPA, not an independent human transcription study.