We bring Geoffrey Khan’s reconstruction of the Tiberian pronunciation tradition of Biblical Hebrew into modern NLP — rule engines, distilled neural models, and interactive demos that map fully pointed Masoretic text (niqqud and teʿamim) to Tiberian IPA.
The Tiberian reading tradition is the oral system behind the vocalization and accent signs of the medieval Masoretes of Tiberias. Khan’s open-access volumes remain our linguistic source of truth; we encode that phonology in software and train models that imitate a carefully curated teacher, not a free-form “guess” at Biblical Hebrew sound.
| Artifact | Role |
|---|---|
tiberianai/tiberian-hebrew-ipa-byt5 |
ByT5 distillation: Masoretic Hebrew → Tiberian IPA |
tiberianai/tiberian-hebrew-ipa-bhs |
Gated parallel corpus (manual approval; Dataset Viewer for approved users) |
Demo Space |
ZeroGPU Gradio app for interactive transcription |
| Rule-based teacher | Khan/Loder-aligned pipeline that labels training data |
Authors: John Locke, Teodor Bors
High automatic metrics mean the model matches the teacher IPA, not an independent human transcription study.