Tiberian AI

We bring Geoffrey Khan’s reconstruction of the Tiberian pronunciation tradition of Biblical Hebrew into modern NLP — rule engines, distilled neural models, and interactive demos that map fully pointed Masoretic text (niqqud and teʿamim) to Tiberian IPA.

The Tiberian reading tradition is the oral system behind the vocalization and accent signs of the medieval Masoretes of Tiberias. Khan’s open-access volumes remain our linguistic source of truth; we encode that phonology in software and train models that imitate a carefully curated teacher, not a free-form “guess” at Biblical Hebrew sound.

What we ship

Artifact Role
tiberianai/tiberian-hebrew-ipa-byt5 ByT5 distillation: Masoretic Hebrew → Tiberian IPA
tiberianai/tiberian-hebrew-ipa-bhs Gated parallel corpus (manual approval; Dataset Viewer for approved users)
Demo Space ZeroGPU Gradio app for interactive transcription
Rule-based teacher Khan/Loder-aligned pipeline that labels training data

Authors: John Locke, Teodor Bors

Intended use

Linguistic foundation

Links

High automatic metrics mean the model matches the teacher IPA, not an independent human transcription study.