TiBLA-RTDETR

Primary checkpoint of TiBLA (Tibetan Book Layout Analysis) — an RT-DETR-l detector for the page layout of modern Tibetan books (headers, text area, footers, footnotes).

  • Base model / provenance: RT-DETR-l (Ultralytics), fine-tuned on the leak-free v4 tam2col split of TiBLAD.
  • License: AGPL-3.0 (inherited from the Ultralytics RT-DETR weights).
  • Dataset: BDRC/TiBLAD
  • Paper: buda-base/papers (papers/2026-tibetan-book-layout) — arXiv link forthcoming
  • Code: github.com/buda-base/tibla

This checkpoint is seed 0. Across five training seeds the paper reports mean canonical F1 0.961 ± 0.009 (unified scorer, per-seed operating point); at the validation-selected operating point used in the table below this seed scores 0.952.

Task

A 4-class detector — header, text-area, footer, footnote — kept as four classes at training time. Evaluation folds them into a 3-class canonical scheme: header+footer are combined into one header-footer class (matched individually, merged losslessly afterwards), text-area is merged to a single page/column envelope as a post-processing step (two boxes only on genuine two-column pages), and footnote is left as-is. All numbers below are in that canonical space, on the leak-free TiBLAD v4 833-page test set, unified scorer (pycocotools bbox mAP 0.50:0.05:0.95; F1 by greedy IoU≥0.5 at the best-mean-F1 operating point).

Inference

# pip install ultralytics
from ultralytics import RTDETR

model = RTDETR("tibetan_book_layout.pt")
# recommended per-class confidence thresholds (see below); predict at the floor
res = model.predict("page.jpg", imgsz=1024, conf=0.25)[0]
TH = {0: 0.60, 1: 0.55, 2: 0.25, 3: 0.60}  # header / text-area / footnote / footer
for cls, conf, xywhn in zip(res.boxes.cls.tolist(), res.boxes.conf.tolist(),
                            res.boxes.xywhn.tolist()):
    if conf >= TH[int(cls)]:
        print(res.names[int(cls)], round(conf, 3), [round(v, 4) for v in xywhn])

A ready-made infer.py (batch, YOLO-format output) is included in this repo.

Recommended confidence thresholds (per-class max-F1 operating points): header/footer ≈ 0.60, text-area ≈ 0.55, footnote ≈ 0.25. Footnote is deliberately kept low (recall-safe): the v4 test has only 38 footnote GT boxes, so a low threshold keeps recall near 1.0. Raising header/footer from 0.25 to 0.60 lifts precision +0.028 for a −0.014 recall cost; raising text-area from 0.25 to 0.55 (native) lifts precision +0.007 at no recall cost. If you prefer one global knob, the single validation-selected best-mean-F1 confidence is 0.64 (costs ≈0.008 mean F1 vs per-class tuning).

Evaluation (TiBLAD v4, 833-page test)

metric TiBLA-RTDETR TiBLA-PP-DocLayout-L TiBLA-RFDETR
license AGPL-3.0 Apache-2.0 Apache-2.0
base model RT-DETR-l (Ultralytics) PP-DocLayout-L (PaddleOCR, RT-DETR-L) RF-DETR-L (Roboflow)
mean F1 (canonical 3-class) 0.952 0.955 0.921
  header-footer F1 0.954 0.953 0.947
  text-area F1 0.999 0.998 0.996
  footnote F1 0.902 0.914 0.821
mean AP@0.50 0.974 0.959 0.925
mean AP@[0.50:0.95] 0.786 0.781 0.667
shared-class mAP@[.50:.95] (DocLayNet-aligned) 0.650 0.641 0.604
Hidden Trespass — header/footer 0.009 0.004 0.021
Hidden Trespass — footnote 0.043 0.037 0.178
COTe (Trespass) 0.975 (0.001) 0.978 (0.000) 0.974 (0.002)
operating confidence 0.64 0.61 0.47

"operating confidence" is the single global best-mean-F1 confidence, selected on the leak-free validation split and frozen for test (no test-set tuning). COCO AP rows are threshold-free (all detections above the fixed 0.05 floor).

Hidden Trespass = peripheral (header/footer/footnote) ground-truth area that survives in the actual OCR body crop C = E \ P, where E is the predicted text-area envelope and P is the union of the predicted peripheral boxes the pipeline subtracts; area-based, micro-averaged over the test set. Lower is better (less peripheral text bled into the OCR region). Formal definition in the paper.

Which checkpoint to pick

checkpoint license mean F1 shared mAP footnote HT
TiBLA-RTDETR (primary) AGPL-3.0 0.952 0.650 0.043
TiBLA-PP-DocLayout-L Apache-2.0 0.955 0.641 0.037
TiBLA-RFDETR Apache-2.0 0.921 0.604 0.178

RT-DETR-l leads on mAP, shared-class mAP and the 5-seed mean F1 (0.961 ± 0.009), but its weights are AGPL-3.0 (Ultralytics). If you need a permissive license, PP-DocLayout-L is an Apache-2.0 match (statistically on par on F1); RF-DETR is a lighter PyTorch-native Apache-2.0 option.

Citation

@misc{tibla2026,
  title        = {TiBLA: Tibetan Book Layout Analysis},
  author       = {Buddhist Digital Resource Center (BDRC)},
  year         = {2026},
  howpublished = {\url{https://github.com/buda-base/tibla}},
  note         = {arXiv link forthcoming}
}
Downloads last month
50
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train BDRC/TiBLA-RTDETR

Collection including BDRC/TiBLA-RTDETR