control-r3 — DETR object detection for news-page layout regions

This model is a fine-tuned DetrForObjectDetection (DETR with a ResNet-50 backbone, starting from facebook/detr-resnet-50) tuned to localize layout regions in historical and modern newspaper pages. It predicts bounding boxes for the 7 classes of the BigLam “Locating Objects Beyond Words” dataset:

class id label
0 Photograph
1 Illustration
2 Map
3 Comics/Cartoon
4 Editorial Cartoon
5 Headline
6 Advertisement

The base model facebook/detr-resnet-50 is released under the Apache 2.0 license, so this fine-tune can be freely shared and used.

Model Details

  • Model type: DETR (DEtection TRansformer), encoder-decoder Transformer with a ResNet-50 backbone.
  • Backbone: ResNet-50 (provided by timm), pretrained on ImageNet.
  • Number of queries: 100.
  • Input: RGB image resized/padded to a square of 600 × 600 after a smallest-max-size resize.
  • Output: per-image box predictions with class logits for the 7 layout classes (plus the implicit no object class).
  • Optimized by: fine-tuned on the training split (2,846 images) of biglam/loc_beyond_words.
  • Base model license: Apache 2.0.

Intended Use

The model is intended for document / newspaper-page layout analysis: given a scan or a page image, it detects coarse layout regions such as headlines, photographs/illustrations, advertisements, maps and comics. It is designed as a layout-region detector and is not meant for fine-grained text recognition or OCR (use OCR/HTR systems for reading text).

Example inference:

from transformers import pipeline

detector = pipeline("object-detection", model="harness-race/control-r3")
results = detector("path/to/newspaper_page.png")
# results: list of {label, score, box: {xmin, ymin, xmax, ymax}}

Training Data

  • Dataset: biglam/loc_beyond_words (BigLam “Locating Objects Beyond Words”, a Library-of-Congress-derived newspaper layout dataset).
  • Training split: 2,846 images.
  • Validation split: 712 images (reported results are on this held-out split).
  • Format: COCO-style bounding boxes [x, y, width, height] in pixel coordinates.
  • The data was repackaged (only renaming annotation fields for compatibility) into harness-race/loc_beyond_words_coco; no annotations were modified.

Training Procedure

The model was fine-tuned end-to-end (all weights trainable) with the Hugging Face Trainer-style loop on a single NVIDIA T4 GPU (fp16 AMP), with the RGB images resized and padded to 600 × 600 and light augmentation (horizontal flip, random brightness/contrast, hue/saturation, random crop with box clipping).

Hyperparameters

  • optimizer: AdamW (betas 0.9/0.999, eps 1e-8)
  • learning rate: 1e-4 with linear schedule + 10% warmup
  • weight decay: 1e-4
  • train batch size: 4
  • eval batch size: 8
  • epochs: 8
  • mixed precision: fp16 (native AMP)
  • image size: 600 × 600 (resize + pad)
  • gradient clipping: 1.0
  • seed: 42

Evaluation is run at the end of every epoch and the checkpoint with the best validation mAP is kept.

Evaluation Results

Reported on the validation split (712 images), using COCO-style metrics (torchmetrics.MeanAveragePrecision, box_format=xyxy). Metrics are in %:

Metric Value
mAP (IoU .5:.95) 24.65
mAP @ IoU 0.50 34.76
mAP @ IoU 0.75 28.39
mAR@100 37.94

Per-class mAP (IoU .5:.95):

Class mAP
Photograph 39.02
Illustration 1.17
Map 0.03
Comics/Cartoon 13.21
Editorial Cartoon 0.00
Headline 59.51
Advertisement 59.64

The model detects Headline and Advertisement regions very well (>59 mAP) and detects Photograph and Comics/Cartoon reasonably. The rare classes (Illustration, Map, Editorial Cartoon) show very low mAP, which is largely a consequence of heavy class imbalance in the dataset (e.g. only ~215 Map and ~293 Editorial Cartoon object instances across the whole dataset vs ~27.9k Headline instances). More data or class-balancing/oversampling for those classes would improve them.

The raw per-epoch metrics are stored in val_metrics.json in this repository.

Bias, Risks, and Limitations

  • Performance is strongly skewed by class imbalance; rare layout classes are unreliable.
  • Detector quality on page layouts/engravings/styles outside the training distribution may degrade.
  • Predictions should be reviewed when used in archival/curation workflows (layout regions may span multiple semantic zones; bounding boxes are approximate).
  • This is a layout-region detector, not an OCR/HTR model.

Technical Notes

  • Architecture: DetrForObjectDetection, 100 queries, 6 encoder + 6 decoder layers, d_model=256.
  • Repository files: config.json, model.safetensors, preprocessor_config.json (DetrImageProcessorFast, size=600), val_metrics.json, train_detr.py (training script).

Acknowledgements

Based on the DETR model (Carion et al., 2020) and the Transformers library. Dataset from BigLam and the Library of Congress newspaper collections.

Downloads last month
31
Safetensors
Model size
41.6M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for harness-race/control-r3

Finetuned
(806)
this model

Dataset used to train harness-race/control-r3

Paper for harness-race/control-r3

Evaluation results