OCR Match Captcha

A TrOCR model fine-tuned to recognize short mathematical captcha expressions such as 26+7=? and 45-6=?.

Handwritten captcha Clean captcha
Distorted captcha Math captcha

Model details

  • Architecture: Vision Encoder-Decoder / TrOCR
  • Base model: microsoft/trocr-base-printed
  • Task: Image-to-text OCR
  • Target geometry: 130x30 RGB images
  • Output: Mathematical expressions without spaces

Training configuration

Parameter Value
Epochs 5
Learning rate 5e-6
Batch size 8
Gradient accumulation 2
Operator-token weight 3.0
Generation beams 4
Maximum output length 32

Usage

from PIL import Image
from transformers import TrOCRProcessor, VisionEncoderDecoderModel

repo_id = "arkhabbazan/ocr-match-captcha"
processor = TrOCRProcessor.from_pretrained(repo_id, token=True)
model = VisionEncoderDecoderModel.from_pretrained(repo_id, token=True)

image = Image.open("captcha.jpeg").convert("RGB")
pixel_values = processor(images=image, return_tensors="pt").pixel_values
generated_ids = model.generate(pixel_values, num_beams=4, max_length=32)
prediction = processor.batch_decode(
    generated_ids, skip_special_tokens=True
)[0]
print(prediction.replace(" ", ""))

Limitations

  • Training data is primarily synthetic.
  • Unseen fonts and layouts may reduce accuracy.
  • The model is intended for short expressions, not document OCR.
  • Usage must be authorized and compliant with applicable policies.
Downloads last month
159
Safetensors
Model size
0.3B params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for arkhabbazan/ocr-math-captcha

Quantized
(4)
this model

Spaces using arkhabbazan/ocr-math-captcha 2