← 开源
jiangxiluning

FOTS.PyTorch

FOTS Pytorch Implementation

Model DevelopmentArchitecturePython
在 GitHub 打开
增长势头
+024 小时新增 Star0.0%
641
Star
191
Fork
+0
本周
3
贡献者
创建于 2018-07-24 · 更新于 2026-08-30 · 今日第 15324 名
主要开发者
README

Introduction

This is a PyTorch implementation of FOTS.

  • [x] ICDAR Dataset
  • [x] SynthText 800K Dataset
  • [x] detection branch
  • [x] recognition branch
  • [x] eval
  • [x] multi-gpu training
  • [x] reasonable project structure
  • [x] wandb
  • [x] pytorch_lightning
  • [x] eval with different scales
  • [ ] OHEM

Instruction

Requirements

conda create --name fots --file spec-file.txt
conda activate fots
pip install -r reqs.txt

cd FOTS/rroi_align
python build.py develop

Training

# quite easy, for single gpu training set gpus to [0]. 0 is the id of your gpu.
python train.py -c pretrain.json
python train.py -c finetune.json

Evaluation

python eval.py 
-c finetune.json
-m 
-i 
--detection    
-o ./results
--cuda
--size "1280 720"
--bs 2
--gpu 1

with --detection flag to evaluate detection only or without flag to evaluate e2e

Benchmarking and Models

Belows are E2E Generic benchmarking results on the ICDAR2015. I pretrained on Synthtext (7 epochs). Pretrained model (code: 68ta). Finetuned (5000 epochs) model (code: s38c).

Name Backbone Scale (W * H) Hmean
FOTS (paper) Resnet 50 2240 * 1260 60.8
FOTS (ours) Resnet 50 2240 * 1260 46.2
FOTS RT (paper) Resnet 34 1280 * 720 51.4
FOTS RT (Ours) Resnet 50 1280 * 720 47

Samples

img_295.jpg img_486.jpg img_497.jpg

Acknowledgement