Abstract
Optical Character Recognition (OCR) plays an important role in automating the extraction of student data from answer sheets in large-scale educational assessments. Tesseract is one of the most widely used open-source OCR engines, yet systematic evaluation of its performance on structured answer-sheet data, together with independent external validation, remains limited in the Indonesian context. This study evaluated the accuracy of Tesseract OCR using 400 answer sheets collected from Mid-Semester Examination (UTS) documents of Universitas Gunadarma students. The documents were acquired using both flatbed scanners and smartphone cameras under varying acquisition conditions. A full factorial experiment was conducted on a stratified subset of 40 documents (5 preprocessing techniques × 4 image resolutions × 2 language models × 3 lighting conditions × 3 repetitions = 14,400 test cases), while the remaining 360 documents were reserved for independent external validation. Combined preprocessing (binarization, denoising, and 2x upsampling) achieved the highest accuracy (94.2%) at 200 DPI on the factorial subset, while external validation on the held-out 360 documents confirmed a comparable accuracy of 91.6% (CER = 8.4%), indicating reasonable generalization beyond tuning. Resolution was the most critical factor, with substantial gains between 100–200 DPI (approximately 19%) but diminishing returns beyond 200 DPI. Character-level analysis identified digit–letter confusion (8?B, 0?O, 1?I) as the dominant error source; Tesseract also showed a competitive accuracy-to-speed trade-off (0.9 seconds per page on CPU) against EasyOCR and PaddleOCR. These findings offer evidence-based guidance for OCR-based answer sheet processing in resource-constrained settings.
Downloads
Galleys
Published
Issue
Section
Copyright (c) 2026 Jurnal Studi Inovasi

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.


