Turn Paper Receipts into a CSV with Free Tesseract OCR
Scan a shoebox of receipts with an open-source OCR engine on your own computer and export one spreadsheet-ready CSV for your budget.
Why this works
Receipts fade, and expense apps charge monthly fees. Tesseract is a free, open-source OCR engine that reads printed text well. A short script and a spreadsheet are enough to turn a folder of photos into a budget.
Step by step
- Photograph flat and well litPlace each receipt on a dark surface, shoot from above, and save as PNG or JPG.
- Install TesseractUse your package manager (for example
brew install tesseractor your Linux repository) and check withtesseract --version. - Extract the textRun
tesseract receipt.png stdout -l engfor each photo and save the output to a text file per receipt. - Pull out the fieldsAsk a local model, or a short script, to extract date, merchant, total, and category into CSV columns. Use the prompt below.
- Reconcile the totalCheck the extracted total against the printed total for every receipt. Fix the ones that differ.
Copy-paste prompt
From this OCR text of one receipt, return one CSV line with the columns: date (YYYY-MM-DD), merchant, total (numbers only, dot as decimal), currency, category (food, transport, household, health, other). If a field is not clear, write UNCLEAR in that cell. Output the header line once, then one data line. OCR TEXT: <paste>
Check your result
- Each total matches the printed receipt
- Dates are in one format
- UNCLEAR cells were fixed by hand
Pitfalls to avoid
- Thermal receipts fade over time; scan them soon after you get them.
- OCR confuses 0/O and 1/l. Check any total that looks wrong.
What we checked: Licence and project status were read from the project's GitHub repository on 8 Oct 2026. Install commands and flags change between releases, so follow the project README for the current version. Source: https://github.com/tesseract-ocr/tesseract.