Every document reading tool looks great in the demo. Clean scan. Flat surface. Perfect lighting. Professional scanner output.
Then you try it on the document you actually have: a receipt photographed on a desk at 40 degrees, under fluorescent light, with a coffee stain in the corner. The tool reads maybe 70% correctly and hands you a wall of text to fix by hand.
This isn't a bug. It's a design choice. Most tools are built for scans, not photos.
Phone cameras shoot from an angle. The top of the document appears narrower than the bottom. Text lines converge. Traditional readers assume parallel lines and uniform character spacing — assumptions that fail immediately on a tilted photo.
Desk lamps create hot spots. Windows cast shadows. Fluorescent tubes produce banding. Glossy paper reflects the light source directly into the lens. Every variation degrades contrast and confuses character detection.
Receipts get crumpled. Contracts get folded. Business cards get bent in pockets. Wrinkles create false edges that look like character boundaries to algorithms expecting flat surfaces.
Photographing a screen introduces interference patterns. The pixel grid of the display interacts with the camera sensor, creating wavy artifacts that overlay the actual content. This is invisible to humans but devastating to automated readers.
Thermal receipts lose contrast over time. A receipt from two weeks ago may have text that's barely visible. Standard contrast thresholds miss it entirely.
The answer isn't "better character recognition." It's fix the photo first, then read it.
shenshen applies a preprocessing pipeline before any text extraction:
Only after these corrections does shenshen attempt to read the document. And it doesn't stop at text — it extracts structured fields (vendor, total, date, line items) so the output is ready to use, not another editing task.
| Photo condition | Traditional reader | shenshen |
|---|---|---|
| 40° angle | Garbled, misaligned rows | Straightened, fields extracted |
| Desk glare | Missing characters in highlight zone | Glare corrected, full extraction |
| Crumpled receipt | False line breaks, split words | Wrinkle compensated, clean output |
| Screen moiré | Wavy text, unreadable | Moiré filtered, accurate reading |
| Faded thermal | Blank or partial | Enhanced, full extraction |
Don't measure accuracy on clean scans. Measure it on the photos your team actually takes.
If your team photographs documents on desks, in cars, at conferences, or off screens — those are the conditions that matter. A tool that scores 99% on flat scans but 60% on phone photos is worse than useless in practice, because the 40% failure rate means someone is still doing the job by hand.
Take the worst document photo in your camera roll right now. Upload it to shenshen. Compare what you get against what your current tool produces. Free to start, no card required.
Try shenshen on your messiest photo →
Related: Receipt Scanner · Invoice Extractor · Business Card Scanner