How OCR runs inside the browser
This page uses Tesseract, the open source recognition engine, compiled to WebAssembly through tesseract.js. When you pick a language and press Extract text, the page starts a worker, downloads the trained data for that language, and reads the image entirely on your device. The trained data is the only thing fetched from the network, and it is cached afterwards; the image itself is never transmitted. Language packs are a few megabytes to around 15 MB each depending on the language, so the first run takes a little longer, and switching to a different language starts a fresh worker with the new pack. The status line shows the download and recognition progress as percentages.
Seven options are offered: English, Russian, Estonian, German, French, Spanish, and a combined English plus Russian mode for screenshots that mix both alphabets. Choosing the right language matters more than most people expect; running an Estonian document through the English pack produces garbage where the letters with diacritics appear.
What the output looks like
Recognised text appears in an editable box, with Tesseract's own line breaks, so a two column layout comes out as one column after the other rather than interleaved. From there you can copy it to the clipboard or download extracted-text.txt. It is plain text; bold, headings and tables are not preserved, and a table becomes rows of words separated by spaces. If the page finds nothing it says so rather than returning an empty file.
As a reference point, a 1920x1080 screenshot of an article page with normal 14 to 16 pixel text is typically read in 3 to 8 seconds on a laptop and comes out with very few errors. A 12 megapixel photo of a printed page takes longer, 10 to 30 seconds, and accuracy depends heavily on how the photo was taken.
Getting good results from photos
Tesseract was trained on scanned documents, so the closer your image is to a flat, evenly lit, high contrast scan, the better it does. Shoot straight on, not at an angle; a tilt of more than a few degrees drops accuracy sharply because the engine does not correct perspective. Fill the frame with the text so the characters are at least 20 to 30 pixels tall; tiny text in a wide shot is the most common cause of gibberish. Avoid shadows across the page and turn off the flash to prevent glare. Crop away everything that is not text on the crop tool; a busy photo border adds false characters. Handwriting is not supported in any practical sense; Tesseract is a printed text engine and will return a few random words from a handwritten note.
What it cannot do
There is no PDF input; take a screenshot of the PDF page or export it as an image first. There is no layout analysis beyond simple lines, so invoices and forms lose their structure. There is no automatic language detection. And there is no correction pass: characters such as 0 and O, or 1 and l, are confused occasionally in poor scans, so proofread anything that matters, especially bank details and reference numbers. Longer documents are better handled a page at a time; a 40 page contract in one image is not realistic.
Because the recognition never leaves the browser, the page is suitable for material you would not paste into an online service: contracts, medical letters, ID documents, internal screenshots. Once you have the text, the image itself may still carry hidden location data if it is a photo; the EXIF remover shows what is inside before you share it. If you later need the images bound into one document, the image to PDF tool does that, though the PDF it produces has no text layer, which is exactly why this OCR page exists.