Turn a scanned Arabic PDF into copyable text
The file is open on your screen, the words are right there, and yet not one letter will highlight and search finds nothing in it. That is not a fault in your reader. The file is pictures rather than text. Drop the scan or the photo of the page into the tool below and the letters get read out of the image itself, so you get text you can copy or a searchable copy of your own file. All of it runs inside your browser on your own device, so nothing is uploaded to a server, we never see what is in the document, and there is no account and no daily cap.

Drop your files here
Files are processed on your device and never uploaded
Runs in your browser. Your file never leaves your device.
You can process several files at once, up to 3.
Why PDF to Word comes back empty on a scan
When you open a scanned file, try to highlight a line, and nothing highlights, you are looking at a file made of pictures rather than text. The scanner or the phone camera took an image of the page and put it inside a PDF wrapper, so the page looks written while the file itself does not know a single letter of what is on it. That is why converting to Word comes back empty on this kind of file: there is no text to carry across, and the tool does not invent text it did not find. The missing step is recognition, which means looking at the picture of the page and pulling the letters out of it one by one, so the text exists where it did not before. The tool on this page does that step inside your browser. And if a Word file is what you actually want, you do not need to come through here at all: drop the scan straight into the PDF to Word tool and it offers to read the file and build the Word document from what it reads, on the page you are already standing on.
How to tell your file is pictures before you try anything
The test takes two seconds and saves you a run that was never going to work. Open the file in any PDF reader and drag the cursor across a line of text the way you would to copy it. If the line highlights, the file has real text in it, and your fastest route is the extract text tool or PDF to Word directly, because both copy text that already exists and that is more accurate than any recognition and far quicker. If nothing highlights, or the whole page highlights at once as though it were a single image, the file is pictures and this page is where it belongs. There is a third case that catches people out: the mixed file, whose first pages are text and whose middle carries scanned pages inserted later, such as a typed contract with a photographed ID or a stamped page appended to it. For that one, set the page range to the scanned pages and read only those, which spares your device the work and gets you the result sooner.
Two different outputs, not one
The tool has an option called Output with two choices that suit genuinely different needs. Text to copy hands you the words of the pages as text, ready to paste into Word, into a message, or into a search box, and when the file has more than one page each page is separated from the next by a line carrying its label, so you can see where any passage came from. Searchable PDF hands you your own file looking exactly as it did, with the images, stamps, signatures and margins where they were, and an invisible text layer over every page that was read, so the document becomes searchable, selectable and copyable in any PDF reader. The first is for someone who wants the words themselves somewhere else. The second is for someone who needs the document to stay the document and simply become findable, which is what archiving wants, and what sending a file to an office that will search it later wants.
Why Arabic search works in the searchable copy
The invisible layer written over the page is ours rather than the recognition engine's, and that difference is not a technicality for programmers. When the engine builds a searchable PDF by itself, it writes Arabic letters in visual order, meaning the order the eye sees them arranged on the page, rather than the logical order in which a word is stored in any piece of text. The result is a file that looks perfectly correct and in which searching finds not one Arabic word, and from which copying yields scattered letters. So we take the word boxes from the engine along with their text in logical order, then draw every word ourselves at its own place on the page in an embedded Arabic font at full transparency. Search finds the Arabic word, and selecting and copying give back sound text you can paste anywhere. Rotated pages travel the same path, so the layer on a page turned ninety degrees lands where it belongs instead of mirrored.
Document language: why to leave Arabic and English alone
The default is Arabic and English together, which is right for most of the paperwork read in Saudi Arabia, so leave it unless you are certain the page is in one language. The reason is that picking a single language is slightly faster but turns everything on the page in the other language into random letters rather than into a blank you could ignore. A bank statement carries an IBAN in Latin characters, an invoice carries the tax number and the system name, a government form carries the platform title in English, and a certificate carries the university name in both. The page you take to be purely Arabic rarely is once you look at its details. The price of both languages together is that the second language model downloads once on the first run, and after that there is no speed difference worth talking about.
Twenty pages per run, and what to do with a longer file
One run reads up to twenty pages. That is a technical ceiling and not a commercial one, and no paid plan lifts it because nothing here is for sale: recognition is heavier by orders of magnitude than anything else on the site, since every page is rendered as an image at double resolution and then swept in full for letters, and all of it happens on your own processor rather than on a distant server. Twenty pages is what a mid range phone finishes without the browser killing the tab halfway through. A longer file has two routes and both are free and uncapped. The first is the page range box in the tool itself: read one to twenty, then run it again on the next range. The second is to split the file first with the split tool and read the parts. And here is a detail worth knowing before you worry about it: when you choose Searchable PDF for a PDF, the whole document comes back to you with all of its pages as they were, and the invisible layer sits on the pages that were read, so no page outside the range is lost.
A phone photo gets cleaned before it gets read
What lowers accuracy most is not the language or the typeface, it is the state of the picture. A page photographed with a phone usually comes out with a grey background, tilted by a few degrees, and carrying the shadow of a hand or of the device along one side. The engine tells ink from paper by the difference between them, so when that difference is a grey gradient rather than a sharp edge, the ends of letters and the diacritics dissolve and one dot becomes two. That is why the right order for a phone photo is to send it through the scan cleanup tool first: it puts the background back to white, brings the ink forward, and corrects a tilt of up to six degrees either way, which is the range a hand introduces while photographing, so recognition receives a page much closer to what a real scanner produces. The cleanup tool does not recognise text and does not claim to. It is a step before recognition, not a substitute for it. A page that is upside down or turned ninety degrees is not a tilt at all, and the rotate tool is what handles that.
Accuracy, honestly, and what you must check
Recognition is not copying, it is inferring a letter from its shape, which is why it is measured in likelihood rather than certainty. Clean printed Arabic or English reads well. Handwriting does not read, and nothing should be built on it. A page that was printed, then photographed on a phone, then sent through a chat app, then saved again loses a little of its ink definition at every step and arrives weaker than it began. And the engine's most dangerous mistakes are in what language itself does not guard: figures, names and account references. A wrong sentence draws attention because its meaning breaks, so it gets corrected. A wrong digit in an amount, an ID number or a date passes in silence and travels to whoever you sent it to. So there is one rule we ask you to keep: check every number and name in the output against the original page before you decide anything on it or send it anywhere.
The first run is slower, and after that nothing leaves your device
The first time you run recognition, the engine and the language model download, a few megabytes in total, and then stay cached in your browser so later runs start straight away. That download comes from our own servers rather than a public delivery network, and that is a deliberate rule rather than an accident: no request should go to a third party of any kind while somebody's document is open in the tab. The file itself is never uploaded under any circumstance. Pages are rendered and read inside the tab on your own processor, they never reach a server, we never see what is in them, and there is no copy on our side to keep or to delete. And you can check that rather than take our word for it: save the offline copy of the site from the work without internet page, cut the network once the engine has downloaded, and run recognition. It still works.
Related pages
Three steps, that's it
Drop the scanned file or the photo of the page into the tray below
Leave the language on Arabic and English, pick Text to copy or Searchable PDF, and set a page range if the file runs past twenty pages
Click extract, then check the numbers and names against the original page before relying on the result
Frequently asked questions
Why does PDF to Word not work on a scanned file?
Because a scanned file is pictures rather than text, so there is nothing in it to carry into Word. Recognition pulls the letters out of the image first, and then the text exists. You do not need two steps for it either: drop the scan into the PDF to Word tool and it offers to read the file and build the Word document from what it reads, on that same page.
How do I know if a PDF is a scan or real text?
Open it in any reader and try to highlight a line by dragging the cursor. If the line highlights, the file has real text and the extract text tool is faster and more accurate for you. If nothing highlights, or the whole page highlights at once, the file is pictures and it needs recognition.
Is my scanned file uploaded to a server to be read?
No. The recognition engine runs inside your browser and the pages are read on your own processor, so the file never leaves the device, we never see what is in it, and there is no copy of it on our side. We even host the engine on our own servers rather than a third-party source, so no request goes to anyone else while your document is being read.
How many pages can be read at once?
Up to twenty pages per run, which is a technical ceiling because recognition happens on your own processor and every page is rendered and swept in full. For a longer file, use the page range box and run it again on the next range, or split the file first with the split tool. No paid plan lifts this limit.
How accurate is Arabic OCR?
High on clean printed text, and noticeably lower on handwriting, tilted photos and low-quality images. Check figures and names in particular, because a wrong sentence gives itself away through its meaning while a wrong digit passes in silence. To improve the result, send a phone photo through the scan cleanup tool before reading it.
Can I get a Word file straight from a scanned PDF?
Yes. Drop the scan into the PDF to Word tool and it offers to recognise it and build a docx from the text it reads, inside your browser, under the same page ceiling. This page gives you the other two shapes instead: text you can copy, or a searchable PDF that keeps the document looking exactly as it does now.
Why does search not find Arabic words in a searchable PDF?
Because most tools write the text layer in visual letter order rather than logical order, so the file looks correct while search finds not one Arabic word in it. Here the layer is drawn by us from the word boxes with their text in logical order, so searching, selecting and copying all work in Arabic.
