OCR to pdf description

How can I attach an txt.file into the description of a document? . I want to use spotlight to find a document easier, like adding metadata into a file. I am extracting OCR of a pdf file as .txt file, using tesseract, to then add it in the info space of the original pdf (manually)

Posted on Mar 26, 2024 4:49 PM

Reply
Question marked as Top-ranking reply

Posted on Mar 27, 2024 8:09 AM

Finder and Preview.app GetInfo readily display PDF:Subject metadata (not XMP-dc:Subject or XMP-dc:Description).


exiftool -overwrite_original '-Subject<=-' cicero.pdf < cicero.txt 

exiftool -a -G1 -s -Subject cicero.pdf                            
[PDF]           Subject                         : Text from cicero.txt
[XMP-dc]        Subject                         : Text from cicero.txt


And also PDF:Author, Title, and Subject:


exiftool -overwrite_original -Author='Charles Dickens' -Title='A Tale of Two Cities' -Subject='A historical novel set during the French Revolution' cicero.pdf

exiftool -a -G1 -s cicero.pdf         
[PDF]           Author                          : Charles Dickens
[PDF]           Title                           : A Tale of Two Cities
[PDF]           Subject                         : A historical novel set during the French Revolution
[XMP-pdf]       Author                          : Charles Dickens
[XMP-dc]        Subject                         : A historical novel set during the French Revolution
[XMP-dc]        Title                           : A Tale of Two Cities


6 replies
Question marked as Top-ranking reply

Mar 27, 2024 8:09 AM in response to VikingOSX

Finder and Preview.app GetInfo readily display PDF:Subject metadata (not XMP-dc:Subject or XMP-dc:Description).


exiftool -overwrite_original '-Subject<=-' cicero.pdf < cicero.txt 

exiftool -a -G1 -s -Subject cicero.pdf                            
[PDF]           Subject                         : Text from cicero.txt
[XMP-dc]        Subject                         : Text from cicero.txt


And also PDF:Author, Title, and Subject:


exiftool -overwrite_original -Author='Charles Dickens' -Title='A Tale of Two Cities' -Subject='A historical novel set during the French Revolution' cicero.pdf

exiftool -a -G1 -s cicero.pdf         
[PDF]           Author                          : Charles Dickens
[PDF]           Title                           : A Tale of Two Cities
[PDF]           Subject                         : A historical novel set during the French Revolution
[XMP-pdf]       Author                          : Charles Dickens
[XMP-dc]        Subject                         : A historical novel set during the French Revolution
[XMP-dc]        Title                           : A Tale of Two Cities


Mar 27, 2024 7:23 AM in response to Matti Haveri

I have a three paragraph text file named cicero.txt in my home directory. I have a PDF on my Desktop named cicero.pdf containing the same text as the text file. Now, I want to use ExifTool to insert the contents of that text file into the XMP:descriptioin field of the PDF.


exiftool -overwrite_original '-Description<=-' cicero.pdf < ~/cicero.txt


One requires a PDF Editor to view this XMP description data, but I have written a shell script that extracts the XMP data from the PDF to stdout using ExifTool. The pipe to awk removes dozens of spaces at the end of the ExifTool output.


#!/bin/zsh

[[ "${1:e:l}" == "pdf" ]] || {echo "not a pdf."; exit 1}

XMPData="$(which exiftool) -XMP -b "${1}" | awk 'NF' "
[[ ${XMPData} ]] || {echo "no XMP data in PDF"; exit 1}
print ${XMPData}
exit 0



Mar 27, 2024 1:07 AM in response to ofranyut

exiftool can pipe tesseract output into file metadata (example below uses english but you can set it to use other OCR languages as well). But tesseract can not read .pdf. And do you want to manually limit what text snippet to use?


tesseract file.jpg - -l eng |exiftool -overwrite_original '-Description<=-' file.jpg


https://exiftool.org/forum/index.php?topic=15872.msg85337#msg85337

Mar 27, 2024 9:33 AM in response to Matti Haveri

Matti,


Fully aware of the standard PDF metadata fields that Finder and Preview can display. Have written tools in Python and Swift to do that as well. I don't see the advantage of sending three paragraphs of text from Cicero.pdf into the Subject metadata field though… 🧐


Yes, those standard metadata fields can be updated by Exiftool, but I think the place for description category content remains in the PDF XMP data container.

Mar 27, 2024 7:46 AM in response to VikingOSX

Or, without ExifTool, how to extract XMP metadata from a PDF (if it exists), and write to a text file on the Desktop. This assumes one has the command line tools for Xcode installed as it uses the /usr/bin/strings command:


#!/bin/zsh

# force absolute path
INFILE="${1:a:s/\~\//}"
OUTFILE="$HOME/Desktop/${INFILE:r:t}_meta.txt"

hasXMPData=$(egrep -a -c -m1 "x:xmpmeta" "${INFILE}")
(( $hasXMPData )) || {echo "PDF XMP data not found."; exit 1}

awk -v out="${OUTFILE}" '/<x:xmpmeta/,/<\/x:xmpmeta/ {print $0 > out}' <<<"$(strings "${INFILE}")"
exit 0


Usage: pdfxmp_meta.zsh ~/Desktop/cicero.pdf
Output: ~/Desktop/cicero_meta.txt



Mar 27, 2024 9:42 AM in response to Matti Haveri

https://medium.com/@vinuthomas/extracting-text-from-a-scanned-pdf-in-mac-bb4dda706119


I'm getting a txt file with all the writing of the scanned documents. I want add the text to original document, comments in the Getinfo menu, in order to do the file more searchable using spotlight, like scanning keyword on a document or mail. there is a way to avoid the limitation of words on comments, on the Get Info menu?

This thread has been closed by the system or the community team. You may vote for any posts you find helpful, or search the Community for additional answers.

OCR to pdf description

Welcome to Apple Support Community
A forum where Apple customers help each other with their products. Get started with your Apple Account.