python - 使用 Python tesseract 从 PNG 图像中提取文本

Question

最近，我接了一个项目。使用 Python tesseract 将扫描的 PDF 转换为可搜索的 PDF/word。

经过几次尝试，我可以将扫描的 PDF 转换为 PNG 图像文件，之后，我很震惊，谁能帮我将 PNG 文件转换为可搜索的 Word/PDF。附上我的一段代码

请找到随附的图像以供参考。

Import os
Import sys
from PIL import image
Import pytesseract
from pytesseract import image_to_string

 Libpath =r'_______' #site-package
 Pop_path=r'_______' #poppler dlls
 Sys.path.insert(0,LibPath)

  from pdf2image import convert_from_path

     Pdfpath=r'_______' # PDF file directory
     imgpath=r'_______' #image output path

     images= convert_from_path(pdf_path = pdfpath, 
         dpi=500, poppler_path= pop_path)
      for idx, of in enumerate (images):
                 pg.save(imgPath+'PDF_Page_'+'.png',"PNG")
                 print('{} page converted'.format(str(idx)))

       try:
          from PIL import image
       except ImportError:
                 import image
         import pytesseract

     def ocr-core(images):
              Text = 
       pytesseract.image_to_string(image.open(images))
       return text
  print(ocr_core("image path/imagename))

就是这样，我已经写了......然后我得到了多个“.PNG”图像......现在我只能将一个 PNG 图像转换为文本。

如何转换所有图像并将其保存为 CSV/word？

score 1 · Accepted Answer

  from PIL import image
  from pdf2image import convert_from_path
  import pytesseract
  import OS
  import sys

   Pdf_file_path = '_______' #your file path

  Images = convert_from_path(Pdf_file_path, dpi=500)

Counter=1
for page in Images:
       idx= "image_"+str(Counter)+".jpg" ##or ".png"
       page.save(idx, 'JPEG')
       Counter = Counter+1

 file=Counter-1
  Output= '_____' #where you want to save and file name
 f=open(output, "w")
 for i in range(1,file+1):
          idx= "image_"+str(Counter)+".jpg" ##or ".png"         
 text=str(pytesseract.image_to_string(Image.open(idx)))
     f.write(text)
     f.close()

python - 使用 Python tesseract 从 PNG 图像中提取文本

1 回答 1

Related

Reference