Hi everyone,
I have thousands of PDF documents that I’d like to use as a knowledge base for my Rovo agents.
I’ve tried importing them by converting the PDFs into pages, but the content is essentially vector-based images. As a result, the conversion produces pages full of images with no usable text. I’ve attempted different approaches (LibreOffice and other methods), but I haven’t been able to extract the text properly.
Would you recommend using OCR in this case? Or are there alternative approaches you’ve found effective?
Not being able to import PDFs as proper Confluence pages is quite limiting and is currently putting my “Single Source of Truth” initiative at risk.
Thanks in advance for any suggestions!
Best regards,
Giuseppe